Open Source and Social Media Analysis
Department of Homeland Security · CBP
Overview
- Development stage
- DeployedSource: c) Deployed – The use case is being actively authorized or utilized to support the functions or mission of an agency.
- High-impact designation
- High-impactSource: a) High-impact
- Impact justification
- Not reported
- Start date
- 2025-01-01 (day precision)Source: 2025-01-01T00:00:00
- Withheld from public reporting
- Not reported
Mission
- Topic area
- Law Enforcement
- Operational functions
- Translation and transcription, Image and video analysis Derived by keyword rules; see methodology
Problem the AI is intended to solve
The AI is intended to solve the problem of efficiently identifying potential threats and admissibility concerns by quickly analyzing vast amounts of open-source and social media data for security risks to enhance U.S. national security. This tool then presents information to a CBP Officer/analyst for manual review, verification and validation for violations of Title 8 and Title 19 or other laws that CBP is sworn to enforce. The output is not used as the sole basis for action or decision making.
Expected benefits
CBP uses this tool to conduct targeted queries to aid CBP in open-source research to monitor potential threats or dangers or identify travelers who may be subject to further inspection for violation of laws CBP is authorized to enforce or administer.
System outputs
This tool utilizes AI modules for Text detection and translation as well as object and image recognition to provide analysts with possible matches to manually review in a single interface versus doing multiple manual queries. The output is not solely used for action or decision making and are used to identify additional Open Source or Social Media of a person or identify additional selectors (such as phone and emails) that are previously unknown to CBP and compared by an analyst against Government systems to identify additional derogatory information.
Technology
- AI classification
- Natural language processingSource: Natural Language Processing: AI that processes, interprets, and shares information in human language.
- System name(s)
- Not reported
- Custom-developed code
- No
- Public source code
- Not reported
Sourcing
- How it was built
- Purchased from vendorSource: a) Purchased from a vendor
- Vendor (as reported)
- NexisXplore
- Vendors (standardized)
- None on the standard list
A vendor is the supplier named by the agency. It does not identify the underlying model or AI technology.
Data and privacy
- Involves PII
- No
- Privacy Impact Assessment
- www.dhs.gov ↗
- Authorization to Operate
- No
- Demographic features
- Not reported
- Training and evaluation data
- Training data was collected from several publicly available, social media, and media outlet sites. This approach ensured the model was trained across several different groups representing an array of possible language types and vernaculars so as not to cause bias toward a specific demographic. Along with the above open-source data, the vendor leverages a mix of proprietary data to ensure the data is representative of real-world conditions and context.
- Federal Data Catalog
- Not reported
Governance
8 of 8 minimum-practice questions answered. A blank answer means the agency reported nothing; it does not mean the practice is absent.
Potential impacts and how they were identified
AI could potentially mis-label an object however all results are reviewed by a law enforcement officer and OSINT results are only one section of data among many when reviewing admissibility.