CBP Translate
Department of Homeland Security · CBP
Overview
- Development stage
- DeployedSource: c) Deployed – The use case is being actively authorized or utilized to support the functions or mission of an agency.
- High-impact designation
- High-impactSource: a) High-impact
- Impact justification
- Not reported
- Start date
- 2019-08-07 (day precision)Source: 2019-08-07T00:00:00
- Withheld from public reporting
- Not reported
Mission
- Topic area
- Law Enforcement
- Operational functions
- Translation and transcription Derived by keyword rules; see methodology
Problem the AI is intended to solve
Assist officers and agents with immediate interpretation needs when human translators are not available.
Expected benefits
CBP Translate enhances efficiency by expediting questioning when immediate interpretation is needed. It ensures clear communication, minimizes misunderstandings, and offers immediate accessibility via mobile and web platforms. This improves operational flexibility and creates a smoother experience for travelers.
System outputs
The outputs of CBP Translate include translated text or audio in the form of chat bubbles, which store each interaction. Additionally, CBPOs can capture images of non-travel documents for text translation, but images of actual travel documents are not taken.
Technology
- AI classification
- Natural language processingSource: Natural Language Processing: AI that processes, interprets, and shares information in human language.
- System name(s)
- CBP Translate
- Custom-developed code
- Yes
- Public source code
- Not reported
Sourcing
- How it was built
- Contract and in-houseSource: c) Developed with both contracting and in-house resources
- Vendor (as reported)
- Aneesh Technologies, 24X7, Ellumen Inc., Deloitte, NiyamIT
- Vendors (standardized)
- Deloitte
A vendor is the supplier named by the agency. It does not identify the underlying model or AI technology.
Data and privacy
- Involves PII
- Yes
- Privacy Impact Assessment
- www.dhs.gov ↗
- Authorization to Operate
- Yes
- Demographic features
- Not reported
- Training and evaluation data
- The models are trained using examples of translated sentences and documents, which are typically collected from the public web. A data miner that focuses more on precision than recall is used, which allows the collection of higher quality training data from the public web.
- Federal Data Catalog
- Not reported
Governance
8 of 8 minimum-practice questions answered. A blank answer means the agency reported nothing; it does not mean the practice is absent.
Potential impacts and how they were identified
The key risks would be the programs inability to accurately translate what was spoken by both sides of the conversation, leading to significant delays in emergency response situations when trying to leverage traditional phone based translation services in areas with limited cell phone reception. Inaccuracy may also lead to longer processing times at Ports of Entry. These were identified via feedback from the end-users and a common understanding regarding LLM language translation models.