Methodology and data quality
How the data was loaded, cleaned and analysed, what each measure means, and what this platform cannot tell you.
Data source and refresh status
- Dataset
- 2025 Federal Agency AI Use Case Inventory (individually reported use cases)
- Publisher
- Office of Management and Budget (OMB) under OMB Memorandum M-25-21
- Official source
- https://github.com/ombegov/2025-Federal-Agency-AI-Use-Case-Inventory
- Source file
- 2025_individually_reported_AI_use_cases.xlsx
SHA-256 62d5baa65fa3205972cbfe59… - Update mode
- Static snapshot. This is not a live feed. OMB publishes inventories once per reporting cycle; this platform shows the validated file listed here.
- Last successful processing
- 2026-10-11 16:12 UTC
- Dataset version
- 2025-inventory@62d5baa65fa3
- Additions or changes
- Not applicable: one snapshot loaded, no prior version to compare.
A new inventory is loaded by running the ingestion pipeline on the new file. The pipeline validates the schema, normalises with versioned rules, and refuses to publish if headers do not match. Version-to-version change logs are designed but not yet built (see docs/ARCHITECTURE.md).
Data-quality summary
- Worksheet
- Consolidated Inventory · 36 columns · 3,613 rows including headers
- Records retained
- 3,611 from 41 agencies
- Blank or repeated header rows
- 0 blank, 0 repeated headers
- Missing agency IDs
- 704 (every record has an internal ID)
- Duplicated agency IDs
- 28 records sharing 13 values
- Start dates
- blank 2209, parsed 1399, out_of_range 3
- Date precision
- month 484, day 879, year 36
- Governance answers flagged
- 75 (wrong-question options or free text)
- Contact emails
- 2496 in source, 0 in this platform
The full report, with completeness for every field and every category distribution, is in docs/DATA_QUALITY_REPORT.md.
Definitions and calculation rules
Use-case record. One row of the inventory. Agencies split systems into use cases differently, so counts compare reporting, not AI capacity.
Lifecycle. From the stage answer. Blank stage is shown as "Status not reported" and never merged into another stage.
High-impact. The agency's own designation under M-25-21, not an independent risk rating.
Disclosure completeness. Share of 12 core fields filled per record, averaged per agency. It measures reporting, not quality of the AI.
Agency comparisons. Agency figures aggregate many records. The agency is the unit of comparison; records are not treated as independent observations of agency behaviour, and no statistical tests are run.
Governance statuses. Reported in place; Reported in progress; Reported not in place; Not applicable; Precluded by law or guidance; Waived by CAIO; Other or ambiguous answer; No answer reported. Blank is never read as negative. Ambiguous answers are kept as "other" and flagged.
Governance denominators. Default scope is records designated high-impact, the population the minimum practices apply to. Every governance view shows its scope and counts for each status. "In place of applicable" divides by scope minus not applicable, precluded and waived.
Operational functions. Keyword rules on name, problem and outputs (18 tags). Rules are in pipeline/normalize.py. They miss some records and catch some wrongly; use them to explore, not to count precisely.
Vendors. Free-text field matched against a list of >60 companies. 78% of records name no vendor.
No governance maturity score. The fields are too sparse and too dependent on designation choices for a defensible composite. The platform shows a multidimensional profile instead.
Cross-agency similarity
Score = 0.85 × text similarity + 0.15 × categorical agreement. Text similarity is TF-IDF cosine on problem (0.4), outputs (0.4), benefits (0.1) and data description (0.1), reweighted over fields both records report. Categorical agreement: same topic (0.4), same AI classification (0.25), shared function tag (0.35). Only pairs from different agencies scoring at least 0.2 are kept, top 10 per record. Bands: strong ≥ 0.50, moderate ≥ 0.35, weak below. Bands were set by reading samples, not by a labelled benchmark. No external model or API is used.
Data dictionary
43 fields| Internal field | Source header | Type | Filled | Normalisation and notes |
|---|---|---|---|---|
| agency_abbr | Agency Abbreviation | text | 100% | Agency abbreviation. Used as the agency key; names are consistent per abbreviation. |
| agency_name | Agency Name | text | 100% | Full agency name. |
| use_case_id | Use Case ID [Agency Abbrev.] – [#] | text | 80.5% | Agency-assigned ID. 704 missing and some duplicated or placeholder values; every record also gets a stable internal ID (rid = 'R' + worksheet row). |
| name | Use Case Name | text | 100% | Use case name as reported. |
| bureau | Bureau/Component | text | 98.4% | Bureau or component. |
| contact_email | Email Address | text | EXCLUDED at ingestion. Never written to processed data, search, exports or analyst answers. | |
| withheld | Should this AI use case be withheld from public reporting? | category | 68.7% | Withheld-from-public-reporting answer: a) no, b) risk to disclosure, c) prohibited by law, d) other. |
| stage | Stage of Development | category | 90.6% | a) pre-deployment, b) pilot, c) deployed, d) retired; blank becomes 'unknown'. |
| high_impact | Is the AI use case high-impact? | category | 88.2% | a) high-impact, b) presumed high-impact but determined not, c) not high-impact; blank becomes 'unknown'. |
| impact_justification | Justification | text | 22.9% | Justification for the high-impact determination. |
| topic | Use Case Topic Area | category | 83.3% | Topic area. 'Administrative functions' merged with 'Administrative Functions'; 'Other – X' variants fold into 'Other' with topic_detail. |
| ai_class | AI Classification | category | 81.7% | AI classification, mapped from the leading phrase of the answer option. |
| problem | What problem is the AI intended to solve? | text | 83.6% | Problem the AI is intended to solve. |
| benefits | What are the expected benefits and positive outcomes from the AI for an agency’s mission and/or the general public? | text | 81.8% | Expected benefits and outcomes. |
| outputs | Describe the AI system’s outputs. | text | 78.8% | Description of system outputs. |
| start_date | Date when AI use case became operational or the pilot’s start date | date | 38.8% | Operational or pilot start date. Mixed formats parsed to ISO with day, month or year precision; original kept; years outside 1980 to 2026 flagged. |
| sourcing | Was the system involved in this use case purchased from a vendor or developed under contract(s) or in-house? | category | 44.4% | a) purchased from vendor, b) in-house, c) both; blank becomes 'unknown'. |
| vendor | Vendor(s) Name | text | 21.9% | Free-text vendor names. Matched to a standard company list (vendors[]); original kept. |
| ato | Does this AI use case have an associated Authorization to Operate (ATO)? | yes/no | 42.7% | Authorization to Operate. |
| system_names | System(s) Name | text | 25.1% | System name(s). |
| training_data | Describe any data used to train, fine-tune, and/or evaluate performance of the model(s) used in this use case. | text | 35.7% | Data used to train, fine-tune or evaluate. |
| data_catalog | If the data is required to be publicly disclosed as an open government data asset, provide a link to the entry on the Federal Data Catalog. | url/text | 5.7% | Federal Data Catalog link; first URL extracted. |
| pii | Does this AI use case involve personally identifiable information (PII) that is maintained by the agency? | yes/no | 41.8% | Involves PII maintained by the agency. |
| pia | If publicly available, provide the link to the AI use case’s associated Privacy Impact Assessment (PIA). | url/text | 6.5% | Privacy Impact Assessment link; first URL extracted. |
| demographics | Which, if any, demographic variables does the AI use case explicitly use as model features? | category list | 30.6% | Demographic variables used as model features; parsed to a list of variable keys plus a status. |
| custom_code | Does this project include custom-developed code? | yes/no | 47.4% | Includes custom-developed code. |
| code_link | If the code is open source, provide the link for the publicly available source code. | url/text | 6.9% | Open-source code link; first URL extracted. |
| gov_testing | Has pre-deployment testing been conducted for this AI use case? Practice: Complete AI Impact Assessment | governance | 7.1% | Pre-deployment testing. Status vocabulary: affirmative, in_progress, negative, not_applicable, precluded, waived, other, blank. |
| gov_impact_assessment | Has an AI impact assessment been completed for this AI use case? Practice: Complete AI Impact Assessment | governance | 6.5% | AI impact assessment completed. |
| impacts_text | What are the potential impacts of using the AI for this particular use case and how were they identified? Subpractice: Complete AI Impact Assessment | text | 6.8% | Potential impacts and how they were identified. |
| gov_independent_review | Has as independent review of the AI use case been conducted? Sub practice Complete AI Impact Assessment | governance | 7.4% | Independent review. 'Yes' by another office, oversight board or CAIO all map to affirmative; CAIO waiver maps to waived. |
| gov_monitoring | Is there a process to conduct ongoing monitoring to identify any adverse impacts to the performance and security of the AI functionality, as well as to privacy, | governance | 7.4% | Ongoing monitoring process. |
| gov_training | Has the agency established sufficient and periodic training for operators of the AI to interpret and act on the its output and managed associated risks? Practic | governance | 7.8% | Operator training. Answers copied from the monitoring question are flagged and classed as 'other'. |
| gov_failsafe | Does this AI use case have an appropriate fail-safe that minimizes the risk of significant harm? Practice: Provide Additional Human Oversight, Intervention, and | governance | 8.1% | Fail-safe that minimises risk of significant harm. |
| gov_appeal | Is there an established appeal process in the event that an impacted individual would like to appeal or contest the AI system’s outcome? Practice: Offer Consist | governance | 8% | Appeal process. 'Law, operational limitations or guidance precludes' maps to precluded. |
| gov_feedback | What steps has the agency taken to consult and incorporate feedback from end users of this AI use case and the public? Practice: Consult and Incorporate Feedbac | governance | 7.3% | Consultation with end users and the public. Lettered options a to c and recognised free text map to affirmative; 'd) Other' and unrecognised free text map to other. |
| rid | (derived) | id | Stable internal record ID from the worksheet row. | |
| source_row | (derived) | integer | Worksheet row number in the source file (provenance). | |
| id_status | (derived) | category | ok, missing, duplicate or placeholder. | |
| vendors | (derived) | list | Standardised company names matched in the vendor text. | |
| functions | (derived) | list | Operational-function tags from keyword rules on name, problem and outputs. | |
| completeness | (derived) | 0 to 1 | Share of 12 core fields filled: stage, high-impact, topic, AI class, problem, benefits, outputs, sourcing, start date, ATO, PII, custom code. | |
| original | (derived) | object | Original source values for every normalised category field. |
Agency disclosure completeness
Mean share of 12 core fields filled| Agency | Records | Completeness |
|---|---|---|
| DOC | 223 | 0.9% |
| TVA | 59 | 16.5% |
| CFTC | 3 | 16.7% |
| ED | 56 | 16.7% |
| NSF | 17 | 25% |
| GSA | 49 | 33.3% |
| PBGC | 17 | 41.7% |
| FCA | 3 | 55.6% |
| VA | 367 | 57.4% |
| NEA | 4 | 58.3% |
| STB | 2 | 58.3% |
| STATE | 60 | 61.8% |
| FDIC | 50 | 62.3% |
| NASA | 425 | 62.6% |
| FCC | 6 | 65.3% |
| SEC | 60 | 66.1% |
| EPA | 29 | 68.4% |
| DOT | 70 | 69.6% |
| DOL | 41 | 71.7% |
| TREAS | 129 | 71.8% |
| DHS | 238 | 73.2% |
| HUD | 11 | 77.3% |
| USDA | 162 | 77.4% |
| HHS | 447 | 77.5% |
| FHFA | 16 | 79.2% |
| DOJ | 314 | 80.8% |
| DOE | 340 | 83% |
| FERC | 6 | 83.3% |
| FRB | 38 | 83.6% |
| NCUA | 6 | 84.7% |
| DOI | 247 | 90.5% |
| NIGC | 5 | 91.7% |
| OSC | 1 | 91.7% |
| EAC | 4 | 93.8% |
| NTSB | 4 | 93.8% |
| SSA | 33 | 99% |
| FTC | 16 | 100% |
| NARA | 14 | 100% |
| NRC | 4 | 100% |
| OSHRC | 1 | 100% |
| SBA | 34 | 100% |
Feature tiers
Preview mode: all features available| Feature | Planned tier |
|---|---|
| Executive overview | Free |
| Use-case explorer | Free |
| Agency profiles | Free |
| Governance analytics | Free |
| Technology landscape | Free |
| Multi-agency comparison | Professional |
| Cross-agency discovery | Professional |
| Cross-tabulation workspace | Professional |
| Printable reports and briefings | Professional |
| CSV export of filtered records | Professional |
| Question-based analyst | Professional |
| JSON API access | Institutional |
| Full-dataset bulk export | Institutional |
There is no login or payment yet. Tiers are configuration only, enforced server-side when switched on.
What this platform does not do
- It does not verify agency answers or audit any system.
- It does not show or infer budgets, spending, accuracy, performance, return on investment or deployment success. The inventory contains none of these.
- It does not label any agency or system compliant, non-compliant, safe or unsafe.
- It does not identify the underlying AI model from a vendor name or classification.
- It does not declare that two systems can be shared; discovery produces candidates for human review.