What I see repeatedly is not a technology problem. The models work. What breaks is everything around the model.
Every agency I talk to has an artificial intelligence pilot running somewhere. Fewer of them have an AI system that has actually made it through a compliance review, into an authority to operate (ATO) package, and out the other side into daily use by a program office. That gap — between “we built a demo” and “we operationalized a system a federal customer can actually rely on” — is where most government AI initiatives quietly stall.
In my experience running an AI implementation firm that works with small and mid-size government contractors, what I see repeatedly is not a technology problem. The models work. What breaks is everything around the model: the audit trail, the data boundary, the explainability requirement, the vendor risk questionnaire that nobody scoped time for, and the assumption — often made by teams who built the pilot on a commercial cloud account — that the compliance review is a formality to handle after the fact rather than a design constraint from day one.
A few patterns show up often enough that they’re worth naming plainly.
First, pilots get built against the wrong environment. A team stands up a proof of concept using a public application programming interface (API) key, a personal or unmanaged cloud tenant, and whatever data is convenient to grab. It works beautifully in a demo. Then someone asks where the data actually lives, whether it crossed a boundary it shouldn’t have, and whether the model provider’s terms of service are even compatible with the agency’s data handling requirements. At that point the “pilot” isn’t six weeks from production — it’s a rebuild. The fix is boring but effective: Build the first version inside the actual authorization boundary, even if that means a slower, less flashy start. A prototype that has to be thrown away because it lived in the wrong environment costs far more time than starting inside the constraints.
Second, explainability gets treated as a documentation task instead of a design requirement. Compliance reviewers, inspector general offices, and program counsel increasingly want to know not just that a system produces a recommendation, but why, and what happens when it’s wrong. If a team bolts an explanation layer on after the model is already built, they usually discover the underlying architecture doesn’t support it — the model is a black box by construction, and there’s no clean way to expose the reasoning a reviewer needs. Systems that survive review are the ones where someone asked, “How will we explain this decision to an auditor?” before the first line of code, not after the first rejected package.
Third, the human-in-the-loop requirement gets satisfied on paper but not in practice. It’s easy to write “a human reviews all outputs before action” into a system design document. It’s much harder to build a workflow where that review is real — where the reviewer has enough context, time, and authority to meaningfully catch a bad output rather than rubber-stamping a queue of AI recommendations because the volume makes genuine review impractical. Reviewers routinely tell me their actual review time per item dropped once volume increased, which quietly turns a real safeguard into a theoretical one. Compliance officers are starting to ask pointed questions about actual review throughput, not just the existence of a review step, and systems need to be built to hold up under that scrutiny.
Fourth, vendor and data-handling documentation gets started too late. FedRAMP status, SOC 2 reports, subprocessor lists, data residency commitments — these take real lead time to assemble, and if a team waits until the compliance office asks for them, the AI capability sits in limbo for months while paperwork catches up to code that was otherwise ready. The teams that move fastest are the ones who treat the compliance package as a parallel workstream from day one, not as a gate at the end.
None of this argues against pursuing AI in government work — quite the opposite. The agencies and contractors I’ve watched succeed are the ones who accepted early that the compliance review is not an obstacle bolted onto the “real” engineering work; it is part of the engineering work. Building for FedRAMP boundaries, explainability, and genuine human review from the start is slower in week one and dramatically faster by month six, because there’s no rebuild waiting at the far end.
For small and mid-size government contracting firms in particular, this is also a competitive question. Prime contractors and agencies are increasingly wary of AI capabilities that look impressive in a sales demo but can’t produce an audit trail on request. A smaller firm that can show up with a system already designed around compliance requirements — rather than one that has to explain why it needs another two quarters to retrofit them — has a real advantage in a crowded field of AI vendors making similar promises.
The honest version of “AI adoption in government” isn’t a story about model capability. The models are, for the vast majority of realistic government use cases, already good enough. The real story is about whether the system around the model — the data boundary, the explanation layer, the review workflow, the paper trail — was built to survive contact with a compliance office. Firms and agencies that internalize that distinction early will spend less time explaining why their pilot never shipped, and more time running systems that actually did.
Patrick Bryant is the founder of Code and Trust.
Copyright
© 2026 Federal News Network. All rights reserved. This website is not intended for users located within the European Economic Area.

