Between the ledger and the lever

Distilling high-level governance frameworks into concrete mechanisms starts with documenting internal principles and putting pen to paper.

While high-level governance frameworks recommend boundaries and actions, federal program managers and technologists are often left to interpret a framework’s key tenets into defined operating procedures. Principles often express organizational intent, but they don’t technically constrain a model’s behavior. For these principles to “have teeth,” they must be paired with guardrails that define the specific, discrete decisions and actions an AI system is authorized to take. 

“Put the guidance document down. It describes principles, and a principle doesn’t constrain anything,” explained Jamal Khan, chief growth and innovation officer at Connection Public Sector Solutions. “The work begins with an inventory of the decisions your program’s AI is going to make, and for each one, there are three questions you should ask: Can it be undone? What happens to a person if it’s wrong? And is there a cheap way to check the outcome?” 

Once those answers are established and mapped to specific priority levels and outcomes, Khan said the guardrails mostly write themselves. “They never write themselves from the memo alone,” he added. “If you cannot say what an assessor would collect to prove the control worked yesterday, you have a principle with a control’s name on it.” 

Redefine the federal ‘sandbox’ 

As recent industry incidents have demonstrated, highly capable agentic swarms are able to bypass environmental constraints, attempt privilege escalation and then alter their evaluation metrics when attempting to solve for assigned tasks. In short, AI models have learned to obfuscate the truth, or more simply put, lie. To account for this, agencies must build guardrails at the speed of the model they’re meant to govern. 

“A control that’s reviewed once a year against a model that changes every few weeks is decorative. If a control cannot be checked continuously, either redesign it until it can or move the decision it was protecting back to a human,” explained Khan. “Also, it’s important to provision governance before granting autonomy to these agents. Start by deploying them in advisory mode and then widen their permissions as your ability to observe them grows.” 

Provisioning also extends to the federal sandbox. It’s common for a sandbox and guardrails to be properly configured on paper only for a leak to occur days or months later.

“Any sandbox holding a capable agent should be evaluated the way you would evaluate a boundary facing an adversary, because from the agent side, that’s what it is,” Khan said. “Verify boundaries by trying to cross them. Do not just read the configuration report that says egress is blocked. Attempt egress from the inside and confirm it fails. Try to reach production data, internal services and the open internet, and write down what succeeded.” 

Following that logic, whatever grades, tests or approves an agent’s output should sit somewhere the agent cannot reach. An agent that can touch its evaluator, will, as evidenced by METR’s investigation into the Hugging Face breach. 

Grade external models for security and alignment 

Building AI models isn’t a solo effort; developers at OpenAI, Anthropic, Mistral and others will partner with trusted allies as they race to build the next model. Similarly, in the federal government, delivering services isn’t a one-horse race. Partnerships with third-party AI tools and companies can be helpful, but they should ultimately start with what Khan calls a “look under the hood.”

“If an agency cannot test, audit or interrogate the model that informs a government decision, then that tool should not be used in high-consequence decision-making, no matter how well it performed during the demonstration. An agency that deploys a model it cannot inspect is substituting the vendor’s assurance for its own ability to verify, and remains legally responsible for outcomes it cannot explain.” 

Another important factor to consider when onboarding a new technology is who retains control over the data. The contract should state in plain terms that the government data will stay under government control, with defined provisions for getting it back or deleting it if necessary. And if federal data moves around an external estate, leaders should know where their data has been, which sub-processor handled it and whether any of that information is used to feed the vendor’s AI model. 

As today’s frontier AI labs race toward artificial general intelligence, it may seem as if there’s a new model announced every week. AI capabilities often arrive inside products that already hold authorization. When working with partners, agencies should require advance notice of any model changes or new AI functions, with the ability to pin or decline a version. 

“Ask for your vendor’s own evaluations, red-team findings and incident history and then run your own tests against the data,” explained Khan. “A benchmark score describes a model plus a particular harness on a particular day, but it doesn’t describe your deployment.” 

The Connection Public Sector team works to bridge this gap between today’s AI vendors and the public sector by providing agencies with the vendor-neutral infrastructure layer needed to manage governance, align AI capabilities and scale AI responsibly. Through offerings like CNXN Helix, Connection Public Sector provides agencies with the information needed to make smart, responsible decisions that are mission-first and mission-focused.

Are you looking to integrate AI into your agency’s operations but unsure of where to start? See how Connection’s team of experts can guide you through the process.

This content is made possible by our sponsor Connection Public Sector. It is not written by and does not necessarily reflect the views of NextGov/FCW's editorial staff.

NEXT STORY: Agentic AI in Government: Moving from Experimentation to Operational Impact