What Agentic AI Deployment Actually Looks Like in Regulated Operations

02.08.26 03:54 AM By Karen Norden

By Malcolm Gill 

  

A city procurement lead I spoke with on a recent project put it bluntly: she didn't need another vendor telling her what the AI could automate. She needed to know what happens to a resident's case file when the system gets it wrong. 

 

That's the question that decide these deployments, and it's rarely the one vendors come prepared to answer.

 

I work on go-to-market for a frontier AI agent platform built for regulated, exception-heavy back office work, the kind where every transaction carries compliance exposure and a mistake doesn't surface until weeks later, usually as a complaint, an appeal, or a number that doesn't reconcile. The sales pitch is almost always about capability: automate the workflow, navigate the system of record, eliminate manual work. None of that is false. It's also not what gets a public sector buyer to sign.

 

Public sector buyers, and the operations leaders who serve them, aren't evaluating performance. They're evaluating accountability. Can the system recognize what it doesn't know. Is there a complete, reviewable record. Who is responsible when the model is wrong, and how fast does a resident, a patient, or a small business find out. These aren't compliance checkboxes. They're the conditions under which a public institution can use AI without breaking the public's trust in it.

 

That distinction matters more in this sector than almost any other. A private company that deploys a flawed automation absorbs the cost internally, maybe a write-off, maybe an awkward client call. A city that deploys one absorbs it as a resident who got denied a permit, a benefit, or a service, with no clear way to find out why or appeal it. I've sat through a demo where the vendor's answer to “what happens on an edge case” was a shrug and a promise to “tune the model.” That's an answer for a company optimizing a marketing funnel. It is not an answer for a system that decides whether someone's housing application gets reviewed this month or next.

 

The stakes here aren't just operational. They're about who gets to participate fairly in the systems a city runs, and that's a different burden than the one most automation vendors were built to carry.

 

I'll say the part that doesn't make for a clean vendor pitch: most automation providers aren't ready for this conversation, because their answer to “what happens when you're wrong” is thinner than their demo suggests. A platform that can't describe, specifically, how it surfaces  and escalates the cases it can't handle hasn't solved the public sector problem. It's solved the enterprise efficiency problem and is hoping nobody asks the harder question.

 

This is also why the conversation about agentic AI in cities can't stay contained to technology teams. Privacy officers, equity and access advocates, and front line staff are the ones who see where automated decisions land on real people. A frontline caseworker knows, in a way a vendor's roadmap never will, which residents are most likely to fall through the cracks of an automated process: the ones without reliable internet access, the ones who don't speak the dominant language fluently, the ones who never get a callback because the system marked their case resolved. The deployments that hold up over time are the ones where those voices were in the room before the contract was signed, not brought in afterward to manage a

 

If you're evaluating a vendor for this kind of work, here's a concrete test: ask them to walk through their last real exception, the case itself, not a hypothetical, and tell you who got notified, how long it took, and what changed afterward. If they can't produce one, that's the answer.

 

I don't think there's a tidy resolution to this, and I'd be cautious of anyone who offers one. Public trust in AI gets built case by case, appeal by appeal, not by a vendor's aggregate accuracy number. The providers worth taking seriously are the ones willing to walk through their failure cases in detail, unprompted, because that's the conversation a city needs to have before anyone touches a resident's file.


About the Author

Malcolm Gill is a Paris-based enterprise AI commercial strategist, advisor and Fractional Chief Revenue Officer (CRO) with more than 20 years of experience scaling enterprise software and AI businesses across Europe, the Middle East, Africa (EMEA) and North America. Having held senior commercial leadership roles with global organizations including Oracle, Hitachi Consulting, Slalom Build and NTT DATA, he specializes in helping AI companies accelerate growth by aligning strategy, go-to-market execution and revenue operations. Malcolm works with founders and executive teams to build scalable revenue architectures, strengthen market positioning and translate AI innovation into measurable business outcomes. A recognised thought leader, speaker and author, he publishes Strategic Signals, a newsletter exploring enterprise AI governance, commercialization and deployment, and regularly shares practical insights on the adoption of artificial intelligence across business and government. 

Learn more at malcolmgill.ai.