Two things often seem true when it comes to revolutionary technologies: Organizations rush to implement them, and their potential initially exceeds their impact. That’s the case so far with artificial intelligence, which many enterprises have aggressively adopted without yet seeing meaningful results. The solution may be stronger controls.
As Peter Senge famously noted in his seminal systems thinking book The Fifth Discipline (Crown Currency, March 2006),1 the optimal rate of growth is significantly less than the fastest possible rate. That’s because pushing a complex system such as that needed to fully deploy AI too hard can create bottlenecks and rework, while embedding the right controls early can increase overall speed.
That’s especially true for the public sector, where leaders are typically forced to choose between caution and speed, because slowing down—through the monitoring and management of AI—tends to happen at the end of processes rather than being embedded throughout. The result is stop-and-start activity: reacting to events rather than catching them early, adjusting, and maintaining momentum.
In this second article of our series on rewiring the public sector, we examine the importance of having strong controls that build the trust organizations need to maintain momentum. Our belief is that the public sector agencies most likely to pull ahead in the next decade through the use of AI may be not the boldest or most careful but the most thoughtful—those that have embedded the controls necessary to maintain momentum and accelerate safely over time.
The paradox of permission
The AI mantra for many major governments in the past two years has moved from “be careful” to “go faster, safely.” It’s the right impulse, and public sector agencies in countries from Australia to Britain to the United States have been given permission to accelerate and expand their use of AI. Yet progress does not necessarily flow naturally from permission. AI adoption often still clusters in back-office pilots, stalling when it comes to the services residents actually use and where the value is greatest.
Part of the reason is that government is not a business. Businesses that get AI wrong may lose customers, money, and reputation; when public agencies get AI wrong, they may risk vital services. The result is the loss of something even harder to rebuild: the public’s permission to act. And because people tend to view the public sector as a single entity—“the government”—and can’t take their business elsewhere, a mistake at one agency can effectively become a mistake for all.
That’s why the real limit on how quickly an agency can act is not technology or budget but trust: the degree to which residents, policymakers, and public sector employees believe it can get AI right. In recent surveys, people express concern about AI and a desire for strong oversight—including those remaining ultimately accountable no matter when, where, and how AI is deployed.2 Trust and confidence are gained when there is transparency about the technology’s use and mechanisms to challenge decisions.3
Beyond trust, the balance of the ability to act is shaped by potential consequences. Get AI right, and the credit is slow and shared. Get it wrong, and the backlash may be quick and individual, especially for agency leaders. Governments around the world have seen automated systems fail in public and paid for it in court or at the ballot box.
Given the high risk relative to reward, it’s no wonder many agencies regard the safe move as implementing a long committee-driven process or a series of hurdles to clear in sequence. Yet while that looks like prudence and diligence, it often ultimately changes very little. And, in many cases, the committee itself becomes the problem rather than the answer by impeding the innovation and iteration it was created to enable.
Designing controls from the start
When it comes to balancing the need to act quickly but safely with AI, a committee at the end of the process is a weak control mechanism. It is typically slow, because everything waits for approval, which often arrives late. The committee process also misses too much. By the time anyone reviews a finished system, the choices that mattered—which data to use, what it may decide, and what recourse residents have—are likely to have been settled months earlier (for leaders wanting to evaluate where their agency stands, see sidebar, “Six questions for the leadership team”).
The alternative is controls designed at the start and carried through the life of the system, with degrees of intensity depending on whether AI is being used for everyday or mission-critical tasks. In our experience, this ultimately delivers faster performance. Problems are surfaced while they are still cheap to fix, with reusable controls and a final approval stage that confirms work already done rather than uncovering work that was not. In this mode, leaders provide oversight not by attending committee meetings but by being actively engaged in the design of the control system to ensure it is tailored to an agency’s mission, ethos, and objectives.
Doing this well requires two components (Exhibit 1). The first is a set of controls following every AI use case from design to launch to ongoing monitoring, in which governance supports an environment that allows AI to scale safely. The second is the shared machinery making those controls practical: people, ways of working, technology, data, and change management—all supported by risk governance. The controls define what must happen. The shared machinery lets an agency do it repeatedly and at scale.
Embedding risk management across the process
In the first part of the process, a single use case travels from idea to live service through five stages, with three approval gates along the way. The trick is to treat each control as an early design choice, not as a box to check at the end.
Design. Before anyone writes code, two questions should shape everything that follows. First, should the agency use AI here at all? The need to adopt the technology should be judged against publicly defined AI ethics principles and whether the same outcome can be achieved more simply or inexpensively in another way (for example, deterministic automation using rules-based systems). Second, how much freedom should the system have to act? The answer depends on the complexity and harm a wrong decision may cause (Exhibit 2). The best early uses of AI are often simple, low-risk, high-volume tasks for which it can handle most of the work. Decisions that materially affect individuals are different. In those cases, a person should remain in control. This is also the time to set clear rules for when the system must refuse, escalate to a person, or rely only on approved sources.
Get data. AI-enabled workflows are only as trustworthy as the information they use. Agencies should use AI governance processes to verify that every source can be legally used, handle personal information carefully, and feed the AI system clean, approved material rather than a messy shared drive.4 Difficult test prompts should also be created early, allowing the system to be tested later against realistic failure modes.
Build and evaluate. At this stage, many controls are engineering choices. An AI-enabled workflow needs clear instructions, limits on the form of answers, tight permissions for any tools used, and filters on what goes in and comes out. The harder step? Proving that the system works well enough for real service. This requires repeatable tests, expert review, stress tests for misuse or data leaks, red-team tests for security and harmful outputs, and checks for fair and equitable performance across affected groups. Gate two in the process, approval to implement, should depend on this evidence, not on the project deadline.
Industrialize. Moving into production requires another set of controls, including locking in the AI-enabled workflow and prompt versions, setting limits on use and cost, adding emergency stop points, and documenting what the system can and cannot do. A sponsoring leader should name an owner, train staff on safe use and escalation, and keep a record of what the system does. Last, the finished use case should be registered and key artifacts retained, including its prompts, data sets, and test results, so that the next team can build on the work rather than starting again. At this point in the process, gate three is reached: approval to go live.
Monitor. Once the system is live, the controls still have to work. In our experience, more than half of all AI oversight work happens after production release. Alerts should flag unsafe behavior and attempts to trick the system. The agency should track who is using the AI solutions, what those solutions are doing, how well they are working, and what they cost to run and maintain. Tests should keep running, because AI-enabled workflows and supplier systems can change in ways that affect behavior. Someone should also ask periodically whether using AI is still appropriate. If the answer has changed, the agency needs to be willing to roll back and update the system. The objective is to create an environment that is both safe and fast, enabling AI applications to scale as efficiently and effectively as possible.
Building capabilities to make the process repeatable
Controls applied to individual use cases will not scale on their own. Public sector agencies need shared capabilities so the same discipline can be applied again and again, faster each time. That is the second part of the process detailed in Exhibit 1: It cannot be bought off the shelf or delegated to a vendor. Rather, it lies at the center of what an agency must build. Three broad capability areas individually and collectively reinforce the control and governance of AI-enabled workflows to strengthen trust.
The technical base. Technological capabilities are critical because overseeing AI use cases with logs in spreadsheets will not scale, especially as systems begin to take actions, call other systems, and link steps together. Agencies should have a common, governed platform connecting use cases with the models, tools, and data on which they rely (Exhibit 3). The platform should control which tools an AI system may use, manage identity and permissions, record activities for auditing, enforce policy, support testing, and maintain a register of approved models. It should also provide reusable building blocks such as hosting, orchestration, and shared agents.
In addition, while newer risks may be less visible, they are just as important. Agencies should protect systems from malicious prompts or corrupted context. They also need to control wasteful resource consumption, which drives costs up as AI systems run. A wide range of innovative start-ups and major cloud providers are already moving in this direction as they enhance their governance layers, and the benefits compound: Once guardrails are built into the platform, every new use case inherits them and the cost curve of control bends with scale, creating greater AI delivery productivity and velocity.
A platform is only as good as the data beneath it. Reliable AI needs current, approved information to draw clear rules for the systems it connects to. It also needs privacy protections built in from the start. In the public sector, data is often sensitive, and misuse can cause serious harm. That means data quality and protection are not just housekeeping but what makes the system work—and acceptable to the public.
The people who run it. The second capability area involves creating truly multidisciplinary teams to build and run AI safely. This includes a product owner who understands the AI-enabled workflow well enough to redesign it and who owns the risk as well as the features; people with skills the public sector is often short of, such as evaluation, red teaming (cybersecurity testing), and AI safety; and emerging roles such as AI trust architects, who ensure that control technologies and techniques are continually advancing and that they are reused to the greatest degree possible. Agencies also need a workforce confident enough to work alongside agentic tools and enhance them, which is more than simple “human in the loop” oversight. Many governments today rely on contractors for these skills, but as AI becomes more important, they may increasingly choose to build expertise in-house. Doing so may require more than hiring; it likely will require rethinking roles, job classifications, operating models, and workforce policies to attract, develop, and retain AI talent.
Product ownership, perhaps the most important choice, is often where agencies go wrong. Making a chief AI officer responsible for trust can recreate the same end-stage checkpoint under a new title. For that reason, accountability should rest with the service leader affected by the AI system. Responsibility lies with the chief AI officer and central AI function, which together provide the platform, standards, shared testing, approved patterns, and scarce technical expertise. Delivery should still sit with the service team that owns the outcome. In practice, this means there should be cross-functional teams that own value and risk and a light central layer for common tools and guardrails. It also means measuring performance not by the number of pilots launched but by outcomes such as accuracy, safety, resolution time, and resident experience.
What keeps it honest and effective. As the use of AI-enabled workflows grows, two things prevent standards from slipping. The first is governance that fits into an agency’s existing risk approach rather than sitting apart from it: clear ownership, clear risk rules, and clear paths for escalation and exceptions. For example, a cross-functional forum can coordinate oversight without becoming a bottleneck. The second is change management, which is critical given how AI alters the nature of work. Leaders should be clear about where AI supports staff and where it removes work, building capability through hands-on coaching. They should also be open about the technology’s limits, model expected behavior, and measure and share real impact, such as time saved, errors avoided, and outcomes improved. While AI’s potential stems from data, algorithms, and talent, it must be supported by trust to achieve real impact. If trust declines, so does impact. Trust defined by strong controls is critical to accelerating adoption and achieving impact from AI investments.
Controls accelerate AI
The public sector plays a unique role in providing critical services to residents, and the cost of failure can be high. But as the opportunity AI presents to drive public sector efficiency and effectiveness becomes clearer, so does the need to find paths to deliver AI-enabled services safely and at speed.
Sustainable speed comes from designing the critical control system to support it. The agencies likely to deliver on the promise of AI in the decade to come will not necessarily be the boldest or the most careful. They will be the ones with the best engineering discipline—the ones that have thoughtfully implemented checks and balances across AI-enabled workflows rather than relying on committees at the end of a process. Strong controls help maximize impact by ensuring issues can be caught earlier, before large sunk costs, and by building a system that can accelerate adoption by mitigating risk, building public confidence, and assuring trust in the ability of governments to deploy AI.
Six questions for the leadership team
Public sector leadership teams wondering where they stand can start by asking six questions. In each case, the desired answer is not just “yes” but “yes—and here is the evidence.”
- Judgment before capability. Do we decide whether using AI is appropriate before we decide whether it is possible?
- Built in, not bolted on. Are the controls designed into how we build each system from the start rather than added at sign-off at the end?
- Demonstrable reliability. Can we show with evidence rather than assurances that each live system still does what we promised?
- Capability to repeat. Are we building shared capability—platform, data, and governance—so the next use case is faster and safer, not built from scratch?
- Future-ready workforce. Are we evolving our workforce, roles, and operating model so the organization can succeed in an AI-enabled world?
- Clear ownership. Does the leader of each affected service own its AI, with the center enabling rather than deciding?


