The AI pilot is over. Now someone has to run it.
For the past two years, most businesses have treated AI adoption as a discovery exercise. Find a willing team. Pick a model. Run a pilot. Collect a few striking examples. Put the results in a presentation.
That phase is ending.
The harder question is no longer whether AI can help. It is whether a company can operate AI-assisted work repeatedly, safely and economically when dozens, hundreds or thousands of people are using it.
That is a very different management problem.
OpenAI’s new Admin plugin for ChatGPT Work and Codex is a useful signal of where the market is heading. It lets authorised administrators inspect adoption and credit use, manage members and groups, change supported settings and confirm what happened, all from a conversation. It does not create new permissions. It works inside the roles, policies and approval requirements already in place.
The product announcement matters less than the operating model it reveals. Enterprise AI is moving out of the innovation corner and into the machinery of the business.
Adoption is not the same as operation
A pilot asks: can the technology do something useful?
Operations asks a longer list of questions. Who is allowed to do it? Which data can the system reach? What happens when a request exceeds a limit? Who approves a high-impact action? Can somebody see what changed? What happens when the tool fails halfway through?
Those questions sound mundane beside a model launch. They are also the difference between a demonstration and a dependable capability.
OpenAI says its own IT team uses AI-assisted workflows to handle employee requests, retrieve context, check policies, complete supported tasks and escalate exceptions. At the time of its report, those workflows resolved about 45% of ticket volume. The company also says support volume roughly doubled while live operational data helped the team eliminate its backlog.
Those figures are OpenAI’s account of its own deployment, not an independent benchmark. Even so, the pattern is instructive: the value did not come from a chatbot answering isolated questions. It came from connecting context, policy, action and escalation inside one managed workflow.
The control layer is becoming the product
When AI only drafts text, governance can look like a usage policy and some training. When AI can read systems, update records, change access, spend credits or trigger a workflow, governance has to become part of the execution path.
That means permissions cannot sit in a PDF nobody reads. Limits cannot live in a spreadsheet reviewed once a quarter. Approval cannot be an informal message sent after the action. The controls have to travel with the work.
This is why the unglamorous features now matter:
permission-aware tools;
clear boundaries between reading and writing;
approval before consequential changes;
structured confirmation after an action;
visible usage and cost data;
escalation when a request falls outside the supported path.
These are not brakes on adoption. They are what allow adoption to spread without creating an invisible layer of risk and waste.
Practitioner discussion already shows the gap. People are not only asking which model is smartest. They are asking why an agent forgot it had access to a tool, where a workflow should be orchestrated, which system owns the hand-off and whether a general assistant adds value beyond a dedicated integration. These complaints are messy, but they point to the same truth: capability is only one part of reliability.
Stop counting seats. Start measuring completed work
Many AI programmes still report adoption through licences, logins and prompt volume. Those numbers can describe activity, but they do not show whether the business is better run.
A more useful scorecard starts with completed work:
What task reached a verifiable end state?
How much human handling did it remove or improve?
How often did it need escalation or rework?
What did a successful outcome cost?
Which permissions and data were used to produce it?
This changes the conversation. A team with fewer AI interactions but a reliable claims-checking workflow may create more value than a team generating thousands of summaries nobody acts upon.
It also exposes where the process, rather than the model, is weak. If an agent repeatedly stalls because ownership is unclear, buying a more capable model will not fix the operating design. If every action requires a manual exception, the workflow has not been properly bounded. If nobody can confirm what changed, the organisation does not have a production process; it has an experiment with better marketing.
Build the boring system now
Leaders do not need to wait for a perfect enterprise platform. They can start with one recurring workflow and make the operating rules explicit.
Choose work with a clear beginning and end. Define the systems the AI may read, the actions it may take and the cases it must escalate. Put one named human in charge of the outcome. Record the evidence required to confirm completion. Track cost per successful result rather than cost per token alone.
Then run it often enough to discover the awkward edges.
The objective is not maximum autonomy. It is dependable delegation. Sometimes that will mean the AI completes the task. Sometimes it will prepare a decision for a person. Sometimes it will stop and ask for approval. A mature workflow knows which mode it is in.
The AI pilot was useful because it proved possibility. The next phase is less theatrical and more valuable: turning possibility into an operating capability the business can trust on an ordinary Tuesday.
Someone has to own that system. The companies that decide who, and give them the controls to run it, will move faster than those still collecting use cases.
Sources
1. OpenAI, “Introducing the Admin plugin for ChatGPT Work and Codex”, 25 August 2026. Primary source for product capabilities, permission model and OpenAI IT deployment figures. https://openai.com/index/introducing-admin-plugin/
2. OpenAI Help Center, “Plugins in ChatGPT and Codex”. Primary product documentation for plugin governance, role management and app permissions. https://help.openai.com/en/articles/20001256-plugins-in-codexOpenAI
3. OpenAI, “The full stack behind abundant intelligence”, 25 August 2026. Primary source for the shift towards measuring useful work and successful outcomes rather than raw compute alone. https://openai.com/index/the-full-stack-behind-abundant-intelligence/