Scale AI published a piece last week asking a question worth sitting with: when automation gets cheap, which work is still worth routing to a human? Their answer is to price it. Multiply the probability of error by the dollar impact of that error, compare it to the cost of a reviewer's time, and you get a threshold. Above the line, escalate. Below it, let the agent run.
It is a good framework, and for a large enterprise with a finance team and a labeled corpus, it is workable. But it assumes something most teams do not have yet: a calibrated confidence score, produced by a pipeline that has already run enough times to know what it is wrong about.
That is the gap we built ActionBoard.ai to close. Before you can price the boundary, you have to know what your actions are, which ones can hurt you, and whether your pipeline has ever finished one cleanly. Most teams are missing all three.
Step one: you can't classify actions you haven't discovered
Ask a ten-person ops team to list the actions their business runs and you will get a list of tools. Ask them which of those actions, if executed wrong at 2am by software, would cost them a customer, and the room goes quiet.
ActionBoard starts in Guided mode. The user drives, the agents assist, and nothing is autonomous. The point of Guided mode is not productivity. It is registration. Every time a team walks through a goal by hand, the sequence of steps, the tools touched, the data read, and the decisions made get captured as a mission pattern in the Action Graph.
After a few weeks you stop having opinions about your actions and start having an inventory. That inventory is the raw material for everything downstream.
Step two: classify by what an error actually costs
Once actions are registered, they get classified. Not by how hard they are, but by what happens when they go wrong. Four dimensions do most of the work.
| Dimension | The question it answers | Why it moves the threshold |
|---|---|---|
| Reversibility | Does a downstream process catch this error, or does only an audit? | Dominant term. A disputed posting is cheap; a silent close costs full value |
| Blast radius | One record or a table? One customer or a segment? Sandbox or production? | Scales the dollar impact of a single wrong run |
| Data sensitivity | Does the action touch regulated data, credentials, or customer PII? | Some actions never go autonomous regardless of score. Policy, not threshold |
| Reasoning fragility | How much depends on the agent inferring intent versus following an explicit instruction? | Inference is where confidence degrades quietly |
The output is a risk tier per action, attached to the action itself in the Action Graph. It travels with the action into every ActionList that uses it.
Step three: the formation gate
Here is the part that differs most from the standard human-in-the-loop design.
In a conventional pipeline, a confidence score is computed at runtime and compared to a threshold. That works when you have production history to calibrate against. When you do not, a confidence score is a number the model made up about itself, and the research is clear that models are systematically overconfident about their own certainty.
So ActionBoard does not gate on a self-reported score at decision time. It gates on demonstrated history.
Five agents. Five gates. The formation runs as orchestrator, data, analysis, action, and defense/audit. Each is a stage in the pipeline, and each stage is scored on every run. An action becomes eligible for autonomous execution only when all three conditions hold.
- It has been executed successfully multiple times, not once
- Every stage in the pipeline clears a 90% success rate across those runs
- The user has driven the same pattern to completion alongside the agents
Miss the 90% formation success rate and the action stays in Guided or ActionList mode. It does not get a warning or a lower autonomy setting. It simply does not execute on its own. The agents can plan it, present it, and wait, but the go/no-go stays with the human.
The third condition is the one people push back on, so it is worth defending. The gate covers the operator, not only the software. If the team cannot reliably drive a mission to a clean result, the agent's 90% is measuring a goal nobody has defined well. The compounding math only works when the user is clear on what they are trying to achieve, running the same goal and the same ActionList until the result stops varying.
Step four: maturity decides the ceiling
Action-level gates answer whether the action is ready. Maturity answers whether the operator is ready.
ActionBoard's AIOps Maturity Model runs L1 Apprentice through L7 Chief AutoAction Officer, and an agent's execution environment is configured to the operator's certified level. An L2 practitioner and an L6 architect working on an identical action do not get identical autonomy. Higher-risk tiers require higher certified maturity before the formation gate is even consulted.
The effect is that autonomy expands along two axes at once. Actions earn trust by succeeding repeatedly. Operators earn trust by demonstrating they can run the system. Neither one alone unlocks execution.
The same feature, shipped two ways
Take a mid-sized enterprise team shipping a new feature. Developers use GitHub Copilot throughout, the pipeline is standard CI/CD, and the code is good. The difference is not the code. It is what the pipeline knows about itself.
| Pipeline stage | Copilot only, no maturity gate | Copilot + ActionBoard maturity gate |
|---|---|---|
| Goal definition | Held in the ticket and the developer's head; varies run to run | Registered as a mission pattern in the Action Graph; same goal, same ActionList across runs |
| Code generation | Suggestion accepted or rejected per keystroke; no record of which patterns hold up | Generation runs inside a registered pattern; per-stage success is scored |
| Review | Human reviews everything at one flat rate, regardless of risk | Actions classified by reversibility and blast radius; review concentrates on high-severity, low-confidence items |
| Test and validate | Coverage measured against code, not against the action's risk tier | Coverage requirement scales with the action's risk tier |
| Deploy decision | Approval based on ticket status and spend threshold | Approval based on demonstrated stage success; below 90%, the action cannot execute autonomously |
| Failure detection | Loud failures caught by monitoring; silent failures caught by audit, quarters later | Defense/audit stage scores every run; silent-failure classes are gated out of autonomy by design |
| Rollback | Reconstructed after the fact from logs and memory | Pattern history shows which stage degraded and on which run |
| Second run of the same pattern | Starts from zero; the learning lives with whoever did it last | Starts from the recorded pattern; success rate compounds or the gate stays shut |
| Audit evidence | Assembled manually when someone asks | Produced as a by-product of the gate, aligned to ISO/IEC 42001 |
Where the money actually moves
The figures below are illustrative. Replace them with your own before they mean anything for your organization. Assume a fully loaded engineer-hour of $120, twelve features per quarter, and a production incident with customer impact costing roughly 40 engineer-hours once you include detection, response, comms, and the fix.
| Cost line | Without the gate | With the gate | Why the difference |
|---|---|---|---|
| Senior review time | Flat review on every change | Review concentrated on high-severity actions | Same capacity, aimed at the decisions where it changes the outcome |
| Rework from escaped defects | Priced at full incident cost | Reduced by the share of silent-failure classes held out of autonomy | Irreversible errors are the expensive ones; the gate blocks those first |
| Token and compute spend | Re-solving the same problem each run | Falls as patterns stabilize | Registered patterns converge; unguided runs restart the reasoning every time |
| Audit and evidence prep | A project each cycle | Near zero marginal cost | Evidence is generated by the mechanism, not reconstructed for the auditor |
| Key-person dependency | Pattern knowledge leaves with the person | Pattern lives in the registry | The expensive failure nobody budgets for |
The headline savings are not in the review hours. They are in the escaped-defect line, because that is where the asymmetry sits. A reversible error costs you an apology and a patch. An irreversible one that nothing downstream catches costs you the full value, repeatedly, until an audit finds it.
The future risks it closes
Three of these matter more than the rest for an enterprise team.
Regulatory exposure. Agentic execution is moving from interesting to evidenced in every AI governance regime being written. A pipeline that can produce a per-stage success record per action is answering a question that will be asked. One that cannot will be assembling the answer under deadline.
Compounding silent error. An unguided pipeline's failure rate is not stable. It drifts with model versions, prompt edits, and staff turnover, and nothing in the pipeline notices. A gated pipeline notices, because the gate is measured continuously and closes when the rate drops.
Autonomy granted by default. The common enterprise pattern is to grant execution rights during a pilot and never revisit them. Maturity-gated deployment inverts that. Autonomy is the thing you earn last, per action and per operator, rather than the thing you configure first and hope about.
Why the slow path is the fast path
The obvious objection is that this is slower than pointing an agent at your business and letting it work.
It is. It is also the only version where week twelve looks different from week one. A pipeline that executes from day one has no baseline, no history, and no way to distinguish a lucky run from a reliable one. A pipeline that spends three weeks in Guided mode has a registry of what worked, a classification of what is dangerous, and a per-stage success record.
At that point you have what Scale's framework actually requires: a real error rate, attached to a real dollar value, per action. You can price your own boundary. Most teams try to compute that threshold before they have earned the inputs.
Cheap automation expands the universe of work worth doing. It does not expand the amount of trust you have established. That still has to be built one successful run at a time.
Where this framework comes from
I did not arrive at maturity gating from theory. I arrived at it from watching regulated operations teams get handed automation they had not earned.
Most of my career has been spent on the operations side of regulated industries: FINRA, the Department of Transportation, Marriott, EMC Documentum, and work across J&J, Pfizer, and Intel. At AWS I was the first Financial Services Enterprise Solutions Architect, and at Capital One I helped stand up the first cloud operations automation practice. Later, as Global CTO at SoftwareOne, I was responsible for delivery across 23 countries. Along the way I built FedRAMP and AWS GovCloud deployments, which is where I learned that evidence you cannot produce on demand is evidence you do not have.
The pattern repeated everywhere. Teams got the tool before they got the practice. The tool worked in the demo and drifted in production, and nobody could say when it had started drifting, because nothing was keeping score.
I founded CloudsCockpit.io in 2022 as one of the first AIOps companies, and ActionBoard.ai is what came out of it. The maturity model behind it is grounded in three years of field research with NGOs in Bangladesh, AIUB, and six university partners, now spanning twelve-plus university AI lab partnerships. In the same period my team built more than thirty generative AI applications for fintech and NGO customers, including twenty-one fintech AI applications and over two hundred financial models. The gate described in this article is the mechanism that kept those deployments from failing quietly.
One more thing worth admitting: the earliest version of this framework was not built for enterprises at all. It was a tool I built to teach my sons goal mastery with AI. The insight held at both scales. You do not get to act autonomously until you have shown, repeatedly, that you know what you are doing.
Jarjis is the founder and CEO of ActionBoard.ai, operating under CloudsCockpit.io Inc. This is Part 2 of the ActionBoard AI Labs series on compounding system design. Part 1 covered why compounding is a multiplication problem.
Reference: Scale AI, "In An Agentic World Where Automation Gets Cheap, Which Work Is Worth Routing to a Human?", September 15, 2026.