How does an AI agent know it did a good job?

In brief: Most companies measure AI progress by the tools they've bought, but the constraint that matters is how many agents one person can actually review. Boris Cherny's "Steps of AI Adoption" model argues that parallel work only becomes possible once a good output and a bad output can be defined precisely enough for a machine to check automatically. This piece looks at what that means in practice — for example, in supplier evaluation — and how a company's contracts, policies, and process design determine how far up the adoption ladder it can climb.

Most organizations measure their progress with AI by the tools they have bought. A more useful question is how many agents one person can review before the reviewing itself becomes the bottleneck. This is the angle behind a model that Boris Cherny, one of the creators of Claude Code, published over the summer under the title Steps of AI Adoption.

Cherny's model looks simple. What sets it apart is its perspective. Most AI plans start with the product and continue with handing out access, while this model asks where the work piles up once the agents are already running. The answer almost always leads to the same place: you can only distribute a task across many agents if someone can state precisely what counts as a good result and what counts as a bad one. Where that definition exists, the work scales. Where it does not, one person stays the bottleneck.

A staircase infographic with four steps rising from bottom to top. Step one, A Pair: one person and one agent working together at a desk. Step two, The Orchestrator: one person directing several agents at once. Step three, Manager of Managers: one agent delegating work to further subagents. Step four, Steering by Intent: the person monitoring only exceptions in front of screens. Labels along the staircase read build verification loops, create dynamic workflows, shift to intent, and personal and organizational efficiency gains.
The steps of agent adoption, from a single person paired with one agent to steering by intent with a thousand agents running.

What the levels mean in daily work

In the assisted stage, one person works alongside a single agent and reads every output before passing it on. The task gets faster, but the shape of the work stays the same, because the colleague sits there watching the agent work. This is the most common state, and many treat it as the finish line of adoption.

One step higher, in parallel work, the same colleague runs around ten agents at once on different cases and reviews only the results, because the system flags anything that falls outside the rules. This is where individual speed starts turning into faster turnaround for the whole organization.

With supervised autonomy the question changes. Instead of "did you review it?" it becomes "what context was the model missing, and how do we supply it next time?" Errors are traced back to the process rather than the single output, and the scale is now around a hundred agents.

At the top step, steering by intent, the agents launch most of the work themselves, scaling well into the thousands, and the person only looks at the exceptions. For most large organizations here this is still a direction rather than everyday reality, but it shows where the staircase leads.

The condition that unlocks parallel work

The most important point in the model lies in the transition between the assisted and the parallel stage. According to Cherny, parallel work has a single precondition: the model checks its own output with automated verification before anyone looks at it. Raising the number of agents solves nothing on its own. You have to build that verification first, and everything else follows from it.

This explains the disparity between different areas. Where a machine can also tell that the output is sound, one person can keep ten sessions running. Where only an experienced colleague can judge whether something is correct, that person's attention stays the bottleneck, and neither incentives nor new tool purchases change it. In software development this benchmark is easier to state; elsewhere it has to be worked out first.

The reason for that difference is straightforward. In software the output can be measured almost immediately, because an expression in a programming language always means the same thing, and the machine can check more easily when something is off. Human language and business decisions do not work that way; the same word can mean something quite different in another situation, so the benchmark has to be spelled out separately. The common thread on both sides is the same, though: an agent cannot do anything that a person was unable to define precisely beforehand. You need a description of what a good and a bad output looks like, with acceptance criteria and examples, and that description is the true precondition, not the power of the model.

What the benchmark is in a business process

If an agent evaluates thirty bids, ranks them and makes a recommendation, on what basis can we claim its work is correct? In software, tests break when something is wrong; in a bid evaluation there is no such automatic signal, and the "we read it and get a feel for it" approach works up to exactly one agent.

The answer already exists inside organizations, just under a different name. The definition of a good output is often written down already: in procedures, internal policies, prequalification criteria, approval thresholds, contract clauses. In a bidding round a machine can decide on its own from these which supplier's prequalification has lapsed, which bid deviates from the agreed payment terms, and which recommendation crosses the signatory's threshold. None of this needs a person, because the condition exists in writing.

What remains is still decided by a person: whether a delivery risk is acceptable with a strategic partner, whether the longer lead time is worth the lower price, and what to do with a supplier who has underperformed for the first time. The two sets can be separated, and the actual work of preparation is exactly that separation.

In fact it is not the outcome of the whole process that is worth defining, but the output of each step, because that is far easier to state precisely. If we can say for every step what counts as a good result, we can trust the outcome as well. Two caveats apply here, though. The separation is never complete: a core remains that a person decides on principle, because due diligence or a partner's reliability cannot be reduced to a list of conditions. And the written benchmark is not enough on its own: the model can misread a situation or be confidently wrong, so every process needs human control, only less of it and more targeted. Writing the process down opens the way toward the higher levels, but it is not enough on its own to walk all the way up.

At a panel discussion in the spring I said that a supply chain is held together first by the contract and only secondarily by the ERP. Alongside agents that idea takes on a technical meaning too, because the contract is one of the places where the conditions for a good output are already written down. Where the process runs ahead of the contract, the benchmark comes from internal policy and standard procedures, the same ones we teach a new colleague. What they share is that the condition is stated. In a company where all of this lives only in people's heads or in free-text annexes, there is nothing to measure against, and no stronger model makes up for that gap.

The order of the work

The order is uncomfortable, because the visible part comes last:

  • Describability. The steps of the process, the decision points and the conditions exist in a stated, consistent form.
  • Verifiability. The output is governed by rules a machine can evaluate on its own, so the system flags anything out of line.
  • Parallelization. From here it makes sense to put ten agents next to one person.

When we designed Fluenta One, we chose the BPMN 2.0 standard and a microprocess-based structure for the sake of adaptability: we wanted the work to be adjustable in small steps rather than in one monolithic system. That same structure is what later makes agent-based work possible. A small, self-contained step is exactly what you can hand to an agent, because at the end you can tell whether it succeeded. In work made of thirty steps that each rely on human judgment you cannot tell that, and such work stays in the assisted stage no matter how many users you add. How we carve out the steps also determines how much can be handed over to agents later. In development this step-by-step benchmark is easier to state; in business we have to formulate it ourselves, and the structure of the process either helps with that or gets in the way.

Where many have not even started

There is a step on the staircase we have not mentioned yet, because it comes before the first one: blocked access. Here no agent runs at all, access depends on approval cycles, and more large organizations in this region stand at this point than would care to admit. For this level Cherny names the old security procedures as the obstacle, along with decisions driven by the pressure to cut token cost instead of by outcomes, and the absence of technical voices in the decision making.

That description fits the kind of negotiation where the conversation is about the monthly fee while at the same company a request for proposal sits in approval for weeks. The two figures are not on the same scale.

Worth asking

Cherny's material is vendor material, its product lists follow Anthropic's range, but the line of thought carries over regardless. The names of the levels say little on their own at a mid-sized company; three questions, however, quickly show where an organization stands:

  • If the caseload doubled tomorrow, on whose desk would it pile up? If one or two names come back, their attention is the limit, and another tool will not help with that.
  • For which of our processes is it written down what counts as a good result? Where this exists in a rule, a contract, a threshold, automation has something to hold on to. Where it lives in an experienced colleague's head, it does not.
  • What would happen if no one looked at the system for a week? The answer shows how much control is built in and how much depends on someone paying attention at the right moment.

The answer to the third question tends to cause the most surprise. At many organizations it turns out that the process is held together not by the system but by the routine of a few people, and as long as that is the case, the number of agents makes little difference to the result.

The sooner you start, the sooner you experience the benefits.