Rootstrap Get Started →

Insights

Human in the Loop for AI Agents: Judgment, Not Verification

Human in the Loop for AI Agents: Judgment, Not Verification

In two weeks, our agentic delivery system ran 976 verification checks. A person ran one. It needed a phone.

"Where exactly is the human in the loop?" comes up more every month, from two directions. Stakeholders ask, clients and engineering leadership alike, because they want to know why they still need people and not AI agents alone. Developers, QA, product and project managers, designers: the people who built careers in software delivery ask because they want to know what will be left of their job. Both get the same answer, and it reassures because it's vague: there's a human in the loop. It promises human oversight somewhere. It doesn't say what for. A recent paper calls this humanwashing: using the presence of a person to imply a safety the system doesn't have.

Look at what people actually do inside an agentic delivery loop and it splits in two. One half is checking work against a standard that already exists. The other is deciding when that standard is missing, ambiguous, or contradicted. Human in the loop (HITL), done right, means the AI agents handle the verification and stop to hand the judgment to a person. This piece measures where that line fell on a real project.

The project is a voice-driven mobile assistant for frontline supervisors: people who manage a team on their feet and can't stop to type. One of our teams, four people, builds it through Rootstrap's multi-agent delivery system. The system instruments the work as it moves through the loop. From 28 August to 11 September 2026 we pulled everything it logged about who checked what and who decided what. Every number below comes from those ledgers: one production project, fourteen days, 108 pull requests.

In two weeks, Rootstrap's delivery system ran 976 verification checks. That means agents reviewing each other's code, plus the spec gates, tests, builds, linters and security scans those agents ran and read. A person performed one. A navigation bug had been reported on someone's phone. An agent fixed it, a second agent reviewed the fix, and the system stopped and asked for a human, because confirming the fix meant holding that phone. Its note: "only a human with the device can run it." Even so, people were needed 120 times in the same two weeks. Once to check. Half the time to judge.

Two kinds of work hide under the word "review"

Verification is checking work against a standard that already exists. Does the pull request meet the story's acceptance criteria? Do the tests pass? Does the code follow the stack's conventions and stay maintainable? Does this new story contradict a decision already on record?

Judgment is what you do when the standard doesn't exist yet, or two valid standards collide. What should the app do when the spec is silent? Which of two working implementations ships, the quick one or the one the team can maintain? Does this integration go synchronous or async? What gets cut when the date holds? Nobody can look those up. Someone has to decide.

Engineers meet this split every day, usually without naming it. Running the test suite is one kind of work. When a test fails, deciding whether the test or the code is wrong is another. Nobody confuses the two in the moment. Teams confuse them in how they organize the work: both go to the same people, in the same meetings, under the same word, review. That's why "human in the loop" sounds like one job. It's two, with opposite futures. The checking is being absorbed by machines. The deciding is becoming the human job.

Verification is already the agents' job

Start with the checking, the half that is harder to let go of. Here is what the agents and their tools did with it on one production project over two weeks. The delivery system logged those 976 verification checks across seven kinds; nothing else was counted. One kind is AI code review: a second agent reading the first agent's diff against a fixed brief covering correctness, security, test strength and architecture, 138 times. The other six are tools the agents run and read.

Every verification check the delivery system logged in two weeks: seven kinds, 976 in total. A second agent reviewed code 138 times; the agents ran and read every other check themselves.

Most teams have had those six tools for years. What's new is who acts on the results. Here the agents run the checks, read what comes back and fix what fails, with no person in between. The buildability gate said no more often than it said yes. Those refusals named 263 specific gaps in specs before a line of code depended on them. Across 108 pull requests, one defect escaped to a later stage than it should have, and nothing was rolled back.

All of that used to be people's work. Reading someone else's pull request. Turning a spec into test cases and running them one by one. Writing the code line by line, and filling in the spec template before any of it. Good people built careers on that work, and many liked it. If that was you, that work is the agents' now, and the checking half of it they already do more thoroughly than a person can, every row, every time. The value was always in the calls the agents keep stopping to ask you for, and it still is.

This is a structural argument, not a measured comparison: the ledger holds no record of people re-reviewing the same work. That's not because a model is a smarter reviewer. Verification is exactly the shape of work where a machine's advantages are structural. It checks every row instead of skimming the familiar names. It applies the same standard to the fiftieth pull request as to the first, and doesn't get tired of saying no to a spec. Because the standard is explicit, every model improvement lands here first and compounds. On this project that showed up as 976 checks to one and a single escaped defect, and the gap will widen. A person re-reading a diff a machine has already reviewed is not a second layer of safety. It's a sample pretending to be a census. And it burns the hours the other half of the job, the deciding, was waiting for.

When should an AI agent escalate to a human? The moment a check turns into a judgment call.

If 976 checks on one production project ran without a person, what were people needed for? Over the same two weeks, the delivery system logged 120 moments where it stopped and asked for one:

  • 61 were human judgment: a gap the spec never answered, two approved documents that couldn't both be right, a review that deadlocked.
  • 32 were sign-offs: human approval attached to work the agents had already verified.
  • 26 were access: things only a person can physically do.
  • 1 was verification: the phone.

The humans stayed in the loop, with a different job. To see what that job looks like, take two of the 61 judgment calls, both from the supervisors' app.

First case. A supervisor on the floor presses the microphone and nothing happens: recording starts late, and nothing on screen says so. The agents traced the bug and stopped short of fixing it. The fix was two lines of on-screen text, in two languages, and by standing instruction words that end users read are not theirs to choose. The agents laid out the options and recommended none. Sixty-six hours later a person made the call. The result was a better product than any option on the table. It covered the words themselves, how formally to address the user, and how long the app should wait before saying anything. The person also set a rule the agents had not thought to propose: the error color must never appear before sound does. Nothing in the spec covered any of that. The call had to come from a person. The agent's closing note:

"This interruption was not redundant."

Second case, for the engineers. An agent running the project's integrity check found four documents on the main branch with a status that doesn't exist. The allowed list has five values, and this wasn't one of them. The fix would take a minute. The agent didn't make it, because picking the replacement changes what those documents mean. A test passing is not the same as a document being approved, and nobody had decided which these were. The agent's note:

"a decision and not a fix I should make unilaterally."

That stop was still waiting when the window closed. A minute to fix. One decision nobody had made.

Both cases began with a check. Verification and judgment were not separate tasks here. The same piece of work turned from one into the other the moment the check found something no standard could settle, and only then did a person get called. In the previous piece in this series I wrote about how the delivery system knows when an agent loop has stopped making progress and needs to escalate. This is the other half of that problem: knowing what deserves a person once it does, and tuning that rule is as much engineering as anything else in the system. In this project the pattern held wherever the ledger has text. Of the 24 judgment calls whose log includes text, 21 were surfaced by an agent's check.

That is the working definition: human in the loop means a person is called in at the moment an agent's check produces a question the standard can't answer. Not before, not after. This differs from classic human in the loop, where a person approves every action, and from human on the loop, where a person watches and intervenes after the fact. Verification runs without a person and leaves a trail. The loop halts for a person when a check turns into judgment, and for the signatures and physical access only a person can give.

Every change a person made was a decision

The ledger records what a person changed each time the loop stopped for one. Of the 120 stops, 55 were closed inside the window. In 45 of them the person accepted exactly what the agents had proposed. Six times they changed something, and all six were decisions. Among them: the microphone wording, a ruling on what counted as valid test data, and a choice to reuse an existing mechanism where the agents had proposed a new field. Not once did a person catch a factual error the agents had let through. (The other four were rejections or an override, three of them the system's own false alarms, covered below.)

Read that against the developer who reviews every agent pull request because they don't trust the agents. Those hours go to verification, where this ledger says a person almost never changes the outcome. Judgment is where those hours were needed: 65 stops were still waiting for a person when the window closed, the phone test among them. In this project the bottleneck was never review capacity. It was decision capacity.

A team that knows how to work with a multi-agent system spends its hours on judgment, not on verification.

Sign-offs are ownership

Of the 120 moments when the loop stopped for a person, one kind looks redundant and isn't: the 32 sign-offs. A person approves a spec so it can be built, or puts their name on a release the agents had already verified. As checks, most of them added nothing. As signatures, none were redundant. An agent can verify a release; it can't answer for it. When something breaks in production, "the agents did it" is no more an answer than "Ruby did it" or "AWS did it" ever was. Someone controls the system, and that someone owns what it ships, whichever hand does the merging. Regulators are converging on the same point. Article 14 of the EU AI Act, which applies to high-risk AI systems, asks for a person who can interpret the system's output, override it and stop it. It does not ask for one who re-reads what a machine already read. Those 32 are the part of the loop that should not shrink, however good the verification gets.

What the ledger doesn't see

These numbers come from one project over two weeks, 28 August to 11 September 2026, not from an average. About a fifth of the runs were driven by hand and log less, so the verification counts are a floor. The ledger has no field for a person re-reviewing a pull request, so 45 of 55 is an indirect measure of that habit. And the system itself was wrong about needing a person at least three times. A stop rule matched the word "review" in a step name, summoned a human to a deadlock that hadn't happened, and withdrew within minutes. Taking humans out of verification only works if the verification leaves a trail, and if the rule that calls them back in gets checked too.

The strongest objection comes from Margaret Mitchell and Avijit Ghosh of Hugging Face and Samir Passi of Data & Society. They argue that extended use of AI systems degrades the cognitive capacities that oversight requires. Two weeks of data cannot show whether that happened here, or whether it will. What the ledger can show is narrower: in this model, people stayed in contact with the work through the judgment calls they answered. Whether that contact is enough to keep the skill intact is a question for a longer window.

The job was never the checking. Put the people where the questions are.

Stakeholders who ask why they still need people and not AI agents alone can read the answer off one project's ledger. The answer is still people, for everything except the checking. The checking is going to the machines, and it gets better without anyone being hired. You need people for the moments a check turns up a question no standard can answer, half of all the stops in this project. You need them for the name on what ships. And you need them for the system itself, because deciding what it checks, when it stops and whom it asks is engineering work nobody has automated.

Staff for judgment, sign for what ships, and stop paying people to check. For every person in your loop, ask what they are there to do: decide, own, or check. Only the third answer should worry you. It means you are paying for a bottleneck dressed as a safety net.

Developers, QA, product managers and designers who ask what will be left of the job get the other half of the answer: the part that was always the point. In this project, 65 stops were still waiting for a person when the window closed. The judgment calls among them look like this: what the app says when the spec is silent, which implementation ships, what a status means, which words a user reads. That is the job now. It is harder than reviewing diffs, because nobody has written the answer down yet. Get fast at it.

Those questions don't respect job titles. The words a user reads used to be the designer's call, the meaning of a status the developer's, whether a feature counts as shipped the QA's. When the agents stop, they ask whoever can decide rather than a role, and in this project the words and the status question landed on the same engineer. What that does to the roles is another article.

We wrote about the four questions worth asking any team making the human-in-the-loop claim, and this is the answer underneath all four. The goal was never a system that stops needing people, but one that knows, exactly, which moments still do. If you're working out where the people go in your own loop, ask us to map it with you. Every stop is a decision, a signature, or a check the agents should already be running.

Frequently asked questions

What does "human in the loop" mean for AI agents?

Human in the loop, for AI agents, means the agents verify the work and hand a person every question the standard can't answer. The work in that loop splits in two. The agents handle verification: checking work against a standard that already exists. A person handles judgment: deciding when that standard is missing or contradicted. A third kind of stop, the sign-off, is ownership rather than either.

Should a human review every pull request an AI agent produces?

On one Rootstrap production project, agents ran 976 verification checks in two weeks and a person ran one. When the system did stop to ask a person, they accepted the agents' proposal 45 times out of 55, and every change they made was a decision, not a correction. On that evidence, no: not when the agents' verification leaves an auditable trail. A person re-reading a diff a machine has already reviewed checks a sample of what the machine checked in full, and spends the hours the decisions were waiting for.

When should an AI agent escalate to a human?

The moment a check produces a question the standard can't answer. That can be a gap the spec never covered, two approved documents that can't both be right, or a wording choice the agents are not allowed to make. On one Rootstrap production project, 21 of the 24 judgment calls with logged text were surfaced by an agent's check, not by a person noticing something. The stop rule itself needs checking too: it called for a person by mistake at least three times.

← All insights