Human in the Loop: Where a Person Still Decides

Human in the Loop: Where a Person Belongs in an AI System and Where They Do Not
Human in the loop means an AI system does part of the work and hands defined cases back to a person. Defined cases, rather than a general promise that somebody is keeping an eye on things. One point in the process, one named person, one kind of decision. This article covers where that point sits, what the person needs to do the job, and when a company can leave it out.
We have described how we built our own AI invoice automation (opens in new tab) from the build side: what the system reads, what it matches, what it posts. Underneath every stage of that build sits a smaller question that has never had an article of its own. Which cases come back to a person, and what is that person expected to do with them?
What human in the loop actually means
The phrase comes from control engineering and survived into AI because the problem did not change. A system runs a process. At defined points it stops and waits for a person, because the next step needs a decision it was not built to make.
Three words in that sentence carry the meaning. Defined, because somebody who might look at anything is standing next to the loop rather than in it. Stops, because a system that carries on and files a report afterwards has put a person in the audit. Waits, because the case has to sit still until an answer arrives.
A rule-based tool has no version of this. It matches the pattern or it throws an error, and an error is not a question. An AI system returns an answer together with how sure it is, and that second number is what makes a loop possible at all. Below a decision threshold you set, the case stops and goes to a person. Above it, the system closes the case itself and nobody sees it.
The three places a person appears in the process

In an AI workflow a person can sit before the system acts, at the cases the system flags, or afterwards on a sample of what went through on its own. These three positions come from systems we have put into production, and the last column is what we have watched go wrong when one of them was left out.
Where the person sits | What they do | What has to be true | What goes missing without it |
|---|---|---|---|
Before the system acts | Approve the rule once, instead of every case forever | The rule is written down and somebody owns it | The process keeps following an assumption from last year |
At the flagged case | Decide the cases the system marked as uncertain | The reason for the flag is on screen, with the source next to it | The queue gets cleared without being read |
After the fact | Read a small sample of what the system closed alone | Somebody is allowed to move the threshold afterwards | A client finds the first wrong pattern before you do |
Sampling is the position that gets dropped first, because nothing is waiting on it and no one complains when it stops. It is also the only one that can tell you whether the threshold is set anywhere near right.
The person is left with the cases nobody could automate
Lisanne Bainbridge set this out in Automatica in 1983, writing about industrial process control. Automating a process, she argued, may expand rather than eliminate problems with the human operator (opens in new tab), because the classic approach leaves the operator responsible for the abnormal conditions.
Read that against your own process. The routine work goes to the system. What reaches the person is the supplier nobody recognises, the figure that does not add up, the request written in a way nobody anticipated. Volume falls and average difficulty rises at the same time, which is the opposite of how the job is usually described when it is being handed over.
Two things follow, and both are design decisions. The checking has to go to somebody who knows the business, rather than whoever has the fewest meetings that week. And the flagged rate has to stay small enough that each case gets read properly. A short queue gets read. One that fills a morning gets cleared.
What the person needs before they can say no
Naming a reviewer is the easy half. The EU AI Act is specific about the other half. Writing about high-risk systems, it requires that deployers assign human oversight to natural persons who have the necessary competence, training and authority, as well as the necessary support (opens in new tab).
A company automating invoices or reports is usually not running a high-risk AI system in that legal sense, so the obligation may never reach it. The list still reads like a catalogue of what goes wrong. Competence, because a reviewer who cannot separate a correct answer from a plausible one adds a delay and no safety. Training, because the system gets updated and the old failure modes are replaced by new ones. Support, because somebody has to be reachable when the problem is the system rather than the case.
Authority is the one that gets skipped. The reviewer is often junior, and the cases that reach them have money or a client attached. When saying no means overruling the person who set the target, the loop is on the diagram and not in the building.
What this looks like in our own invoice system
Our invoice system reads the structure of a document rather than the history of your company. A supplier that has never billed you before is processed the first time the document arrives, with no template prepared in advance.
That capability is also why the loop is needed. Reading a document the system has never seen produces a level of confidence, and the cases under the line come back with their reason attached: this field was unreadable, this supplier matches two records, these lines do not add up to the total. The person answers that question in seconds instead of re-entering the document.
The same shape turns up in reporting, where one company name written three ways in three systems either matches on its own or waits for somebody. We went through that cycle in why the monthly report still takes a week (opens in new tab).
When you do not need a human in the loop
Three cases, and the third one runs against everything above.
The step is reversible and cheap to redo. A draft nobody sends, a translation one person reads, a summary before a meeting. When a wrong output costs a minute, checking every output costs more than the mistakes do. A sample afterwards is the entire loop worth building.
One person does the work and the checking. In a small company the reviewer is the accountant, and the accountant is the whole of accounts. A separate oversight step around one desk produces a document rather than a control.
The decision was always a business decision. Whether to keep a supplier, what discount to give, who absorbs the disputed hours. When an AI system is being put in front of a question like that, the thing to fix is the design. A recommendation that should never have been generated does not become safer by adding a reviewer.
Where to start
Take one step you have already automated, or one you are about to, and answer three questions about it. What does the system do when it is not sure. Who sees that case. What is that person allowed to change afterwards.
If the first question has no answer, the threshold was never set and the system is closing everything on its own. If the third has no answer, you have a reviewer and no loop.
Get in touch (opens in new tab) and we will go through one of your processes and say where the person belongs in it.

Justas Česnauskas
CEO | Founder
Builder of things that (almost) think for themselves
Connect on LinkedIn
