Brief

A human in the loop needs a job description

Specifying review as a control, rather than naming one

Claire FausettPublished Last reviewed 8 min read

Bottom line

“A human reviews the output” appears in a great many control descriptions, and on its own it describes nothing. It names a person and calls that a safeguard. A reviewer becomes a control once somebody has written down what they are meant to catch, what they can see when they try, when they are obliged to act, what their disagreement is worth, and how any of it improves the thing being reviewed. Short of that, the word “human” is doing the work of a design.

1. Start from what the reviewer is meant to catch

A review with no defined failure mode is an inspection with no specification. Four questions settle it:

  • Which specific error is this person positioned to detect?
  • Is a human better placed to see that error than the system is?
  • What does the failure look like from where they sit, on an ordinary day, at the volume they work at?
  • What happens to the case once they see it?

“Anything that looks wrong” fails the first question and predicts the answers to the rest.

Volume decides more than most control descriptions admit. A reviewer holding ninety seconds per case can catch a category of error that announces itself, and will miss one that requires a second source. Both are legitimate control designs. Only one of them survives being written down honestly.

2. Give them enough to decide with

  1. The decision the system reached
  2. The evidence it used
  3. The evidence it discarded
  4. Its confidence, and what that number is worth
  5. Their own authority to act

A reviewer shown only the output is asked to ratify rather than to review, and ratification is what an approval queue produces when its reviewer has no purchase on the decision.

The third line is the one most often missing. A system that scored a case as low risk did so partly by setting something aside, and the set-aside evidence is exactly where the human advantage lives. Show it, and the reviewer has a job. Withhold it, and the queue measures throughput.

Confidence deserves its own scrutiny. A score carried to three decimal places invites a reviewer to treat it as a measurement. Whether it is one depends on calibration work somebody either did or skipped, and the reviewer deserves to be told which.

3. Separate the roles

Role Holds
System owner The model or rule, its thresholds, and its performance
Reviewer The individual decision in front of them, and the authority to change it
Control owner Whether review is working as a control at all

The second and third roles collapsing into the first is the common failure. Somebody reviewing their own system is grading their own homework, and supervisory guidance on model risk has made this point for well over a decade: challenge has to come from somewhere holding both the standing and the incentive to deliver it.

4. Write down what an override is worth

An unrecorded override teaches the organization nothing and protects the reviewer least. Record four things: what was shown, what was decided, what changed, and the reason given at the time. Reasons reconstructed months later during an examination are worth close to nothing.

A free-text box collects “looks wrong” and yields a dataset nobody can classify. The workable middle is a short fixed taxonomy — wrong entity, missing evidence, stale data, policy judgment, system error — with free text for the case that fits none of them. Five categories with counts attached will tell a system owner more in a quarter than a year of prose.

5. Close the loop, or admit it is open

Reviewer outcomes either feed back into thresholds, training data, and testing, or they stay a queue that absorbs errors quietly and reports a healthy rate. Both configurations exist in the world. One of them is a control.

The open loop deserves describing precisely, because from above it looks like health: the reviewer catches errors, the errors get fixed case by case, the reported rate falls, and the underlying system keeps producing the same mistake at the same rate for years. The queue has become a maintenance cost the organization has learned to budget for. Say which configuration is in place, and where the loop is open, put that in the control description rather than in a footnote.

A worked hypothetical

The following scenario is hypothetical and is included only to illustrate the framework.

A payments firm deploys a model that scores incoming screening alerts and routes the lowest-scoring band to automatic closure. A person reviews a five percent sample of that band. The control description reads: “Auto-closure is subject to human review.”

1. What the reviewer catches. The band holds cases the model already judged benign, so the reviewer is positioned against exactly one error: the model’s overconfidence. That is a well-chosen target, and it carries a consequence the control description skipped — a five percent sample finds a rare error slowly. Written down honestly, the control reads: this detects a systematic miss within a quarter, and a rare one within a year.

2. What they see. The queue shows the alert, the score, and a Close button. Adding the attribute that matched, the attributes that scored low, and the list entry’s alias set turns the reviewer’s job from ratification into review.

3. The roles. The analytics team owns the model, chose the five percent sample rate, and reads the sample results. Three roles, one owner. Moving the sample rate and the reporting to a control owner outside that team restores the challenge.

4. Overrides. In one quarter the reviewer reopens nine of four hundred sampled cases, each recorded as free text. Six of the nine share a cause, visible only once somebody reads them together: the list entry held an alias the customer record was structurally incapable of matching, and the model scored the residual match low because that field was sparse. A reason taxonomy would have surfaced it in week two.

5. The loop. The reopened cases reach the model team as a quarterly note. Attaching them to the next threshold review, and adding all nine to the regression set run before any model change ships, closes it.

The control description after the exercise: a control owner outside the model team samples five percent of auto-closed alerts monthly; reviewers see the matched and unmatched attributes and the alias set; overrides are coded to a fixed taxonomy; and every override enters the regression set for the next release. Same person, same queue, same five percent. It is now a control.

6. Measures that tell you whether review is real

  • Override rate, and its trend after every model change
  • Time spent per review against the time the task plausibly requires
  • Agreement rate between reviewers shown the same case
  • Share of overrides carrying a recorded reason from a fixed taxonomy
  • Errors caught by review against errors caught downstream of it
  • Cases where the reviewer lacked the evidence to decide either way

The second is the most uncomfortable and the most informative. Where the average review takes forty seconds and the task plausibly needs four minutes, the control description and the operating reality have parted company, and the fix belongs to staffing or to scope rather than to the reviewer.

What this changes operationally

A model owner writes the reviewer’s brief as part of shipping the model, treats overrides as a labeled dataset, and hands the sampling decision to somebody else.

A compliance lead stops accepting “reviewed by a human” as a control description and asks for the five lines: target error, evidence shown, authority held, override handling, feedback path. A control that fails to supply them is an intention.

An operations manager measures time-per-review against task difficulty and raises the gap as a control finding rather than as a resourcing complaint.

A second-line control owner picks up the question currently held by nobody: whether review is working as a control at all, evidenced by override trends rather than by queue completion.

Limitations

This describes how to specify review as a control. It settles nothing about which decisions ought to be automated in the first place, which is a risk-appetite question belonging to governance rather than to design. A well-specified reviewer sitting behind a decision that should have stayed with a person is still the wrong architecture, carefully documented.

It assumes, too, that the reviewer holds authority, capacity, and independence. Where any of the three is missing, the honest control description says so, and the gap belongs on a risk register rather than in a paragraph describing the review.

Primary sources

  • Board of Governors of the Federal Reserve System and OCC. Supervisory Guidance on Model Risk Management, SR 11-7 / OCC Bulletin 2011-12 (2011). Effective challenge, independence of validation, and the documentation expected of a model’s controls.
  • NIST. Artificial Intelligence Risk Management Framework (AI RMF 1.0) (2023) and the accompanying Playbook. Human-AI configuration, and accountability structures for oversight.
  • European Union. Regulation (EU) 2024/1689, the AI Act, Article 14 on human oversight, and what oversight of a high-risk system is expected to include.
  • FFIEC. Bank Secrecy Act/Anti-Money Laundering Examination Manual, on suspicious activity monitoring and on independent testing. Alert handling, quality assurance, and the role of the audit function.
  • Institute of Internal Auditors. The IIA’s Three Lines Model (2020). Separating management of a risk from oversight of it.