Issue #257

More Sign-Offs Don't Mean More Actual Review

When AI output multiplies, someone still has to read it—so approval counts, review time, and execution risk all need sizing up front.

BusinessMore Sign-Offs Don't Mean More Actual Review

What Comes After “A Human Signs Off”

It’s easy to feel reassured once you’ve required a human to approve whatever AI produces. But naming who’s responsible doesn’t finish the job of review. You still have to decide how many items that person has to read each day, what exactly they’re supposed to check, and whether they actually have the authority to stop a bad result.

Say the volume of output needing review multiplies tenfold after AI adoption, while the number of reviewers and the time they have stays the same. What happens then? Even if you designate one more approver, there may simply not be enough time to actually read the content. And the pressure to process requests quickly raises the risk that checks get skipped altogether.

I think that when you bring in an agent, you have to design generation speed and review capacity together. That means defining which actions require human approval, and then calculating how often those actions are actually going to occur.

Managers Also Worry About Who Owns the Decision

The OECD surveyed more than 6,000 middle managers across France, Germany, Italy, Japan, Spain, and the United States. The “algorithmic management” covered in its 2025 report refers to software that supports or automates tasks like assigning work, monitoring performance, and evaluating employees. The survey wasn’t limited to generative AI.

Managers who use these tools said they help with decision-making, but they also flagged concerns. 28% said it’s unclear who’s responsible when a decision goes wrong; 27% each said the logic behind decisions is hard to follow, and that protections for workers’ physical and mental health are insufficient. Accountability came up most often, but the gap between it and the other concerns wasn’t large.

We shouldn’t read this as “managers aren’t worried about accuracy.” It’s closer to saying that preventing bad decisions and determining who responds when errors occur need to happen together. Nor does the survey establish what’s driving the increase in approval steps, or how effective those steps actually are.

What caught my attention in this survey is the problem with trying to resolve anxiety about accountability through approval procedures alone. The fact that an approver exists and the fact that this person can actually exercise meaningful oversight are two different things.

Approval records can miss what was actually reviewed

Approval processes serve two purposes: catching errors, and recording that someone with authority signed off on an action. For both functions to work together, there has to be information to review — and time to review it.

Catching errors means reading what’s changing, what basis the judgment rests on, and who’s affected. If necessary, you also have to check the original source data. But a system can log an approving account and a timestamp regardless of whether any of that content was actually read.

This is why counting approvals alone doesn’t tell you whether real review took place. A person handling the queue might approve a batch of items at once just to keep throughput up — or, conversely, a request might sit pending for a long time precisely because someone is reviewing it carefully. In either case, the approval screen alone won’t tell you which is which.

Nor is it safe to assume responsibility automatically falls entirely on the single approver. The role of whoever configured the permissions, the organization operating the system, and whoever manages the source data are all tangled together. An approval log is material for reconstructing what happened — it doesn’t substitute for actually allocating responsibility.

NIST’s 2024 Generative AI Risk Management Profile also addresses the risks that arise from how human and AI roles are arranged. It proposes distinguishing oversight roles and building risk-proportionate evaluation. It’s a voluntary reference, not a mandate, but it’s useful precisely because it asks you to look at whether oversight actually functions under real conditions — not just count how many approval buttons got clicked.

Even if you lower the approval-target ratio, the review volume can still rise

The following is a hypothetical calculation meant to illustrate a relationship. It’s not a measurement from any real organization.

If 5 staff members each make 4 requests a day, a manager reviews 20 cases. Once an agent is attached and each person is assumed to propose 50 actions a day, the total becomes 250 cases. Even if only 20% of those require pre-approval, the manager still receives 50 cases.

ItemBefore adoptionAssumption after adoption
Actions generated per day20250
Share requiring pre-approval100%20%
Review requests per day2050
At 3 minutes per case60 min150 min

The share requiring pre-approval went down, but the number of cases to review became 2.5x. If the time per review stays the same, the time required is also 2.5x.

To review only 20 out of 250 cases, the ratio would have to be 8%. But that doesn’t mean 8% is the right approval threshold. If there are 50 risky actions, you can’t just eliminate 30 reviews to hit your time budget. You need to either limit execution volume, add reviewers, or design risky actions to occupy a narrower scope.

I find it easier, at first pass, to think of it this way:

Required review time ≈ number of actions × share requiring pre-approval × time per review + time for re-review and exception recovery

This is a simple formula for estimating workload. In practice, review time varies by action, and sometimes it’s more efficient to batch multiple cases together. You also need to measure the time spent re-checking rejected results or recovering from actions that were executed incorrectly.

If this figure exceeds the available time of the person in charge, throughput needs to be adjusted. Don’t judge the impact of adoption just from a report saying “we lowered the approval rate” — you need to check the actual remaining requests and wait times.

More auto-approvals doesn’t mean human oversight disappears

Anthropic’s February 2026 usage study turned up an interesting finding. Among new Claude Code users, roughly 20% of sessions used full auto-approval, while among experienced users that figure topped 40%. At the same time, experienced users also interrupted tasks mid-execution at a higher rate.

The researchers read this as a shift in how oversight works — rather than approving every action in advance, users watch execution unfold and step in when needed. The rise in auto-approvals alone doesn’t mean review has been abandoned, and it doesn’t prove safety has been established.

In the same study, the more complex the task, the more often the agent asked clarifying questions and the more often users interrupted — and the agent’s question rate rose faster than the interruption rate. Neither the question count nor the interruption count actually measures how carefully humans read the content.

To me, this looks less like evidence that “human intervention isn’t increasing” and more like a case for examining how oversight is being exercised. Post-hoc intervention only works if reviewers can see intermediate results, notice when something’s going wrong, and stop it immediately.

You can review the work plan before a single action

In a separate post in April, Anthropic explained that repeated approval requests create friction for users, and that friction can lead people to wave requests through without really looking. One of the fixes it proposed was Claude Code’s plan mode: the system lays out its execution plan first, the user reviews and revises it, and only then does the work begin.

This is a mindset organizations can apply too. Take a task like reading source material and drafting an internal report — you can review the purpose, the materials to be read, and where the output will be stored before any work starts. If the work later requires data access or external sends that fall outside that plan, it should trigger a fresh judgment call rather than proceeding automatically.

That said, agreeing to a plan doesn’t mean every subsequent action is pre-approved. Tool access permissions, spending limits, and approvals needed for external delivery all need to be configured separately. There also needs to be a way to notice — and stop — any action that exceeds the originally granted scope.

Changing the unit of approval means showing more context for review at once. It does not mean that the original approval keeps authorizing execution even after the risk profile has changed midway through.

Oversight and documentation each need their own check

Article 34 of Korea’s AI Basic Act sets out the obligations of businesses that provide high-impact AI or products/services using it. Item 4 of Paragraph 1 covers human management and oversight; Item 5 covers preparing and retaining documentation that can verify safety and reliability measures.

The mere existence of an approval log doesn’t establish that either obligation has been met. You have to look separately at what management and oversight actually took place, and whether the safety and reliability measures are adequately captured in the documentation. A record that only shows a click timestamp can’t be treated as fulfilling Item 5, either.

What falls under the law’s scope isn’t determined simply by “internal tool vs. external service.” You need to examine both whether the tool is used in a domain the law enumerates and whether it carries a meaningful risk of serious harm to life, physical safety, or fundamental rights. A general document draft and a document used for hiring evaluations don’t get the same treatment just because both are labeled “internal documents.”

If a sector has its own separately prescribed procedures, those must be followed too. A design meant to cut down review time isn’t grounds for skipping a legally required procedure.

Set approval criteria by weighing risk against review capacity

In practice, it helps to start by listing out the actions themselves. Look at what changes, how bad the damage would be if something goes wrong, whether it’s reversible, and who gets affected. Even internal tasks can be risky if they send personal data outside the organization or overwrite important data.

The table below is meant to kick off that discussion. The same action can require different levels of control depending on the data involved and the scope of authority.

Example actionWhat to check firstExample controls
Internal draft, querying approved materialsWhether it touches sensitive info or sends data externallyNarrow access permissions, sample review of results
Internal ticket changesScope of the change and whether it’s recoverableLimits on scope/frequency, change logs and exception approval
Customer notices, price changesRecipients, amounts, external impactVerify content and targets before execution, change caps
Large payments, account deletion, major HR decisionsScale of potential damage and related proceduresRestrict autonomous execution authority, require expert review and approval

The approval screen should show what’s changing, the reasoning behind it, and who’s affected — all together. It also needs a link to check the raw source data and a way to reject or halt the action. If a reviewer has to hunt across multiple systems every single time, cutting the number of approval requests won’t actually save much time.

Approval process and review records

When measuring outcomes, you can add the following metrics alongside the automation rate:

  • Daily review requests, plus time spent per review and wait times
  • Rate of errors discovered or corrections made after approval
  • Cases of batch-approving multiple items, and how that review was actually done
  • Time spent on rework after rejection and on exception recovery
  • Tasks people gave up using because of approval burden, and why

That last item is hard to capture from system logs alone. You have to ask the people involved why they stopped using it. And it’s not just the reviewers’ time and workload you need to check — it’s the requesting staff’s, too.

Oswarld’s Lens

I see “a human makes the final check” as the starting point of safety design. From there, you need to decide which actions to allow, how to supply the information needed for review, and how much volume the person in charge can actually handle.

Treating every action as subject to approval can feel, for the moment, like it lets you avoid deciding the scope of authority. But that decision simply gets pushed onto whoever receives the request each time. As requests pile up, each individual approver has to interpret the same standard over and over. I think it’s better for the organization to set the criteria that can be decided in advance, and send only the judgment-requiring exceptions to a reviewer.

Back in March, this newsletter emphasized trust as a constraint on agent adoption. Now I think that account needs to be joined with reviewing capacity. Users who’ve grown comfortable with a product do show a willingness to allow more autonomy—but that alone doesn’t mean the organization’s capacity to handle the resulting workload has grown too.

To increase throughput, you need to secure time for review, reduce risky executions, or make repeated checks easier. It’s hard to expect that simply adding another approval step will resolve this decision.

If there’s one thing to check this week, I’d suggest looking at the list of actions currently requiring approval alongside the actual review time available. Once you know whether there’s enough time secured to actually read through things, you can discuss where approval should stay in place and where the scope of execution should be narrowed.

💬 Have review requests increased since you adopted automation tools? What criteria have you changed to actually secure time to read through the content?

Your take shapes the next issue

What resonated most in this issue, or where has your experience been different?

Any registered reader can comment for free.