AI governance

Human-in-the-loop only works if the human is actually deciding

Every AI vendor promises a human in the loop. Almost none design for what happens when that human faces four hundred items a week — which is that they approve everything, the control becomes theatre, and the audit trail records consent that was never really given.

How much would reach a person?

Send a quarter of bills and we will show you the review volume at each authority level.

1 / 3
Batched by exceptionReasoning attachedRejection as fast as approval

The design

Six decisions that keep approval meaningful.

Fewer, better decisions

The goal is not to route everything to a person. It is to route the small number of things where a person adds judgement, so that approving means something.

Reasoning attached

Every item arrives with what the agent concluded, why, what it compared against, and its confidence. Approving without that is rubber-stamping with extra steps.

Difference-first presentation

What is unusual about this item relative to the last forty like it, shown before the detail. A reviewer should not have to derive the exception themselves.

Rejection is one action

Rejecting is as fast as approving and always carries a reason, which becomes a labelled example. Where rejection is slower than approval, approval becomes the default.

Your matrix, not ours

Amount, department, vendor, and account thresholds read from your existing approval structure, with delegation for absence and escalation for stalling.

Approval is recorded as an act

Who, when, on what basis, under which policy version, with the item as it stood at that moment. An approval you cannot reconstruct is not evidence.

Approval fatigue is the real failure mode

The risk people worry about with AI in finance is the agent doing something wrong unsupervised. The risk that actually materialises is subtler: the agent does everything correctly, routes all of it to a person for approval, and that person — facing several hundred items a week that have been right every time — starts approving in batches without reading.

At that point the control has inverted. The audit trail records a human approval on every transaction, which looks excellent in a controls test, and no human judgement was applied to any of them. It is worse than no approval step, because it manufactures evidence of oversight that did not occur.

A hundred percent approval rate on four hundred weekly items is not a control operating. It is a control that has been optimised into a formality.

So we measure the approver, not just the agent

Approval rate, time spent per item, and rejection rate are reported per approver. If someone is approving ninety-nine percent of items in under three seconds each, that is surfaced — not as a performance criticism, but as a signal that the routing threshold is wrong and too much is reaching them.

The correct response is almost always to raise the policy’s confidence bar so fewer, genuinely uncertain items arrive. Counter-intuitively, routing less to a person usually produces more actual review.

Difference-first, not detail-first

A reviewer given a full invoice has to work out what is unusual about it. A reviewer given "this vendor is normally coded to 6420; this one is proposed as 6310, because the line description mentions consulting" is making a decision immediately.

Presenting the difference before the detail is a small interface decision with a large effect on whether review is real. The detail is one click away for the cases where it matters.

Rejection has to be cheap

In many systems approving is one click and rejecting requires a reason, a routing choice, and a comment. That asymmetry has a predictable consequence: under time pressure, people approve.

Rejection here is a single action with a reason picked from a short list, and each rejection becomes a labelled example that improves the agent. The person doing the reviewing is, in effect, training the thing that will bother them less next month — which is the only incentive alignment that survives contact with a busy week.

What we will not do

Bulk approve. There is no select-all on a review queue, at any authority level. If a queue is large enough that bulk approval feels necessary, the routing threshold is wrong and the fix is upstream.

Questions

What controllers ask.

Can we bulk approve?
No, at any level. If a queue is large enough that bulk approval feels necessary, too much is being routed and the fix is to raise the policy threshold rather than to make consent cheaper.
How many items should reach a person?
For accounts payable at a settled Level 2, typically ten to twenty percent of bills. If it is more than a third, the policy is too tight; if under five percent, it is probably too loose.
Do you report on approver behaviour?
Yes — approval rate, time per item, and rejection rate per approver. Not to police anyone, but because a ninety-nine percent approval rate at three seconds an item means the routing is wrong.
What happens if an approver is away?
Delegation rules you configure, with escalation after a threshold. An item stalling on someone’s holiday is the most common cause of a close slipping a day.
Does rejecting improve the agent?
Yes. Every rejection with a reason becomes a labelled example specific to your business, which is why rejection is deliberately as cheap as approval.

See how much would actually reach you.

A quarter of bills is enough to model review volume at every authority level before you commit to one.