Skip to content
Back to Insights
Agentic AIBy KE Engineering Team

The Approval Gate

The Approval GateAgentic AI cover for The Approval Gatereqgateguardrail before actionAGENTIC AIThe Approval Gate// FIVE SECONDS TO DECIDE

"A human approves before the agent acts" sounds like safety. In practice it often degrades into a person clicking approve forty times an hour because everything looks fine and the queue is long. Show a reviewer a hundred proposed actions where ninety-eight are correct, and by the fiftieth they are pattern-matching on the shape of the request without reading it. The two that were wrong sail through. This isn't a training problem. It's how attention works under repetition, and no amount of "be careful" fixes it.

Route by risk

Not every action needs the same gate. Classify actions by reversibility and blast radius, and route accordingly.

typescript
type ActionClass =  | "auto"          // reversible, low blast radius: execute and log  | "sample"        // reversible, medium: execute, sample 10% for review  | "approve"       // irreversible or high blast radius: block until reviewed  | "escalate";     // outside policy: route to a senior reviewer function classify(action: ProposedAction, ctx: Context): ActionClass {  if (!action.reversible && action.blastRadius > ctx.threshold) return "approve";  if (action.touchesTaintedInput) return "approve";       // tainted input always gets a human  if (action.reversible && history.approvalRate(action.type) > 0.98) return "auto";  return "sample";}

The auto path keys off a measured signal: the class's historical approval rate, a validator result, or an eval score on similar cases. Never the model's stated confidence, which is narration. Tainted input always goes to a person; Prompt Injection: Defense in Depth covers why. The reviewer sees the "approve" queue in full and a sample of the rest. The sample keeps them calibrated on what normal looks like, so they recognize abnormal when it arrives in the approval queue.

The five-second test

A reviewer should look at a request and, within five seconds, know: what will change, from what to what, why the agent proposed it, and what happens if it's wrong. If any of those takes longer to find than to decide, the interface has failed. That means no raw JSON, no "record 8813," no paragraph of model reasoning to parse. The request is rendered as the reviewer would think about it.

┌──────────────────────────────────────────────────────────┐│  Update customer contact                    IRREVERSIBLE  │├──────────────────────────────────────────────────────────┤│  Acme Corp · account #4471                                ││                                                           ││  email     jane@acme.com  →  jane.doe@acme.com            ││  phone     (no change)                                    ││                                                           ││  Because   Customer replied on ticket #8813:              ││            "please update my email to jane.doe@..."       ││                                                           ││  Undo      Not available after 24h (billing sync)         │├──────────────────────────────────────────────────────────┤│  [ Approve ]   [ Edit ]   [ Reject: wrong target ▾ ]      │└──────────────────────────────────────────────────────────┘

Show the input

A model's stated reasoning is a plausible story about why it did something. The actual input it acted on is evidence. Show the evidence: the ticket text, the email, the field value. A reviewer can judge whether "please update my email" justifies the action. They can't judge a paragraph of confident narration. Chain of Thought Is Not an Audit Trail makes the longer argument.

Edit over reject

Most wrong proposals are almost right: the email is correct but the phone shouldn't have been touched. Rejecting forces the agent to start from scratch. Editing lets the reviewer fix the one field and approve, and the edit becomes a labeled correction that feeds evals. Make edit a first-class action.

Rejection reasons as data

A rejection with no reason is a lost signal. Offer four or five one-click reasons: wrong target, wrong value, should not have been proposed, needs more context, other. Those categories, aggregated, tell you exactly where the agent is weak, and feed back into evals as labeled failures.

Queue design

Don't present an undifferentiated list. Group by risk class. Show irreversible actions first and separately, with a count so the reviewer knows the scope of the session. Cap the session, for example by pausing after 20 actions or the first failure, because reviewer accuracy drops off as a session runs long. It's better to leave items pending, with the run parked as Durable Agents describes, than to rubber-stamp them.

Measure the gate

Review time per decision is the tell: track it per action class. A high, stable approval rate that took a normal amount of time to reach means the agent is reliable for that class and the gate can loosen. A high approval rate reached with very short review times means the human stopped reading, and the response is to reduce their queue.

Some action classes should never leave the approval queue no matter how good the rate looks: anything irreversible above a blast radius threshold, anything touching money, access, or deletion at scale. These are standing controls. Watch their approval rate as a quality metric. Everything else graduates on evidence, and can be demoted on evidence when the rate drops.

Also track how often reviewers edit versus reject, and how often an approved action gets undone later. Those two numbers show whether the gate is catching the right things.