An AI agent that has processed twelve thousand refund requests without an error is still not qualified to approve the twelve-thousand-and-first one for $50,000. A long clean track record changes how much you trust an agent’s judgment. It does not change whether some actions should require a human regardless of judgment. Those are two different governance questions, and most teams only ever build controls for the first one.
The second question is the one that matters when something goes wrong. Not “was the agent usually right,” but “was there a ceiling it could not cross no matter how right it had been.” If the answer is no, you don’t have governance — you have a threshold that erodes every time the agent performs well.
Confidence is exactly what erodes the threshold
Here’s the common failure mode. A team deploys an agent with an approval gate — say, refunds over $200 need a human sign-off. The agent performs well for a few months. Someone reasonably raises the threshold to $1,000 to cut down on approval fatigue. It keeps performing well, so it goes to $5,000. Six months in, the gate that once caught almost everything now catches almost nothing, and nobody made a decision to remove it — the threshold just drifted upward one comfortable increment at a time.
Then a bad week happens. A vendor integration returns malformed order data, or a prompt injection buried in a customer email convinces the agent that a return is warranted when it isn’t, and the agent processes it at $4,800 a time, well within its now-generous approval band, until someone notices the total. The agent didn’t get worse. The fence around it just moved until it wasn’t a fence anymore.
The same pattern shows up with permissions, not just dollars. An agent that’s allowed to request broader tool access “when the task needs it” will eventually request broader access for a task that doesn’t actually need it, and there’s no way to tell the difference from inside a single well-reasoned-sounding justification. Approval thresholds that scale with confidence are useful. Approval thresholds that are the entire governance model are not.
Some lines shouldn’t move no matter what the confidence score says
The fix isn’t a better threshold. It’s a second, separate category of control that doesn’t respond to performance data at all — a short list of actions an agent is never authorized to take autonomously, full stop, regardless of how many times it’s gotten the equivalent action right before.
Where that list usually lands: irreversible financial actions above a fixed ceiling (not a ceiling that ratchets up with trust — a number decided once, by a human, that requires a deliberate policy change to alter). Changes to the agent’s own credentials, scopes, or permissions — an agent should never be able to grant itself capability it didn’t already have. Actions that touch a different customer’s or employee’s data than the one it’s currently working on. And anything that deletes or overwrites a system of record with no undo path.
This is where identity and governance have to work together rather than as separate concerns. A scoped identity that limits what an agent can technically call is necessary but not sufficient — the policy engine still needs to know that some of those calls are permitted-but-gated and others are never-without-a-human, and treat them differently. In deployments we build, that distinction is enforced at the identity layer itself, not just in a prompt instruction the agent could talk itself out of: the credential the agent runs under simply doesn’t have the scope to modify its own permissions, so there’s no threshold to erode because there’s no path to erode. Everything else — approval bands, dollar thresholds, escalation rules — stays tunable, and every decision the agent makes within that tunable range gets written against the policy and the reason at the moment it happens, so the two categories stay visibly distinct in the audit trail rather than blurring together.
Drawing the list too wide defeats the point of automating
The honest difficulty is that this list has to stay short to be worth anything. Put too much on it and you’ve rebuilt full manual review with extra steps — every meaningful action needs a human anyway, and the agent isn’t saving anyone time. Put too little on it and the first incident makes the list obviously incomplete in hindsight.
There’s also a harder problem underneath: enumerating every path to a prohibited outcome is genuinely difficult. An agent blocked from issuing a refund over $2,000 can still, without violating any rule as written, issue two refunds of $1,900 each if nobody thought to close that gap. A hard prohibition is only as good as the specificity of what it actually prohibits, and that specificity takes real domain judgment to get right — it’s not something you can template once and forget.
If you’re working through where those lines belong in your own setup, a diagnostic call is a reasonable place to start — it’s usually faster to find the gaps with a second set of eyes than to wait for an incident to find them for you.
The question to sit with
Pick the one action in your business that, done wrong, wouldn’t be fixed by tomorrow’s cleanup — a large payment, a permission change, a deletion with no backup. Now ask honestly: is there something that technically prevents an agent from doing that action alone, or does your current setup just log it well after the fact?