Read-only access should be the first stage of every AI agent deployment, not a concession you make while waiting for approval to do more. In this stage the agent runs against live data, produces the decisions it would have made, and changes nothing. A person compares those decisions to what actually should have happened.
The most useful thing this stage finds is rarely a reasoning error. It’s an agent that reasons correctly from the wrong place. That kind of mistake can’t be seen in a demo, because a demo uses data someone already chose for it.
A correct answer from the wrong source is still wrong
Here is an illustrative example of the pattern, not a specific client engagement. An agent is asked to decide whether a renewal can auto-extend under the customer’s contract terms. It reads the contract, finds the notice clause, applies it, and reaches a conclusion that is logically sound.
The contract it read was a superseded draft. The document store held three versions of that agreement, and the signed one lived in the CRM. Nothing in the agent’s reasoning was broken. Its retrieval had simply found the best-matching document instead of the authoritative one.
If that agent had write access on day one, the first symptom would be a customer notified of terms that were never in force. It’s hard to explain that as “the model made a mistake,” because the model didn’t. The system around it did.
Demos and pilots hide this because someone curated the inputs
Most agent pilots run on a cleaned-up slice of data. Someone picked the folder, deleted the duplicates, and pointed the agent at the one system they trusted. The pilot works, the budget gets approved, and the agent is then connected to the real environment, where nobody ever tidied anything.
Real environments have two systems that each claim to be the source of truth for the same field. They have exports that stopped refreshing months ago and spreadsheets that became load-bearing by accident. An agent doesn’t know which of these to distrust unless someone has told it, and nobody has told it, because the people who know are carrying it in their heads.
Observe-only mode turns tacit knowledge into a written source map
In an observe-only phase, the agent gets a scoped read identity, with access to the systems it needs and nothing else. This is the Identity principle doing its job early: the credential is narrow enough that “read-only” is a property of the account, not a promise in a prompt. Alongside it, every record the agent reads and every conclusion it reaches is written down with the reason attached.
Then a person who actually does the work reviews a sample of those conclusions against what they would have decided. When the two disagree, the question is never only “why did the model say that?” It’s “which source did it use, and which source should it have used?”
Those answers become the Knowledge layer: an explicit map of which system is authoritative for which question, and which ones the agent must never cite. This is where we spend most of the time in this phase. It’s less glamorous than prompt tuning, and it’s the part that determines whether the agent is trustworthy later.
The same phase gives you the approval policy. Once you’ve watched the agent make a few hundred would-be decisions, you can see which ones were routine and which ones a person would have wanted to look at. That’s far better evidence for a policy than a workshop guess. If you haven’t settled who writes that policy, who should own the policy your AI agent enforces is worth settling before this stage ends, not after.
The record from this phase is the baseline for everything after it
An observe-only run produces something a demo never does: a dated record of what the agent would have done, next to what actually happened. That record is what lets you say, with evidence, that the agent’s decisions match human decisions on a given class of work.
It also forces the logging question early. Plain application logs rarely hold the source, the reason, and the policy in one place, which is why “we have logs” isn’t the same as an audit trail. If the observe-only record can’t answer “why did it conclude that?” for a specific case, the write-enabled record won’t either.
Write permissions then arrive in steps, not all at once. Low-consequence, reversible actions come first, behind an approval gate. The agent earns wider scope on its record, not on its pitch. The boundary conditions still apply the whole way through, and what an AI agent should never be allowed to do doesn’t change just because it has performed well.
Where read-only stops being enough
Observe-only has real limits. An agent that only reads can’t show you how downstream systems react to its actions, so write-side problems such as rate limits, duplicate records, and unexpected triggers in other tools only appear once it can write. Some of those can be tested in a sandbox. Others can’t, and the first writes need to be small and watched.
It also costs time, and there’s a failure mode on the other side: a phase that never ends. If reviewers stop looking at the samples, or the review becomes a rubber stamp, you’re paying for the delay without getting the evidence. The phase needs an exit criterion agreed up front, such as what agreement rate with human decisions, on which classes of work, counts as ready.
And the review depends on someone with real knowledge of the work having time to do it. If that person is already your busiest operator, the phase will drag. That’s a staffing problem, and no amount of tooling fixes it.
The question to sit with
If you connected an agent to your systems tomorrow, could you say which one is the authoritative source for each decision it would make? And would the people who know that answer be consulted before the agent acts, or after it has already acted on the wrong one? If you’d like to work through it against your own systems, Book a diagnostic call.