Most companies find out they weren’t ready for an AI agent after it’s live, not before. The signal isn’t a technical failure — the model usually works fine. It’s an operational one: nobody can say who’s accountable when it’s wrong, what it’s actually allowed to do without asking first, or how to reconstruct why it took a specific action three weeks after the fact. Those are governance questions, not engineering ones, and they’re answerable before you write a line of code.
The six questions below are close to what a real readiness assessment asks. None of them are about which model to use or how good the demo looked.
The demo hides the questions that actually matter
A pilot usually gets built by one engineer, with one shared API key, against one narrow task, with a person watching every output before anything happens downstream. That setup proves the model can do the task. It proves almost nothing about whether the business can run it unsupervised.
The gap shows up the moment the agent needs to act on its own — send an email, update a record, approve a refund — without someone reviewing every step. At that point, questions that didn’t matter during the pilot suddenly do: whose credentials is it using, what stops it from doing something out of scope, and if it does something wrong, how does anyone find out before a customer does. Teams that skip past these in the excitement of a working demo are the ones who quietly shut the pilot down two months later.
What a diagnostic actually asks
These six hold up across almost every deployment we’ve scoped, and they map to the operational prerequisites, not the technical ones.
1. Who owns this agent if it’s wrong at 2am? Not “the AI team” — a named person whose job includes noticing and responding. If no one currently owns the manual process the agent would replace, that’s the first thing to fix.
2. Can you state, in one sentence, what it’s allowed to do without asking a human first? If the honest answer is “we’ll figure that out as we go,” you don’t have a governance policy — you have an agent running on hope.
3. Does it have a real system of record to work from, or does the truth live in someone’s inbox? An agent grounded in your actual documents and systems behaves predictably. One improvising from general knowledge doesn’t — and you won’t know which one you built until it’s already wrong.
4. Does it need its own identity, or is it going to borrow someone else’s? A shared login can’t be scoped, rotated, or revoked without affecting everyone else using it. That’s not a security nice-to-have — it’s the difference between being able to shut off one agent and having to shut off a person’s entire access.
5. How many systems does this touch, and who’s coordinating between them? One agent calling one tool is simple. One agent touching your CRM, your billing system, and your inbox needs actual orchestration, or the failure in system A quietly becomes a bad decision in system B.
6. If someone needs to know why it did something specific, three weeks from now, can you answer in minutes? Server logs tell you what happened. They rarely tell you why, under what policy, or who was accountable. That gap is exactly what an audit trail exists to close.
We write every action an agent takes against the policy and the reason behind it, at the moment it happens, inside the customer’s own infrastructure — not as a report generated after someone asks. That’s the difference between an audit trail and a log file someone has to reverse-engineer under pressure. If you want a straight read on where your business stands against these six, book a diagnostic call.
Some of these questions don’t have clean answers yet
Not every “no” is disqualifying, and not every “yes” means you’re ready. A narrow, low-stakes agent with a loose approval policy might be fine — the cost of being wrong is small, and an airtight governance framework would be more overhead than the task deserves. A wide-reaching agent with the same loose policy is a different problem entirely.
The honest limit here is that these questions require judgment, not a checklist score. Getting the answer wrong in either direction — over-governing a trivial task or under-governing a consequential one — is a real failure mode, and no framework removes the need to think about which situation you’re actually in.
Before you evaluate another agent vendor or demo, ask which of these six questions you couldn’t answer cleanly right now — and whether that’s because the answer doesn’t exist yet, or because no one’s had to give it.