Ask a general-purpose AI model how PTO approval works, and it will give you a confident, plausible answer: submit a request, your manager approves it, HR logs it. That answer is wrong for your company the moment your actual policy diverges from the generic template — a second-level approval above five consecutive days, a blackout period during quarter-end close, a carryover cap that changed in this year’s handbook revision. The model isn’t guessing badly. It’s answering a different, more generic question than the one you asked.
This matters more than it looks like it should, because the failure is silent. The agent doesn’t say “I don’t know your policy.” It says something that sounds exactly like something your HR team would say, with the same tone and structure as a correct answer. An employee or manager reading it has no signal to double-check.
The model answers from averages, not from your wiki page
An LLM’s training data is a statistical blend of thousands of companies’ HR documentation, employment law summaries, and generic “best practice” content scraped from the internet. When you ask it about PTO approval, it produces the most statistically common version of that policy — which is a reasonable draft for a company that has no policy yet, and actively wrong for one that does.
The actual policy lives somewhere specific: a Confluence page updated last March, a field in BambooHR, a PDF in the shared drive that HR revised after last year’s audit. The model never saw that page. It has no way to know it exists, let alone that it supersedes the generic version baked into its training data.
The result is an agent that sounds authoritative while being confidently out of date. Teams often don’t catch this until someone acts on the wrong answer — approving a request that should have been escalated, or telling an employee their carryover doesn’t expire when it does. By then the cost isn’t a bad chatbot response, it’s a policy violation with a paper trail leading back to the agent.
Grounding means the agent reads your source, not a snapshot of it
This is what the Knowledge principle in agent deployment actually requires: the agent’s answers come from your company’s own current documents and systems, retrieved at the time of the question, not from what a model happened to memorize during training. That’s a meaningfully different architecture than “upload your handbook once and move on.”
A one-time upload solves the wrong problem. It gets the agent past the “never seen your policy” gap on day one, but policies change — PTO rules, expense thresholds, vendor approval limits, security exceptions. If the agent is grounded in a PDF from six months ago, it’s now confidently wrong in a different way: it knows your policy, just not the current one. The failure mode looks identical to the employee on the other end.
Real grounding tracks provenance and freshness, not just content. The agent needs to know which version of the PTO policy it’s citing, when that version was last confirmed current, and where to check if it’s uncertain. In practice that means connecting the agent directly to the source system — the live wiki page, the HR platform’s API, the document management system — rather than a static copy, so that when HR updates the policy, the agent’s next answer reflects it without anyone re-training or re-uploading anything.
This is also where orchestration and governance intersect with knowledge: an agent that’s unsure whether it has the current version of a policy should be routed to check with a source system or escalate to a human, not guess. When we scope a deployment, we design that retrieval path explicitly — which system is authoritative for which policy, how staleness is detected, what happens when two sources disagree — before the agent is allowed to answer questions that depend on it.
Where this gets genuinely hard
Not every company has a single source of truth to point the agent at. Policy often lives in fragments — part in the handbook, part in an old email thread everyone still references, part in tribal knowledge that was never written down at all. Grounding an agent well sometimes means someone doing the unglamorous work of writing the real policy down for the first time before the agent can be trusted with it.
There’s also a real cost to over-indexing on freshness. Re-verifying every answer against a live source on every query adds latency and complexity that isn’t always worth it for policies that rarely change. The judgment call is deciding which policies are volatile enough to need live grounding and which are stable enough that a periodic refresh is fine — and being honest with yourself about which category you’re actually in, rather than assuming everything is stable until an agent proves otherwise.
If this is a live question for your business, a diagnostic call is a reasonable next step before you commit to a deployment.
The question to sit with
If you asked your AI agent a policy question right now, would you actually know which document it answered from — and whether that document is still the current one?