Every AI Agent Has a Risk Tolerance Slider: Enterprise IT Inherited Its Default
The configuration management database (CMDB) is the record IT leaders treat as ground truth. The IT service management (ITSM) platform sitting on top of it routes tickets, approvals, and escalations based on what that record says. Both assumptions hold as long as a person reads the record before acting on it. They stop holding the moment an AI agent reads the same record and acts on it directly. Trust in that record fails the same way institutional trust fails everywhere else, through confident action standing in for an honest pause.
When Hallucinations Are Public, Someone Outside Can Check
Deloitte and EY hit this exact failure in 2025 and 2026, and their versions became public because a citation is checkable by anyone with a browser. Deloitte’s AUD 440,000 review for the Australian government carried a fabricated legal quote alongside invented academic references. Months later, GPTZero‘s May 2026 investigation into an EY Canada cybersecurity report found 16 of its 27 citations were fabricated, misattributed, or led to broken pages. Both firms pulled or corrected the reports. Both failures reached daylight because an outside reader with nothing but curiosity and a search bar could check the work.
Inside the Firewall, the Same Failure Stays Quiet
An agent acting on a stale configuration item (CI) inside a private enterprise network operates behind a wall no outside reader can see through. Checking that record falls entirely to people inside the company, whether the question is a decommissioned server still marked live or a change record closed before the change actually finished. The mistake gets caught in a post-incident review, if it gets caught at all, and it stays inside the building either way. The evidence here comes from two surveys showing the same pattern running at scale, quietly, rather than a single dramatic scandal. An Ivanti survey of 1,500 IT professionals found that 68% have personally seen AI produce hallucinations with potential operational impact, and 16% say those errors reached production before anyone caught them. Separately, Gravitee‘s survey of 750 executives found that accountability claimed clearly in December 2025 turned out, when the same organizations were asked four months later how that accountability actually gets exercised, to be informal or undefined. That absence of a headline signals the mechanism working exactly as expected. Harm that stays invisible at first still accumulates, quietly, until it surfaces at a cost far larger than the original mistake.
This pattern runs on every platform currently giving agents write access to production infrastructure. A team granting an agent permission to restart a service, close a ticket, or approve a low-risk change is running the same experiment Deloitte and EY ran with citations, at a scale and speed no post-incident review cycle was built to catch.
What an IT Agent Hallucinates Is Different From an LLM Citation


An LLM hallucinates a citation because it predicts the next plausible token with no filter to distinguish plausible from accurate. An IT agent hallucinates something different: it queries whatever the CMDB, the monitoring tool, or the ticket history returns, and treats that response as complete by default. Its filtering stops at plausibility. A record refreshed an hour ago and one stale for three weeks return through the same query, indistinguishable to the agent reading them. The agent’s confidence describes the quality of its reasoning. It says nothing about the completeness of what it was handed to reason with.
That gap is why AI agents in ITSM break when the data layer is treated as optional. The failure is not only model quality. It is unverified state accepted as complete.
Grounding Cuts Hallucinations, and Raises the Rate of Saying “I Don’t Know”
Whether this kind of failure is avoidable has already been tested, just not in enterprise IT. A study published in JMIR Cancer by researchers at Japan’s National Cancer Center and the University of Tokyo found that grounding GPT-based chatbots in a curated, verified information source cut hallucination rates from close to 40% down to near zero. The same study found the chatbots refused to answer far more often once that grounding was in place. Constraining the model for accuracy raised the rate at which it admitted uncertainty instead of guessing. That tradeoff is not specific to medicine. It shows up in enterprise IT as an agent that escalates to a human more often instead of acting, and it forces every team deploying agents to decide out loud whether they want the one that pauses or the one that always has an answer.
Every Layer of the Stack Rewards Speed Over Verified Accuracy
The reason this constraint shows up in oncology chatbots and rarely yet in enterprise IT operations is incentive: every layer of the stack rewards a different outcome than verified accuracy. Platform and automation engineers get rewarded for broader automation coverage and shorter resolution times. Ops teams granting agent permissions get rewarded for lower ticket volume. Product leadership gets rewarded for shipping autonomy as a headline feature. Executive leadership gets rewarded for headcount efficiency and speed to market. All four layers reward speed, completion, and coverage. Pausing to check a record before acting sits outside what any of them measures. The emergent outcome is the one all four layers produce together: agents that complete the action with confidence, inherited rather than chosen, running against real infrastructure instead of a chat window.
Closed Tickets Can Reinforce a Wrong CI
What happens next compounds the problem instead of resolving it. An agent acts on a stale record with full confidence. The ticket closes as resolved. That closure reinforces the record instead of correcting it, treating a closed ticket as proof the CI was accurate rather than a trigger to verify it. The next agent that queries that same configuration item inherits the same wrong assumption, now with one more closed ticket behind it as evidence the record was fine. Call it the infrastructure’s own version of model collapse: a CMDB carrying fewer verified facts and more inherited assumptions with every cycle, each one accepted with full confidence by whatever queries it next.
Teams that still treat the CMDB as a source of truth without freshness and provenance checks feed that loop. Stale data does not stay still once agents write back into the same system of record.
Governance Answered “Was the Agent Allowed?” Not “Was the World Still True?”
The industry’s response so far has addressed one half of this. In a 2026 prediction on agentic AI in infrastructure and operations, Gartner projected that over 40% of agentic AI projects will be canceled by the end of 2027, citing inadequate risk controls as a named driver alongside cost and unclear return. The governance tools built in response, identity verification, permission scoping, audit trails, answer a real and necessary question: was this agent allowed to do what it did. They leave a second question fully open: whether the world the agent was reasoning about was still true when it acted. Permission and truth are two separate checks, and enterprise IT has built infrastructure for exactly one of them.
That is the same gap covered when arguing that a CMDB alone is not AI governance: a record store without live provenance, blast radius, and policy-aware context cannot carry the full control plane agents need.
The Risk Tolerance Slider: Match Error Budget to What the Agent Can Touch


The original framing of this problem was a tolerance for uncertainty, set somewhere between an agent that always defers and one that always acts. In enterprise IT, that tolerance changes by what the agent is allowed to touch. A knowledge-summary agent can carry a wide margin for unverified state, since a wrong summary gets corrected in conversation. A ticket-routing agent needs a narrower one. A change-approval agent needs less still. An agent restarting infrastructure autonomously needs almost none, because the cost of acting on a wrong assumption compounds the moment the action executes. Most organizations inherited whatever tolerance came bundled with the access they granted, tier by tier, without setting any of it on purpose.
Enterprise AI risk tolerance is that slider: how much unverified state each tier of agent may act on before it must stop and ask a person. Leaving it at the vendor or platform default is still a decision. It is just an unowned one.
Even Labs With Evaluation Infrastructure Miss Failures for Months
What happens when that tolerance gets left at its default is already documented, at the most safety-focused end of the industry that builds these systems. Virima’s reporting on Anthropic and OpenAI traced it directly: the earliest of three cybersecurity-evaluation incidents at Anthropic dated back to April 2026 and sat inside its own systems for roughly three months, until an unrelated disclosure from OpenAI forced the retrospective review that found it. Anthropic runs dedicated evaluation infrastructure built specifically to catch this category of failure. The mistake still needed someone else’s mistake, at a different company, to surface it.
Set the Slider on Purpose Against an Independent Record
That is the decision every enterprise running agentic IT operations is already making, named or not: how much unverified state each tier of agent gets to act on before it stops and asks a person instead. Making that decision on purpose, tier by tier, requires an independent record the agent can be checked against, separate from its own account of what it did: a record of what’s actually true in the environment, timestamped, sourced, and kept current through discovery that runs on a scheduled, high-frequency cycle rather than an audit-cycle snapshot. A CMDB for AI agents built this way turns that decision into something an organization sets on purpose, rather than something it discovers only after something breaks.
Frequently Asked Questions
Why do AI agents fail on CMDB data that humans used for years?
Humans re-read a record, notice staleness, and pause. An AI agent treats the same query result as complete and acts. The CMDB did not suddenly get worse. The consumer of the record stopped applying judgment between read and write. Enterprise AI risk tolerance collapses when write access is granted without a freshness or provenance check on the CI the agent used.
What is the difference between agent permission controls and runtime truth?
Permission controls answer whether the agent was allowed to take an action. Runtime truth answers whether the world the agent reasoned about was still accurate when it acted. Identity, scopes, and audit trails cover the first question. Timestamped, discovery-sourced configuration and dependency state cover the second. Most agent governance stacks still only build the first half.
How does a closed ticket make a wrong configuration item worse over time?
When an agent acts on a stale CI and closes the ticket as resolved, the closure is often treated as proof the record was fine. The next agent inherits the same wrong assumption with one more closed ticket as supporting evidence. That loop is infrastructure model collapse: fewer verified facts, more inherited confidence, each cycle harder to unwind.
What does Virima mean by an independent record for agentic IT?
An independent record is environment state the agent did not invent from its own action log: what exists, how it is connected, what changed, and who owns it, kept current through multi-source discovery on a scheduled high-frequency cycle. That record is what you set the risk tolerance slider against when agents gain write access to production.
How does high-frequency discovery change the error threshold for AI agents?
High-frequency discovery shortens how long a wrong or missing CI can sit unmarked. Agents still need tiered autonomy, but the unverified window they act inside gets smaller. Combined with provenance and service context, that is how teams move the enterprise AI risk tolerance slider on purpose instead of inheriting full trust by default.






