Dev.to Security 🔐 Cybersecurity 👁 0 📖 18 min read

When AI Hallucination Becomes Action

SGAEIA Research Series — Article 13 Aridio Silva · Independent Researcher, Brazil · ORCID When a language model produces an unsupported statement, the immediate problem may be answer quality. When an AI agent can pres

When AI Hallucination Becomes Action

SGAEIA Research Series — Article 13

Aridio Silva · Independent Researcher, Brazil · ORCID

When a language model produces an unsupported statement, the immediate problem may be answer quality. When an AI agent can preserve that statement in memory, use it in planning, delegate it to other agents, invoke tools, or exercise authority, the same factual failure can produce a real external effect.

This technical edition preserves the complete research argument and references published on the SGAEIA homepage and Medium while preparing the navigation, metadata, and image delivery for developers, architects, security practitioners, and the DEV Community audience.

Cover — When AI Hallucination Becomes Action

When AI Hallucination Becomes Action. A conceptual representation of unsupported model output crossing memory, planning, tools, and authority boundaries before producing a real-world effect. © 2026 Aridio Silva | Project SGAEIA | CC BY 4.0.

Contents

  • Abstract
  • 1. Hallucination is useful shorthand, but an incomplete diagnosis
  • 2. Why fluent models can generate false claims
  • 3. Agency changes the risk category
  • 4. Five surfaces of agentic hallucination
    • 4.1 Knowledge and grounding
    • 4.2 Planning and causal inference
    • 4.3 Tool use and execution claims
    • 4.4 Persistent memory
    • 4.5 Multi-agent coordination
  • 5. Why another model is not automatically an independent verifier
  • 6. A layered control model
  • 7. A formal risk model
  • 8. The SGAEIA contribution
    • What this article adds
  • 9. Detection Contracts for claim-dependent action
  • 10. Evaluation without manufactured assurance
  • Conclusion
  • Bibliography / References
  • About the Author
  • Research and project resources
  • Figures and public-disclosure status
  • License and status

Abstract

A hallucination in a language model may produce a false statement, a nonexistent citation, or a plausible explanation unsupported by evidence. In an AI agent, the same class of failure can pass through memory, planning, tools, credentials, and other agents until it changes external state. The problem is therefore no longer limited to answer quality. It becomes a systems problem involving authority, execution, propagation, evidence, and accountability.

This article develops a precise and practical account of that transition. It distinguishes hallucination from adjacent failure modes, identifies where unsupported claims can enter agentic workflows, explains why a second model is not automatically an independent verifier, and proposes a layered control model. Through the public SGAEIA research framing, the article argues that no assertion made by an agent about facts, authorization, execution, or compliance should grant authority or serve as its own proof.

When an answer can cause an action, factuality must become a governed system property.

1. Hallucination is useful shorthand, but an incomplete diagnosis

The NIST Generative AI Profile uses confabulation for confidently presented erroneous or false content, including outputs that diverge from the prompt or contradict earlier statements in the same context. NIST emphasizes the downstream risk created when people or systems believe the content and act upon it [1]. OWASP LLM09:2025 treats misinformation as a core vulnerability in applications that rely on language models and identifies hallucination as one major cause, while recognizing that bias and incomplete information can also create misinformation [2].

The term should not become a catch-all label for every undesirable AI outcome. A result may be wrong because a source is obsolete, retrieval selected the wrong document, a tool failed, a policy was incomplete, persistent context was poisoned, or an attacker manipulated the workflow. These failures can look similar at the interface, but they involve different causal mechanisms, controls, and evidence.

An operational taxonomy should therefore distinguish at least seven cases: factual error, unsupported assertion, invalid inference, retrieval or tool failure, memory failure, deliberate deception, and adversarial compromise. The distinctions matter because a factuality detector does not repair a broken tool, a citation requirement does not neutralize poisoned memory, and a second model cannot authenticate an external action that never occurred. Precision in terminology is a precondition for meaningful measurement and incident response.

The broader research literature reinforces this need for taxonomy. Huang et al. organize LLM hallucination research around causes, detection methods, benchmarks, mitigation, retrieval-augmented systems, and unresolved questions about model knowledge boundaries [15]. For SGAEIA, the practical implication is diagnostic: the observed false statement is an outcome, while the causal surface may lie in model knowledge, retrieval, context, memory, tool state, coordination, or authorization.

2. Why fluent models can generate false claims

Language models learn to generate probable sequences from large datasets. Fluency does not imply that every proposition was compared with an authoritative record. NIST describes confabulation as a natural consequence of generative models approximating the statistical distribution of their training data [1].

OpenAI research adds an incentive-based explanation. Many training and evaluation procedures reward correct guesses but do not adequately distinguish prudent abstention from confident error. When attempting an answer can improve a benchmark score while saying “I do not know” guarantees no credit, the system is encouraged to guess [3]. Reducing hallucination therefore depends not only on the model but also on the objectives and metrics that define successful behavior.

The distinction becomes clearer when evaluation records three outcomes instead of one accuracy score: correct answer, error, and abstention. In OpenAI's SimpleQA example, gpt-5-thinking-mini recorded 22% correct answers, 26% errors, and 52% abstentions, while the older o4-mini recorded 24% correct answers, 75% errors, and 1% abstention [14]. Accuracy alone would rank the second result slightly higher even though its factual error rate was almost three times as large. These figures belong to a specific benchmark and model snapshot; they illustrate a scoring failure mode rather than a universal ordering of models.

This three-outcome view changes the engineering question. An evaluator should specify the relative cost of a false assertion, a justified abstention, and a correct answer for the intended use. In a low-impact brainstorming task, abstention may carry a usability cost; in a high-impact agentic action, a confident unsupported claim can be far more costly than delay, clarification, or escalation. SGAEIA therefore treats abstention policy as part of the evaluation envelope and connects it to authority, consequence, and recovery rather than rewarding answer volume alone.

Anthropic's mechanistic interpretability research suggests a more specific mechanism in one studied model. Features associated with known entities could inhibit a default refusal tendency; when a partially recognized entity activated the “known” signal without adequate supporting knowledge, the model could continue and confabulate details [4]. This is useful empirical evidence, but it remains model- and method-specific. It should not be presented as a universal causal account of every hallucination in every architecture.

Anthropic's 2026 Claude Academy explainer translates this research problem into practical warning conditions. It identifies obscure, niche, or recent subjects; little-known people and places; citations and statistics; and exact dates, names, and numbers as situations where verification deserves particular attention. It also describes three classes of mitigation used or recommended by Anthropic: training the model to acknowledge uncertainty, testing it with questions designed to reward appropriate abstention, and asking users to require sources and check whether those sources actually support the claims [12].

These tactics are useful, but their evidentiary weight differs. Permission to say “I do not know” changes the incentive to guess; citation checking tests claim-to-source entailment; and cross-referencing can introduce an external basis of trust. Asking the same model for confidence, a second opinion, or a critique may expose inconsistency, but it remains self-evaluation rather than independent validation. Anthropic's earlier calibration research found encouraging self-evaluation results while also reporting that estimates of whether the model “knows” struggled to remain calibrated on new tasks [13].

3. Agency changes the risk category

A conversational model ordinarily produces content for a person to read. An agent can treat generated content as operational state, use it to select tools, fill parameters, delegate work, and modify external systems. The risk chain can therefore move from an unsupported proposition to a decision, a tool call, a state change, persistent reuse, inter-agent propagation, and an external consequence.

Figure 1 — From Unsupported Claim to External Effect
Figure 1 — From Unsupported Claim to External Effect. A model output becomes materially riskier when it crosses planning, authority, tool, memory, and coordination boundaries without independent grounding. © 2026 Aridio Silva | Project SGAEIA | CC BY 4.0.

Severity is not determined by factual error alone. It depends on the authority available to the agent, the reversibility of the proposed action, the number of downstream components that may inherit the claim, and the probability that the system detects the error before impact. A false statement with no authority may be corrected cheaply; the same statement coupled to broad credentials and automatic execution can alter records, send communications, move assets, or contaminate persistent memory.

OWASP's Agentic AI — Threats and Mitigations 1.1 formalized T5 as Cascading Hallucination Attacks, focusing on propagation through reasoning, memory, and component interactions [5]. Later OWASP artifacts also discuss the broader category of cascading failures. The exact version and title matter: research should not combine evolving taxonomies as if they were a single stable standard.

4. Five surfaces of agentic hallucination

4.1 Knowledge and grounding

An agent may invent a fact, identity, policy, citation, or state of the world. Retrieval can reduce this risk, but retrieval is not proof: a source may be irrelevant, outdated, incorrectly segmented, maliciously inserted, or interpreted outside its scope. A claim should remain linked to the evidence that supports it, including the source version and the limitations of that evidence.

4.2 Planning and causal inference

An agent may construct a coherent plan on a false premise or infer that an action will produce a particular effect without adequate support. Visible step-by-step reasoning does not resolve the problem. Research shows that reported chains of thought can omit material influences or produce convincing explanations that do not faithfully represent the model's actual process [6].

4.3 Tool use and execution claims

An agent may report that it queried a database, checked a payment, ran a command, or completed a transaction when the action did not occur. Confirmation must come from the tool, observed system state, or a verifiable receipt. The agent's narrative is a claim about execution, not execution evidence.

4.4 Persistent memory

A false statement stored as memory can return in future tasks and gain apparent authority through repetition. Confidentiality does not establish integrity, provenance, freshness, authorization, or fitness for purpose. Private memory may still be wrong; authentic memory may be stale; accurate memory may still be unauthorized in the current context.

4.5 Multi-agent coordination

In a serial workflow, an early error can become an assumption for every later step. In a hierarchy, a coordinator's mistaken interpretation can spread to several workers. Parallel diversity may reveal disagreement, but it can also create false consensus when agents share training data, architecture, retrieval sources, or systematic bias. A 2026 preprint on multi-agent cascades observed a decrease in its normalized hallucination measure across revision stages together with a small loss in factual accuracy; the result illustrates a possible trade-off, but requires independent replication before broader generalization [7].

5. Why another model is not automatically an independent verifier

A second LLM can catch some errors, particularly when it performs a narrow task and receives the original source. Yet two models may share blind spots, accept the same false premise, or reinforce one another's confidence. A probabilistic model judging the plausibility of another model's answer remains a probabilistic control.

Meaningful independence requires material separation in one or more dimensions: evidence source, decision mechanism, identity and authority, model or provider, available context, logging channel, and the ability to block or revoke an action. Merely assigning a system prompt that says “you are the verifier” does not create independence if the verifier relies on the same evidence, the same failure mechanism, and a writable record controlled by the executor.

The verification chain should not regress indefinitely into “LLM validates LLM.” Where possible, it should reach a primary source, deterministic rule, reproducible calculation, schema, authenticated query, or accountable human decision. These mechanisms are not infallible: code can contain defects, policies can be incomplete, and humans can err. Their value is that their failure conditions can be specified, tested, and audited differently from open-ended generation.

Self-consistency methods still have operational value when interpreted correctly. SelfCheckGPT samples multiple responses and uses divergence as a black-box signal of possible non-factuality [17], while Chain-of-Verification separates a draft, verification questions, separately answered checks, and a revised response [18]. Both can reveal instability or discipline review, but neither converts correlated model output into authoritative evidence. In SGAEIA terms, they inform detection; they do not independently authorize effect.

Figure 2 — Verification Requires a Different Basis of Trust
Figure 2 — Verification Requires a Different Basis of Trust. Additional models may contribute useful evidence, but high-impact decisions should ultimately reach an independently controlled source, rule, observation, or human authority. © 2026 Aridio Silva | Project SGAEIA | CC BY 4.0.

6. A layered control model

No single technique eliminates hallucination. A defensible architecture combines five layers. First, it reduces the probability of unsupported output through better context, retrieval with provenance, uncertainty-aware training, and appropriate evaluation incentives. Second, it detects signals through schema validation, source comparison, semantic consistency, controlled resampling, specialized detectors, and adversarial testing.

Third, the architecture contains consequences through least authority, typed tools, allowed targets, time and cost ceilings, sandboxing, and mediation before execution. Fourth, it requires evidence such as tool receipts, actor identity, source identifiers, policy versions, parameters, and observed results. Fifth, it supports interruption and recovery through abstention, fail-closed behavior, memory quarantine, revocation, rollback, and proportional human escalation.

Retrieval must also be treated as a fallible subsystem rather than a truth oracle. Self-RAG shows how adaptive retrieval and critique can improve factuality and citation quality [19], while Corrective RAG evaluates retrieved-document quality and triggers corrective retrieval when evidence is weak or irrelevant [20]. These methods strengthen grounding, but remain dependent on corpus quality, retrieval coverage, evaluator behavior, and claim-to-source alignment. A defensible design records what was retrieved, why it was selected, which claim it supports, and what occurred when retrieval confidence was inadequate.

Methods based on semantic entropy can detect an important subset of confabulations by estimating uncertainty over meanings rather than surface wording [8]. Benchmarks such as HaluEval evaluate whether models can recognize generated content that conflicts with sources or cannot be verified against factual knowledge [9]. These approaches provide useful evidence within their evaluation envelopes; neither establishes that a complete agentic system will remain correct across long trajectories, changing context, tools, and real-world feedback.

7. A formal risk model

For governance purposes, the expected impact of an unsupported proposition can be expressed conceptually as:

Agentic Hallucination Risk =
P(unsupported proposition)
× P(non-detection before action)
× available authority
× propagation potential
× consequence severity
× recovery difficulty

This is not a calibrated universal equation. It is a decomposition that prevents an organization from treating model accuracy as the whole problem. A modest error rate can remain unacceptable when non-detection, authority, propagation, and consequence are high; a less accurate model may be usable when authority is narrow, actions are reversible, and independent checks are effective.

The same decomposition clarifies control ownership. Model teams may reduce the first term, evaluation teams estimate detection performance, security architecture constrains authority and propagation, operational governance defines consequence thresholds, and incident-response design reduces recovery difficulty. Hallucination management is therefore a cross-functional control problem rather than a prompt-engineering task.

8. The SGAEIA contribution

SGAEIA distinguishes capability from authority. An agent may know how to perform an operation without being legitimately permitted to execute it. Bounded and revocable authority constrains identity, scope, resource, action, context, time, and delegation [10]. The public Governed Execution Boundary framing separates model-generated intent from external effect through identity, policy, risk evaluation, enforcement, evidence, and revocation [11].

What this article adds

Most hallucination research concentrates on model output, factuality metrics, uncertainty signals, retrieval, or self-correction. This article connects those results to a different question: what happens when an unsupported proposition is converted into operational state and crosses memory, delegation, authority, tools, and external systems. Its unit of concern is therefore not only the false sentence, but the complete path from claim to effect.

The article contributes four linked propositions: verification must reach a differently controlled basis of trust; abstention must propagate into downstream control rather than disappear at the next agent or tool; claim-level evidence must remain connected to the action that depended on it; and recovery must be designed alongside detection. The Detection Contract provides the public conceptual structure for observing that path, while the Governed Execution Boundary limits what uncertain information may cause. These are SGAEIA research contributions derived from the literature, not claims that the underlying detection methods were invented here.

Applied to hallucination, this framing yields three candidate properties:

An agent's assertion about facts, authorization, execution, or compliance must not serve as its own proof.

Material uncertainty in a safety-relevant condition should reduce authority or interrupt execution, never expand it.

An unverified claim should not cross an agent handoff as trusted state.

These properties shift attention from asking a model to be more careful toward governing how uncertain information can influence authority and effect. A consequential decision can be represented at a public conceptual level as a progression from proposition, source and provenance, confidence and limitations, current authority, applicable policy, impact and reversibility, decision, and observed outcome evidence. This is a research proposal, not a claim of production implementation or certified effectiveness.

9. Detection Contracts for claim-dependent action

A Detection Contract makes the failure observable before the action occurs. For a high-impact operation based on generated content, the contract should define the observed objects, such as the proposition, source, tool, memory item, target, and action. It should also define signals including missing evidence, contradiction, low confidence, disagreement across controlled runs, stale state, or failed external confirmation.

The threshold should be tied to consequence and reversibility rather than a single generic confidence score. The observation window may cover one response, a full trajectory, or a distributed multi-agent chain. Attribution should preserve the model, agent, version, context, tool, and delegating principal, while the response can allow, restrict, request evidence, escalate, deny, revoke, or quarantine. The resulting record should connect the decision and policy to the data consulted and the effect actually observed.

This formulation is stronger than the instruction “do not hallucinate.” It defines what the system observes, when uncertainty becomes material, which component may decide, and which evidence must survive for later review.

10. Evaluation without manufactured assurance

A single hallucination rate rarely supports deployment or governance decisions. Evaluation should preserve the complete envelope: model and version, language, domain, task, tools, source corpus, context length, budget, abstention policy, benchmark snapshot, and judging method. Results from short factual questions should not be transferred automatically to long-running trajectories with persistent memory and tools.

Useful measures include factual error, unsupported-claim rate, citation correctness, calibration, appropriate abstention, detector precision and recall, false blocking, recovery after failure, and propagation across stages. For action-capable systems, evaluators should also measure blocked unauthorized attempts, time to detection, revocation success, reversibility, evidence completeness, and the difference between the proposed action and the observed effect.

Long-form evaluation should operate at claim level. FActScore decomposes generated text into atomic facts and measures the proportion supported by a reliable source [16], preventing a mostly correct passage from hiding a small but consequential unsupported claim. RAGChecker adds a complementary pipeline view by diagnosing retrieval and generation modules separately [21]. Together, these approaches support an evidence record that identifies the exact proposition, its source, retrieval path, generated transformation, and downstream decision.

Reporting accuracy without the accompanying error and abstention rates can manufacture assurance. The evaluation record should therefore preserve the full outcome distribution and the scoring rule, including how uncertainty, clarification, refusal, and false confidence are valued. A model-level abstention must also propagate into the control plane: downstream agents and tools must not reinterpret “unknown” as permission to guess, substitute defaults, or continue execution.

The defensible objective is not a promise of zero hallucination. It is to prevent unrecognized uncertainty from crossing silently from language into consequence, and to produce evidence about how the system behaved when its assumptions failed.

Conclusion

Hallucination remains a model-level challenge, but agency turns it into a systems-security and governance problem. Risk emerges from the interaction of plausible language, operational authority, memory, tools, coordination, and human trust. Better models help; controlled consequences remain necessary.

A responsible architecture does not ask only whether an answer appears correct. It asks what evidence supports the claim, which authority depends on it, what action will be permitted, how the effect will be observed, and how the system can interrupt, revoke, or recover the operation. In agentic AI, factuality is no longer only a content quality. It is part of the control plane.

The model may propose. Evidence must support. Policy must decide. Architecture must bound the effect.

Bibliography / References

[1] NIST. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1, 2024. https://doi.org/10.6028/NIST.AI.600-1

[2] OWASP GenAI Security Project. LLM09:2025 Misinformation. https://genai.owasp.org/llmrisk/llm092025-misinformation/

[3] Kalai, A. T. et al. Why Language Models Hallucinate. 2025. https://cdn.openai.com/pdf/d04913be-3f6f-4d2b-b283-ff432ef4aaa5/why-language-models-hallucinate.pdf

[4] Anthropic. Tracing the Thoughts of a Large Language Model. 2025. https://www.anthropic.com/research/tracing-thoughts-language-model

[5] OWASP Agentic Security Initiative. Agentic AI — Threats and Mitigations 1.1, T5 — Cascading Hallucination Attacks. 2025. https://genai.owasp.org/download/45674/

[6] Anthropic. Reasoning Models Don't Always Say What They Think. 2025. https://www.anthropic.com/research/reasoning-models-dont-say-think

[7] Jamshidi, S.; Moradi Dakhel, A.; Nafi, K. W.; Khomh, F. Hallucination Cascade: Analyzing Error Propagation in Multi-Agent LLM Systems. arXiv:2606.07937, 2026. https://arxiv.org/abs/2606.07937

[8] Farquhar, S. et al. Detecting Hallucinations in Large Language Models Using Semantic Entropy. Nature 630, 625–630, 2024. https://doi.org/10.1038/s41586-024-07421-0

[9] Li, J. et al. HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models. EMNLP 2023. https://aclanthology.org/2023.emnlp-main.397/

[10] Silva, Aridio. Why Autonomous AI Agents Need Bounded and Revocable Authority. 2026. https://medium.com/@aridiosilva/why-autonomous-ai-agents-need-bounded-and-revocable-authority-da75656098f7

[11] Silva, Aridio. From Model Capability to Governed Action: An Architecture for Secure Agentic AI. SGAEIA Research Series, 2026. https://medium.com/@aridiosilva/from-model-capability-to-governed-action-an-architecture-for-secure-agentic-ai-de0bf29d0cdc. SGAEIA research artifact: https://doi.org/10.5281/zenodo.22557796

[12] Anthropic. Why Do AI Models Hallucinate? Claude Academy, 2026. https://academy.claude.com/tutorials/why-do-ai-models-hallucinate

[13] Anthropic. Language Models (Mostly) Know What They Know. Alignment Research, 2022. https://www.anthropic.com/research/language-models-mostly-know-what-they-know

[14] OpenAI. Why Language Models Hallucinate. 2025. https://openai.com/index/why-language-models-hallucinate/

[15] Huang, L. et al. A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions. arXiv:2311.05232, 2023. https://arxiv.org/abs/2311.05232

[16] Min, S. et al. FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation. arXiv:2305.14251, 2023. https://arxiv.org/abs/2305.14251

[17] Manakul, P.; Liusie, A.; Gales, M. J. F. SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. arXiv:2303.08896, 2023. https://arxiv.org/abs/2303.08896

[18] Dhuliawala, S. et al. Chain-of-Verification Reduces Hallucination in Large Language Models. arXiv:2309.11495, 2023. https://arxiv.org/abs/2309.11495

[19] Asai, A. et al. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. arXiv:2310.11511, 2023. https://arxiv.org/abs/2310.11511

[20] Yan, S.-Q. et al. Corrective Retrieval Augmented Generation. arXiv:2401.15884, 2024. https://arxiv.org/abs/2401.15884

[21] Ru, D. et al. RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation. arXiv:2408.08067, 2024. https://arxiv.org/abs/2408.08067

About the Author

Aridio Silva is an independent researcher based in Brazil working on the architecture, security, governance, and trustworthiness of autonomous and distributed artificial intelligence systems.

His research focuses on Agentic AI, Multi-Agent Systems, Edge AI, AI Security, Zero Trust, Security-by-Design, AI Governance, Spec-Driven Development, and continuous security assurance.

He is the creator and lead researcher of SGAEIA — Secure Governed Autonomous Edge Intelligence Architecture, an open research initiative investigating architectural foundations for secure, governed, auditable, and trustworthy autonomous AI systems operating across distributed edge-cloud environments.

Research and project resources

Figures and public-disclosure status

The cover is unnumbered and Figures 1–2 are numbered sequentially. All three images are the original public homepage assets and carry the SGAEIA attribution and CC BY 4.0 license information.

The images communicate the path from an unsupported claim to external effect and the need for a differently controlled basis of trust. They remain within the public-disclosure boundary by avoiding private protocols, algorithms, state machines, policy logic, operational pipelines, or reconstruction-enabling implementation details. No C2PA Content Credentials claim is made.

License and status

Except where otherwise noted, the text and original conceptual illustrations are licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). The SGAEIA software research artifact remains subject to its separately stated Apache License 2.0.

This DEV Community draft is a technical edition of the same public research work. It is not a new study, implementation certification, legal-compliance determination, accredited standard, or production guarantee.

© 2026 Aridio Silva | Project SGAEIA | CC BY 4.0

Autonomous AI. Governed by Design. Trusted by Evidence.

📰 Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.