Dev.to AI 🤖 Ai 👁 0 📖 8 min read

Who Decides What Is True for AI Agents? The Source of Truth Problem

Originally published at getunblocked.com on October 7, 2026. Priya has been a staff engineer on the billing platform for nine years, and this is the fourth time today someone has pinged her with a question an agent coul

Originally published at getunblocked.com on October 7, 2026.

Priya has been a staff engineer on the billing platform for nine years, and this is the fourth time today someone has pinged her with a question an agent could not settle. The runbook says the payment retry policy is three attempts. The code says five, with jitter. A Jira ticket from March says the change to five was reverted. A Slack thread from April says the revert was itself reverted after an incident. The coding agent working on a refund fix found all four, listed them, and stopped. The engineer driving it did what everyone does: asked Priya.

Who decides what is true for AI agents? In most engineering organizations today, it is Priya. The source of truth for AI agents is whoever has been around long enough to know which of four contradicting records is current, and that person is answering the same questions a dozen times a week. The job she is doing has a name, arbitration, and it belongs in the retrieval path at answer time, with provenance attached, rather than in a human's head or a hand-curated wiki. It is a separate problem from AI output contaminating your knowledge base. That post covers agent-written content getting indexed. Here the records are human-written, already indexed, and still in conflict.

Why is managing "what is true" for our AI tools becoming a full-time job?

Because the people doing it were already the team's source of truth before the agents arrived. The 2026 Stack Overflow Developer Survey asked developers which sources they trust for specific answers: 72% said the people they work with, ahead of the code (63%) and documentation (60%). Agents inherit that dependency and multiply the request volume.

An agent does not read a Slack thread the way Priya does. She knows the April thread outranks the March ticket because she sat in the incident review. The agent sees four records with similar relevance scores and no signal that says which one won. So it guesses, or it escalates to the same person already fielding questions from new hires. With 73% of developers who use AI coding assistants now using them daily, the number of askers has gone up while the number of people who can settle a conflict has stayed flat. The job did not get harder. It got more frequent, and it now has a queue.

In brief: When code, docs, tickets, and chat disagree, someone picks a winner before an agent acts. That is arbitration. It needs to happen at answer time, with the losing sources shown and permissions enforced, or the work lands on one tenured engineer who becomes the bottleneck for every agent in the company.

Is there tooling for this, or is it a wiki with extra steps?

There is tooling, and most of it is a wiki with extra steps. Curated knowledge bases, agent-maintained wikis, and memory consolidation features all move the arbitration earlier, to the moment someone writes the page. Only tools that reconcile sources at query time can resolve a conflict the curator never saw. We compared six of them in a separate roundup.

The appealing version is the self-maintaining wiki. A 2026 arXiv paper on knowledge compounding measured an agent-maintained knowledge layer against plain retrieval and reported 84.6% fewer tokens across four sequential queries. Anthropic's Dreaming feature for Managed Agents, announced in May 2026, works between sessions: it reviews earlier transcripts, extracts recurring patterns, and restructures the memory store so it stays high-signal.

Both compress what the agent already concluded. Neither can tell you whether the March ticket or the April thread is current, because that fact was never in a transcript. It was in Priya's head. A wiki, hand-written or agent-written, records the arbitration that happened when the page was last touched. The next conflict arrives after that.

What is arbitration, and why does it have to happen at answer time?

Arbitration is the act of deciding which of several conflicting records is authoritative for this question, and saying why. DORA's 2026 report on the ROI of AI-assisted development, as summarized by InfoQ, found that returns depend on "the quality of the internal platform, the clarity of workflows, and the alignment of teams." Arbitration is where those three meet.

Truth in an engineering org has a timestamp. A record is correct relative to a moment, and the moment that matters is when the agent acts. A wiki treated as the source of truth freezes one moment and hopes nothing changes.

This is the job a context engine exists to do. A context engine is a system that retrieves across code, pull requests, chat, tickets, docs, and incidents, ranks the results by recency and authority, resolves the conflicts between them, and returns an answer with sources attached and the asker's permissions applied. The output is decision-grade context: synthesized, conflict-resolved, permission-enforced. Provenance keeps it honest: the answer that says five retries also says why the April thread won.

Our agent always needs the one human who has been here forever. How do we capture that?

You capture what the person does, which is arbitrate, and you stop trying to capture what they know, which is unbounded. A March 2026 paper accepted in Information Sciences defines the bus factor as the number of people whose sudden unavailability would stall a project, and proves that computing it exactly is NP-hard. You will not document your way out.

The instinct is a knowledge transfer program: sit Priya down, record everything, write it up. Given four contradicting records, she checks who wrote each one, whether the change shipped, and whether anyone senior pushed back. That is a method. Yesterday's facts do nothing for tomorrow's conflict.

James Stanier made the organizational version of this point in LeadDev in July 2026: "if only one person can explain a system, that belongs on the risk register, not the list of your stars." Asked which context they most need AI to consider, 44% of developers named prior decisions and rationale. The rationale is the part that lives in Priya. Capturing her means encoding how she weighs sources. Software can hold that.

What does a retrieval system do when sources disagree?

Left alone, it picks badly. A January 2026 study of 13 open-weight models found that LLMs prefer institutionally corroborated sources over individuals and social posts, and that this preference "can be reversed by simply repeating information from less credible sources" (arXiv 2601.03746).

The stale retry policy appears in the runbook, in a Confluence page copied from it, and in a PR description an agent wrote from the runbook. The correction appears once, in a Slack reply. Retrieval that rewards repetition hands the agent the wrong answer, and the first plausible hit ends the search (satisfaction of search).

Research is converging on arbitration before generation. ConflictRAG, published in May 2026, detects conflicts between retrieved documents, scores source credibility, and resolves the conflict before generation, with a reported 5.3 to 6.1% correctness gain over the strongest baseline. An August 2026 paper, PURPOSE, shows the resolver itself is an attack surface: injected content framed as an update rather than a contradiction slips through. Provenance defends against both. An arbiter that cannot show its sources cannot be audited, and will eventually be gamed.

Frequently asked questions

Is a single source of truth for AI agents a realistic goal?

No, and chasing it is where the full-time job comes from. Code, tickets, docs, and chat each record a different moment in a decision's life, and none back-propagates to the others. The realistic goal is a single arbiter: one place where conflicts get resolved at query time, with the losing records still visible.

How is this different from a knowledge graph or an agent memory layer?

A knowledge graph stores relationships someone asserted. Agent memory stores conclusions the agent reached in past sessions. Both are inputs to arbitration, and neither performs it, so neither is a source of truth on its own. When two edges disagree, or a memory note contradicts a new PR, something still has to rank them with a reason. See a context engine versus a knowledge graph.

What about permissions? Our senior engineer can see things most people can't.

That is a feature of the human oracle people forget to replicate: Priya answers differently depending on who asks, because she knows who is allowed to know about the incident. An arbiter that resolves a conflict using a document the asker cannot read has leaked it. Permission filtering has to run before ranking.

Does showing provenance slow the agent down?

It adds a few hundred tokens of citations and removes the round trip to a human. In the 2026 Stack Overflow survey, 93% of developers said source attribution matters at least somewhat when deciding whether to trust an AI answer. Without it, the answer is one more record for Priya to adjudicate.

What does decision-grade context look like in practice?

It looks like an answer that names a winner, shows the losers, and respects who is asking. At Codat, engineers describe tribal knowledge spread across teams that used to cost three or four hours of code reading per question. Unblocked is the context engine that does that arbitration for them across code, PRs, Slack, Jira, and Confluence.

Ask about the retry policy and the answer says five attempts with jitter, cites the April Slack thread and the PR that followed the incident, marks the runbook and the March ticket as stale, and omits anything the asker lacks permission to see. Priya appears in the citations as the person who made the call. The agent proceeds. Nobody pings anyone.

"There's a lot of tribal knowledge spread across teams. Unblocked gave us a way to self-serve answers without spending three or four hours reading code."

— Matt Thompson, Staff Software Engineer, Codat

Codat's other observation matters here: an obscure build error resolved by surfacing a Slack thread from two and a half years earlier with the same issue. That thread was never going to be in a wiki. Only something that ranks every source the team writes to, and says where the answer came from, was going to find it.

Where to put the arbitration instead

Take it out of one person's head and put it in the retrieval path. Three moves this quarter, each of which shrinks Priya's queue without asking her to write anything down.

  1. Inventory the known conflicts. Ask your longest-tenured engineers which questions they answer more than once a week. Each one is a place where records disagree and the arbitration happens in a DM.
  2. Require provenance from every agent answer on those topics. If a tool cannot show which source won and which lost, it is retrieving, and someone downstream is still arbitrating. Being the context engine yourself never scaled, and agents make it worse.
  3. Connect the arbitration to every source, with permissions: code, PRs, chat, tickets, docs, incidents. The correction to any record can live in any of the others.

The source of truth for AI agents should be an arbiter that can show its work, and the human who used to be that arbiter should get to go back to engineering. To see it run against your own repos, Slack, and tickets, try Unblocked.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.