By Seth Earley, Founder & CEO, Earley Information Science
Published: September 29, 2026
Last Updated: October 7, 2026 | Version 1.0
────────────────────────────────────
Who This Is For: Content operations leaders, Chief Data Officers, information architects, AI/ML leaders, and Enterprise Architects building AI-ready content environments. Also valuable for KM leaders and anyone responsible for RAG accuracy and content governance.
Prerequisites: Basic familiarity with retrieval-augmented generation (RAG) and enterprise content management.
────────────────────────────────────
Ask six groups inside a software company to describe the same product feature and you will get six descriptions.
Engineering documents what was built. QA documents what broke. Support documents what customers actually encountered. Technical publications documents what the user needs to do. Marketing documents what the feature means in the market. Product management documents what it is for.
None of these are wrong. They describe the same thing through different lenses, and each lens captures something real. The problem is that overlap, redundancy, and contradiction across these six perspectives makes retrieval unreliable. Point a RAG pipeline at all six sources and it will retrieve a version it determines is correct. The problem is that it may not be the best answer — and there is nothing in the system to tell it otherwise.
The standard prescription is to designate a source of truth. Lock it down and feed it from the other systems. That is reasonable when an answer is unambiguous and specific. The problem is that content needs context, and different sources can provide different contexts. What the enterprise actually needs is not a source of truth. It needs a source of authority.
The six-lens challenge is not a case of five groups being wrong and one being right. Marketing's description of a feature is not a defective version of engineering's — it answers a different question for a different reader with a different tolerance for technical detail. Support's description captures real problems that never appear in a specification because nobody predicted them. Collapsing all six into one document does not produce clarity. It produces a document that serves no audience well.
What is missing is adjudication. In most enterprises, nothing decides which lens the organization will stand behind when two of them disagree about a specific fact. In organizations where content governance has not matured, ownership may not be clearly assigned at all. When that is the case, the retrieval system decides — by whichever chunk scores highest — and humans will not know the result was wrong unless a measurement process is in place to catch it.
Consider two contributors to the same knowledge base.
Joe is an engineer who tested a behavior and documented it when it shipped. His documentation has provenance: a named person, a verified build, a timestamp. When an AI cites Joe's documentation, the answer can be defended.
Sue works in support. A customer reported a problem that Joe's specifications said was not supposed to occur, but Sue found a workaround that resolved it for that customer in that environment.
Sue's fix is true in the sense that it solved a real problem. But it has never been tested by engineering. It worked in one configuration for one customer and may not generalize. Sue also has no authority to change what the organization asserts about its own product. Her input is valuable, but it is not authoritative. The path forward runs back through engineering: Sue identifies the issue, engineering tests and validates the fix, engineering documents the result. At that point the answer is verifiable — because the organization is prepared to stand behind it.
This distinction between "true" and "verified and authorized" is the distinction between a source of truth and a source of authority. The retrieval system cannot tell the difference on its own.
There is no stable source of truth because truth about a product changes every time the product does. Versions ship, features deprecate, workarounds get superseded. A statement that was correct in version 4.2 is wrong in version 4.3, and in most content environments nobody retracted it. It sits in the corpus — well written, correctly tagged by every standard the organization had at the time of publication — and completely wrong now. Multiply that by a decade of releases and thousands of topics and you have the actual condition of most enterprise content: not false when written, false now, with nothing in the record to mark when it stopped being true.
This is why searching for a single authoritative document fails. The unit that needs authority is not the document. It is the assertion, bound to a version, a platform, a configuration, and a date. A document can be current and still contain three statements that expired at different times.
A source of authority is different from a source of truth in a critical way. A source of truth implies a repository — something you can point to. A source of authority is a process with an owner: inputs arrive from support tickets, field observations, QA runs, and customer feedback; a designated person reviews, tests, approves, and publishes. That is an assigned human role, not a system configuration.
A content operations assessment conducted inside a large enterprise networking company found a pattern that is consistent across industries. Roughly 60% of respondents said every major content set had a named owner with defined responsibility for accuracy and currency. By that measure, ownership appeared to be solved. But only a small number of respondents said retrieval and answer quality was measured against defined benchmarks, or that the full set of content an AI could retrieve from was inventoried and held to a known standard.
Naming an owner addresses custodial responsibility. It does not address accountability — who decides whether a specific assertion inside a document is still correct, and who answers for it when a customer or field team acts on it and it is wrong. Those are different jobs, and most organizations have defined the first while assuming it covers the second.
Two failure modes surface when authority is absent. In the first, sources conflict and the model resolves the conflict by retrieval score rather than by designated authority, producing a confident answer that nobody in the organization formally endorsed. In the second, no source has been designated to address a particular question at all, so the system returns nothing — and the resulting support case or compliance escalation costs far more than the documentation gap that caused it. Ungoverned AI produces wrong answers when sources conflict and no answers when no source has been designated to speak. Both are expensive.
AI is well-suited to the collation step of the authority process. A model can look across engineering specifications, release notes, support cases, and QA reports simultaneously — more sources than a person could hold in working memory — and report what those sources collectively appear to say. For research and drafting, that efficiency is real.
What AI cannot do is adjudicate. It can propose that two topics are functionally the same; it cannot decide whether they must remain separate because one applies to a product variant about to diverge. It can propose chunk boundaries; it cannot determine what is important enough to codify, because knowing everything is not the same as knowing what matters. It can draft a reconciliation of five contradictory sources; it cannot warrant the result.
Giving a model adjudicative authority is not an architecture decision. It is an accountability gap. Authority is an accountability relationship, and a model cannot be held accountable. There is also a practical constraint: models are probabilistic. The same question submitted an hour later, with no inputs changed, will likely produce different phrasing. That variability is acceptable for exploration. It is not acceptable for an authoritative answer. The architecture needs a certified snapshot: this assertion is verified, this is what the organization asserts today, and the model retrieves from it rather than revising it. Everything outside the snapshot is open for remediation; everything inside it is settled until a designated human unsettles it. Without that boundary, drift becomes the system's default behavior.
Building a functioning source of authority requires decisions that most content governance programs have not yet made explicit.
First, name the authority per content domain rather than per system or role. Identify who has the standing to bless a specific assertion about a specific product area. Second, distinguish inputs from authority in the architecture. Support tickets, customer voice data, and field observations are inputs that require a defined intake path so they are not lost — but they should not write directly to the authoritative record. Third, certify snapshots as verified, dated, and attributed. This is the official version; the retrieval pipeline treats it as read-only. Fourth, bind assertions to their conditions in structured metadata: version, platform, configuration, effective date. These belong in fields the retriever can use, not in narrative text the retriever may drop or misread. Fifth, scope what the AI agent is authorized to do: it can collate, propose, and draft, but certification remains a human responsibility. Sixth, keep provenance retrievable. The answer should be traceable not to "the system says" but to "verified by engineering on build X on date Y." That is what a customer, a regulator, or a field team requires in order to trust and act on the answer.
Every organization working to make its content AI-ready is solving what appears to be a retrieval problem. Underneath it is a governance question that predates AI by decades and that AI has made impossible to defer: who is empowered to say what is true here?
There is no AI without information architecture, and there is no information architecture without someone empowered to maintain it. Provide the designated authority and the structured metadata that supports it, and the retrieval system has the data it needs to produce answers the organization can stand behind. Without it, the agent will respond with information that is inaccurate, outdated, or inapplicable — delivered with the same confidence it would show for a verified answer.
Read the original version of this article on VKTR.