
Boards have started asking a pointed version of the same question: why does the demo always work and the rollout never does. Two years of pilots have impressed everyone in the room and then failed to survive contact with the rest of the business. Is our data that unique? That complex? That bad?Â
The model is rarely what breaks. Models are good, they improve every few months, and they broadly get cheaper. If raw capability were the constraint, the pilot would have failed too, and the pilot is the part that works. What breaks at scale is that the system doesn’t know what the company’s own data means once it butts up against the edge of its nicely-scoped and defined proof of concept environment.Â
Two different things sit inside that. The first is technical: where a number comes from, which table feeds which report, how a figure was calculated. The second is institutional, and it’s the harder one. Which of the four columns called revenue is the one finance closes the books on. That the EMEA figures always land three days late. That “active customer” means something specific at your company that it doesn’t mean at your competitor. None of that is written down in the systems. It’s carried by people.Â
The artifact starts drifting the day you finish itÂ
For twenty years the answer has been to write it down. First data dictionaries, then catalogs, then business glossaries and semantic layers, then ontologies. The vocabulary turns over every few years but the bet underneath never does. Build the right place for people to record how the business works, and they will record it and keep it current. Â
They do record it. Keeping it current is where it collapses.Â
A team renames a metric and the reports downstream keep the old label. A pipeline gets rewritten and the definition that described it no longer describes what runs. Someone builds a new dashboard because the existing one was almost right for what they needed, and now the company has two, slightly different, both in use. None of this is negligence. It’s a normal week and that’s why it’s so hard to build into demos.Â
So this isn’t a discipline problem, however much it looks like one. No review cadence can keep documentation current, because the review cadence is not what sets the pace. The business changes continuously and the documentation is revisited periodically, and so the gap between those two rates only widens.Â
Why people could live with this and agents can’tÂ
For twenty years the drift was survivable, because the thing reading the documentation was a person, and people are very good at working around a stale catalog. An analyst notices a number doesn’t match what finance reports, asks the colleague two desks over, and gets told which version the CFO actually uses. The correction takes ninety seconds and is recorded nowhere because the analysts always work around it.Â
So the documentation was never really doing the work. The people were doing the work, and the documentation was a hint that helped them do it.Â
An agent has no one to ask. Handed a stale definition it will use it confidently, because it has no independent sense of what the right answer looks like and no idea that someone down the hall could have fixed it in ninety seconds.Â
That is most of what separates a pilot from a rollout. In a pilot, a small team hand-feeds the system everything it needs to know about one well-groomed slice of the business, doing deliberately what analysts have always done in passing. It doesn’t survive being asked ten thousand times a day across every domain in the company.Â
Anthropic recently published figures from its own internal analytics work that put a number on the decay. Documentation covering a data model that changed daily went stale within weeks, and without active maintenance, offline accuracy fell from roughly 95% to 65% over a single month. That is a company with every possible advantage in scaling this kind of effort, measuring itself honestly.Â
Deriving instead of declaringÂ
The question my team has spent the past two years on is whether the premise can be inverted. Rather than asking people to declare what the data means and then maintain the declaration, how much of it can be worked out from what the organization already does?Â
For the technical half, the answer is ‘most of it’. How data moves, what feeds what, which reports are used and by whom: that is a byproduct of operating, and it can be reconstructed from the systems themselves. Because it’s derived rather than written by hand, it stays current without anyone maintaining it.Â
The institutional half is harder, and it’s the half that decides whether any of this is trustworthy. That knowledge is encoded in patterns rather than in any single authoritative record. A transformation, a certification, an access grant, a deprecation: each says something faint about what a thing means and how far it can be trusted, and any one of them alone will mislead you. Read together they corroborate each other, and from the pattern you can infer ownership, reliability and preferred use without needing any single declaration to be correct.Â
The way to get this badly wrong is to let the result become opaque. If the inferred knowledge disappears into model weights where nobody can inspect it, it will be harder to trust than the hand-written catalog ever was: you can’t anticipate what a change will do, you can’t localize a correction, and every repair risks breaking something unrelated. Software has already been through this. AI writes the code, but the code, not the prompt, is what the company keeps, reviews and repairs. Institutional knowledge has to work the same way.Â
Where people belong in thisÂ
There is a version of this argument circulating that ends with the people removed: the systems come to understand the business on their own and the data organization shrinks. I’d be careful with anyone selling it.Â
Judgment is where trust gets established. Resolving genuine ambiguity, drawing organizational boundaries, signing off on consequential changes, stepping in when the evidence is thin: none of that delegates well. What does delegate is the maintenance. Noticing what changed, proposing the update, checking what the update did.Â
The measure of success isn’t whether people are in the loop. It’s whether the attention they have to spend stays flat while the number of agents and domains and workflows keeps climbing.Â
For a decade the industry has argued about which artifact to build and who should own it. The more useful question is how much of it anyone needed to write down in the first place, and what keeps it true once the writing stops.Â


