Enterprise AI

Why Enterprise AI Agents Fail at the Data Layer

By Dileep Mundakkapatta

A governed architecture for schema mapping, transformation, and human approval 

Sooner or later, every large enterprise has to move data it has scattered across dozens of systems into one place — a migration, a consolidation, a new target model that a pile of legacy applications is supposed to feed. And before a single row moves, someone has to work out how each source field maps onto that target model: what it means, what it becomes, which rules apply. That mapping is slow, manual, expert work — and it is exactly the kind of thing teams now want to hand to an AI agent. 

It’s a reasonable instinct, and it fails in a specific, dangerous way — not because the model isn’t smart enough. It fails because the data layer being mapped is decades of inconsistent definitions, half-true documentation, and field names that quietly lie about the values beneath them. An agent that maps on those names produces answers that sound authoritative and are wrong. 

So this isn’t a general complaint about LLM data quality. It’s about one high-stakes job — mapping source systems onto an enterprise data model — and why capable agents keep getting it wrong. One concrete example runs through this piece: a mapping a model proposes with real confidence, and the simple check that shows it’s wrong. That gap — between sounds right and is right — is the whole problem in miniature. 

The real problem isn’t intelligence, it’s representation 

Go one level deeper into that mapping job. The source systems typically carry fields like: 

They look related. Sometimes they even are. But device_type might mean the manufacturer’s hardware model in one system, the operational role of the device in another, and a set of locally invented abbreviations — documented nowhere — in a third. Fields with identical names carry different meanings; fields with different names carry the same one. 

Before any data moves, each source attribute has to be mapped to the right target attribute, and then transformation rules have to be written for formats, enumerated values, relationships, mandatory fields, and identifiers. Traditionally this is slow, manual work done by analysts, SMEs, and architects reading data dictionaries and interviewing whoever still remembers how the 2011 provisioning system worked. 

It looks like a perfect job for an AI agent. It is — but only inside a governed architecture. To see why the naive version fails, follow one field. 

A mapping that sounds right — and isn’t 

Inventory-A has a column called device_type. Its data dictionary defines it, reasonably enough, as “Type or role of the network device.” The target model has an attribute called network_function — an operational role, with four allowed values: 

 Map on the name and description alone — which is exactly what a documentation-driven agent does — and the answer writes itself. Ask a frontier model to propose the mapping, and it returns something like this: 

A confident, well-reasoned mapping. It is also wrong. Here is what the column actually contains: 


Those aren’t roles. They’re hardware models. The documentation calls the field a role; the data says model, and the data is the ground truth. 

This is the single most common shape of enterprise data debt — and a bigger model doesn’t fix it, because a better-worded lie is still a lie. The mapping is semantically convincing and operationally wrong, and nothing in a name-based pipeline can tell the two apart. 

The fix isn’t a smarter prompt. It’s to stop treating the mapping as a single question sent to a single model, and start treating it as a controlled decision process — one that keeps the model where it’s genuinely useful and wraps it in things it cannot provide. 

Surround the probability with certainty 

The idea is simple to state: let the model do the one thing it’s good at — reading ambiguous schemas and proposing candidates with reasoning — and surround that probabilistic core with real profiling of the data, deterministic rules it cannot overrule, and a human who is accountable for the calls that matter. 

The seven stages of a governed mapping workflow. Colour marks what governs each stage — probabilistic, deterministic, human, or stored knowledge. Image generated by the author. 

Seven stages carry the weight. I’ll walk them against the field above. 

  1. Profile the data, not the documentation

The first stage never trusts a field name. It reads representative rows and reports what is actually there — null rates, distinct values, patterns, candidate keys. For our field it reports four short alphanumeric codes, none of which resembles an operational role. This is where “type or role of the network device” gets confronted with ASR9000. 

  1. Retrieve what’s already known

Data dictionaries, previously approved mappings, transformation standards, naming conventions, known exceptions. Retrieval-augmented generation belongs here: it gives the proposal step context. But retrieval only supplies evidence — it never guarantees the answer, and a lot of “agentic” designs quietly let it stand in for validation. 

  1. Generate candidates, plural

Ask for more than one mapping, each with its target attribute, transformation logic, supporting evidence, and reasoning. Forcing a single confident answer out of a probabilistic system throws away exactly the uncertainty you most want to see. 

  1. Validate with rules the model can’t overrule

This is the stage that catches our example. It is deterministic and independent of the model, and two plain questions are enough: 

 Neither question consults the model’s confidence; both consult only the data. The second is the sturdier of the two — it inspects the raw values directly, so no amount of confident relabelling by the model can suppress it. On our field, both fire. 

  1. Score confidence and risk separately

A model’s self-reported confidence is not a risk assessment. Risk is a different axis: a mapping can be clean and highly confident and still deserve review because it touches a business-critical identifier. Collapse the two and you auto-approve exactly the mappings that most deserve a second pair of eyes. 

6 & 7. Human approval, then a registry that remembers 

Low-risk, well-supported mappings can auto-approve. Ambiguous or high-impact ones go to a subject-matter expert who sees everything — sample values, evidence, validation results, alternatives — and can approve, reject, or edit. Every approved decision lands in a versioned registry with lineage and effective dates, and future proposals retrieve those decisions as examples. The system grows more consistent over time because it is grounded in approved organizational knowledge instead of starting cold for every new source. 

Trace the example through the two paths and the difference is stark. The naive agent applies device_type → network_function and moves on. The governed workflow profiles the values, both checks fire, and the mapping is rejected and routed to a human, with the reasons attached. 

And the workflow isn’t merely a rejector. A genuinely clean field like equipment_role — whose values PE, CE, AGG, CORE map straight onto the target roles — passes validation cleanly, yet because network_function is business-critical it is still held for sign-off rather than auto-applied. Three fields, three outcomes — applied blindly, rejected, held — and only the governed path gets them right. 

Where these systems actually break 

Bad knowledge that reinforces itself 

Suppose a mapping is approved under deadline pressure and later turns out to be wrong. If it just sits in the registry, future agents retrieve it as trusted evidence and faithfully reproduce the mistake across every new system. Continuous learning without continuous governance is just an efficient way to spread an error. Mappings need owners, versions, expiration, and a way to be deprecated — the registry has to distinguish proposed from approved from deprecated from conditional. 

Reaching for an agent when a function would do 

Reformatting a date does not need an autonomous agent. Deterministic transformation code is cheaper, faster, testable, and more predictable. Save the agent for what actually has semantic ambiguity: fragmented documentation, competing interpretations, judgment calls. Using probabilistic reasoning where deterministic engineering suffices is how you get a system that is both more expensive and less reliable than the thing it replaced. 

From concept to production 

None of the above is a finished product, and it would be dishonest to pretend two checks and a diagram add up to one. The example uses a target model with a handful of attributes; a real one has hundreds, which turns “which attribute does this field map to?” into a retrieval-and-ranking problem in its own right. A production validator needs far more than the two rules shown here — type compatibility, referential integrity, uniqueness, conditional cross-field rules, real transformation logic rather than relabelling. And above all it needs an evaluation harness that measures mapping accuracy against known-correct answers, so you can tell whether the agent is actually good rather than merely confident. 

The contribution here is not a tool. It’s a shape: where to put the model, where to put deterministic rules, and where to put people — so that a plausible-but-wrong mapping is caught by something that reads the data instead of the documentation. 

The practical lesson 

Enterprise data transformation is not a prompt. It’s a controlled decision process, and the strongest version of it assigns each part to whatever handles it best: the model interprets ambiguous schemas and drafts candidates; deterministic rules enforce the technical and business constraints; humans resolve genuine uncertainty and own the high-impact calls. 

Enterprise AI agents don’t fail at the data layer because they aren’t smart enough. They fail because the data layer is a sediment of years of inconsistent definitions, undocumented decisions, and fragmented knowledge — and no amount of model intelligence reads a value’s true meaning off a field name that lies about it. Agentic AI can genuinely help with that mess. But only when the probability at its core is wrapped in deterministic checks, trusted knowledge, and a human who signs their name to the result. 

Related Articles

Back to top button