AI Business Strategy

Your agents are failing at the parts of the job nobody wrote down

By Chase W. Hughes, three-time founder, ProAI

Salesforce AI Research built a benchmark that does something most agent evaluations avoid: it drops leading models into a realistic sandbox of a system businesses actually run on, with nineteen expert-validated tasks across sales, service and quoting. The results are worth reading as three numbers rather than one. On tasks the researchers class as Workflow Execution, the best agents cleared 83% in a single turn. Overall single-turn success was around 58%, and when the task required multiple turns of conversation with a simulated user, it fell to 35%. 

The usual reading of a spread like that is a capability story: agents are not clever enough yet, and the next model will close the gap. I think it is a specification story, and the difference matters because one of those problems gets solved by waiting and the other does not. Workflow Execution is the category where somebody had already written the job down. The rest of the work was being held in a person’s head, and retrofitting an agent into that seat hands it the job minus the half nobody ever had to say out loud. 

The gap between 83% and 35% is a documentation gap 

Look at what actually changes across those three numbers. The task domain is constant, the models are constant, the platform is constant. What varies is how much of the job existed in writing before the agent arrived — and the more the agent has to extract from a person mid-conversation, the worse it does. 

The same paper reports a third finding that makes the point harder to argue with. The agents showed “near-zero inherent confidentiality awareness”, improvable with prompting “but often at a cost to task performance”. Nobody had written down what this role must not say, because the rule had always lived in an employee who understood it without being told. The agent did not fail a capability test there; it faithfully executed a job description with a hole in it. 

One caveat, said plainly because the number deserves it: this is a single benchmark, on a single platform, using synthetic data. It is evidence about where agents struggle, not a measurement of your company. What makes it useful is that the three figures come from the same environment, so the spread between them is telling you something about the tasks rather than about the models. 

Agent-first design starts with an org chart, not a model choice 

Here is what changes when you design for agents from the beginning instead of bolting them onto a product built for people. The unit of design stops being the model call and becomes the role: what this agent is for, what it is allowed to do, who checks its work. That sounds like management because it is management, and the most useful published example I know reads like a staffing plan. 

Anthropic’s engineering write-up of its Research feature describes an orchestrator-worker pattern — a lead agent that plans, then delegates to specialised subagents running in parallel. What is striking is how prescriptive they had to be about effort. Their prompts encode explicit scaling rules: “Simple fact-finding requires just 1 agent with 3-10 tool calls, direct comparisons might need 2-4 subagents with 10-15 calls each, and complex research might use more than 10 subagents with clearly divided responsibilities.” That is a headcount table, and they wrote it because early versions spawned fifty subagents for simple questions. 

The payoff and the price are both reported. On their own internal research evaluation, the multi-agent configuration outperformed a single strong agent by 90.2%, and multi-agent systems in their data burn about 15 times the tokens of a chat interaction. So the decision to run a hierarchy is a staffing decision with a budget attached, justified only where the task is worth the spend. It is also a decision about which role gets which model — and if any of your agents run on edge hardware, as some of mine do, the cheap role is not a preference but a constraint you design the whole hierarchy around. 

Delegation failures in that system look exactly like delegation failures anywhere. Their write-up describes a lead agent giving instructions vague enough that “one subagent explored the 2021 automotive chip crisis while 2 others duplicated work investigating current 2025 supply chains”. I filed a patent-pending multi-agent research system in early 2023, before this pattern had a widely used name, and the work that mattered was never the model. It was the brief. 

Where the management metaphor breaks, and what that teaches 

The metaphor gets you to the org chart, then it starts lying to you, and the place it breaks is the most useful finding in this whole area. In June 2025, Cognition’s Walden Yan published a piece called Don’t Build Multi-Agents, arguing that agents must “share context, and share full agent traces, not just individual messages”, because “actions carry implicit decisions, and conflicting decisions carry bad results”. His illustration is a Flappy Bird clone split between two subagents: one builds a background that looks like Super Mario Bros, the other builds a bird that moves wrong, and the final agent inherits the job of combining two miscommunications. His verdict at the time was that running multiple agents in collaboration “only results in fragile systems”. 

Ten months later he published a follow-up saying the opposite had become true in a narrower shape: multi-agent systems that actually work, where “multiple agents contribute intelligence to a task while writes stay single-threaded”. Read that constraint again, because it is doing enormous work. Extra agents may read anything, reason about anything and advise on anything; one thread does the writing. 

Then comes the result that no manager would ever design. Their code-review agent catches an average of two bugs per pull request written by their own coding agent, roughly 58% of them severe — and it works best when the coding and review agents share no context beforehand. Clean context makes the reviewer sharper, because it has to reason backwards from the implementation and because attention degrades over long contexts. Withholding the specification from your reviewer is career-limiting advice for a human team and it is the correct architecture here. 

That is the honest shape of the metaphor. It is right that you should be designing roles, delegation and review rather than prompts. It is wrong about why: human teams need shared context to stay aligned, while agents often get better when you deny it to them. The units of an agent org chart are context, write authority and verification — not seniority, not trust, and not anything you would put in a job ad. 

“Won’t better models make this go away?” 

It is the fair objection and it deserves a real answer, because part of it is correct. Yan’s own reversal is evidence that models got better at exactly the coordination he said they were bad at. Anything I write about today’s ceilings has a shelf life of about a year, and I would not bet against the next generation dissolving several of these problems. 

But watch what replaced the old ceiling in his follow-up. Their experiment pairing a fast cheap model with an expensive one failed for a specific reason: “the quality ceiling was set by the primary, and the primary wasn’t strong enough” — a weak model does not know when it is out of its depth. Where it worked, between two frontier models, the interesting part was that delegation “becomes a capability router rather than a difficulty escalator”. Both of those are organisational problems, not capability problems, and better models reshaped them rather than removing them. 

Anthropic is candid about the same limit from the other side, noting that most coding tasks involve fewer genuinely parallelisable pieces than research, and that agents “are not yet great at coordinating and delegating to other agents in real time”. The pattern that keeps showing up is that scaling moves the frontier and the failure mode moves with it, from “can it do the task” to “who decides, who writes, and who checks”. No model release answers that question for you. 

Three questions that make it an architecture 

Before any of this is a design rather than an experiment, I would want three things written down. First: can you state a role’s brief precisely enough that its output can be checked without replaying the whole conversation? If not, you do not have a role, you have an ambiguity you are about to run in parallel. 

Second: who writes? Keep the write single-threaded and explicit, and let the other agents contribute judgement rather than actions — that constraint is what turned a sceptic’s position around, and it costs you very little. Yan notes that most multi-agent setups in the wild are read-only subagents that “mostly resemble tool calls rather than true multi-agent collaboration”, which is worth knowing before you tell your board you run a multi-agent system. 

Third: what does done mean, in a form something other than a person can check? Cognition make a sharp observation about the spectacular agent demos — the ones that build a browser or a compiler — that they “all share a property most real software doesn’t: a simple, verifiable success criterion”. Where you cannot write that criterion, you do not have an autonomous step; you have a draft and a review, which is a perfectly good product as long as you price it that way. 

The retrofit question worth asking first 

None of the three questions above is exotic, and that is rather the point. What makes them hard is that answering them forces a company to write down work it has never had to specify, because the specification lived in whoever was doing it. The 35% and the near-zero confidentiality score are what it looks like when an agent inherits that job as-is. 

So before the architecture diagram, ask the retrofit question about any seat you are considering: what did the person doing this know that is not written anywhere? Answer it and you have most of an agent design, plus a document your human team probably needed anyway. Skip it and you will ship something that behaves exactly like a new hire who was handed a login and no induction — and, unlike that new hire, never asks.  

 

Related Articles

Back to top button