
Ask Saurab Dhir for the strongest argument against everything he believes, and he tells you machine learning’s founding parable better than his opponents would. For years, researchers built careful, hand-curated datasets and clever, specialized methods; then the internet-scale approach arrived and flattened all of it. “The long history of methods, like raw internet beating every carefully constructed NLP dataset, proves human judgment to be obsolete,” he said, reciting the doctrine that hardened out of that history: scale wins, curation loses, and human judgment about data is a bottleneck to automate away.
Then comes his turn, delivered with the timing of someone who has rehearsed it against himself. “That pattern stops at the physical world, where data can’t be scraped for free. It has to be made by paid humans, and people are incentivized to game the system for reward. Ambient collection, the way Tesla does it, isn’t in place here.”
The turn is the whole argument, and the claim follows: model architecture is becoming a commodity, solved-ish and shared, while the durable edge in AI belongs to whoever keeps their training data honest and high-quality at scale. The moat, on this view, was never the model. It is the data, and the verified kind specifically.
Commodity, in his usage, does not mean worthless. It means available: strong architectures increasingly ship in the open, get copied within months, and stop separating the leaders from everyone else. When the same model design is within reach of every serious team, the thing that still differs is what you feed it, and whether you can trust what you fed it. That is where the argument points.
His vantage point is worth naming, because it grants authority and complicates it at once. Dhir leads data and machine-learning work at Mecka AI, a company whose whole business is selling verified physical-world training data, so his thesis, believed, enriches his employer. It also means he spends every working day at the junction the thesis describes, deciding which recordings of human work deserve to become training signal. Pundits theorize about data quality. Dhir adjudicates it, thousands of clips at a time. The conflict is real, and so is the firsthand view; a fair reader holds both.
The evidence from the language-model era cuts his way more than the parable admits. The headline gains of recent years owe an underappreciated debt to data work: deduplication, filtering, curation, and human feedback quietly drove leaps that architecture papers took credit for. The bitter lesson was always subtler than its bumper-sticker form. Scale won, but the winners scaled clean data, and the labs learned to guard their data pipelines as crown jewels while saying in public that architecture was the game.
Physical AI tilts the board further. Language data was a found resource, free and nearly infinite. Demonstrations of physical work are manufactured, one paid contributor at a time, and because they are produced rather than captured passively, the way a Tesla fleet records the road, they both cap supply and invite gaming: every payment gives someone a reason to shade the work. In that world the compounding asset is a corpus you can prove is real. Whoever holds it can trust its own benchmarks, train without fear of quiet poison, and sell confidence rather than footage. Whoever lacks it is scaling a rumor.
It matters what verified means, because the word gets used loosely. Labeled data tells you what a clip claims to show. Verified data tells you the claim survived interrogation: the task physically happened, the label matches the pixels, the contributor’s history does not reek, the batch passed checks nobody who produced it controlled. That interrogation is expensive, unglamorous engineering, which is exactly why Dhir thinks it compounds into a moat. Cheap advantages get copied. Expensive, boring ones get conceded, and he has built his career on the expensive, boring side of the line.
If he is right, the effects land on budgets before leaderboards. The physical-AI companies acting on the thesis will spend on data operations the way the last generation spent on compute: verification teams staffed like security teams, provenance treated as a first-class feature, acquisitions chosen for data discipline over demos. Watch the job postings, Dhir suggests, and the argument settles itself early.
What would prove him wrong? He named a test, then sharpened it so it could actually fire. “A leading physical-AI system trained mostly on synthetic or scraped data, verified human demonstration only a small fraction of the corpus, that works on real hardware outside a lab demo. If that lands within two years, I’m wrong.” What he expects instead is the opposite tape: leaders pulling ahead on proprietary, verified data, and every synthetic or in-the-wild pipeline that works turning out, on inspection, to have high-quality curated human data at the bottom of it.
There is something clarifying about an argument that names a condition under which it dies. Most theses in this industry are built to survive any outcome; Dhir’s, in its sharpened form, can fail, on schedule, in public. The commodity-model claim will irritate architecture researchers, and the verified-data claim will read, to skeptics, as a data vendor calling his own inventory the future. Both can be true at once. The parable that opened this piece was also told, in its day, by the people who happened to own the compute. The lesson still held.



