The cost of producing a translation fell by three orders of magnitude. The cost of verifying one did not move at all.
Translation is now one of the cheapest things an enterprise can buy. Raw machine output is priced at roughly $10 to $20 per million characters, which works out at something under 15 cents for a thousand words. A professional human translation of the same thousand words, translated, edited and proofread, still costs $150 to $300.
That is a gap of roughly three orders of magnitude, and over the past three years most organisations have quietly walked across it. Marketing content, support documentation, product listings, internal policy, contract summaries and patient-facing instructions have all migrated to machine output, often without a formal decision being taken anywhere.
What has not migrated is the review layer. The cost of generating a translation collapsed. The cost of checking one stayed exactly where it was, because checking still requires somebody who reads both languages. The research on how to make AI translation more reliable keeps arriving at the same conclusion from different directions: the bottleneck has moved. It is no longer generation. It is verification.
This report sets out what the 2026 evidence base actually shows about machine translation accuracy, where quality breaks down, and why the organisations most exposed to translation risk are usually the ones least aware they are carrying it.

Figure 1. The 2026 translation price ladder. Raw output converted from per-character pricing at approximately six characters per word. The only variable that meaningfully changes across these three tiers is how much human verification is attached. Alt text: log-scale bar chart showing raw machine translation at under 15 cents per 1,000 words versus $150 to $300 for full human translation.
FINDING 01
Accuracy is not a single number, and the average hides the failure
The most quoted statistic in this market is some version of “machine translation is now over 90% accurate.” It is not wrong. It is just measured at the wrong resolution.
A study published in the Journal of General Internal Medicine took 20 commonly used emergency department discharge instructions and ran them through a mainstream engine into seven widely spoken languages, then had native speakers score the output. The accuracy rate ranged from 94% in Spanish down to 55% in Armenian. Same tool. Same source text. Same afternoon.

Figure 2. Accuracy of a mainstream translation engine on identical emergency discharge phrases, scored by native speakers. Seven languages were tested; six published discrete accuracy rates. Source: Journal of General Internal Medicine. Alt text: bar chart of AI translation accuracy by language, Spanish 94% to Armenian 55%, with threshold lines at 80% and 90%.
A 39-point spread inside a single product is not a rounding error. It is the difference between a tool you can deploy with light oversight and a tool that should not be deployed at that task at all. The researchers concluded plainly that output was inconsistent between languages and should not be relied on for patient instructions.
Earlier work published in JAMA Internal Medicine put a sharper edge on it. Analysing 100 sets of discharge instructions, the authors found 92% accuracy in Spanish and 81% in Chinese, and then measured what the failures actually did: 2% of the Spanish errors and 8% of the Chinese errors carried the potential to cause clinically significant harm. In one case, an instruction to hold a medication was rendered as an instruction to keep taking it.
The headline accuracy figure and the harm rate are different measurements, and only one of them is on the dashboard.
Most enterprise buyers procure on the first number. Almost none of them measure the second.
FINDING 02
A wrong translation gives off no signal
Software failure is usually loud. A malformed API response throws an error. A broken deployment fails a health check. A pricing bug shows up in the numbers by Friday.
Translation failure is silent by construction. The output is grammatical, fluent, confident, and formatted correctly. It arrives in a language the person who requested it cannot read, and it is delivered to a person who has no way to compare it against the source. Every party in the chain is looking at something that appears complete.
This is the same structural problem The AI Journal has covered elsewhere in enterprise data and agentic systems, where the most dangerous failures are the ones that look exactly like the right answer. Translation is arguably the purest expression of it, because the entire point of the transaction is that the buyer cannot evaluate the product.

Figure 3. The structural asymmetry in machine translation deployment. Neither party to the transaction is positioned to detect an error. Alt text: diagram of the AI translation pipeline showing the verification step bypassed between output and reader.
FINDING 03
Detecting errors is harder than making translations
The intuitive fix is to automate the checking too. If a model can translate, surely a model can grade the translation.
The 2026 evidence is more sobering than that. At the tenth Conference on Machine Translation, the shared task on automated evaluation systems reported that accurate error detection, and balancing precision against recall, remain persistent challenges. Reference-based baseline metrics still outperformed large language models at the segment level, which is precisely the level at which an individual sentence goes wrong in a contract or a dosage instruction. Robustness across a broad diversity of languages was flagged as a major unresolved problem across all three subtasks.
The organisers of the general translation task that year titled their findings paper, without much diplomacy, “Time to stop evaluating on easy test sets.”
Read those two results together and the picture is uncomfortable for anyone running an automated multilingual pipeline. Generation has raced ahead. Evaluation has not kept pace. The system producing your translation is materially more capable than any system currently available to tell you whether it worked.
What did work
One result from the same conference points at the practical answer. In the error-correction subtask, the winning submission did not attempt to detect and repair errors inside a single output. It generated multiple candidate translations from different models and selected the strongest one. That approach beat the automatic post-editing method, which had a known tendency to overcorrect and degrade output that was already fine.
The operational lesson generalises well beyond a research leaderboard. Where independent engines converge on the same rendering, confidence is reasonable. Where they diverge, something in the source is ambiguous, idiomatic, domain-specific or poorly covered in training data, and that divergence is a usable risk signal available before anyone has read the output. It is not the same as a human review. It is dramatically cheaper than one, and it is the only signal in this pipeline that arrives automatically.
That matters commercially as well as technically, because users are demanding more transparency, not less, as these systems get more capable. A translation that shows its confidence is a different product from one that simply asserts an answer.
FINDING 04
The market has already repriced, and most buyers have not noticed
The industry data tells the same story from the supply side. The 2026 Nimdzi 100 records the top ten language service providers growing 3.6% and the top fifty growing 2.0% in 2025, while providers ranked 51 to 100 contracted by 4.3%, the first decline in that segment since 2021.
That is the signature of a market where commodity volume has moved to machines and the surviving value has concentrated at the top, where verification, liability, domain expertise and certification live. The mid-market squeeze is not a story about translators losing work. It is a story about what buyers are now willing to pay for, and the answer is increasingly: assurance, not words.
TABLE 1 · WHAT EACH TIER ACTUALLY BUYS IN 2026
| Tier | Indicative cost | Who verifies the output | Where the liability sits |
| RAW MACHINE | Under $0.15 per 1,000 words | Nobody | Entirely with the deploying organisation |
| MULTI-ENGINE + FLAGGING | Low multiple of raw | Automated agreement check, human review on flagged segments only | Shared, with an audit trail of what was flagged |
| POST-EDITED | $50 to $150 per 1,000 words | One linguist, reviewing machine output | Shared with the provider, subject to scope |
| FULL HUMAN (TEP) | $150 to $300 per 1,000 words | Translator, editor and proofreader | Contractually with the provider, often certified |
FINDING 05
Regulation is closing the gap faster than procurement is
Until recently, the argument for translation governance was reputational. It is becoming statutory.
The EU AI Act’s Article 50 transparency obligations took effect on 2 August 2026, while the heaviest Annex III high-risk obligations have been provisionally extended to December 2027. Deployers of high-risk systems are required to assign meaningful human oversight and retain automatically generated logs. Neither requirement is satisfiable if your multilingual pipeline has no record of which content was machine translated, which engine produced it, and whether anyone competent looked at it.
The practical exposure is straightforward. Machine-translated safety instructions, employment terms, clinical guidance, financial disclosures and consumer contracts all sit inside regulated categories in one jurisdiction or another. For readers tracking this, The AI Journal’s legal and compliance coverage follows the timelines in detail. The relevant question for a governance lead is narrower and more awkward: could you currently produce a log of every language your organisation published in last quarter, and say who signed off on each one.
Most organisations cannot, because translation was never treated as an AI deployment. It was treated as a utility.
A workable model: tier the content, not the tool
The failure mode in most localisation programmes is a single global policy. Either everything goes through human review, which is unaffordable and therefore quietly abandoned, or nothing does, which is what actually happens.
The alternative is to stop asking “is AI translation good enough” and start asking “good enough for which content.” Risk is a property of the document, not of the engine.
TABLE 2 · CONTENT RISK TIERING AND MINIMUM VERIFICATION
| Risk tier | Content type | Minimum verification | Failure cost |
| CRITICAL | Clinical instructions, drug and device labelling, legal contracts, safety warnings, regulatory filings | Qualified human review, certified where required, full audit log | Physical harm, litigation, regulatory action |
| ELEVATED | Employment terms, financial disclosures, technical documentation, terms of service, complaints handling | Multi-engine agreement check with human review of all divergent segments | Disputes, mis-selling exposure, remediation cost |
| COMMERCIAL | Product listings, campaign copy, brand messaging, UX strings | Automated confidence scoring plus native-speaker spot check | Lost conversion, brand damage |
| ROUTINE | Internal comms, ticket triage, gisting, first-pass research | Raw output acceptable, labelled as machine translated | Minor rework |
Three things make this model work in practice, and all three are architectural rather than procedural, which is the point The AI Journal has made repeatedly about agentic systems: governance has to be designed into the architecture rather than bolted on as policy.
- Classify at ingestion, not at review. The risk tier has to be attached when content enters the pipeline. Deciding afterwards means nobody decides.
- Treat model disagreement as a routing rule. Where independent engines diverge on a segment, that segment escalates automatically. This turns verification from a fixed cost applied to everything into a variable cost applied where the evidence says it is needed.
- Log the language, not just the job. Which engine, which version, which tier, which reviewer. This is the artefact regulators will ask for, and it costs almost nothing to capture at the time and almost everything to reconstruct later.
The bottom line
Cheap translation is not the risk. Cheap translation with no verification layer, deployed into content nobody classified, in languages nobody on the team reads, is the risk.
The 2026 data is consistent on this point. Accuracy varies enormously by language and by domain. Automated error detection is measurably weaker than automated generation. The failure mode is silent by design. And the regulatory framework now expects an organisation to know which of its outputs were machine-produced and who was accountable for them.
The organisations that get this right will not be the ones that spend the most on translation. They will be the ones that know which 5% of their content actually needed checking.
That question is answerable today, with tooling that already exists, at a fraction of the cost of reviewing everything. What it requires is the decision to treat translation as an AI deployment rather than a utility bill.
Frequently asked questions
How accurate is AI translation in 2026?
There is no single figure. Peer-reviewed testing of a mainstream engine on identical source text found accuracy ranging from 94% in Spanish to 55% in Armenian. Accuracy depends heavily on the language pair, the domain, and how much training data existed for that combination. Aggregate accuracy claims average across all of this and conceal the low end.
What is the biggest risk of using machine translation in business?
That errors are undetectable at the point of use. Machine translation output is fluent and correctly formatted regardless of whether it is accurate, and the person reading it usually cannot compare it with the source. Unlike most software failures, there is no error message.
Can AI check its own translations?
Only partially. Findings from the 2025 Conference on Machine Translation reported that accurate error detection remains a persistent challenge and that reference-based metrics still outperformed large language models at the segment level. Automated quality estimation is useful as a triage signal, but it is currently weaker than the generation systems it is meant to police.
What is multi-engine verification?
Running the same source text through several independent translation models and comparing the outputs. Agreement across models is a reasonable confidence signal; disagreement flags a segment as ambiguous or high-risk and routes it for human review. At WMT25, selecting the best candidate from multiple models outperformed automatic post-editing of a single output.
Does the EU AI Act apply to machine translation?
It can, depending on use. Article 50 transparency obligations took effect on 2 August 2026, and translation embedded in a high-risk use case, such as employment, essential services, or medical devices, brings human oversight and logging duties for the deploying organisation. The heaviest Annex III obligations have been provisionally extended to December 2027.
Should we stop using AI translation?
No. The economics are decisive and the quality is genuinely strong on high-resource language pairs and routine content. The recommendation is to tier content by risk, apply verification proportionally, and maintain a log of what was machine translated and by whom it was approved.



