Conversational AI

One Language Is Not Enough: How AI Video Localisation Is Changing the Economics of Content

Forty per cent of consumers will not buy from a website that is not in their language. That is not a marketing blog. That is CSA Research’s “Can’t Read, Won’t Buy” study, run with Kantar across 8,709 consumers in 29 countries. Seventy-six per cent prefer products with information in their own language. Sixty-five per cent prefer their own language even when the translation quality is poor. Localisation is not a polish layer. It is the gatekeeper to the market.

For video, the economics have always been brutal. A conventional studio dub takes six to twelve weeks per language. A creator making one video in English who wants Spanish, Portuguese and Mandarin versions used to face a production cycle longer than the news cycle they were trying to ride. Most did not bother. Platforms and advertisers simply left money on the table in every non-English market.

That is changing fast. AI dubbing is compressing turnaround from months to days. A 2026 Research and Markets report on AI dubbing for OTT cited a Deepdub deployment with AWS for Paramount’s Ananey Studios that reduced localisation turnaround by more than 70%. Korea’s Ministry of Science and ICT funded AI dubbing of 1,200 K-content titles totalling 1,400 hours into English, Spanish and Portuguese, reaching 100 million cumulative viewers across 22 countries within five months. Deepdub’s live product, announced in April 2025, reports first-token audio latency of 125 milliseconds.

The platforms have noticed. Amazon Prime Video now says it can offer up to 22 dubbed languages and 36 subtitle tracks across 240 countries and territories. The economics of adding a language track are shifting from “is this title worth it?” to “can the pipeline handle one more?”

The creator side is moving too. Goldman Sachs Research estimates the creator economy will grow from $250 billion to $480 billion by 2027, with roughly 50 million creators globally but only about 4% earning more than $100,000 a year. Adobe’s 2026 Creators’ Toolkit Report, based on 16,000 creators across eight countries, found that 93% say AI output is faster, but 57% say it usually needs moderate or heavy editing before it is publishable. Faster is not finished. But faster means a solo creator can afford to fail four times and keep the best take.

The cost curve for that is now concrete. One second of lip-sync from Sync.so costs $0.025 on the hobbyist tier. Full text-to-video generation ranges from $0.08 for Kling O3 to $0.70 for Sora 2 Pro at 1080p. A sixty-second video therefore costs somewhere between $1.50 and $42 before retries. Add a four-take selection pipeline and you are paying for four seconds of output for every one you keep.

That spread turns localisation from a post-production decision into a routing decision. You do not dub the whole video with the most expensive model. You dub the close-up talking head with a lip-sync model, the B-roll with a cheaper generator, and the title cards with a text overlay. The model that wins on lips does not win on background plates. The platform that knows how to route between them wins on budget.

There is a regulatory clock running in parallel. AI-generated video of people is now required to be marked in most major markets.

China’s AIGC labelling measures and the mandatory standard GB 45438-2025 have been in effect since 1 September 2025. They require both visible and implicit labels. For video, the visible label must display for at least two seconds and the text height must be at least 5% of the frame’s shortest side. The EU’s AI Act Article 50 transparency obligations apply from 2 August 2026, with machine-readable marking for synthetic output and disclosure for deepfakes; penalties reach €15 million or 3% of global turnover. India’s IT Amendment Rules 2026 took effect on 20 February 2026, requiring platforms to label synthetically generated information and embed provenance metadata. California’s AI Transparency Act, in force from 1 January 2026, requires providers with over a million monthly users to embed provenance metadata and offer a free public detection tool.

None of these laws say “do not generate synthetic video.” They say “mark it and be able to prove where it came from.” That is a pipeline requirement, not a creative one. It means every generated clip needs to carry provenance metadata from the moment of creation, because retrofitting it later is expensive and error-prone.

The practical implication is that the winning workflow is not a better single generator. It is a pipeline: generate, translate, dub, label, review, publish. The review step is non-negotiable. Adobe’s report found that 85% of creators believe final creative decisions must stay with the creator, and 44% want the ability to review, edit or undo at any time. The public is not asking for autonomous agents that replace creators. It is asking for agents that do the boring language and lip-sync work and then hand the interesting choices back.

This is why we think localisation is the killer use case for AI video tools. A unified AI video tool that can generate, translate, dub and watermark in one pipeline does not just save time. It changes which markets are profitable to reach. The one-language creator becomes the ten-language creator without hiring nine translators.

The companies that get this right will not be the ones that chase the best lip-sync benchmark. They will be the ones that build a reliable, reviewable, compliant pipeline. Because the audience is there. Forty per cent of them will not wait for you to speak English.

Author:

Related Articles

Back to top button