AI Dubbing Tools: What a Minute Actually Costs
In short
Nobody publishes a price per dubbed minute, so you derive it. Doing that from vendors' own pricing pages in July 2026 gives a spread of roughly 25× across credible options. HeyGen's Creator plan ($29/month, 600 credits, 5 credits per lip-synced minute) works out near $0.24 per minute. Rask AI's entry Creator plan ($60/month for 25 minutes) works out at $2.40, and its enhanced lip-sync bills 3 quota minutes per video minute, pushing the effective rate past $7. YouTube's auto-dubbing is free and on by default, but only inside YouTube, with no lip-sync and no voice control. Three things wreck naive comparisons: whether you are billed per target language, whether lip-sync carries a multiplier, and whether unused quota expires. Rolling your own with Whisper, an open TTS model and LatentSync makes the compute nearly free and the engineering expensive.
How AI dubbing works end to end
Every tool here runs the same four stages, and they tell you where both the price and the quality come from.
- Transcribe and diarise. Speech-to-text plus speaker separation. Whisper or a derivative underneath almost everywhere.
- Translate. The transcript goes to an LLM or an MT model. Where idioms, jargon and proper nouns die.
- Synthesise. TTS in the target language, usually with a voice cloned from the original speaker. It has to fit the translated line into the original's time slot — the source of most audible weirdness.
- Lip-sync (optional). A video model rewrites the mouth region to match the new audio.
Stages 1–3 are cheap. Stage 4 is why price lists have multipliers. If you do not need lip-sync — podcasts, voiceover, screen recordings, anything not shot head-on in close-up — you can skip the expensive part, and most vendors let you.
Tools compared by price per minute
Derived from each vendor's official pricing page, checked July 2026. Effective rate = plan price ÷ minutes the plan yields, at monthly billing unless noted.
| Tool | Entry paid plan | What it yields | Effective $/min | The catch |
|---|---|---|---|---|
| YouTube auto-dubbing | Free | 24 source languages, into 20+ | $0 | YouTube only; no lip-sync, no voice choice, no re-render control |
| HeyGen Creator | $29/mo, 600 credits | 120 min lip-synced (5 cr/min) | $0.24 | Credits shared with avatar generation; Avatar IV/V burns 20 cr/min |
| HeyGen Creator, audio only | $29/mo, 600 credits | 300 min (2 cr/min) | $0.10 | No lip-sync at this rate |
| HeyGen Pro | $49/mo, 1,000 credits | 200 min lip-synced | $0.25 | Same credit-sharing, bigger bucket |
| Rask AI Creator | $60/mo, 25 min | 25 output minutes | $2.40 | Annual billing drops it to ~$1.32/min — a 45% penalty for staying monthly |
| Rask AI Creator Pro | $150/mo, 100 min | 100 output minutes | $1.50 | Enhanced lip-sync (beta) burns 3 quota minutes per video minute → ~$4.50/min |
| Descript Creator | $24/mo | Dubbing in 30 languages, bundled | No per-minute meter | Gated to Creator and above; Hobbyist gets captions only |
| DIY (Whisper + open TTS + LatentSync) | GPU time | Unlimited | Cents of compute | RunPod lists an RTX 4090 at $0.34/hr community, $0.69/hr secure — the cost is your week, not the GPU |
ElevenLabs also sells dubbing on a credit meter, billed per target language; its pricing page is region-gated and we could not read it, so we are not quoting unverified numbers.
The pattern, plainly: credit-metered tools are roughly an order of magnitude cheaper per minute than minute-metered ones, because credits are shared across the whole product and the vendor is betting you spend them on something with a fatter margin. A real discount if dubbing is all you want from the plan; a trap if it is not.
The three multipliers that break the price list
Per-language billing. Rask's pricing page states that "1 minute equals 1 minute of final translated video/audio". Dub a 10-minute video into three languages and you have spent 30 minutes, not 10. HeyGen's credits do the same implicitly. Read every per-minute figure anyone quotes as per output minute.
Lip-sync multipliers. Rask's standard lip-sync costs 1 quota minute per video minute; the enhanced beta costs 3. HeyGen charges 5 credits per lip-synced minute against 2 for audio-only. Decide up front — this is most of your bill.
Expiry. HeyGen rolls unused credits over one extra month on monthly plans, and accumulates them until renewal on annual. Rask's annual plans front-load the year's minutes. If your output is bursty, annual billing on a minute-metered tool is the cheapest structure available.
Where quality collapses
Failure modes are consistent across vendors, and Google's auto-dubbing help page is unusually candid about them.
Timing mismatch. Translated lines are rarely the same length as the original, so the system speeds up delivery or clips the pause after it. Eight minutes of that and the dub sounds subtly panicked.
Proper nouns, jargon and numbers. Google lists mispronunciations, accents, dialects, background noise, proper nouns, idioms and jargon as known error sources. For developer content that is close to fatal: library names, CLI flags, version numbers and acronyms are exactly the tokens that break. Budget for proofreading by hand.
Music and effects bleed. Good pipelines separate voice from the mix and re-lay the dub over the original bed. Weaker ones re-synthesise over everything or duck the music inconsistently. Test with a music-bed clip first.
Overlapping speakers. Two people talking at once is where diarisation breaks and a line gets dropped or given to the wrong voice.
On-screen text stays put. Nothing here translates burned-in captions, slide titles or UI in a screen recording. If your video is mostly screencast — as most developer content is — dubbing the audio gets you halfway and the mismatch is jarring. Localised visuals are a separate job, closer to our faceless video guide than to anything the dubbing vendor sells.
Hard rejections. YouTube refuses videos over 120 minutes, with minimal speech, with undetectable or unsupported source language, with rapid pacing, or carrying Content ID claims. Other tools fail more quietly on the same inputs.
Voice cloning and consent
Every tool here clones the original speaker's voice by default, because that is the selling point. Two things follow.
If the voice is not yours, get written consent naming AI synthesis and the target languages. A talent release signed before 2023 almost certainly does not cover having that person speak Portuguese. Right-of-publicity law has caught up: Tennessee's ELVIS Act, in force since July 2024, covers voice explicitly and even reaches providers of tools whose primary purpose is unauthorised voice replication.
If you distribute in the EU, Article 50 of the AI Act applies from 2 August 2026: providers of systems generating synthetic audio must mark output machine-readably, and deployers of deepfakes must disclose it. A dubbed video of a real person saying words they did not say is squarely in scope. The practical answer is one line in the description plus the consent on file.
A worked example
A 10-minute talking-head video into Spanish, German and Portuguese, with lip-sync. That is 30 output minutes.
| Route | What you pay | Note |
|---|---|---|
| YouTube auto-dubbing | $0 | No lip-sync; locked to YouTube |
| HeyGen Creator ($29/mo) | 150 of 600 credits ≈ $7.25 | Room for three more such jobs that month |
| HeyGen Creator, audio only | 60 credits ≈ $2.90 | Same, without lip-sync |
| Rask Creator Pro ($150/mo) | 30 of 100 min ≈ $45 | The $60 Creator plan cannot hold this job — 25 min allowance |
| Rask, enhanced lip-sync | 90 of 100 min ≈ $135 | One video eats the month |
| DIY on a rented 4090 | GPU-hours at $0.34/hr | Plus a week building the pipeline |
Same job, roughly 6× between the two commercial tools and 18× with enhanced lip-sync on. That is not a quality gap of the same size. Watch the output before assuming the expensive option earned it.
For the DIY route the parts are all open: Whisper for transcription, any LLM for the translation pass, Chatterbox or another open TTS for the voice, LatentSync (Apache-2.0, 8 GB VRAM for v1.5, 18 GB for v1.6) for the mouth. Buildable in a weekend if you already run local models; not worth it at two videos a month. The honest threshold is near 100 output minutes a month, or a hard requirement that source audio never leaves your infrastructure. Our AI video tools round-up covers the rest of the stack, and OpenCut handles assembly for free.
FAQ
How much does AI dubbing cost per minute?
Between $0 and roughly $7 per output minute in July 2026, depending entirely on billing model: free on YouTube for YouTube, about $0.10–$0.25 on credit-metered tools like HeyGen, about $1.30–$2.40 on minute-metered tools like Rask AI at entry tiers, and around $4.50 with Rask's enhanced lip-sync. Always calculate from output minutes — source length times number of target languages.
Is AI dubbing good enough to publish?
For monologue content with clean audio, no music bed and few proper nouns: usually yes, with a script proofread. For interviews, panels, anything with crosstalk, or technical content dense with library names and version numbers: not without human review of the translation. The tell is rarely voice quality — it is timing drift and mangled jargon.
Do I need lip-sync?
Usually not, and it is most of the cost. Skip it for podcasts, voiceover, screencasts and tutorials, and anything not framed head-on in close-up for long stretches. Turn it on for direct-to-camera marketing and course intros. On HeyGen, 2 credits per minute against 5 is a 60% saving for something nobody will notice.
Can I dub a video into a language I don't speak and just publish it?
You can, and it is the most common way people embarrass themselves with this technology. Run the translated script past a native speaker at least once per language before publishing the first video in it — after that you will know which failure modes your content triggers. Mistranslated product names are the ones that reach your support inbox.
Is it legal to clone a speaker's voice for dubbing?
Your own voice, yes. Someone else's requires consent that specifically covers AI synthesis — an old talent release does not. Tennessee's ELVIS Act names voice directly and reaches tool providers too, and other jurisdictions are following. From 2 August 2026 the EU AI Act also requires synthetic audio to be marked machine-readably and deepfakes disclosed. Get consent in writing, keep it, label the dub.