Purchasing Guides / Video Translation & Dubbing

Video Translation & Dubbing

A purchasing guide to buying a workflow that translates your videos, adds a new narration and prepares subtitles for your review.

Yearly running cost

Assuming 600 language versions a year

By hand
$15.8K / year
450 hours of work
Hosted PAYG
$6.7K+ / year
This workflow
$392 / year
Machine usage only; setup, hosting and review are extra.
Which videos do you want to translate?

Translate narrated tutorials, product demos and explainers, with an editable script and a licensed replacement voice.

One language version includes up to 5 source-video minutes and 1 target language and 6,000 translated characters.

Review every language version before sharing. No automatic publishing, voice imitation or lip synchronisation. The default uses a licensed replacement voice, not the original speaker's identity.

What is included

Upload an authorised narrated video. Review its transcript and translation, then generate a licensed replacement narration and export the video with subtitles.

Includes a source transcript, glossary-aware translation, script approval, timed narration, subtitle files and a video export. Start with English-to-Spanish narration. Agree other language pairs with your provider. Supply a clean music/effects track if you want it retained; the workflow does not recreate mixed background audio.

Narrators per source: up to 1 narrator.

Configured language pairs: 1 language pair.

Glossary setup: up to 100 terms.

Maximum source height: up to 1,080 pixels.

Maximum frame rate: up to 30 frames per second.

Maximum source file: up to 500 MiB.

Narration boundary tolerance: up to 500 milliseconds.

Subtitle boundary tolerance: up to 200 milliseconds.

Reference worker capacity: 2 physical CPU cores.

Reference worker memory: 4 GiB.

What should make the translation ready to use?

A good language version preserves the message, sounds clear and fits the video without rushing or dropping words.

How these standards are measured

Use unseen videos and bilingual reviewers. Check preserved meaning, names, numbers, glossary terms, pronunciation and timing against the source. Count omitted segments, failed renders and unusable outputs. AI back-translation alone is not an acceptance test.

Targets for your selected standard
What is checkedTarget
Segments preserve the original meaningIndependent bilingual reviewers compare every source segment with the translation using an agreed error rubric. Missing segments count as failures. Report critical errors separately; do not rely only on automatic similarity or back-translation.≥95%
Glossary terms used correctlyCheck every required glossary occurrence in both the script and generated speech. Report textual correctness and pronunciation separately; either failure makes that occurrence fail. Preserve names, quantities and negation.≥98%
Narration accepted by native listenersNative listeners score intelligibility, pronunciation and natural delivery blind to preset. Count all generated segments, including silence, clipped words and failed generations. Do not use voice similarity as a substitute for clarity.≥90%
Narration fits its agreed time slotsCompare each narration segment with its agreed source time slot. Do not rush speech or delete meaning to pass. Subtitles must match the generated audio within their separate tolerance; this is not mouth-motion matching.≥95%
Exports pass technical checksValidate decodable video, complete audio, duration, subtitle encoding and timing for every requested file. Preserve the original visuals. Failed exports and missing segments remain in the denominator.≥100%

Incorrect names, quantities, warnings or meaning must be corrected before sharing, even when average scores pass. A fluent voice is not proof of an accurate translation.

How quickly do you need a language version?

Choose how soon the translation and rendered narration should be ready. Your approval time is separate.

Timing details

Measure both machine stages: source transcription and translation, then speech generation, mixing and export after script approval. Include queues and retries. Add those processing times; report upload and human waiting time separately. Later script changes are another run.

The target applies to at least 95% of agreed test runs, with 2 in progress at a time.

How much do you want to spend per language version?

Choose the machine budget for translation, narration and rendering in each target language.

Cost details

Includes speech recognition, text models, voice generation, rendering, retries and shared hosting with a commercial voice-plan allowance. Count each language and every regeneration. Bilingual review, premium voice licences and the calling agent are separate.

Reference machine cost: $1.17 – 1.97 per language version at 50 language versions a month. The selected cap is a target to test, not a replacement for this estimate.

Where do you want it to run?

Run in your cloud or on a machine you control. Choose whether audio and scripts can go to approved AI APIs.

Data and access details

Runs in a cloud account you control, with access controls and logs.

Audio and scripts go to approved speech and translation APIs. Confirm commercial voice rights and retention.

Confirm rights to the source and the selected voice. Local deployment with API access still sends content outside your machine. Offline mode needs different private speech models and a new voice-quality pilot.

How do you want to use it?

Use a review web page, connect your existing agent or run a dedicated localisation agent.

Anything else your provider should know?

Optional. Your choices are included automatically.

Common questions

When should I use a hosted dubbing tool instead?

Use a hosted tool if its languages, voices and editor fit your work. Buy a workflow when you need your own approval rules, glossary, integration or deployment.

Will it sound exactly like the original speaker?

Not by default. This scope uses a licensed replacement narrator. Voice matching needs explicit permission, suitable recordings, model access and separate evaluation.

Can I correct the translation?

Yes. Edit and approve the script before generating narration. Later changes create a new version and consume more generation resources.

Does it change lips or text inside the video?

No. Lip synchronisation and translated on-screen graphics are different capabilities. Agree them separately if you need them.

Can I add more languages?

Yes, but each language needs a suitable voice, glossary and native-language tests. Each output language is a separate run; expanding the validated language set can change setup cost.

Download RenX

Get the app.

Install, sign up, and start on your free plan with welcome credit included. No credit card required.

On a platform not listed? Leave your email and we'll notify you when a build is available.