Subtitle-based dubbing
Dub videos from their subtitles
If your video has an SRT file, you have everything needed for a dub. VocaTTS translates the subtitles, reads them in the voices you assign to each speaker, and fits every line to its original timing.
The dubbing workflow
- 1
Start from the source SRT
Export subtitles in the original language from your editor or captioning tool and upload the .srt file.
- 2
Translate and review
Turn on translation, choose the target language and quality, then edit any line that needs a better wording.
- 3
Voice and export
Assign a voice per speaker, generate timeline-fitted audio and download it with the translated SRT.
Dubbing workspace
SRT to Voice Dubbing
Upload SRT, edit the content, map each speaker to a voice, then generate audio in a simpler step-by-step flow.
Usage details
Source
Input ModeVoice
Auto-detect [A], [B], then map each speaker to a voice.
Generate
Details that make a dub usable
- Glossary dictionaries keep product names, people and terms translated the same way every time.
- Speaker tags [A], [B] let each person in the video keep a consistent voice.
- A speed range keeps translated lines inside their time slot; set min and max equal for a fixed pace.
- Gemini voices accept emotion tags, useful for reactions and emphasis.
- Your cloned or designed voices can be assigned to speakers like any stock voice.
Related tools
Frequently Asked Questions
Q: Can I upload a video to dub?
A: Not in VocaTTS. The workflow starts from an SRT file in the original language, which most editors and captioning tools can export.
Q: Will the dubbed audio stay in sync?
A: Each translated line is fitted to the original subtitle's time slot within the speed range you set, so the audio follows the video's timeline.
Q: Can I keep the same voice for each person in the video?
A: Yes. Tag their lines [A], [B] and so on, then assign one voice per speaker.
Q: How do I keep names translated correctly?
A: Add them to a translation dictionary. Signed-in users can create glossaries and apply them to subtitle translations.
Q: Is there lip-sync?
A: No. The output is an audio track timed to the subtitles, not a lip-synced video.