Dialogue builder
Multi-speaker audio from a single script
Build conversations line by line. Each line belongs to a speaker preset — a voice plus its pitch, speed and volume — and VocaTTS joins every line, with the pauses you set, into one finished track.
Usage details
No voices found. Start by creating a voice preset (Language → Speed → Pitch).
Script Composer
| # | Time | Voice / Blocks | Status | Actions |
|---|---|---|---|---|
| No data | ||||
How a dialogue comes together
- 1
Create speaker presets
Set up one preset per character: choose a stock, cloned or designed voice and adjust pitch, speed and volume.
- 2
Write the lines
Add a block for each line, pick who says it, and set how long to pause before the next line.
- 3
Export one track
Preview individual lines, then generate the whole conversation as a single audio file.
What to make with it
Two-host podcasts
Script a conversation between hosts with distinct voices and natural gaps.
Language-learning dialogues
Model everyday conversations, with pauses for listeners to repeat.
Audio stories
Give the narrator and each character their own voice.
Training scenarios
Customer and agent role-plays for onboarding and support training.
Tips for natural dialogue
- Keep each block to one speaker turn; long monologues sound better split into a few blocks.
- Use slightly different speeds for speakers so they are easier to tell apart.
- Short pauses (a few hundred milliseconds) feel conversational; longer ones work for scene changes.
- Your work is saved as a draft in the browser while you write.
Requirements and limits
- Script
- At least two blocks; the maximum number of blocks depends on your plan
- Per block
- Speaker preset (voice, pitch, speed, volume) and pause after the line
- Voices
- Stock voices plus your cloned and designed voices
- Output
- One combined audio track
- Credits
- Charged per character at the rate of each voice used
Frequently Asked Questions
Q: How many speakers can a dialogue have?
A: As many presets as you need. The limit is on the number of blocks per generation, which depends on your plan.
Q: Can I mix my cloned voice with stock voices?
A: Yes. A preset can use any voice from the picker, including those under “My voices”.
Q: Do I get separate files per speaker?
A: No. The tool exports the whole conversation as one track. For one file per segment, use multi-segment mode in Text to Speech.
Q: Is my script saved?
A: A draft is kept in your browser while you work, so a refresh does not lose it.
Already have subtitles with speaker tags?
Tag lines [A], [B] in an SRT file and assign a voice per speaker in SRT to Audio — the timing then follows your video.