Dialogue builder

Multi-speaker audio from a single script

Build conversations line by line. Each line belongs to a speaker preset — a voice plus its pitch, speed and volume — and VocaTTS joins every line, with the pauses you set, into one finished track.

0/0 credits. Resets: N/A.

Usage details

Max Characters / Block: 2,000
Max Blocks: 0/0
+Credits / Block: 0

No voices found. Start by creating a voice preset (Language → Speed → Pitch).

Script Composer

0 blocks • 0 charactersEstimated duration: ~1 sec
No data

How a dialogue comes together

  1. 1

    Create speaker presets

    Set up one preset per character: choose a stock, cloned or designed voice and adjust pitch, speed and volume.

  2. 2

    Write the lines

    Add a block for each line, pick who says it, and set how long to pause before the next line.

  3. 3

    Export one track

    Preview individual lines, then generate the whole conversation as a single audio file.

What to make with it

  • Two-host podcasts

    Script a conversation between hosts with distinct voices and natural gaps.

  • Language-learning dialogues

    Model everyday conversations, with pauses for listeners to repeat.

  • Audio stories

    Give the narrator and each character their own voice.

  • Training scenarios

    Customer and agent role-plays for onboarding and support training.

Tips for natural dialogue

  • Keep each block to one speaker turn; long monologues sound better split into a few blocks.
  • Use slightly different speeds for speakers so they are easier to tell apart.
  • Short pauses (a few hundred milliseconds) feel conversational; longer ones work for scene changes.
  • Your work is saved as a draft in the browser while you write.

Requirements and limits

Script
At least two blocks; the maximum number of blocks depends on your plan
Per block
Speaker preset (voice, pitch, speed, volume) and pause after the line
Voices
Stock voices plus your cloned and designed voices
Output
One combined audio track
Credits
Charged per character at the rate of each voice used

Frequently Asked Questions

Q: How many speakers can a dialogue have?

A: As many presets as you need. The limit is on the number of blocks per generation, which depends on your plan.

Q: Can I mix my cloned voice with stock voices?

A: Yes. A preset can use any voice from the picker, including those under “My voices”.

Q: Do I get separate files per speaker?

A: No. The tool exports the whole conversation as one track. For one file per segment, use multi-segment mode in Text to Speech.

Q: Is my script saved?

A: A draft is kept in your browser while you work, so a refresh does not lose it.

Already have subtitles with speaker tags?

Tag lines [A], [B] in an SRT file and assign a voice per speaker in SRT to Audio — the timing then follows your video.