Flagship feature

Clone your voice with AI

Give VocaTTS a short recording of how you speak and get a synthetic version of your voice. Use it to narrate scripts, voice subtitles and play a part in multi-speaker audio — without recording every take yourself.

Create your cloned voice

Sign in with a Basic plan or higher. You will make two recordings: a natural reference sample and a consent statement.

Sign in to create your personal voice

Personal voices are saved to your account and can be used in Text to Speech, Multi-Voice and SRT to Speech.

Sign in

How cloning works

  1. 1

    Record a reference sample

    Talk for 10–30 seconds the way you want the AI voice to sound. Energy and intonation carry over, so avoid a flat, careful reading.

  2. 2

    Record the consent statement

    The same person reads a fixed sentence confirming they own the voice. It must be read word for word in one of 30 supported languages.

  3. 3

    Generate and use it

    When the voice is ready it appears under “My voices” in every voice picker, next to the stock voices.

Getting a good sample

  • Use a headset or phone microphone about a hand's width from your mouth rather than a laptop microphone.
  • Pick a small, quiet room. Echo, fans and background music end up in the voice.
  • Aim for 20–30 seconds and keep talking naturally if you finish the suggested text early.
  • A clean WAV or MP3 from a real microphone usually beats recording in the browser.

Where your cloned voice works

Good fits

  • Creators who publish often and want a consistent voice across videos
  • Course authors updating lessons after a product or syllabus changes
  • Product teams refreshing tutorial narration when the interface changes
  • Podcasters fixing a line without setting up the microphone again

Requirements and limits

Reference sample
10–30 seconds of natural speech, recorded or uploaded
Consent recording
2–30 seconds, verbatim statement, same speaker, 30 supported languages
Plan
Basic plan or higher (3 custom voice slots; Premium and Studio: 5)
Storage
Kept for up to one year; delete it any time to free the slot
Credits
Speech generated with a cloned voice uses the Gemini voice rate
Availability
Not offered in the UK, the EEA, Switzerland, Illinois or Texas

What to expect

  • A clone follows the recording: a noisy or monotone sample produces a noisy or monotone voice.
  • Very expressive acting, singing and shouting are not reliable targets for a clone.
  • The original recordings are converted and sent to Google Gemini to build the voice; VocaTTS does not keep the raw files.

Frequently Asked Questions

Q: How long does the recording need to be?

A: The reference sample must be 10–30 seconds. Around 20–30 seconds of natural, energetic speech gives the most complete voice.

Q: Why do I have to record a consent statement?

A: Google Gemini, which builds the voice, requires the speaker to confirm ownership by reading a fixed statement. It protects people from having their voice cloned without permission.

Q: Can I clone someone else's voice?

A: Only with their explicit permission, and they must record the consent statement themselves. Cloning a voice you have no right to use is not allowed.

Q: Can I use my cloned voice with SRT files?

A: Yes. Select it under “My voices” in SRT to Audio and the subtitle file is read in your voice with timing fitted to each line.

Q: Can I delete my cloned voice?

A: Yes, at any time from the tool or from “My voices”. Deleting removes it from Gemini and frees the slot on your plan.

Q: How long does a cloned voice last?

A: Gemini stores it for up to one year. After that you can create it again from a new recording.

Cloning not available to you?

If you are in a region where cloning is not offered, or prefer not to use a real voice, Voice Design creates an original voice from a written description.