Flagship feature
Clone your voice with AI
Give VocaTTS a short recording of how you speak and get a synthetic version of your voice. Use it to narrate scripts, voice subtitles and play a part in multi-speaker audio — without recording every take yourself.
Create your cloned voice
Sign in to create your personal voice
Personal voices are saved to your account and can be used in Text to Speech, Multi-Voice and SRT to Speech.
Sign inHow cloning works
- 1
Record a reference sample
Talk for 10–30 seconds the way you want the AI voice to sound. Energy and intonation carry over, so avoid a flat, careful reading.
- 2
Record the consent statement
The same person reads a fixed sentence confirming they own the voice. It must be read word for word in one of 30 supported languages.
- 3
Generate and use it
When the voice is ready it appears under “My voices” in every voice picker, next to the stock voices.
Getting a good sample
- Use a headset or phone microphone about a hand's width from your mouth rather than a laptop microphone.
- Pick a small, quiet room. Echo, fans and background music end up in the voice.
- Aim for 20–30 seconds and keep talking naturally if you finish the suggested text early.
- A clean WAV or MP3 from a real microphone usually beats recording in the browser.
Where your cloned voice works
Text to Speech
Narrate scripts, lessons and voice notes. Edit the text and regenerate instead of re-recording.
Open Text to SpeechSRT to Audio
Read a subtitle file in your voice with timing fitted to each subtitle line.
Voice a subtitle fileMulti-Speaker
Cast yourself as one speaker in a dialogue, alongside stock or designed voices.
Build a dialogue
Good fits
- Creators who publish often and want a consistent voice across videos
- Course authors updating lessons after a product or syllabus changes
- Product teams refreshing tutorial narration when the interface changes
- Podcasters fixing a line without setting up the microphone again
Requirements and limits
- Reference sample
- 10–30 seconds of natural speech, recorded or uploaded
- Consent recording
- 2–30 seconds, verbatim statement, same speaker, 30 supported languages
- Plan
- Basic plan or higher (3 custom voice slots; Premium and Studio: 5)
- Storage
- Kept for up to one year; delete it any time to free the slot
- Credits
- Speech generated with a cloned voice uses the Gemini voice rate
- Availability
- Not offered in the UK, the EEA, Switzerland, Illinois or Texas
What to expect
- A clone follows the recording: a noisy or monotone sample produces a noisy or monotone voice.
- Very expressive acting, singing and shouting are not reliable targets for a clone.
- The original recordings are converted and sent to Google Gemini to build the voice; VocaTTS does not keep the raw files.
Frequently Asked Questions
Q: How long does the recording need to be?
A: The reference sample must be 10–30 seconds. Around 20–30 seconds of natural, energetic speech gives the most complete voice.
Q: Why do I have to record a consent statement?
A: Google Gemini, which builds the voice, requires the speaker to confirm ownership by reading a fixed statement. It protects people from having their voice cloned without permission.
Q: Can I clone someone else's voice?
A: Only with their explicit permission, and they must record the consent statement themselves. Cloning a voice you have no right to use is not allowed.
Q: Can I use my cloned voice with SRT files?
A: Yes. Select it under “My voices” in SRT to Audio and the subtitle file is read in your voice with timing fitted to each line.
Q: Can I delete my cloned voice?
A: Yes, at any time from the tool or from “My voices”. Deleting removes it from Gemini and frees the slot on your plan.
Q: How long does a cloned voice last?
A: Gemini stores it for up to one year. After that you can create it again from a new recording.
Cloning not available to you?
If you are in a region where cloning is not offered, or prefer not to use a real voice, Voice Design creates an original voice from a written description.