How to Clone Your Voice with AI (and What Makes a Good Sample)

Published · 6 min read

Thirty seconds of audio decides how your clone sounds for the next year. Here is how to record those seconds well.

An AI voice clone is only as good as the recording it learns from. On VocaTTS the reference sample is short — between 10 and 30 seconds — which means every second carries weight. A noisy room, a flat reading or a clipped microphone all end up baked into the voice you will use for months. The good news is that you do not need a studio. You need a quiet corner, a reasonable microphone and a few minutes of preparation.

What the clone actually needs from you

You will make two recordings. The first is the reference sample: natural speech that shows how you sound. The second is a consent statement: a fixed sentence, read word for word by the same person, confirming that the voice belongs to them. Google Gemini, which builds the voice, checks that both recordings come from the same speaker. The statement is available in 30 languages, so read it in the one you speak most comfortably.

  • Reference sample: 10–30 seconds. Aim for 20–30 — more material gives a fuller voice.
  • Consent statement: 2–30 seconds, read exactly as shown.
  • A paid plan (Basic or higher). Each clone uses one custom voice slot.
  • Cloning is not available in the UK, the EEA, Switzerland, Illinois or Texas. If you are there, voice design is the alternative.

Pick the right room before the right microphone

Echo hurts a clone more than a cheap microphone does. Hard, empty rooms bounce your voice back into the mic, and the model learns that reverb as part of your sound. A small room with soft furnishings — a bedroom with curtains, a closet full of clothes, even a car parked somewhere quiet — usually beats a bare home office.

  • Turn off fans, air conditioning and anything with a hum.
  • Silence notifications on your phone and computer.
  • Sit 15–20 cm from the microphone, slightly off-axis so breaths do not pop.
  • Record a few seconds of silence first and listen back. If you can hear the room, the clone will too.

What to say in the reference sample

The model learns from how you speak, not what you say, so choose material that brings out your natural delivery. Reading a dry paragraph in a monotone gives you a dry, monotone clone. Talking about something you actually care about gives the model pitch movement, emphasis and rhythm to work with.

A sample that works

“Last weekend I finally fixed the old bike in the garage. The chain was rusted solid, so I soaked it overnight — and honestly, I didn't expect it to work. But the next morning it spun like new, and I rode it all the way to the river.”

  • Speak at the energy you want the clone to have. If you want a lively narrator, be a little livelier than usual.
  • Include a question and an exclamation so the model hears your range.
  • Avoid long pauses. Twenty seconds of continuous speech beats thirty with gaps.
  • Stay in one language for the whole sample.

The consent statement is not a formality you can rush. It must match the shown text word for word, and it must be the same voice as the sample. Read it at a normal pace in the same room, with the same microphone and the same distance. If the system rejects it, the most common causes are a skipped word, background noise or a different person reading it.

Common mistakes and how to avoid them

  1. Using a clip from a video or podcast. Music, room tone and compression all leak into the clone. Record fresh audio.
  2. Recording someone else's voice. You may only clone your own voice, or the voice of a person who agrees and records the consent statement themselves.
  3. Whispering or shouting. Extremes do not generalise well; your normal speaking voice gives the most usable clone.
  4. Clipping. If the waveform hits the top, the peaks are distorted. Lower the input level or move back a little.

After the clone is ready

Your new voice appears under “My voices” in Text to Speech, SRT to Audio and Multi-Speaker. Test it with a short paragraph before committing to a long project. Gemini voices also respond to emotion tags such as <excited> or <calm> at the start of a sentence, which is the quickest way to add variety without re-recording. If you are not happy with the result, delete the voice — the slot frees up immediately — and record again with what you have learned.

The finished voice is stored for up to a year and billed at the normal Gemini credit rate when you generate speech with it. Your raw recordings are not kept after the voice is created.

Ready to record?

The cloning tool walks you through both recordings and checks them before creating the voice.

← All guides