Writing Voice Design Prompts That Actually Work

Published · 5 min read

A voice design prompt is a casting brief, not a script. Five traits get you most of the way.

Voice design creates a brand-new synthetic voice from a written description. There is no recording and no real person behind it, which makes it the right tool when you need a specific character — a calm meditation guide, a gruff old sailor, a bright product presenter — that you could not or would not record yourself. The catch is that the model can only build what you describe. Vague prompts give generic voices.

Think casting brief, not script

The most common mistake is writing what the voice should say instead of how it should sound. “Welcome to our channel, today we’re talking about…” tells the model nothing about the speaker. Describe the person as if you were briefing a voice actor at an audition. On VocaTTS you have between 20 and 1,500 characters for this, which is plenty — most good prompts are two or three sentences.

The five traits that matter most

  1. Age and role. “A woman in her late twenties who teaches online classes” anchors pitch, vocabulary and attitude in one phrase.
  2. Timbre. Deep, bright, breathy, crisp, nasal, gravelly. Pick one or two, not five.
  3. Pace and energy. Slow and unhurried, brisk and upbeat, measured with pauses.
  4. Accent or region, only if your audience will notice. “Neutral American” or “southern Vietnamese” is enough.
  5. The listener’s feeling. Reassured, excited, focused, amused. This often does more than any technical adjective.

Gender and the main language are set separately in the form — male or female, and one of 15 main languages — so you do not need to repeat them in the prompt. Use the prompt for everything else.

Examples by use case

E-learning

A warm, patient teacher in her early thirties. Clear articulation, steady medium pace, smiles slightly when explaining. Listeners should feel it is fine to make mistakes.

Documentary narration

A man around sixty with a low, resonant voice and a hint of gravel. Speaks slowly, lets sentences land, never sounds rushed. Authoritative but kind.

Short-form video

An energetic presenter in his mid-twenties, bright and punchy, fast but still easy to follow. Sounds like he is letting you in on something exciting.

Audiobook, children’s story

A gentle, slightly breathy female storyteller. Soft volume, slow rhythm, warm and playful, as if reading at a child’s bedside.

What to leave out

  • Names of real people or characters. “Sounds like a famous actor” is not something you can rely on, and imitating real people is not allowed.
  • Contradictions. “Calm yet extremely energetic” forces the model to guess.
  • Long lists of adjectives. Past five or six traits, extra words dilute the important ones.
  • Sound effects or music. The design describes a voice, not a production.

Iterate one trait at a time

If the first voice is close but not right, change a single trait and generate again. Adjust age first, because it moves everything else; then pace; then timbre. Changing three things at once makes it impossible to tell which change helped. Keep your prompts in a note — a voice you like today is easy to recreate later if you still have the description.

Putting the voice to work

A designed voice behaves like any other voice in “My voices”: use it for narration in Text to Speech, for subtitle voiceover in SRT to Audio, or give each character in a Multi-Speaker script its own designed voice. Because it is a Gemini voice, emotion tags such as <cheerful> or <serious> at the start of a sentence still work, so one well-designed voice can cover several moods.

Try your prompt

Write a description, pick gender and language, and hear the voice before you save it.

← All guides