Skip to content

ElevenLabs v3 (eleven_v3) Prompting Principles

Reference notes for directing the ElevenLabs eleven_v3 text-to-speech model, which is the model all TTS in this system targets. Sourced from the official ElevenLabs v3 prompting best-practices docs (https://elevenlabs.io/docs/overview/capabilities/text-to-speech/best-practices).

For our own empirically-tested findings (single-quote emphasis, ? for volume, separate brackets, letter-repetition for shouting, word-joining for speed, etc.), see the companion doc learnings_11labs_prompting_v3.md. Both feed the eleven_v3 dialogue-tagging prompt in src/promptTemplates/workbenchV2/tagDialogueForElevenV3.template.ts.


Core idea

eleven_v3 is directed through inline audio tags (bracketed cues), punctuation, and capitalization — not SSML. A tag colors the text that follows it. Text structure strongly influences the output.

Audio tag catalog (by category)

Emotional / delivery direction

  • [laughs], [laughs harder], [starts laughing], [wheezing], [laughing hysterically]
  • [whispers]
  • [sighs], [exhales], [exhales sharply], [inhales deeply], [frustrated sigh], [happy gasp]
  • [sarcastic], [curious], [excited], [crying], [snorts], [mischievously]
  • [happy], [sad], [angry], [annoyed], [appalled], [thoughtful], [surprised]
  • [dismissive], [cute], [deadpan], [professional], [sympathetic], [questioning], [reassuring], [panicking]

Non-verbal sounds

  • [giggles], [chuckles], [clears throat], [swallows], [gulps]

Environmental sound effects

  • [gunshot], [applause], [clapping], [explosion]

Pacing / pauses

  • [short pause], [long pause], [pauses]
  • v3 does not support SSML <break> tags — use audio tags, ellipses, and text structure for pacing instead. Speed is also directed via tags.

Multi-speaker / dialogue

  • [overlapping], [interrupting] (assign a distinct voice per speaker, e.g. "Speaker 1:", "Speaker 2:")

Unique / experimental

  • [strong <X> accent] (replace <X>, e.g. [strong French accent]), [sings], [singing quickly], [robotic voice], [binary beeping], [woo], [fart]

Many effective tags exist beyond this published list — descriptive states often work. Experimental tags can be inconsistent; test before production.

Punctuation & capitalization

  • Ellipses (…) add pauses and weight.
  • Capitalization increases emphasis.
  • Standard punctuation provides natural speech rhythm.
  • Example: It was a VERY long day [sigh] … nobody listens anymore.

Stability setting (the most important v3 setting)

  • Creative — more emotional/expressive, but prone to hallucinations.
  • Natural — closest to the original voice recording; balanced and neutral.
  • Robust — highly stable, but less responsive to directional prompts (audio tags).
  • For maximum expressiveness with audio tags, use Creative or Natural; Robust reduces responsiveness to tags.

Voice selection

  • The chosen voice is the single most important parameter — it must match the intended delivery (a shouting voice won't respond to [whispers]).
  • Neutral voices are more stable across languages.
  • PVCs aren't fully optimized for v3 — prefer IVC or designed voices.

IPA pronunciation

  • v3 natively supports IPA across 70+ languages — wrap transcriptions in forward slashes, e.g. /fəˈnɛtɪks/; include stress markers (ˈ primary, ˌ secondary). No XML tags needed. Reliability is roughly 80–90%.

Dos & don'ts

  • Do match tags to the voice's character — a professional voice may not respond well to playful tags.
  • Do combine multiple tags for complex emotional delivery; experiment.
  • Do use natural speech patterns and clear emotional context — text structure strongly influences v3 output.
  • Don't rely on experimental tags in production without testing.
  • Don't over-tag; most lines need only a few tags.

Example snippets

  • [whispers] I never knew it could be this way, but I'm glad we're here.
  • [applause] Thank you all for coming tonight! [gunshot] What was that?
  • [strong French accent] "Zat's life, my friend — you can't control everysing."

Known follow-ups (this repo)

Done

  1. Prod wiring. DialogueTaggingService (src/workbench/dialogueTagging.service.ts):
  2. generateScenesV2 awaits tagScenesInMemory after dialogue extract.
  3. Manual addDialogue / updateDialogueText / getStoryVideo are not tagged (a load-time backfill used to restore tags on refresh after a user stripped them).
  4. testDebug workbenchV2/tagDialogueForElevenV3 remains for scripts/eval.
  5. TTS. Tagged text already flows through dialogue audio gen → ElevenLabs (dialogue.text).

Still open

  1. Eval not in test_runner. dialogue_tagging stays unregistered in python_eval/tests/test_definitions/__init__.py (opt-in; G-Eval needs OPENAI_API_KEY). Uncomment the import + merge there to run it. Manual: python_eval/scripts/tag_scene_dialogue.py.
  2. Shot distribution is char-index based. DialogueUtils.redistributeDialogueAcrossShots (src/shared/dialogue.utils.ts) slices dialogue.text by raw character index when splitting a line across shots. With tags inlined in text, a [tag] could be split across a shot boundary, and tags would surface in on-screen caption portion.text. A follow-up should make distribution tag-aware (don't split inside a [...]; strip tags from caption text).
  3. Note: timing is already safe — DialogueUtils.calculateTextDuration strips [...] via DIALOGUE_BRACKET_TAG_PATTERN (src/shared/dialogue.constants.ts) and maps [pause]/[long pause] to fixed silence before computing duration.
  4. voiceDirection field fate. The tagging step does not consume voiceDirection (see "Design decision" below). Decide whether to keep or deprecate it on SceneDialogue (TODO on the field in src/shared/dialogue.types.ts).
  5. Scene Guru newDialogues. AI lines added via insert_shot are origin: GENERATED and are not auto-tagged today (same origin as manual addDialogue).

Design decision: the tagging step does NOT use voiceDirection

The eleven_v3 tagging step (tagDialogueForElevenV3.template.ts) deliberately ignores the per-line voiceDirection field. Its inputs are: each speaker's vibePersonality (inlined verbatim per line), the scene context (title / description / script), and deliverySpeed (kept as a pacing constraint). Reasons:

  • Same context, weaker source. voiceDirection is produced upstream by the Scene-breakdown extraction (generateScenesV2.template.ts) from the script — the same context this step already has — but by a prompt that has no access to vibePersonality. Re-consuming it would be a lossy round-trip: re-expanding a terse note that was itself a lossy compression of context this step reads directly.
  • Often a placeholder. For migrated, guru-added, and manually-added dialogue, voiceDirection is stamped with the literal default 'normal delivery' (see dialogue.migration.ts, workbenchSceneGuru.ts, workbench.service.ts), which carries no signal and would flatten tagging toward "normal."
  • Scene context is fused directly, not via voiceDirection. Dropping voiceDirection does not drop scene context — title/description/script are independent top-level inputs the system prompt explicitly fuses with each speaker's vibe (vibe = stable baseline, scene = situational modulation).

deliverySpeed is retained because it is functional, not stylistic: it feeds DialogueUtils.calculateTextDuration (timing). The tagging prompt treats it as a constraint to stay consistent with, not as intent to re-encode.