ElevenLabs v3 (eleven_v3) Prompting Principles¶
Reference notes for directing the ElevenLabs eleven_v3 text-to-speech model, which is the
model all TTS in this system targets. Sourced from the official ElevenLabs v3 prompting
best-practices docs
(https://elevenlabs.io/docs/overview/capabilities/text-to-speech/best-practices).
For our own empirically-tested findings (single-quote emphasis, ? for volume, separate brackets,
letter-repetition for shouting, word-joining for speed, etc.), see the companion doc
learnings_11labs_prompting_v3.md. Both feed the
eleven_v3 dialogue-tagging prompt in
src/promptTemplates/workbenchV2/tagDialogueForElevenV3.template.ts.
Core idea¶
eleven_v3 is directed through inline audio tags (bracketed cues), punctuation, and
capitalization — not SSML. A tag colors the text that follows it. Text structure strongly
influences the output.
Audio tag catalog (by category)¶
Emotional / delivery direction
[laughs],[laughs harder],[starts laughing],[wheezing],[laughing hysterically][whispers][sighs],[exhales],[exhales sharply],[inhales deeply],[frustrated sigh],[happy gasp][sarcastic],[curious],[excited],[crying],[snorts],[mischievously][happy],[sad],[angry],[annoyed],[appalled],[thoughtful],[surprised][dismissive],[cute],[deadpan],[professional],[sympathetic],[questioning],[reassuring],[panicking]
Non-verbal sounds
[giggles],[chuckles],[clears throat],[swallows],[gulps]
Environmental sound effects
[gunshot],[applause],[clapping],[explosion]
Pacing / pauses
[short pause],[long pause],[pauses]- v3 does not support SSML
<break>tags — use audio tags, ellipses, and text structure for pacing instead. Speed is also directed via tags.
Multi-speaker / dialogue
[overlapping],[interrupting](assign a distinct voice per speaker, e.g. "Speaker 1:", "Speaker 2:")
Unique / experimental
[strong <X> accent](replace<X>, e.g.[strong French accent]),[sings],[singing quickly],[robotic voice],[binary beeping],[woo],[fart]
Many effective tags exist beyond this published list — descriptive states often work. Experimental tags can be inconsistent; test before production.
Punctuation & capitalization¶
- Ellipses (
…) add pauses and weight. - Capitalization increases emphasis.
- Standard punctuation provides natural speech rhythm.
- Example:
It was a VERY long day [sigh] … nobody listens anymore.
Stability setting (the most important v3 setting)¶
- Creative — more emotional/expressive, but prone to hallucinations.
- Natural — closest to the original voice recording; balanced and neutral.
- Robust — highly stable, but less responsive to directional prompts (audio tags).
- For maximum expressiveness with audio tags, use Creative or Natural; Robust reduces responsiveness to tags.
Voice selection¶
- The chosen voice is the single most important parameter — it must match the intended delivery
(a shouting voice won't respond to
[whispers]). - Neutral voices are more stable across languages.
- PVCs aren't fully optimized for v3 — prefer IVC or designed voices.
IPA pronunciation¶
- v3 natively supports IPA across 70+ languages — wrap transcriptions in forward slashes, e.g.
/fəˈnɛtɪks/; include stress markers (ˈprimary,ˌsecondary). No XML tags needed. Reliability is roughly 80–90%.
Dos & don'ts¶
- Do match tags to the voice's character — a professional voice may not respond well to playful tags.
- Do combine multiple tags for complex emotional delivery; experiment.
- Do use natural speech patterns and clear emotional context — text structure strongly influences v3 output.
- Don't rely on experimental tags in production without testing.
- Don't over-tag; most lines need only a few tags.
Example snippets¶
[whispers] I never knew it could be this way, but I'm glad we're here.[applause] Thank you all for coming tonight! [gunshot] What was that?[strong French accent] "Zat's life, my friend — you can't control everysing."
Known follow-ups (this repo)¶
Done¶
- Prod wiring.
DialogueTaggingService(src/workbench/dialogueTagging.service.ts): generateScenesV2awaitstagScenesInMemoryafter dialogue extract.- Manual
addDialogue/updateDialogueText/getStoryVideoare not tagged (a load-time backfill used to restore tags on refresh after a user stripped them). - testDebug
workbenchV2/tagDialogueForElevenV3remains for scripts/eval. - TTS. Tagged
textalready flows through dialogue audio gen → ElevenLabs (dialogue.text).
Still open¶
- Eval not in test_runner.
dialogue_taggingstays unregistered inpython_eval/tests/test_definitions/__init__.py(opt-in; G-Eval needsOPENAI_API_KEY). Uncomment the import + merge there to run it. Manual:python_eval/scripts/tag_scene_dialogue.py. - Shot distribution is char-index based.
DialogueUtils.redistributeDialogueAcrossShots(src/shared/dialogue.utils.ts) slicesdialogue.textby raw character index when splitting a line across shots. With tags inlined intext, a[tag]could be split across a shot boundary, and tags would surface in on-screen captionportion.text. A follow-up should make distribution tag-aware (don't split inside a[...]; strip tags from caption text). - Note: timing is already safe —
DialogueUtils.calculateTextDurationstrips[...]viaDIALOGUE_BRACKET_TAG_PATTERN(src/shared/dialogue.constants.ts) and maps[pause]/[long pause]to fixed silence before computing duration. voiceDirectionfield fate. The tagging step does not consumevoiceDirection(see "Design decision" below). Decide whether to keep or deprecate it onSceneDialogue(TODO on the field insrc/shared/dialogue.types.ts).- Scene Guru
newDialogues. AI lines added viainsert_shotareorigin: GENERATEDand are not auto-tagged today (same origin as manualaddDialogue).
Design decision: the tagging step does NOT use voiceDirection¶
The eleven_v3 tagging step (tagDialogueForElevenV3.template.ts) deliberately ignores the
per-line voiceDirection field. Its inputs are: each speaker's vibePersonality (inlined
verbatim per line), the scene context (title / description / script), and deliverySpeed
(kept as a pacing constraint). Reasons:
- Same context, weaker source.
voiceDirectionis produced upstream by the Scene-breakdown extraction (generateScenesV2.template.ts) from the script — the same context this step already has — but by a prompt that has no access tovibePersonality. Re-consuming it would be a lossy round-trip: re-expanding a terse note that was itself a lossy compression of context this step reads directly. - Often a placeholder. For migrated, guru-added, and manually-added dialogue,
voiceDirectionis stamped with the literal default'normal delivery'(seedialogue.migration.ts,workbenchSceneGuru.ts,workbench.service.ts), which carries no signal and would flatten tagging toward "normal." - Scene context is fused directly, not via
voiceDirection. DroppingvoiceDirectiondoes not drop scene context — title/description/script are independent top-level inputs the system prompt explicitly fuses with each speaker's vibe (vibe = stable baseline, scene = situational modulation).
deliverySpeed is retained because it is functional, not stylistic: it feeds
DialogueUtils.calculateTextDuration (timing). The tagging prompt treats it as a constraint to stay
consistent with, not as intent to re-encode.