Selections strategy and metrics for voices
Roadmap and Strategy¶
In the UI, we have a list of voices in the "voice settings" section.
After obtaining metrics for some of the library voices from the experiment, how do we make use of them?
- We have already preselected the most used voice from 11labs (in the last year). We can show it at the top of the list in the voices-selection left-view in "Voice Settings".
- We order the voices by score, in the following descending order: excellent → good → good but limited emotional range → unsuited. These tags are found in the mongodb collection
evaluatedVoicesin the fieldoverall_rating. - After displaying all evaluated voices, you can show the remaining 11labs voices. Avoid overlaps by matching on voice_id.
This is the first step (should have its own PR).
Second Step¶
We also make use of the per-emotion alignment scores. These are found within the fields eval.emotions.<emotion-name>.emotional_alignment.avg_score within the evaluatedVoices collection. An avg_score of 4.0 is considered good, 5.0 is the maximum.
To reach this emotion-based suggestion we first infer the character’s most important emotions, then match them against the annotated voices in evaluatedVoices. This logic is now implemented (see PR #623):
- A single LLM call reads the character’s "vibe & personality" (
vibePersonality), visual identity (coreVisualIdentity), and tagged dialogue lines, and returns a weighted priority over the six macro emotions, plus an inferred voice gender and age group. - The candidate voice pool is hard-filtered by gender (and softly by age) before emotion scoring, so a character is never matched to an opposite-gender voice.
- Each remaining voice is scored as a priority-weighted mean of its per-emotion
avg_score, minus penalties for emotions the voice isunsuitedon or missing. Every voice in the pool is ranked, but only the top 9 are stored on the character assuggestedVoices(best fit first). - Trigger via
POST /projects/:projectId/characters/suggest-voices(fire-and-forget, returns 202; optional body{ characterId, force }—characterId= characterassetId). WithoutcharacterId, fills every character missing suggestions; with it, only that character if suggestions are absent; withforce, recomputes regardless (see the Third Step below). On completion, publishes JaduSpineassetUpdatedwith{ result: { project } }on the project channel.
Third Step — voice tags on the character, and keeping the list fresh¶
PEA-2626. The user can now say how a character should sound, and the suggested list reacts to it.
voiceTags— free-text descriptors on the character asset (projects.characters[].properties.voiceTags, e.g.['deep', 'wise', 'soft tone']), set in the character edit view next to the other description fields. Normalized on write (trimmed, whitespace-collapsed, de-duplicated case-insensitively, max 12 tags × 40 chars) bysrc/shared/voiceTags.ts, which is the only writer.- They feed the match on two paths, because emotion scores cannot express timbre:
- the derivation prompt reads them as deliberate casting intent — the strongest signal it has — and lets them raise emotion weights (
"menacing"→ anger,"warm"→ joy) and, only when a tag names one outright, override gender/age. Timbre words ("deep","raspy") are explicitly not gender or age evidence. - the matcher matches them literally against each voice's own
name + description + labels, adding up toTAG_AFFINITY_BONUS = 0.5scaled by the fraction of tags echoed. Sized as a tie-breaker on purpose: it reorders emotionally-comparable voices but cannot outrank theUNSUITED_PENALTY = 2.0a voice pays for being unsuited on a needed emotion. Tag matching ignores stopwords ("voice","tone","sound"appear in nearly every ElevenLabs description), matches any significant token of a multi-word tag, and prefix-matches from 4 characters ("rasp"→"raspy") while requiring whole words at 3. - Signal priority for the emotion weighting:
voiceTags(when they bear on emotion) → dialogue tags →vibePersonality→coreVisualIdentity. Appearance counts, just slightly less than personality, and loses to it on conflict — a "frail, elderly-looking man" whose vibe is "fiery, quick to rage" is an angry voice. The one exception is gender/age, where a stated sex or age incoreVisualIdentityis authoritative. - Recalculation is wired into the write, not left to the client. Editing a character's
voiceTags,vibePersonalityorcoreVisualIdentityschedules a forced recompute (suggestVoicesForProject({ characterId, force: true })) after the DB write and publishesassetUpdatedwhen it lands; regenerating a character's details does the same. Withoutforcethe existingsuggestedVoicesshort-circuit the recompute and the edit would silently change nothing. Edits that the derivation prompt never sees (clothingAndAccessories,background, title) and no-op re-saves (same tags reordered/re-cased, whitespace-only reflow) do not recompute — each recompute is an LLM call. - Overlapping edits are coalesced. A recompute takes an LLM call plus a rank, so adding three tags in a row would start three overlapping runs, and whichever finished last — not whichever read the newest tags — would decide the stored list. Requests arriving while a run is in flight for the same project + character are collapsed into one rerun afterwards (forced if any collapsed request was), which re-reads the character and therefore sees the final state. The guard is process-local (
ProjectsService.inFlightVoiceSuggestions): across instances behind a load balancer two edits can still overlap, which is the unguarded behaviour, not a worse one. - In the UI (studio-frontend, branch
feat/pea-2626-voice-tags-on-character-asset): - The tags are edited as chips in the character edit modal (
voiceTagsField.tsx), below the description fields. Every add/remove saves immediately — that save is what triggers the recompute — andapp/_shared/voiceTags.tsmirrors the backend's normalization rules so a duplicate or a 13th tag is refused in the editor instead of being accepted and silently dropped on write. The editor is hidden while a variant is being edited, because a variant save writes to the variant's own properties, which voice matching never reads. - The right-panel filter chips (gender, age, accent, language, category) now apply to the suggested section too (
matchesVoiceFilters). The marketplace list is filtered server-side byfilterEvaluatedVoices, but a character's suggestions are held client-side, so the same rules — including "a voice with no value for an active chip does not match" — are mirrored there; otherwise the same voice would be hidden from the Library section and kept in Suggested.
How the metrics were calculated during the Evaluation¶
- The scoring pipeline lives outside this repo, in
Scenarix/audio-tools(checked out locally asobtain-relevant-audio). Nothing in studio-backend computes these ratings — it only reads what that repo wrote intorenderboard.evaluatedVoices. Go there to change a threshold, re-score a voice, or add new voices to the picker: src/eval/qualispeech/resumable_runner.py— the two-pass eval run: P1 is a blind multimodal audio judge (qwen3-omni-flash) and P2 a text-only alignment scorer (claude-haiku-4-5);p2_eais the 1–5 score the ratings below are derived from.src/eval/qualispeech/metrics_per_emotion.pyandmetrics_per_voice.py— the authoritative implementation of the per-emotion and per-voice rating rules restated in this section. These two files, not this doc, decide whatoverall_ratinga voice gets.scripts/seed_evaluated_voices.py+docs/seeding_evaluated_voices.md— how documents land inevaluatedVoices. Insert-only by default;--allow-updateis required to touch an existing voice.docs/eval_sets.md— the versioned sentence sets.eval.eval_seton each document records which one produced its scores, so ratings are only comparable within a version.docs/rational_select_voices_for_prepopulation.md— how the candidate voices were chosen from the ElevenLabs shared library (English accents, ranked by 1-year character usage, balanced across gender × age).docs/selection_logic_voices.md— the origin of the ordering rules in the Roadmap section above; this doc is the polished restatement of it.
Each voice is tested on 6 emotions (joy, sadness, anger, fear, disgust, surprise), each with 9 sentences (3 intensities × 3 sentences). Every sentence is scored 1–5 for emotional alignment by the judge.
Per voice (aggregating the 6 emotion ratings) the overall rating is: - excellent — every emotion is excellent - good — at least 4 emotions are good and at least 1 more is limited-or-better - unsuited — the majority of emotions are unsuited - good but limited emotional range — anything else
The average emotional-alignment score used for tie-breaking is the stored emotional_alignment.avg_score for the relevant emotion(s) — the per-sentence scores themselves are not persisted in evaluatedVoices, only their per-emotion average, rating, sample count, and below-threshold count.
Per-emotion Metrics¶
Per-voice metrics are based on per-emotion metrics. For now we don't use these metrics singularly in isolation, but we might down the road to find best characters-to-voice fits based on emotion.
Per emotion (9 scores), the rating is: - excellent — all 9 scores are 5 - good — at most 1 score below 4.0 - good but limited emotional range — more than 1 but ≤ 30% of scores below 4.0 - unsuited — more than 30% of scores below 4.0
(For the 9-sentence set this works out to: limited range = exactly 2 scores below 4.0, unsuited = 3 or more below 4.0.)