Skip to content

Selections strategy and metrics for voices

Roadmap and Strategy

In the UI, we have a list of voices in the "voice settings" section.
After obtaining metrics for some of the library voices from the experiment, how do we make use of them?

  1. We have already preselected the most used voice from 11labs (in the last year). We can show it at the top of the list in the voices-selection left-view in "Voice Settings".
  2. We order the voices by score, in the following descending order: excellent → good → good but limited emotional range → unsuited. These tags are found in the mongodb collection evaluatedVoices in the field overall_rating.
  3. After displaying all evaluated voices, you can show the remaining 11labs voices. Avoid overlaps by matching on voice_id.

This is the first step (should have its own PR).

Second Step

We also make use of the per-emotion alignment scores. These are found within the fields eval.emotions.<emotion-name>.emotional_alignment.avg_score within the evaluatedVoices collection. An avg_score of 4.0 is considered good, 5.0 is the maximum.

To reach this emotion-based suggestion we first infer the character’s most important emotions, then match them against the annotated voices in evaluatedVoices. This logic is now implemented (see PR #623):

  • A single LLM call reads the character’s "vibe & personality" (vibePersonality), visual identity (coreVisualIdentity), and tagged dialogue lines, and returns a weighted priority over the six macro emotions, plus an inferred voice gender and age group.
  • The candidate voice pool is hard-filtered by gender (and softly by age) before emotion scoring, so a character is never matched to an opposite-gender voice.
  • Each remaining voice is scored as a priority-weighted mean of its per-emotion avg_score, minus penalties for emotions the voice is unsuited on or missing. Every voice in the pool is ranked, but only the top 9 are stored on the character as suggestedVoices (best fit first).
  • Trigger via POST /projects/:projectId/characters/suggest-voices (fire-and-forget, returns 202; optional body { characterId, force } — characterId = character assetId). Without characterId, fills every character missing suggestions; with it, only that character if suggestions are absent; with force, recomputes regardless (see the Third Step below). On completion, publishes JaduSpine assetUpdated with { result: { project } } on the project channel.

Third Step — voice tags on the character, and keeping the list fresh

PEA-2626. The user can now say how a character should sound, and the suggested list reacts to it.

  • voiceTags — free-text descriptors on the character asset (projects.characters[].properties.voiceTags, e.g. ['deep', 'wise', 'soft tone']), set in the character edit view next to the other description fields. Normalized on write (trimmed, whitespace-collapsed, de-duplicated case-insensitively, max 12 tags × 40 chars) by src/shared/voiceTags.ts, which is the only writer.
  • They feed the match on two paths, because emotion scores cannot express timbre:
  • the derivation prompt reads them as deliberate casting intent — the strongest signal it has — and lets them raise emotion weights ("menacing" → anger, "warm" → joy) and, only when a tag names one outright, override gender/age. Timbre words ("deep", "raspy") are explicitly not gender or age evidence.
  • the matcher matches them literally against each voice's own name + description + labels, adding up to TAG_AFFINITY_BONUS = 0.5 scaled by the fraction of tags echoed. Sized as a tie-breaker on purpose: it reorders emotionally-comparable voices but cannot outrank the UNSUITED_PENALTY = 2.0 a voice pays for being unsuited on a needed emotion. Tag matching ignores stopwords ("voice", "tone", "sound" appear in nearly every ElevenLabs description), matches any significant token of a multi-word tag, and prefix-matches from 4 characters ("rasp" → "raspy") while requiring whole words at 3.
  • Signal priority for the emotion weighting: voiceTags (when they bear on emotion) → dialogue tags → vibePersonality → coreVisualIdentity. Appearance counts, just slightly less than personality, and loses to it on conflict — a "frail, elderly-looking man" whose vibe is "fiery, quick to rage" is an angry voice. The one exception is gender/age, where a stated sex or age in coreVisualIdentity is authoritative.
  • Recalculation is wired into the write, not left to the client. Editing a character's voiceTags, vibePersonality or coreVisualIdentity schedules a forced recompute (suggestVoicesForProject({ characterId, force: true })) after the DB write and publishes assetUpdated when it lands; regenerating a character's details does the same. Without force the existing suggestedVoices short-circuit the recompute and the edit would silently change nothing. Edits that the derivation prompt never sees (clothingAndAccessories, background, title) and no-op re-saves (same tags reordered/re-cased, whitespace-only reflow) do not recompute — each recompute is an LLM call.
  • Overlapping edits are coalesced. A recompute takes an LLM call plus a rank, so adding three tags in a row would start three overlapping runs, and whichever finished last — not whichever read the newest tags — would decide the stored list. Requests arriving while a run is in flight for the same project + character are collapsed into one rerun afterwards (forced if any collapsed request was), which re-reads the character and therefore sees the final state. The guard is process-local (ProjectsService.inFlightVoiceSuggestions): across instances behind a load balancer two edits can still overlap, which is the unguarded behaviour, not a worse one.
  • In the UI (studio-frontend, branch feat/pea-2626-voice-tags-on-character-asset):
  • The tags are edited as chips in the character edit modal (voiceTagsField.tsx), below the description fields. Every add/remove saves immediately — that save is what triggers the recompute — and app/_shared/voiceTags.ts mirrors the backend's normalization rules so a duplicate or a 13th tag is refused in the editor instead of being accepted and silently dropped on write. The editor is hidden while a variant is being edited, because a variant save writes to the variant's own properties, which voice matching never reads.
  • The right-panel filter chips (gender, age, accent, language, category) now apply to the suggested section too (matchesVoiceFilters). The marketplace list is filtered server-side by filterEvaluatedVoices, but a character's suggestions are held client-side, so the same rules — including "a voice with no value for an active chip does not match" — are mirrored there; otherwise the same voice would be hidden from the Library section and kept in Suggested.

How the metrics were calculated during the Evaluation

  • The scoring pipeline lives outside this repo, in Scenarix/audio-tools (checked out locally as obtain-relevant-audio). Nothing in studio-backend computes these ratings — it only reads what that repo wrote into renderboard.evaluatedVoices. Go there to change a threshold, re-score a voice, or add new voices to the picker:
  • src/eval/qualispeech/resumable_runner.py — the two-pass eval run: P1 is a blind multimodal audio judge (qwen3-omni-flash) and P2 a text-only alignment scorer (claude-haiku-4-5); p2_ea is the 1–5 score the ratings below are derived from.
  • src/eval/qualispeech/metrics_per_emotion.py and metrics_per_voice.py — the authoritative implementation of the per-emotion and per-voice rating rules restated in this section. These two files, not this doc, decide what overall_rating a voice gets.
  • scripts/seed_evaluated_voices.py + docs/seeding_evaluated_voices.md — how documents land in evaluatedVoices. Insert-only by default; --allow-update is required to touch an existing voice.
  • docs/eval_sets.md — the versioned sentence sets. eval.eval_set on each document records which one produced its scores, so ratings are only comparable within a version.
  • docs/rational_select_voices_for_prepopulation.md — how the candidate voices were chosen from the ElevenLabs shared library (English accents, ranked by 1-year character usage, balanced across gender × age).
  • docs/selection_logic_voices.md — the origin of the ordering rules in the Roadmap section above; this doc is the polished restatement of it.

Each voice is tested on 6 emotions (joy, sadness, anger, fear, disgust, surprise), each with 9 sentences (3 intensities × 3 sentences). Every sentence is scored 1–5 for emotional alignment by the judge.

Per voice (aggregating the 6 emotion ratings) the overall rating is: - excellent — every emotion is excellent - good — at least 4 emotions are good and at least 1 more is limited-or-better - unsuited — the majority of emotions are unsuited - good but limited emotional range — anything else

The average emotional-alignment score used for tie-breaking is the stored emotional_alignment.avg_score for the relevant emotion(s) — the per-sentence scores themselves are not persisted in evaluatedVoices, only their per-emotion average, rating, sample count, and below-threshold count.

Per-emotion Metrics

Per-voice metrics are based on per-emotion metrics. For now we don't use these metrics singularly in isolation, but we might down the road to find best characters-to-voice fits based on emotion.

Per emotion (9 scores), the rating is: - excellent — all 9 scores are 5 - good — at most 1 score below 4.0 - good but limited emotional range — more than 1 but ≤ 30% of scores below 4.0 - unsuited — more than 30% of scores below 4.0

(For the 9-sentence set this works out to: limited range = exactly 2 scores below 4.0, unsuited = 3 or more below 4.0.)