AssetGen Module¶
Voice Settings¶
Character voice assignment via ElevenLabs. Voices are stored as lean VoiceEntry[] on the project asset -- display data is resolved at runtime from source collections.
Generated Voice Flow¶
- User describes a voice ->
POST /elevenLabs/text-to-voice/design-> returns 3 previews withaudio_base_64 - User picks one ->
POST /elevenLabs/text-to-voice-> permanent voice created in ElevenLabs + saved tovoicescollection
Library Voice Flow¶
- User picks a marketplace voice ->
POST /elevenLabs/addVoice-> voice saved tovoicescollection - TTS job runs with character's dialog text ->
POST /createJob-> custom audio stored inAssetGenJob.outputLinks.result
Pitch Variation Flow (Remix)¶
- User selects pitch (high/low) + strength ->
POST /elevenLabs/remix-> returns 3 remixed previews - User picks one ->
POST /elevenLabs/text-to-voice-> permanent voice saved tovoicescollection
Library Voice Pitch Variation (IVC bridge = Instant Voice Cloning )¶
Library (marketplace) voice ids are not owned by our ElevenLabs account, so remix fails directly. For source: 'library':
- FE sends
assetGenJobId(TTS preview job from library voice assignment) with the remix request - BE loads
AssetGenJob.outputLinks.result(mp3), creates a temporary IVC viaPOST /v1/voices/add - BE runs remix on the temp owned
voice_id, returns 3 previews - BE deletes the temp voice from ElevenLabs only (no Mongo write)
- Confirm step is unchanged:
POST /elevenLabs/text-to-voice+PATCH .../voices/selectwithvariation
If assetGenJobId is missing, BE returns 400 — re-select the library voice to generate preview audio.
Data Resolution (on voice settings open)¶
| Endpoint | Source | Returns |
|---|---|---|
POST /voices/getByIds |
voices collection |
name, preview_url, user_id, created_at |
POST /voices/noticePeriods |
voiceNoticePeriods collection |
notice_period, disable_at_unix (deletion risk) |
POST /getJobsByIds |
AssetGenJob collection |
outputLinks.result (library voice audio) |
POST /users/resolveNames |
users collection |
userId -> display name map |
All batch endpoints capped at max 50 per request (Zod validated). Frontend chunks larger arrays to match.
Audio Resolution Logic¶
- Generated voice ->
voicescollectionpreview_url(custom dialog audio from design step) - Library voice ->
AssetGenJoboutputLinks.result(custom dialog TTS on B2) - Variation voice ->
voicescollectionpreview_url(remix audio)
Collections¶
| Collection | Purpose |
|---|---|
voices (MongoDB) |
Local cache of ElevenLabs voice metadata (premade + user voices) |
voiceNoticePeriods (MongoDB) |
Cached ElevenLabs sharing terms per voice_id, for deletion-risk warnings |
AssetGenJob |
TTS job results -- library voice custom audio lives here, and other jobs also |
Project asset voices[] |
Lean VoiceEntry array: voiceId, selected, source, assetGenJobId?, variations? |
Voice Deletion Risk¶
Library voice owners can remove a voice after its notice_period (days, 30–730) or with no notice at all, so the studio warns users before they commit to a risky voice.
- Cache:
voiceNoticePeriods, keyed byvoice_id. Separate fromvoicesbecause that collection is keyed(voice_id, user_id)while sharing terms are a property of the voice itself. - Source:
ElevenLabsService.getVoiceSharingInfo(GET /v1/voices/{id}->sharing). Per-voice on purpose: batchvoices.search({ voiceIds })silently drops ids outside our workspace. - Reads:
VoiceNoticePeriodService.resolveanswers from the cache and schedules a bounded (10 per call) background refresh for missing/stale (>7 day) ids. It never blocks the picker on an upstream call. - Fail open:
toRiskFieldsomits both fields for an unknown or non-shared voice, so the UI reads "no data" rather than "no notice period". Anotice_periodofnullon the wire means the voice genuinely has none. - Backfill:
npx ts-node scripts/backfillVoiceNoticePeriods.tsseeds every evaluated + stored library voice (one-off).--skip-cachedresumes an interrupted run;--from-csv <path>seeds from avoiceId,noticePeriodDaysexport with no upstream calls (but carries nodisable_at_unix). - Risk band: the rule lives in the frontend (
app/_shared/voiceDeletionRisk.ts); the backend only supplies raw terms. 90–180 days gets the "At risk" badge; below 90 days (and no notice period at all, and any scheduled removal) the voice is hidden from the picker entirely, though a voice already assigned to a character still warns. Above 180 days is treated as safe.disable_at_unixis cached for the follow-up that emails collaborators when a voice in use is scheduled for deletion.
The non-evaluated listSharedVoices branch needs no enrichment — ElevenLabs' shared-voices payload already carries noticePeriod.
Voice Slots¶
Our plan has 160 voice slots, and only voices we create occupy one: Voice Design (generated) and instant voice clones (cloned). Voices saved from the shared library cost nothing, and neither do premade ones — so the workspace total (several hundred) says nothing about remaining capacity.
- Check:
ElevenLabsService.ensureVoiceSlotAvailablereadsvoiceSlotsUsed/voiceLimitfromGET /v1/user/subscription.SharedConstants.ELEVEN_LABS_MAX_VOICES_LIMIT_CHECK(160) is only the fallback when that call fails. - Auto-add on use: rendering TTS with a shared voice id makes ElevenLabs add that voice to the workspace itself, as
professionaland costing no slot. We never call the shared-voice add endpoint —addSharedVoiceToMyLibraryonly writes our Mongo cache. So the workspace total grows with library usage whilevoiceSlotsUseddoes not. - Where it belongs: only on paths that create a voice —
createVoiceFromPreviewandremixLibraryVoice. TTS/STS renders audio from an existing voice id and consumes no slot, so it does no check. Paid plans can render a library voice without it being saved to the workspace. - When full: it reclaims leaked
temp-library-remix-*clones (temporariesremixLibraryVoicenormally deletes itself), oldest first. If there are none it throws rather than guess, because any other voice may still be referenced by a project.
Dialog Preview¶
Each character has a dialogPreview string (min 100 chars) generated on character creation. Template: "My name is {title} and I'm the {role} in {project}." Padded with character description if under 100 chars. Used as the spoken text for all voice previews.
Voice Selection (Assign / Restore)¶
After creating a voice via the endpoints above, the FE selects it on the character via a dedicated endpoint in the projects module:
PATCH /projects/:projectId/assets/:assetId/voices/select
Body: { voiceId, source?, assetGenJobId?, variation?: { voiceId, pitch, strength } }
The BE determines the action automatically:
source present? |
variation present? |
voiceId exists? | Action | Activity Logged |
|---|---|---|---|---|
| yes | no | — | Add new VoiceEntry, select it | ASSIGN_CHARACTER_VOICE |
| no | no | yes | Select existing voice | RESTORE_CHARACTER_VOICE |
| — | yes (new) | — | Add new Variation on parent | ASSIGN_CHARACTER_VOICE_VARIATION |
| — | yes (existing) | yes | Select existing variation | RESTORE_CHARACTER_VOICE_VARIATION |
Activity logging is fire-and-forget. Display data (voiceName, audioUrl) is resolved from source collections before logging:
| Voice type | voiceName from | audioUrl from |
|---|---|---|
| Generated | VoicesModel.getVoicesByIds() -> name |
-> preview_url |
| Library | AssetJobModel.getJobsByIds() -> modelConfig.modelTitle |
-> outputLinks.result |
| Variation | VoicesModel.getVoicesByIds() -> name |
-> preview_url |
Key files: src/projects/projects.service.ts (selectVoice, handleBaseVoice, handleVariation), src/projects/projects.validator.ts (validateSelectVoiceRequest), src/projects/projects.controller.ts (selectVoice). See src/projectActivities/README.md for activity logging conventions.