Skip to content

AssetGen Module

Voice Settings

Character voice assignment via ElevenLabs. Voices are stored as lean VoiceEntry[] on the project asset -- display data is resolved at runtime from source collections.

Generated Voice Flow

  1. User describes a voice -> POST /elevenLabs/text-to-voice/design -> returns 3 previews with audio_base_64
  2. User picks one -> POST /elevenLabs/text-to-voice -> permanent voice created in ElevenLabs + saved to voices collection

Library Voice Flow

  1. User picks a marketplace voice -> POST /elevenLabs/addVoice -> voice saved to voices collection
  2. TTS job runs with character's dialog text -> POST /createJob -> custom audio stored in AssetGenJob.outputLinks.result

Pitch Variation Flow (Remix)

  1. User selects pitch (high/low) + strength -> POST /elevenLabs/remix -> returns 3 remixed previews
  2. User picks one -> POST /elevenLabs/text-to-voice -> permanent voice saved to voices collection

Library Voice Pitch Variation (IVC bridge = Instant Voice Cloning )

Library (marketplace) voice ids are not owned by our ElevenLabs account, so remix fails directly. For source: 'library':

  1. FE sends assetGenJobId (TTS preview job from library voice assignment) with the remix request
  2. BE loads AssetGenJob.outputLinks.result (mp3), creates a temporary IVC via POST /v1/voices/add
  3. BE runs remix on the temp owned voice_id, returns 3 previews
  4. BE deletes the temp voice from ElevenLabs only (no Mongo write)
  5. Confirm step is unchanged: POST /elevenLabs/text-to-voice + PATCH .../voices/select with variation

If assetGenJobId is missing, BE returns 400 — re-select the library voice to generate preview audio.

Data Resolution (on voice settings open)

Endpoint Source Returns
POST /voices/getByIds voices collection name, preview_url, user_id, created_at
POST /voices/noticePeriods voiceNoticePeriods collection notice_period, disable_at_unix (deletion risk)
POST /getJobsByIds AssetGenJob collection outputLinks.result (library voice audio)
POST /users/resolveNames users collection userId -> display name map

All batch endpoints capped at max 50 per request (Zod validated). Frontend chunks larger arrays to match.

Audio Resolution Logic

  • Generated voice -> voices collection preview_url (custom dialog audio from design step)
  • Library voice -> AssetGenJob outputLinks.result (custom dialog TTS on B2)
  • Variation voice -> voices collection preview_url (remix audio)

Collections

Collection Purpose
voices (MongoDB) Local cache of ElevenLabs voice metadata (premade + user voices)
voiceNoticePeriods (MongoDB) Cached ElevenLabs sharing terms per voice_id, for deletion-risk warnings
AssetGenJob TTS job results -- library voice custom audio lives here, and other jobs also
Project asset voices[] Lean VoiceEntry array: voiceId, selected, source, assetGenJobId?, variations?

Voice Deletion Risk

Library voice owners can remove a voice after its notice_period (days, 30–730) or with no notice at all, so the studio warns users before they commit to a risky voice.

  • Cache: voiceNoticePeriods, keyed by voice_id. Separate from voices because that collection is keyed (voice_id, user_id) while sharing terms are a property of the voice itself.
  • Source: ElevenLabsService.getVoiceSharingInfo (GET /v1/voices/{id} -> sharing). Per-voice on purpose: batch voices.search({ voiceIds }) silently drops ids outside our workspace.
  • Reads: VoiceNoticePeriodService.resolve answers from the cache and schedules a bounded (10 per call) background refresh for missing/stale (>7 day) ids. It never blocks the picker on an upstream call.
  • Fail open: toRiskFields omits both fields for an unknown or non-shared voice, so the UI reads "no data" rather than "no notice period". A notice_period of null on the wire means the voice genuinely has none.
  • Backfill: npx ts-node scripts/backfillVoiceNoticePeriods.ts seeds every evaluated + stored library voice (one-off). --skip-cached resumes an interrupted run; --from-csv <path> seeds from a voiceId,noticePeriodDays export with no upstream calls (but carries no disable_at_unix).
  • Risk band: the rule lives in the frontend (app/_shared/voiceDeletionRisk.ts); the backend only supplies raw terms. 90–180 days gets the "At risk" badge; below 90 days (and no notice period at all, and any scheduled removal) the voice is hidden from the picker entirely, though a voice already assigned to a character still warns. Above 180 days is treated as safe. disable_at_unix is cached for the follow-up that emails collaborators when a voice in use is scheduled for deletion.

The non-evaluated listSharedVoices branch needs no enrichment — ElevenLabs' shared-voices payload already carries noticePeriod.

Voice Slots

Our plan has 160 voice slots, and only voices we create occupy one: Voice Design (generated) and instant voice clones (cloned). Voices saved from the shared library cost nothing, and neither do premade ones — so the workspace total (several hundred) says nothing about remaining capacity.

  • Check: ElevenLabsService.ensureVoiceSlotAvailable reads voiceSlotsUsed / voiceLimit from GET /v1/user/subscription. SharedConstants.ELEVEN_LABS_MAX_VOICES_LIMIT_CHECK (160) is only the fallback when that call fails.
  • Auto-add on use: rendering TTS with a shared voice id makes ElevenLabs add that voice to the workspace itself, as professional and costing no slot. We never call the shared-voice add endpoint — addSharedVoiceToMyLibrary only writes our Mongo cache. So the workspace total grows with library usage while voiceSlotsUsed does not.
  • Where it belongs: only on paths that create a voice — createVoiceFromPreview and remixLibraryVoice. TTS/STS renders audio from an existing voice id and consumes no slot, so it does no check. Paid plans can render a library voice without it being saved to the workspace.
  • When full: it reclaims leaked temp-library-remix-* clones (temporaries remixLibraryVoice normally deletes itself), oldest first. If there are none it throws rather than guess, because any other voice may still be referenced by a project.

Dialog Preview

Each character has a dialogPreview string (min 100 chars) generated on character creation. Template: "My name is {title} and I'm the {role} in {project}." Padded with character description if under 100 chars. Used as the spoken text for all voice previews.

Voice Selection (Assign / Restore)

After creating a voice via the endpoints above, the FE selects it on the character via a dedicated endpoint in the projects module:

PATCH /projects/:projectId/assets/:assetId/voices/select
Body: { voiceId, source?, assetGenJobId?, variation?: { voiceId, pitch, strength } }

The BE determines the action automatically:

source present? variation present? voiceId exists? Action Activity Logged
yes no — Add new VoiceEntry, select it ASSIGN_CHARACTER_VOICE
no no yes Select existing voice RESTORE_CHARACTER_VOICE
— yes (new) — Add new Variation on parent ASSIGN_CHARACTER_VOICE_VARIATION
— yes (existing) yes Select existing variation RESTORE_CHARACTER_VOICE_VARIATION

Activity logging is fire-and-forget. Display data (voiceName, audioUrl) is resolved from source collections before logging:

Voice type voiceName from audioUrl from
Generated VoicesModel.getVoicesByIds() -> name -> preview_url
Library AssetJobModel.getJobsByIds() -> modelConfig.modelTitle -> outputLinks.result
Variation VoicesModel.getVoicesByIds() -> name -> preview_url

Key files: src/projects/projects.service.ts (selectVoice, handleBaseVoice, handleVariation), src/projects/projects.validator.ts (validateSelectVoiceRequest), src/projects/projects.controller.ts (selectVoice). See src/projectActivities/README.md for activity logging conventions.