Skip to content

Batched shot image generation

Generates N images for a shot in a single user action. Used in three flows:

  1. Initial generation (sketch → image, via POST /workbench/generateShotImage)
  2. Edit (via the shot-image-edit guru chat, edit_shot_image tool)
  3. Restart from sketch (via guru or direct api, restart_from_sketch tool)

count == 1 keeps the existing single-image code paths byte-identical. count >= 2 opts into the batch path. Capped at SHOT_IMAGE_GEN_MAX_BATCH_COUNT (sharedConstants.ts).

How count gets to the backend

Flow Source Path
Initial-gen count and/or imageGenerationCounts (per-model, keyed by ModelConfigType.id) on POST /workbench/generateShotImage body workbench.validator.ts
Edit / restart-from-sketch referenceData.imageGenerationCount on the user chat message guruChat.validator.ts

For the chat flows the LLM never sees the count — it's read from the latest user message's referenceData (shotImageEditGuru.getLatestImageGenerationCounts) so a tampering tool call can't change it. FE owns entitlement gating.

Which model runs each slot

Two registries, one per flow, keyed by ModelConfigType.id (the key the FE sends counts under):

Flow Registry Default composition
Edit (edit_shot_image) shotImageEditModels.ts FE default: 2 × Nano Banana Pro + 1 × GPT Image 2.5 Sunburst + 1 × GPT Image 2.5 Flare
Sketch → image (initial-gen, restart-from-sketch, picker reject-and-regenerate) shotImageGenModels.ts SHOT_IMAGE_GEN_SLOT_ORDER: Nano, Nano, Flare, Sunburst

Initial-gen with imageGenerationCounts builds slots on the request's own model from the request's modelConfig (keeping FE-chosen inputs and any pre-gen-review prompt) and the other slots from the registry with the same prompt and references. Restart-from-sketch honours only the FE total and fills slots in SHOT_IMAGE_GEN_SLOT_ORDER. OpenAI models are described once in openAIImageModels.ts.

Two emission patterns

Flow When backend HTTP returns Socket events
Initial-gen batch Immediately, with placeholder shot One SHOT_IMAGE_BATCH_STARTED, then one consolidated SHOT_UPDATED after all N slots settle
Edit / restart-from-sketch batch Immediately, tool result includes { status: 'pending', batchId, batchSize, placeholderJobIds, shot } One SHOT_IMAGE_BATCH_STARTED, then N progressive SHOT_IMAGE_BATCH_PROGRESS events, followed by one authoritative SHOT_UPDATED

Initial-gen is consolidated because most providers are slow (Replicate / Fal can be 20–60s) and the FE renders one place. Edit-batch is progressive because edits are fast (Gemini, ~5–10s parallel) and live in chat where progressive feedback feels better.

Socket event payloads

SHOT_UPDATED (existing event, unchanged shape — used by initial-gen):

{ shot, job, jobs }   // job=jobs[0] for legacy handlers; jobs=full list

SHOT_IMAGE_BATCH_STARTED (edit/restart and picker regeneration):

{ storyId, sceneId, shotId, shot, batchId, batchSize, initiatedByUserId }
This is the authoritative project-level delivery of the placeholder shot. Every tab signed in as initiatedByUserId registers the picker; collaborators apply the shot snapshot without opening one. If the originating tab already has an unbound HTTP-owned operation for the shot, it binds that operation to the batch.

SHOT_IMAGE_BATCH_PROGRESS (edit/restart only):

{ batchId, batchIndex, batchSize, jobId, status, images?, error? }
Slot events carry only images[] (not the full shot). The full shot is delivered through SHOT_IMAGE_BATCH_STARTED before any slot events arrive, with the synchronous tool/HTTP result retained as an initiating-client fallback.

After every progressive slot settles, edit/restart batches publish one final SHOT_UPDATED containing the authoritative full shot. This reconciles any missed progress delivery and gives every batch flow the same terminal event.

Both events are wrapped in the standard JaduSpine envelope: { topic, data: { event, status, isSuccess, message, data: <payload>, timestamp } }.

How batched images are stored

Each batched image is an ImageMedia row in shot.images[]. Cohort identity lives in imageParams:

imageParams: {
  assetGenJobId: string;
  assetGenJobStatus: JobStatus;     // 'processing' | 'completed' | 'failed'
  batchId: string;                  // shared across the cohort
  batchSize: number;                // total expected
  batchIndex: number;               // 0..N-1, stable order
  // ...optional fields like generationRefs from sketch-to-shot
  // model config lives on the AssetGenJob (via assetGenJobId), not here
}

For count <= 1 none of batchId / batchSize / batchIndex are written — backward-compatible with FE that doesn't know about batches.

FE groups by batchId to render the cohort and sorts by batchIndex for stable display order. While slots are pending, imageParams.assetGenJobStatus === 'processing' and imageURL === '' — render those as skeleton tiles.

Selection lifecycle

The contract: isSelected: true means displayable image. Empty placeholders are never selected.

  1. Placeholder insert — addShotImageBatch inserts N entries with isSelected: false. Any pre-existing displayable selection (image with non-empty URL) is preserved.
  2. First slot to land successfully — patchShotImageJobResult runs a guarded atomic update: if no slot in this batch is already selected, deselect everything else and select this slot. Concurrent landings can't race because the guard is in the doc-level filter.
  3. Subsequent slots — leave selection alone. User can swap deliberately via the existing updateShotImageSelection endpoint.

This matches the single-image gen UX: a freshly generated image takes over selection from the prior one.

Concurrency

N slot runs happen in parallel. patchShotImageJobResult uses Mongo arrayFilters keyed on imageParams.assetGenJobId so each runner writes only its own slot. The selection promotion in step 2 above is also atomic — see shotImageBatch.service.test.ts and the storyVideosModel concurrency tests for the regression guards.

Activity logging

One GENERATE_SHOT_IMAGE (or EDIT_SHOT_IMAGE) project activity row per slot — the activity feed reflects every generated image, not just one. Dispatch creates N rows in PROCESSING (each keyed to a placeholder job id); each slot's runner flips its own row to COMPLETED / FAILED when its result lands. Cohort grouping is available via imageParams.batchId on the resulting shot images if a consumer wants it.

Defensive validation

Three layers enforce the count cap: 1. Validator (SHOT_IMAGE_GEN_MAX_BATCH_COUNT = 4) 2. getLatestImageGenerationCount falls back to 1 for any out-of-range value 3. The dispatch*Batch methods throw if count is outside [2, MAX] — last-resort guard if the upper layers ever regress.

Future work

  • Conditional batch dispatch based on edit magnitude (skip batch for small touch-ups). See @todo on runBatchedShotImageEdit.
  • Per-slot retry on transient provider failures.
  • Per-model counts on restart-from-sketch (today it honours the total only and fills slots in SHOT_IMAGE_GEN_SLOT_ORDER).