Skip to content

Scene Preflight

Preflight prepares scenes for sketch generation by analyzing assets and shots before anything is rendered: what does each character/prop/environment look like, which assets appear in which shot, and which environment plate should back each shot.

This module is being brought into production incrementally from the original preflight prototype (PR #657). This doc describes what is live today and what is not yet ported. When you add a phase, update this doc.


What exists today

   IMAGE ADDED (upload / select from library / generation)
        │
        ▼
   workbench.service.ts — triggerAnalysisForAssetImage (fire-and-forget)
        │
        ▼
   ┌────────────────────────────────────────────────────────┐
   │ VLM ANALYSIS (one per image, stored on the image)      │
   │  characterAnalysis.service.ts   — GPT-5.4 vision       │
   │  propAnalysis.service.ts        — GPT-5.4 vision       │
   │  environmentAnalysis.service.ts — GPT-5.4 vision, 2    │
   │    passes (objects+bboxes, then script-alias mapping)  │
   └────────────────────────────────────────────────────────┘
        │  writes ImageMedia.imageAnalysis (atomic per-image, projects collection)
        ▼

   ON DEMAND (separate APIs, user- or tool-invoked)
        │
        ▼
   ┌────────────────────────────────────────────────────────┐
   │ scenePreflight.service.ts — orchestrator               │
   │                                                        │
   │  linkSceneAssets            NER: shot text → asset     │
   │   (assetLinking.service)    links, GPT-5-mini          │
   │                                                        │
   │  selectSceneVariants        best env plate per shot +  │
   │   (variantSelection.service) missing-variant sugges-   │
   │                             tions, GPT-5.4             │
   │                                                        │
   │  applySceneVariantSelections  apply stored recommenda- │
   │                               tions to shots later     │
   │                                                        │
   │  computeSceneBlocking       character blocking plan +  │
   │   (blocking.service,        180° screen projection,    │
   │    screenProjection.service) Claude Opus — ASYNC,      │
   │                             result via SCENE_UPDATED   │
   │                                                        │
   │  generateSketch             mannequin blocking sketches│
   │   (sketchGeneration.service) per shot — prerequisite-  │
   │                             gated, ASYNC, result via   │
   │                             SHOT_UPDATED (image gen    │
   │                             runs as AssetGenJob batch) │
   │                                                        │
   │  getShotSketchContext       read-only preview: report +│
   │                             exact generation inputs    │
   └────────────────────────────────────────────────────────┘
        │  surgical per-shot writes (patchShotsFields) +
        │  scene.preflight.{assetLinking,variantSelection,blocking}
        ▼
      MongoDB

1. Asset image analysis (automatic)

Every character/prop/environment image added to a project asset (variant or top-level) triggers a fire-and-forget VLM analysis. Hooks live in workbench.service.ts on all three image-add paths: user upload / select-from-library (uploadAssetImage), async generation completion (addAssetGenJobAssetImage), and synchronous generation (generateAssetImage). Analysis failures are logged, never thrown — an image add must not fail because analysis did.

  • Characters (characterAnalysis.service.ts): prompt-ready description, physical profile (incl. estimatedHeightCm, age-anchored, bounded 30–400cm), clothing, accessories, ordered mustPreserve features, script aliases, bio traits not visible in the reference.
  • Props (propAnalysis.service.ts): description + mustPreserve + script aliases. Deliberately minimal; no real-world scale (invented props have no reliable size prior).
  • Environments (environmentAnalysis.service.ts): two passes — (1) spatial zones, key landmarks, variant profile (viewpoint / time of day / weather / camera orientation), located objects with normalized bounding boxes; (2) same image + bboxes re-sent to map script aliases onto objects. Zero-size bboxes are filtered.

Results are stored on the image itself: ImageMedia.imageAnalysis — a single union field (CharacterAnalysisResult | EnvironmentAnalysisResult | PropAnalysisResult), mirroring Asset.properties. Writes are atomic per image (ProjectsModel.setImageAnalysis, $elemMatch + arrayFilters) so concurrent analyses can't clobber each other. Script context comes from scriptContext.ts (combineScriptsForAnalysis, whole-script drops at a 150k-char cap).

Backfill for images that predate the triggers: scripts/preflight/backfillAssetAnalysis.ts (see scripts/preflight/README.md). Analyzes each variant's selected-or-latest image plus the asset's top-level selected-or-latest image; idempotent (skips analyzed images unless --force); --batchSize for concurrency.

2. NER asset linking (on demand)

POST /workbench/linkSceneAssets {storyId, sceneId}

One GPT-5-mini call per scene extracts entities from shot text (description, acting, frame instructions) and resolves them to known story assets — handling misspellings, nicknames, and category references ("the phone" → prop "Mobile Phone"). Rules baked into the prompt: environment fixed-objects are not props; characters only link if physically present ("would an animator need to draw them?"); existing links are never removed, only added; an existing shot.environment is never overridden. LLM-returned links are validated against the actual asset pools so hallucinated assetIds can't land on shots. shot.characters is kept a superset of shot.charactersVisible.

Shots that already have final images are skipped (their links are settled). Persistence is a surgical per-shot patch of only the link fields on only the changed shots.

Every run is recorded on scene.preflight.assetLinking ({computedAt, changesCount, inputsFingerprint}) — including zero-change runs, so "ran and found nothing" is distinguishable from "never ran". Surfaced in the sketch prerequisite report as assetLinking: ok|missing|stale — informational only, never gating (hand-linked scenes are equally valid).

3. Environment variant selection (on demand, record-then-apply)

POST /workbench/selectSceneVariants {storyId, sceneId, apply?, includeMissingVariants?}

Per environment used in the scene, an LLM picks the best plate per shot from the analyzed candidates: every variant whose selected/latest image has imageAnalysis, plus the asset's analyzed top-level image as the "Default" candidate (recommendedVariantId: null — applying it clears shot.environment.variantId, and image resolution falls back to the asset's own images). Recommendations are returned for all shots, each with currentVariantId, so "your current plate isn't the best fit" is visible even when nothing is applied.

A second LLM call per environment (detectMissingVariants, gated by includeMissingVariants, default on) suggests new variants to create when no existing plate serves some shots — label + guidance on what image to upload + which shots need it.

Every run is recorded on scene.preflight.variantSelection ({computedAt, selections, missingVariants}). Applying is separate:

  • apply: true on the select call applies immediately, or
  • POST /workbench/applySceneVariantSelections {storyId, sceneId, shotIds?} applies the stored recommendations later (optionally a reviewed subset).

Apply semantics: writes only shot.environment.variantId (set, or $unset for default — labels are display-only snapshots, never applied); shots with sketches or images are never touched (changing the plate under a rendered shot silently invalidates it); no-op recommendations are skipped; applied selections get an appliedAt stamp in the stored state.

4. Character blocking + screen projection (on demand, async)

POST /workbench/computeSceneBlocking {storyId, sceneId, force?} → 202 {status: 'computing'}

Groups the scene's shots into physical setups (contiguous shots where characters occupy the same space) and anchors every character to concrete environment geometry from the stored plate analyses — camera-independent, object-relative language only ("at the far end of the table, closest to the door"). Claude Opus via Bedrock, with the distinct plates (variant images, or the environment's default top-level plate for shots without a variantId) plus up to 10 existing shot sketches/images attached (downloaded and compressed via imagePreprocessing.ts) so the plan never contradicts existing visuals. Beats/filler strategy provide setup-boundary hints when confirmed.

The call validates synchronously (edit permission, active shots, characters, no run already in progress), marks scene.preflight.blocking = {status: 'computing'}, and returns 202. The async body then:

  1. Computes an initial plan, or updates an existing one preserving setupIds (so downstream references stay valid). force: true recomputes from scratch.
  2. Runs screen projection (180° rule) per setup with ≥2 characters: GPT-5-mini maps each character's anchor-object bbox x-position to screen-left/center/right, stored as setup.screenProjection.
  3. Reviews every shot's plate against the blocking (reviewPlatesAfterBlocking, one LLM call per shot, with all candidate plate images attached): variant selection ran before blocking, so this is the first point where the model can see where the characters actually ended up. Each shot comes back with decision: keep | update plus noPlateAdequate — "the plate I picked is still the wrong one, a new plate is needed". An update is applied straight to shot.environment.variantId for shots with no sketch/image yet; on a re-trigger an edited shot only gets shot.environment.suggestedVariant for the user to accept. Every verdict is stored on scene.preflight.plateReview.shots[] (plateAdequate, the review's reason, and the plate to create), so a UI can warn on a shot that ended up on the least-bad plate. noPlateAdequate shots are also grouped per blocking setup into plateReview.missingVariants — the plates the user should create. The per-shot neededPlateLabel prefers step 3's own variantSelection.missingVariants entry for that shot (neededPlateSource: 'requested'): those labels are LLM-written per plate and short enough for a sentence — "overhead table tableau". The setup-grouped list is only a fallback ('derived'), because its label is assembled in code as zoneLabel — setupLabel, so it names the place and the action but never the camera, and four shots in one setup all get the same string. A shot the review approved never carries a plate request, even if step 3 asked for one before blocking. Note: a shot that already has a sketch or image is not reviewed (one LLM call per shot), so it has no entry. Absence is not a pass.
  4. Persists scene.preflight.blocking = {status: 'ready'|'failed', plan?, error?} and publishes SCENE_UPDATED (JaduSpine, project channel) with the fresh scene.

Blocking is record-only — nothing on shots to apply; sketch generation consumes the plan later.

5. Sketch generation (on demand, async, prerequisite-gated)

POST /workbench/generateSketch {storyId, sceneId, shotId, count?, allowMissingPrerequisites?}

Generates blocking-review sketches for one shot: characters drawn as plain colored mannequins (featureless, one fixed flat color per character, assigned in code and stable scene-wide so a later T2I prompt can say "the green mannequin is Stan"). Two phases per run: (1) per candidate, a GPT-5.4 prompt-writer call (sees the plate + character refs, decides a crop region by shot size, translates room-space blocking into screen-space placement) — kept outside the job so a prompt failure costs nothing; then (2) all candidates run as one AssetGenJob batch (createJobBatch, config Preflight Sketch Gen, candidates alternating GPT Image 2.5 Flare / Sunburst (i2i-gpt-image-2.5-flare / i2i-gpt-image-2.5-sunburst), 16:9 2560x1440, quality high) — charged, tracked, auto-refunded on failure, results B2-uploaded by the job pipeline, jobs carry storyLinks + jobMetaData.isSketchGeneration. Consumes everything upstream: the plate (variant or default), environment analysis (named objects + anchor bboxes), the blocking setup (positions, setup continuity arc, transitions), screen projection (180° lock), character heights (median across analyzed images → relative-height guidance), and prop refs. Candidate count scales with physically-present characters (photo-only characters are never drawn), and a shot with none of them still generates — a plate-and-props composition check at the low candidate count. Each sketch appends to shot.sketches[] (GENERATED_SKETCH) with imageParams deliberately minimal: assetGenJobId (the job is the provenance record — prompt/config/status live there), promptReasoning (exists only in the prompt-writer response, so this is its one persistent home), and generationRefs (exactly which env/character/prop variant images + cropped-or-full plate fed it).

Also logs a generate_shot_sketch project activity — the only preflight flow that does: the upstream flows (NER, variant selection, blocking) are pipeline preprocessing, not user-facing deliverables, and are intentionally unlogged for now.

Read-only preview: POST /workbench/getShotSketchContext {storyId, sceneId, shotId} returns the prerequisite report, the exact shotContext the prompt LLM would receive, the resolved plate/character/prop refs with mannequin colors, this shot's recorded variant-selection recommendation, and matching missing-variant suggestions. No LLM, no writes — powers the pre-generation review UI.

Prerequisite gate. Every call first builds a per-shot report (SketchPrerequisiteReport) covering plate, environment analysis, variant selection, blocking, screen projection, per-character analyses — each ok/missing/ stale/computing — plus two informational statuses that never gate: assetLinking (NER is one way to link, not a requirement) and propAnalysis (sketch generation doesn't read it — props reach the prompt as name + reference URL):

  • Default: fail closed. Any non-ok entry → 422 with the report; nothing generated.
  • allowMissingPrerequisites: true: degrade loudly. Generation proceeds without the missing artifacts, and the report is returned in the 202 AND attached to the final SHOT_UPDATED payload — degradation is never silent.
  • Hard floor in both modes: no plate image → no generation.

Characters-expected gate. A shot with no linked character is NOT skipped — an empty charactersVisible is ambiguous, because nobody types that field: the shot breakdown writes the prose and asset linking attaches the assetIds from it. So empty means either a genuinely characterless shot (a ferry on open water, a treeline) or a linking miss, and counting refs cannot tell them apart. The prompt-writer call — which already receives the shot description — is asked directly, returning charactersExpected + charactersExpectedReason (the question is only added to the prompt when nothing is linked). false → generate the characterless sketch. true → throw CHARACTERS_EXPECTED_BUT_NOT_LINKED (422) naming the reason, so the fix is "re-run asset linking", not a peopleless sketch of a shot that needs people. Only an explicit true gates: the fallback-prompt path leaves the field undefined. planSketch returns both fields so the pre-generation review can show it before an image job is spent.

The gate also runs on the precomputedPlan path. It has to: the generate modal loads a plan at step 3 and forwards its prompt on every generation, so a gate scoped to the no-plan branch would never fire in the product. When nothing is linked, the prompt writer runs even with a plan present — purely for the verdict. The plan still wins on prompt and crop, and the plate gate stays off that path, because the user saw the plate in the preview and chose to generate.

A body-part insert (a hand, a boot, bare skin) answers false even though the description names whose hand it is — there is no whole figure to stage, so a mannequin would say nothing and the sketch is a detail composition instead. Live-checked on 14 real characterless shot descriptions: 13/14 as expected; the one miss ("doors slam shut behind the procession") answered false, i.e. it generates rather than blocks — the safe direction.

Staleness is fingerprint-based (preflightFingerprints.ts): asset linking, variant selection, and blocking each store an inputsFingerprint — a hash of only the inputs they consumed (shot text fields, character links, plate identities + analyzedAts, asset pools). At sketch time the hash is recomputed and compared: mismatch → stale. Unconsumed writes (favorites, dialogue, our own variant-apply) cannot invalidate; a missing fingerprint (pre-fingerprint states) conservatively reads as stale.

Storage map

What Where Written by
Image analysis projects.{characters,environments,props}[].images[].imageAnalysis and ...variants[].images[].imageAnalysis ProjectsModel.setImageAnalysis
Shot asset links storyvideos.scenes[].shots[].{charactersVisible,characters,props,environment} StoryVideosModel.patchShotsFields
Shot plate choice storyvideos.scenes[].shots[].environment.variantId StoryVideosModel.patchShotsFields
Selection record storyvideos.scenes[].preflight.variantSelection StoryVideosModel.setSceneVariantSelectionState
Blocking state storyvideos.scenes[].preflight.blocking StoryVideosModel.setSceneBlockingState
Plate review verdicts storyvideos.scenes[].preflight.plateReview StoryVideosModel.setScenePlateReviewState
NER run record storyvideos.scenes[].preflight.assetLinking StoryVideosModel.setSceneAssetLinkingState
Generated sketches storyvideos.scenes[].shots[].sketches[] (GENERATED_SKETCH) StoryVideosModel.addShotImage
Sketch image jobs assetGenJobs (config Preflight Sketch Gen, jobMetaData.isSketchGeneration) AssetGenJobService.createJobBatch

All shot writes go through patchShotsFields — per-shot, per-field arrayFilters patches, never a whole-array replace. This matters because these flows are slow (LLM latency): a whole-array write would clobber concurrent shot edits with stale data.

File map

File Role
scenePreflight.service.ts Orchestrator: permissions, flow, persistence. No WorkbenchService dependency — a future pipeline can call it directly.
characterAnalysis.service.ts / propAnalysis.service.ts / environmentAnalysis.service.ts Pure VLM analysis + persist helpers
assetLinking.service.ts Pure NER: story+scene in, updated shots + changes out
variantSelection.service.ts Pure selection: recommendations + missing variants + blocking-driven swap suggestions out
blocking.service.ts Pure blocking: setups + character anchors out (Claude Opus, plates + shot visuals attached)
screenProjection.service.ts Pure 180° projection: screen-left/right per character per setup
sketchGeneration.service.ts Sketch gen: prerequisite report, context preview, mannequin candidates (image gen via AssetGenJob batch)
preflightFingerprints.ts Input fingerprints (staleness detection for asset linking / variant selection / blocking)
../../assetGenModelConfigs/workbenchV2/sketchGenConfig.ts AssetGenJob config for sketch image candidates
imagePreprocessing.ts Download + compress images for vision payloads (sharp)
scriptContext.ts Script combining for prompts (150k-char cap)
scenePreflight.types.ts All result/state types
../../promptTemplates/scenePreflight/ All prompts (repo convention: prompts never inline in services)

LLM analytics events: SCENE_PREFLIGHT_{ENV,CHAR,PROP}_ANALYSIS, SCENE_PREFLIGHT_NER, SCENE_PREFLIGHT_VARIANT_SELECTION, SCENE_PREFLIGHT_BLOCKING, SCENE_PREFLIGHT_SCREEN_PROJECTION, SCENE_PREFLIGHT_SKETCH_PROMPT — every call is logged with model, input/output, and token usage via AnalyticsLogger.logLLMRequest, so model provenance is queryable there (and deliberately not duplicated on stored results).


Not yet ported (from the PR #657 prototype)

Held back deliberately; each needs its own design pass:

  • T2I prompt consumption of image analysis — the prototype injected CHARACTER/PROP IDENTITY NOTES (the mustPreserve lists) into shot-image-generation prompts (promptGenerator.template.ts + analysis fields threaded through shotImageGeneration.service.ts). Today imageAnalysis is produced and stored but nothing reads it yet — the analyses exist so downstream phases land without a data-population step.
  • Automatic orchestration — the prototype ran NER + variant selection + blocking on every shot mutation (refreshScenePreflight hooked into upsert/delete/reorder). We expose the flows as explicit APIs instead. The orchestrator would slot into scenePreflight.service.ts.
  • HITL sketch overrides (PR #662) — the read-only preview (getShotSketchContext) is SHIPPED; still missing are the ephemeral per-generation variant overrides (swaps only, never persisted). prepareSketchContext is already structured for it.
  • T2I from a chosen sketch (PR #662) — generateFromSketch(sketchImageUrl) override so image generation can use a specific preflight sketch instead of the selected one.
  • Character scaling v1 (PR #662 backlog doc) — numeric environment anchoring (door as a height ruler). Hard blocker documented there: locatedObjects.region coords are normalized to the UNCROPPED plate while sketches use a cropped one.
  • Missing-variant → placeholder creation — the PR let the scene guru create placeholder variants from missingVariants suggestions. We return suggestions in the API response and store them on scene.preflight, but nothing creates variants yet.