Scene Preflight¶
Preflight prepares scenes for sketch generation by analyzing assets and shots before anything is rendered: what does each character/prop/environment look like, which assets appear in which shot, and which environment plate should back each shot.
This module is being brought into production incrementally from the original preflight prototype (PR #657). This doc describes what is live today and what is not yet ported. When you add a phase, update this doc.
What exists today¶
IMAGE ADDED (upload / select from library / generation)
│
▼
workbench.service.ts — triggerAnalysisForAssetImage (fire-and-forget)
│
▼
┌────────────────────────────────────────────────────────┐
│ VLM ANALYSIS (one per image, stored on the image) │
│ characterAnalysis.service.ts — GPT-5.4 vision │
│ propAnalysis.service.ts — GPT-5.4 vision │
│ environmentAnalysis.service.ts — GPT-5.4 vision, 2 │
│ passes (objects+bboxes, then script-alias mapping) │
└────────────────────────────────────────────────────────┘
│ writes ImageMedia.imageAnalysis (atomic per-image, projects collection)
▼
ON DEMAND (separate APIs, user- or tool-invoked)
│
▼
┌────────────────────────────────────────────────────────┐
│ scenePreflight.service.ts — orchestrator │
│ │
│ linkSceneAssets NER: shot text → asset │
│ (assetLinking.service) links, GPT-5-mini │
│ │
│ selectSceneVariants best env plate per shot + │
│ (variantSelection.service) missing-variant sugges- │
│ tions, GPT-5.4 │
│ │
│ applySceneVariantSelections apply stored recommenda- │
│ tions to shots later │
│ │
│ computeSceneBlocking character blocking plan + │
│ (blocking.service, 180° screen projection, │
│ screenProjection.service) Claude Opus — ASYNC, │
│ result via SCENE_UPDATED │
│ │
│ generateSketch mannequin blocking sketches│
│ (sketchGeneration.service) per shot — prerequisite- │
│ gated, ASYNC, result via │
│ SHOT_UPDATED (image gen │
│ runs as AssetGenJob batch) │
│ │
│ getShotSketchContext read-only preview: report +│
│ exact generation inputs │
└────────────────────────────────────────────────────────┘
│ surgical per-shot writes (patchShotsFields) +
│ scene.preflight.{assetLinking,variantSelection,blocking}
▼
MongoDB
1. Asset image analysis (automatic)¶
Every character/prop/environment image added to a project asset (variant or
top-level) triggers a fire-and-forget VLM analysis. Hooks live in
workbench.service.ts on all three image-add paths: user upload / select-from-library
(uploadAssetImage), async generation completion (addAssetGenJobAssetImage), and
synchronous generation (generateAssetImage). Analysis failures are logged, never
thrown — an image add must not fail because analysis did.
- Characters (
characterAnalysis.service.ts): prompt-ready description, physical profile (incl.estimatedHeightCm, age-anchored, bounded 30–400cm), clothing, accessories, orderedmustPreservefeatures, script aliases, bio traits not visible in the reference. - Props (
propAnalysis.service.ts): description +mustPreserve+ script aliases. Deliberately minimal; no real-world scale (invented props have no reliable size prior). - Environments (
environmentAnalysis.service.ts): two passes — (1) spatial zones, key landmarks, variant profile (viewpoint / time of day / weather / camera orientation), located objects with normalized bounding boxes; (2) same image + bboxes re-sent to map script aliases onto objects. Zero-size bboxes are filtered.
Results are stored on the image itself: ImageMedia.imageAnalysis — a single union
field (CharacterAnalysisResult | EnvironmentAnalysisResult | PropAnalysisResult),
mirroring Asset.properties. Writes are atomic per image
(ProjectsModel.setImageAnalysis, $elemMatch + arrayFilters) so concurrent
analyses can't clobber each other. Script context comes from scriptContext.ts
(combineScriptsForAnalysis, whole-script drops at a 150k-char cap).
Backfill for images that predate the triggers:
scripts/preflight/backfillAssetAnalysis.ts (see scripts/preflight/README.md).
Analyzes each variant's selected-or-latest image plus the asset's top-level
selected-or-latest image; idempotent (skips analyzed images unless --force);
--batchSize for concurrency.
2. NER asset linking (on demand)¶
POST /workbench/linkSceneAssets {storyId, sceneId}
One GPT-5-mini call per scene extracts entities from shot text (description, acting,
frame instructions) and resolves them to known story assets — handling misspellings,
nicknames, and category references ("the phone" → prop "Mobile Phone"). Rules baked
into the prompt: environment fixed-objects are not props; characters only link if
physically present ("would an animator need to draw them?"); existing links are never
removed, only added; an existing shot.environment is never overridden. LLM-returned
links are validated against the actual asset pools so hallucinated assetIds can't land
on shots. shot.characters is kept a superset of shot.charactersVisible.
Shots that already have final images are skipped (their links are settled). Persistence is a surgical per-shot patch of only the link fields on only the changed shots.
Every run is recorded on scene.preflight.assetLinking
({computedAt, changesCount, inputsFingerprint}) — including zero-change runs, so
"ran and found nothing" is distinguishable from "never ran". Surfaced in the sketch
prerequisite report as assetLinking: ok|missing|stale — informational only, never
gating (hand-linked scenes are equally valid).
3. Environment variant selection (on demand, record-then-apply)¶
POST /workbench/selectSceneVariants {storyId, sceneId, apply?, includeMissingVariants?}
Per environment used in the scene, an LLM picks the best plate per shot from the
analyzed candidates: every variant whose selected/latest image has imageAnalysis,
plus the asset's analyzed top-level image as the "Default" candidate
(recommendedVariantId: null — applying it clears shot.environment.variantId, and
image resolution falls back to the asset's own images). Recommendations are returned
for all shots, each with currentVariantId, so "your current plate isn't the best
fit" is visible even when nothing is applied.
A second LLM call per environment (detectMissingVariants, gated by
includeMissingVariants, default on) suggests new variants to create when no
existing plate serves some shots — label + guidance on what image to upload + which
shots need it.
Every run is recorded on scene.preflight.variantSelection
({computedAt, selections, missingVariants}). Applying is separate:
apply: trueon the select call applies immediately, orPOST /workbench/applySceneVariantSelections{storyId, sceneId, shotIds?}applies the stored recommendations later (optionally a reviewed subset).
Apply semantics: writes only shot.environment.variantId (set, or $unset for
default — labels are display-only snapshots, never applied); shots with sketches or
images are never touched (changing the plate under a rendered shot silently invalidates
it); no-op recommendations are skipped; applied selections get an appliedAt stamp in
the stored state.
4. Character blocking + screen projection (on demand, async)¶
POST /workbench/computeSceneBlocking {storyId, sceneId, force?} → 202 {status: 'computing'}
Groups the scene's shots into physical setups (contiguous shots where characters
occupy the same space) and anchors every character to concrete environment geometry
from the stored plate analyses — camera-independent, object-relative language only
("at the far end of the table, closest to the door"). Claude Opus via Bedrock, with the
distinct plates (variant images, or the environment's default top-level plate for shots
without a variantId) plus up to 10 existing shot sketches/images attached (downloaded
and compressed via imagePreprocessing.ts) so the plan never contradicts existing
visuals. Beats/filler strategy provide setup-boundary hints when confirmed.
The call validates synchronously (edit permission, active shots, characters, no run
already in progress), marks scene.preflight.blocking = {status: 'computing'}, and
returns 202. The async body then:
- Computes an initial plan, or updates an existing one preserving
setupIds (so downstream references stay valid).force: truerecomputes from scratch. - Runs screen projection (180° rule) per setup with ≥2 characters: GPT-5-mini maps
each character's anchor-object bbox x-position to screen-left/center/right, stored as
setup.screenProjection. - Reviews every shot's plate against the blocking (
reviewPlatesAfterBlocking, one LLM call per shot, with all candidate plate images attached): variant selection ran before blocking, so this is the first point where the model can see where the characters actually ended up. Each shot comes back withdecision: keep | updateplusnoPlateAdequate— "the plate I picked is still the wrong one, a new plate is needed". Anupdateis applied straight toshot.environment.variantIdfor shots with no sketch/image yet; on a re-trigger an edited shot only getsshot.environment.suggestedVariantfor the user to accept. Every verdict is stored onscene.preflight.plateReview.shots[](plateAdequate, the review'sreason, and the plate to create), so a UI can warn on a shot that ended up on the least-bad plate.noPlateAdequateshots are also grouped per blocking setup intoplateReview.missingVariants— the plates the user should create. The per-shotneededPlateLabelprefers step 3's ownvariantSelection.missingVariantsentry for that shot (neededPlateSource: 'requested'): those labels are LLM-written per plate and short enough for a sentence — "overhead table tableau". The setup-grouped list is only a fallback ('derived'), because its label is assembled in code aszoneLabel — setupLabel, so it names the place and the action but never the camera, and four shots in one setup all get the same string. A shot the review approved never carries a plate request, even if step 3 asked for one before blocking. Note: a shot that already has a sketch or image is not reviewed (one LLM call per shot), so it has no entry. Absence is not a pass. - Persists
scene.preflight.blocking = {status: 'ready'|'failed', plan?, error?}and publishesSCENE_UPDATED(JaduSpine, project channel) with the fresh scene.
Blocking is record-only — nothing on shots to apply; sketch generation consumes the plan later.
5. Sketch generation (on demand, async, prerequisite-gated)¶
POST /workbench/generateSketch {storyId, sceneId, shotId, count?, allowMissingPrerequisites?}
Generates blocking-review sketches for one shot: characters drawn as plain colored
mannequins (featureless, one fixed flat color per character, assigned in code and
stable scene-wide so a later T2I prompt can say "the green mannequin is Stan"). Two
phases per run: (1) per candidate, a GPT-5.4 prompt-writer call (sees the plate +
character refs, decides a crop region by shot size, translates room-space blocking into
screen-space placement) — kept outside the job so a prompt failure costs nothing; then
(2) all candidates run as one AssetGenJob batch (createJobBatch, config
Preflight Sketch Gen, candidates alternating GPT Image 2.5 Flare / Sunburst
(i2i-gpt-image-2.5-flare / i2i-gpt-image-2.5-sunburst), 16:9 2560x1440, quality high) — charged,
tracked, auto-refunded on failure, results B2-uploaded by the job pipeline, jobs carry
storyLinks + jobMetaData.isSketchGeneration. Consumes everything upstream: the plate
(variant or default), environment analysis (named objects + anchor bboxes), the blocking
setup (positions, setup continuity arc, transitions), screen projection (180° lock),
character heights (median across analyzed images → relative-height guidance), and
prop refs. Candidate count scales with physically-present characters (photo-only
characters are never drawn), and a shot with none of them still generates — a
plate-and-props composition check at the low candidate count. Each sketch appends to shot.sketches[]
(GENERATED_SKETCH) with imageParams deliberately minimal: assetGenJobId (the job is
the provenance record — prompt/config/status live there), promptReasoning (exists only
in the prompt-writer response, so this is its one persistent home), and generationRefs
(exactly which env/character/prop variant images + cropped-or-full plate fed it).
Also logs a generate_shot_sketch project activity — the only preflight flow that does:
the upstream flows (NER, variant selection, blocking) are pipeline preprocessing, not
user-facing deliverables, and are intentionally unlogged for now.
Read-only preview: POST /workbench/getShotSketchContext {storyId, sceneId, shotId}
returns the prerequisite report, the exact shotContext the prompt LLM would receive,
the resolved plate/character/prop refs with mannequin colors, this shot's recorded
variant-selection recommendation, and matching missing-variant suggestions. No LLM, no
writes — powers the pre-generation review UI.
Prerequisite gate. Every call first builds a per-shot report
(SketchPrerequisiteReport) covering plate, environment analysis, variant selection,
blocking, screen projection, per-character analyses — each ok/missing/
stale/computing — plus two informational statuses that never gate: assetLinking
(NER is one way to link, not a requirement) and propAnalysis (sketch generation
doesn't read it — props reach the prompt as name + reference URL):
- Default: fail closed. Any non-ok entry → 422 with the report; nothing generated.
allowMissingPrerequisites: true: degrade loudly. Generation proceeds without the missing artifacts, and the report is returned in the 202 AND attached to the finalSHOT_UPDATEDpayload — degradation is never silent.- Hard floor in both modes: no plate image → no generation.
Characters-expected gate. A shot with no linked character is NOT skipped — an empty
charactersVisible is ambiguous, because nobody types that field: the shot breakdown
writes the prose and asset linking attaches the assetIds from it. So empty means either a
genuinely characterless shot (a ferry on open water, a treeline) or a linking miss, and
counting refs cannot tell them apart. The prompt-writer call — which already receives the
shot description — is asked directly, returning charactersExpected +
charactersExpectedReason (the question is only added to the prompt when nothing is
linked). false → generate the characterless sketch. true → throw
CHARACTERS_EXPECTED_BUT_NOT_LINKED (422) naming the reason, so the fix is "re-run asset
linking", not a peopleless sketch of a shot that needs people. Only an explicit true
gates: the fallback-prompt path leaves the field undefined. planSketch returns both
fields so the pre-generation review can show it before an image job is spent.
The gate also runs on the precomputedPlan path. It has to: the generate modal loads a plan
at step 3 and forwards its prompt on every generation, so a gate scoped to the no-plan branch
would never fire in the product. When nothing is linked, the prompt writer runs even with a
plan present — purely for the verdict. The plan still wins on prompt and crop, and the plate
gate stays off that path, because the user saw the plate in the preview and chose to generate.
A body-part insert (a hand, a boot, bare skin) answers false even though the description
names whose hand it is — there is no whole figure to stage, so a mannequin would say nothing
and the sketch is a detail composition instead. Live-checked on 14 real characterless shot
descriptions: 13/14 as expected; the one miss ("doors slam shut behind the procession")
answered false, i.e. it generates rather than blocks — the safe direction.
Staleness is fingerprint-based (preflightFingerprints.ts): asset linking, variant
selection, and blocking each store an inputsFingerprint — a hash of only the inputs
they consumed (shot text fields, character links, plate identities + analyzedAts,
asset pools). At sketch time the hash is recomputed and compared: mismatch → stale.
Unconsumed writes (favorites, dialogue, our own variant-apply) cannot invalidate; a
missing fingerprint (pre-fingerprint states) conservatively reads as stale.
Storage map¶
| What | Where | Written by |
|---|---|---|
| Image analysis | projects.{characters,environments,props}[].images[].imageAnalysis and ...variants[].images[].imageAnalysis |
ProjectsModel.setImageAnalysis |
| Shot asset links | storyvideos.scenes[].shots[].{charactersVisible,characters,props,environment} |
StoryVideosModel.patchShotsFields |
| Shot plate choice | storyvideos.scenes[].shots[].environment.variantId |
StoryVideosModel.patchShotsFields |
| Selection record | storyvideos.scenes[].preflight.variantSelection |
StoryVideosModel.setSceneVariantSelectionState |
| Blocking state | storyvideos.scenes[].preflight.blocking |
StoryVideosModel.setSceneBlockingState |
| Plate review verdicts | storyvideos.scenes[].preflight.plateReview |
StoryVideosModel.setScenePlateReviewState |
| NER run record | storyvideos.scenes[].preflight.assetLinking |
StoryVideosModel.setSceneAssetLinkingState |
| Generated sketches | storyvideos.scenes[].shots[].sketches[] (GENERATED_SKETCH) |
StoryVideosModel.addShotImage |
| Sketch image jobs | assetGenJobs (config Preflight Sketch Gen, jobMetaData.isSketchGeneration) |
AssetGenJobService.createJobBatch |
All shot writes go through patchShotsFields — per-shot, per-field arrayFilters
patches, never a whole-array replace. This matters because these flows are slow (LLM
latency): a whole-array write would clobber concurrent shot edits with stale data.
File map¶
| File | Role |
|---|---|
scenePreflight.service.ts |
Orchestrator: permissions, flow, persistence. No WorkbenchService dependency — a future pipeline can call it directly. |
characterAnalysis.service.ts / propAnalysis.service.ts / environmentAnalysis.service.ts |
Pure VLM analysis + persist helpers |
assetLinking.service.ts |
Pure NER: story+scene in, updated shots + changes out |
variantSelection.service.ts |
Pure selection: recommendations + missing variants + blocking-driven swap suggestions out |
blocking.service.ts |
Pure blocking: setups + character anchors out (Claude Opus, plates + shot visuals attached) |
screenProjection.service.ts |
Pure 180° projection: screen-left/right per character per setup |
sketchGeneration.service.ts |
Sketch gen: prerequisite report, context preview, mannequin candidates (image gen via AssetGenJob batch) |
preflightFingerprints.ts |
Input fingerprints (staleness detection for asset linking / variant selection / blocking) |
../../assetGenModelConfigs/workbenchV2/sketchGenConfig.ts |
AssetGenJob config for sketch image candidates |
imagePreprocessing.ts |
Download + compress images for vision payloads (sharp) |
scriptContext.ts |
Script combining for prompts (150k-char cap) |
scenePreflight.types.ts |
All result/state types |
../../promptTemplates/scenePreflight/ |
All prompts (repo convention: prompts never inline in services) |
LLM analytics events: SCENE_PREFLIGHT_{ENV,CHAR,PROP}_ANALYSIS, SCENE_PREFLIGHT_NER,
SCENE_PREFLIGHT_VARIANT_SELECTION, SCENE_PREFLIGHT_BLOCKING,
SCENE_PREFLIGHT_SCREEN_PROJECTION, SCENE_PREFLIGHT_SKETCH_PROMPT — every call is logged
with model, input/output, and token usage via AnalyticsLogger.logLLMRequest, so model
provenance is queryable there (and deliberately not duplicated on stored results).
Not yet ported (from the PR #657 prototype)¶
Held back deliberately; each needs its own design pass:
- T2I prompt consumption of image analysis — the prototype injected CHARACTER/PROP
IDENTITY NOTES (the
mustPreservelists) into shot-image-generation prompts (promptGenerator.template.ts+ analysis fields threaded throughshotImageGeneration.service.ts). TodayimageAnalysisis produced and stored but nothing reads it yet — the analyses exist so downstream phases land without a data-population step. - Automatic orchestration — the prototype ran NER + variant selection + blocking on
every shot mutation (
refreshScenePreflighthooked into upsert/delete/reorder). We expose the flows as explicit APIs instead. The orchestrator would slot intoscenePreflight.service.ts. - HITL sketch overrides (PR #662) — the read-only preview (
getShotSketchContext) is SHIPPED; still missing are the ephemeral per-generation variant overrides (swaps only, never persisted).prepareSketchContextis already structured for it. - T2I from a chosen sketch (PR #662) —
generateFromSketch(sketchImageUrl)override so image generation can use a specific preflight sketch instead of the selected one. - Character scaling v1 (PR #662 backlog doc) — numeric environment anchoring (door as
a height ruler). Hard blocker documented there:
locatedObjects.regioncoords are normalized to the UNCROPPED plate while sketches use a cropped one. - Missing-variant → placeholder creation — the PR let the scene guru create
placeholder variants from
missingVariantssuggestions. We return suggestions in the API response and store them onscene.preflight, but nothing creates variants yet.