Skip to content

I2V pre-generation conflict system — implementation handover

Audience: a coding agent (Claude Code / Cursor / Codex) implementing or porting this feature, working with Yashasewi Singh.

Reference implementation: Scenarix/studio-backend PR #847, branch f-i2vconflict, base staging. 64 files, +21,785 / −44. Validator version at handover: 0.8.0. Frontend reference (naive, illustrative only): Scenarix/studio-frontend PR #814, branch f-i2vconflict. Jeet owns the real spec.

Everything below was read off that branch. Where a number is measured, the script that measured it is named. Where something is unbuilt or unverified, it says so. Do not treat any statement here as a substitute for reading the file it cites.


0. Ground rules for the agent

  1. Read the file headers. Every module in src/workbench/sequenceValidation/ opens with a long comment explaining why it exists and which real product failure produced it. Those headers are the design record. Do not delete or "tidy" them.
  2. Do not change a prompt without a live run. Unit tests only assert what text was sent to the model. They cannot prove what the model answered. Every prompt change in this feature that was believed correct from tests alone was wrong at least once. Use a throwaway script under scripts/i2v/_*.ts that calls the real analyzer against a real story (dotenv/config, read-only), and keep it untracked.
  3. No regex in new code in this repo (project rule). normalizeCameraValue in shotFieldFix.ts uses two .replace() calls with regex literals and predates the rule — do not copy the pattern.
  4. Never write to Mongo from an exploratory script. Reads only.
  5. A model's output is untrusted input. Every stage passes model output through deterministic filters. That is the core architectural decision of this feature: the same input must produce the same decision, rather than "did the model happen to read that line".

1. What the feature does

The I2V pre-generation screen sends a sequence — an ordered selection of 1–5 shots — to a video model (Seedance 2.x). Each shot carries text (description, acting, three frame-instruction fields, camera dropdowns) and one picture (a rendered frame, or a sketch). Only one shot in the selection is required to be a finished image.

When the user presses Generate Prompt, we check the selection before spending the generation, and show conflicts the user must answer. Three checks:

Stage id Name Model Question
references Reference checks none (deterministic) What will the request actually send as pictures?
shot_analysis Analysis A 1 vision call per shot Does this shot's picture match this shot's own text?
cross_shot Analysis B 1 vision call, all pictures attached Do the selected shots hold together as one continuous clip?

Analysis A does not do text-vs-text inside one shot. The single-shot T2I checker (src/workbench/preGenValidation/) already does that and already shows the user a verdict; re-deciding it means the user answers the same question twice.

Analysis B is the part that exists nowhere else in the product.


2. The call graph

POST /workbench/validateSequenceForGeneration
  └─ WorkbenchValidator.validateSequenceForGenerationRequest   (zod, .min(1).max(5))
  └─ SharedHelpers.enableHeartBeat(res)                        ← REQUIRED, see §5
  └─ SequenceValidationService.validateSequence
       ├─ readSequenceInputs()  → scene, shotsById, timeline, assets, identity,
       │                          selectionKey, inputHash, stored
       ├─ if (!force && isStillTheAnswer(stored, inputHash))  → return stored   ← cache hit, 0 calls
       ├─ runReferenceChecks(timeline, assets)                 ← deterministic, ~0ms
       ├─ await Promise.all([                                  ← ✱ PARALLEL, NEVER CHAINED ✱
       │      ShotAnalysisStage.run(...)   // Analysis A, fan-out one call per shot
       │      CrossShotStage.run(...)      // Analysis B, one call with every picture
       │    ])
       ├─ adjudicateLookCandidates(...)    ← B's verdict decides a look conflict
       ├─ mergeStages([references, A, B])  → dedupe(), betterEvidenced()
       ├─ applySettledFieldCarryOver(...)  ← DELETES  (sketch-chat decisions)
       ├─ applyT2ICarryOver(...)           ← STAMPS   (T2I dismissals)
       ├─ carryForwardUserActions(...)     ← RE-APPLIES the user's own prior answers
       ├─ deriveStatus() + canGenerate()
       └─ upsertSequenceValidation()       ← failure is logged, never thrown

sequenceValidation.service.ts:200-222 is the Promise.all. Do not chain A into B. Feeding A's facts into B was measured at roughly 70 seconds; the budget is 15–25s. Both stages are independently useful and B reads the pictures itself.

Each stage's .catch() degrades instead of failing the request: A returns {conflicts: [], timing: {ran: false}}; B returns an infrastructure conflict (§8).


3. File map — src/workbench/sequenceValidation/

23 files. Line counts at handover, so you can tell a stub from the real thing.

File Lines What it owns
sequenceValidation.service.ts 1269 Orchestrator, cache, gate, dismiss/resolve, all writes
stages/crossShotStage.ts 1056 Analysis B: prompt assembly, the 8 filters, option building
referenceChecks.ts 635 Deterministic group F checks, look candidates, variant swap effects
stages/shotAnalysisStage.ts 595 Analysis A fan-out, its filters, persistSketchTypes
sequenceIssueCatalog.ts 529 The 27 codes: group, shortMessage, default severity, gates
sequenceStageSchemas.ts 487 OpenAI strict JSON schemas for both stages
sequenceTimeline.ts 397 Shots → timeline / seams / presence grid
sequenceValidation.types.ts 355 Conflict, effects union, editable/camera field lists
shotAnalysis/thinShotAnalyzer.ts 334 The default Analysis A implementation (replaceable)
shotAnalysis/shotAnalysis.port.ts 223 The seam: ShotVisualFacts, setShotVisualAnalyzer
identityRoster.ts 211 Who is who in the picture (mannequin colour + sketch labels)
shotFieldFix.ts 168 Validates a model-proposed field fix → an applicable effect
promptNotes.ts 124 Resolved conflicts → directives for the video prompt
shotAnalysis/priorKnowledge.ts 120 What is already known about a shot; suppressed codes
settledFields.ts 113 Sketch-chat settled fields carry-over
buildSequenceConflict.ts 81 The ONLY constructor of a conflict; severity clamp
t2iCarryOver.ts 74 T2I dismissal carry-over
sequenceAssets.ts 92 assetId+variantId → label, reference image URL, "does this version exist"
sequenceSelection.ts 39 sequenceSelectionKey, SEQUENCE_VALIDATIONS_KEPT

Plus sequenceFlow/ (the prompt generator this feeds into): sequenceFlow.utils.ts 397, buildSequenceUserPrompt.ts 387, sequenceFlow.types.ts 186, sequenceFlow.service.ts 159.

Prompts live in src/promptTemplates/sequenceValidation/: - shotVisualFacts.prompt.ts (237 lines) — Analysis A - crossShotContinuity.prompt.ts (344 lines, 17 sections) — Analysis B

Tests (11 files, all in the PR):

File Lines
tests/workbench/sequenceValidationService.test.ts 1151
tests/workbench/crossShotStage.test.ts 1048
tests/workbench/shotAnalysisStage.test.ts 884
tests/workbench/sequenceTimeline.test.ts 316
tests/workbench/sequenceValidationEndpoints.test.ts 254
tests/workbench/promptNotes.test.ts 253
tests/workbench/identityRoster.test.ts 220
tests/workbench/sequenceValidationStorage.test.ts 181
tests/workbench/t2iCarryOver.test.ts 162
tests/workbench/settledFields.test.ts 143
tests/workbench/uvSequence.fixture.ts 119 (fixture)

Also modified outside the folder: src/agenticGuru/sketchAnalysisGuru.ts (+53, writes settled fields), src/shared/cameraInstructions.ts (+57), src/shared/sketchType.ts (new, 64), src/shared/media.types.ts (SketchType enum), src/scenePreflight/characterMannequinColors.ts (new, 45), src/scenePreflight/sketchGeneration.service.ts (+36, labels its own sketches).


4. Documentation already in the branch

docs/i2v/ holds it all. The one to read before touching the codes is docs/i2v/SEQUENCE_CONFLICT_TAXONOMY.md (698 lines) — what each code means, which group it belongs to, and why it exists. docs/i2v/evidence/ holds the measurement write-ups the taxonomy rests on. The brief, the architecture proposal and the superseded taxonomy drafts are planning material and are not in the repo.


5. The four endpoints

All in src/workbench/workbench.router.ts, all POST, all behind the existing workbench auth:

workbenchRouter.post('/validateSequenceForGeneration', SharedHelpers.handleIfError(WorkbenchController.validateSequenceForGeneration));
workbenchRouter.post('/peekSequenceValidation',        SharedHelpers.handleIfError(WorkbenchController.peekSequenceValidation));
workbenchRouter.post('/dismissSequenceConflict',       SharedHelpers.handleIfError(WorkbenchController.dismissSequenceConflict));
workbenchRouter.post('/resolveSequenceConflict',       SharedHelpers.handleIfError(WorkbenchController.resolveSequenceConflict));

Request schemas (workbench.validator.ts, zod):

validateSequenceForGenerationRequest  { storyId, sceneId, shotIds: z.array(z.string()).min(1).max(5), masterShotId?, force? }
validatePeekSequenceValidationRequest { storyId, sceneId, shotIds: .min(1).max(5), masterShotId? }
validateDismissSequenceConflictRequest{ storyId, sceneId, shotIds, issueId, reason: z.nativeEnum(DismissConflictReason), note? }
validateResolveSequenceConflictRequest{ storyId, sceneId, shotIds, issueId, optionId, note?, overrideValue?, overrideVariantId? }

overrideValue is the option's suggested wording after the user edited it in the panel — see §14. Same 20,000-character ceiling as the properties panel's own text fields, so this route is not a way around it.

overrideVariantId is the character version the user picked when none of the offered ones is right — see §14. z.string().trim().min(1).nullable().optional(): null is a real value and means the asset's own picture (the UI's "Default"), so absent ("apply the option as offered") and null ("use Default") are two different requests and must not be collapsed on the way in. Send the id alone; the shot writes are rebuilt from the live shots on the backend.

The 5-shot cap is real and deliberate — it matches validateGenerateSequencePromptRequest. A selection the generator would reject must not be checkable.

Responses (workbench.controller.ts):

Route Body Notes
validate { validation } enableHeartBeat(res) first — without it the proxy closes the socket on a 20–40s answer. traces stripped.
peek { validation } or { validation: null } null is a normal answer, not an error. Never runs a model, never writes. No heartbeat.
dismiss { validation }
resolve { validation, updatedShots } updatedShots is required by the client, see §12.

forSequencePanel() (controller line 32) strips description from every conflict on the way out. description is the diagnostic record (asset ids, compared fields, the model's justification) kept for Mongo, logs and evals. message is the only sentence a user reads. Wording rules for message live in src/promptTemplates/preGenValidation/conflictLanguage.prompt.ts: one line, plain enough for a child, "what seems wrong, then Consider ".

canGenerate is derived and never stored. The dismiss and resolve controllers recompute it via SequenceValidationService.canGenerate(validation.conflicts) so the panel does not reimplement the gate.


6. Storage

Path: storyvideos → scenes[] → sequenceValidations[] (src/models/storyVideos.model.ts, StoryVideoScene.sequenceValidations?: StoredSequenceValidation[], ~line 508).

Model methods: getSequenceValidation (~4144), upsertSequenceValidation (~4163).

Key = sequenceSelectionKey(shotIds) = first 16 hex chars of sha1(shotIds.join('|')) (sequenceSelection.ts).

Two things about that key:

  • Not keyed by sequenceId. There is no sequenceId at Generate Prompt time — it is minted later by generateSequenceVideoInitial (sequenceId || uuidv4()). StoredSequenceValidation.sequenceId exists but nothing fills it today.
  • ORDER IS PART OF THE KEY. 8→11→14 and 14→11→8 are different sequences, because half the checks are about what happens between one shot and the next. Never sort the selection.

upsertSequenceValidation behaviour, and this is the data-loss surface (§17): - existing key → $set of scenes.$[scene].sequenceValidations.$[validation] = {...validation, updatedAt: now} — whole-record overwrite - new key → $push with $sort: {updatedAt: 1}, $slice: -SEQUENCE_VALIDATIONS_KEPT - SEQUENCE_VALIDATIONS_KEPT = 8. The cap exists because the episode document is the largest in the database and 5-of-23 shots is a lot of combinations; the 16MB limit is the real constraint.

Stored shape (sequenceValidation.types.ts):

interface StoredSequenceValidation {
  selectionKey: string;        // primary key within the scene
  sequenceId?: string;         // never set today
  inputHash: string;
  status: 'pass' | 'pass_with_warnings' | 'blocked' | 'validation_failed';
  validatedAt: Date;
  validatorVersion: string;
  shotIds: string[];           // send order
  masterShotId?: string;       // needed to rebuild the timeline on restamp
  conflicts: SequenceConflict[];
  diagnostics?: { timings: SequenceStageTiming[]; dropped: Array<DroppedFinding & {stage, shotId?}> };
  createdAt?: Date; updatedAt?: Date;
}

diagnostics.dropped is persisted on purpose: without it, a clean result and a result whose only finding was gated away are indistinguishable in the data. traces (raw model output) is not persisted — several KB per run, and it is in the logs.


7. Caching — and the loop this prevents

computeInputHash(timeline, shots, validatorVersion, assets, identity)   // service:1042
  = sha256(JSON.stringify({ timeline, priorKnowledge, referenceImages, validatorVersion }))
  • priorKnowledge = buildShotPriorKnowledge(shot, timelineShot, identity) per shot. The roster is in the hash because running sketch analysis genuinely changes what both stages are told.
  • referenceImages = for every subject, ${assetId}:${variantId}=${assets.referenceImage(...)}. In the hash because the timeline only stores assetId:variantId — someone selecting a different image on the asset changes what the request sends without touching a shot.
  • Deliberately NOT the raw shot documents. A comment, a rename or a new video must not re-run a ~$0.25 / 20s check.
isStillTheAnswer(stored, inputHash)
  = !!stored && stored.inputHash === inputHash && stored.status !== 'validation_failed'

Bumping SEQUENCE_VALIDATOR_VERSION invalidates every stored record (it is inside the hash). The changelog for 0.1.0 → 0.8.0 is in the constant's comment; keep appending to it.

restampInputHash — read this before changing resolutions

service.ts:437. A resolution that rewrites a shot changes the timeline, which changes the input hash, which makes the stored record stale — so the next Generate Prompt re-runs the check, gets different conflicts (this model is not deterministic on identical input), and the user answers forever without reaching a prompt.

So after applying effects, the service rebuilds the timeline and re-stamps the hash against stored.validatorVersion, not the current constant — otherwise a version bump would be silently absorbed by our own write.

Invariant: two presses of Generate Prompt, always. One to get the questions, one to get the prompt.


8. The gate and the status

isOpen(conflict)     = !conflict.dismissed && !conflict.resolution && conflict.code !== INFRASTRUCTURE_CODE
canGenerate(list)    = list.every(c => !isOpen(c))        // public, used by the controller
deriveStatus(list)   → 'validation_failed' | 'blocked' | 'pass_with_warnings' | 'pass'
(service.ts:775-801.)

Every severity counts toward the gate, not just blockers. Disha's rule: the prompt is auto-generated only when there are no conflicts, or every conflict has been actioned. A minor conflict still has to be answered, because answering it is what changes the prompt.

The backend reports the gate; the frontend owns the button. Same contract as T2I. The backend does not refuse to generate.

The infrastructure conflict. SEQUENCE_VALIDATION_INFRASTRUCTURE_FAILED (service.ts:1063) is raised when a stage throws. It is: - excluded from isOpen, so it never gates; - the only non-dismissible code — isSequenceCodeDismissible(code) returns code !== 'SEQUENCE_VALIDATION_INFRASTRUCTURE_FAILED'; - status becomes validation_failed while canGenerate stays true.

Note this differs from T2I, where blockers are not dismissible. In I2V every conflict except the infrastructure one is dismissible, whatever its severity.


9. The conflict object

SequenceConflict extends Omit<ValidationConflict, 'stage'|'fixTarget'|'proposedFix'> — deliberate reuse so the review panel renders an I2V conflict with the same component as a T2I one.

Inherited from preGenValidation.types.ts:78: issueId, code, severity, shortMessage, message, description, suggestion?, fieldPaths, evidence, confidence, dismissible, dismissed?: {at, byUserId, reason, note?}.

New here:

stage: 'references' | 'shot_analysis' | 'cross_shot';
shotIds: string[];                        // every shot it is about, in send order
subject?: { type: 'character'|'prop'|'environment'|'light'|'effect'|'camera'|'reference'; name: string; assetId? };
resolutionOptions: SequenceResolutionOption[];   // 2–3, never undefined
resolution?: SequenceConflictResolution;         // mutually exclusive with `dismissed`
carriedOverFromT2I?: { shotId: string; issueId: string };

buildSequenceConflict() in buildSequenceConflict.ts is the only constructor. It fills issueId, the templated shortMessage, the catalog severity, dismissible, and guarantees resolutionOptions is an array. Never build a conflict literal.

Severity clamp: clampSeverity(reported, ceiling) — a model may LOWER a catalog default and never RAISE one. Severity is the gate; one over-confident call must not create a blocker.

DismissConflictReason = NOT_ACTUALLY_A_CONFLICT | MINOR_CAN_IGNORE | DONT_UNDERSTAND | OTHER.


10. The code catalog

sequenceIssueCatalog.ts, 27 codes, each with group, shortMessage, defaultSeverity, and gate flags (needsThreeShots). shortMessage is templated from the code, so a mislabelled code produces a banner one-liner that lies — hence filter ④ in §14.

Group Code Default severity
people PERSON_LEAVES_UNEXPLAINED blocker
people PERSON_POSE_CHANGES_UNEXPLAINED blocker
people PERSON_IN_TEXT_NOT_IN_CAST blocker
people PERSON_IN_PICTURE_NOT_IN_CAST blocker
people PERSON_LOOK_CHANGES blocker
people PERSON_APPEARS_UNEXPLAINED risky
people SAME_PERSON_LISTED_TWICE risky
objects OBJECT_STATE_GOES_BACKWARDS blocker
objects OBJECT_APPEARS_IN_HAND risky
objects OBJECT_LOOK_CHANGES risky
objects OBJECT_NEVER_CHANGES minor
place_light PLACE_CHANGES_WITHOUT_TRAVEL blocker
place_light ENVIRONMENT_TAG_CONTRADICTS_FRAME_TEXT risky
place_light LIGHT_OR_TIME_CHANGES risky
place_light EFFECT_STATE_JUMPS risky
place_light SET_LAYOUT_CONTRADICTS risky
action_arc ACTION_GOES_BACKWARDS blocker
action_arc ACTION_NEVER_ADVANCES risky
action_arc EMOTION_ROUND_TRIP minor
screen_geometry SCREEN_SIDE_FLIPS risky
screen_geometry SUBJECT_SCALE_JUMPS risky
references NO_LOCATION_REFERENCE_SENT blocker
references SHOT_HAS_NO_VISUAL blocker
references SEQUENCE_VALIDATION_INFRASTRUCTURE_FAILED blocker
references SAME_OBJECT_SEVERAL_ASSETS risky
references REFERENCE_LABEL_IS_PLACEHOLDER risky
references CHOSEN_VARIANT_NOT_SENT risky — deliberately unimplemented

Analysis A does not use these codes. It emits seven shipped T2I codes (SHOT_ANALYSIS_CODES in shotVisualFacts.prompt.ts): SKETCH_POSE_ACTION_MISMATCH, SKETCH_POSITION_MISMATCH, SKETCH_ENTITY_COUNT_MISMATCH, SKETCH_CAMERA_MISMATCH, SKETCH_ANGLE_MISMATCH, SKETCH_PROP_STATE_MISMATCH, SKETCH_ENVIRONMENT_MISMATCH.

That reuse is load-bearing. The T2I carry-over matches on the code, so a new vocabulary for Analysis A would silently break "a T2I dismissal must suppress the same I2V conflict".

SKETCH_MOMENT_MISMATCH is excluded on purpose: the T2I checker decides it with more context (it sees the sketch chat's agreed scope), and re-deciding it would contradict a verdict the user has already seen.


11. Analysis A — shot_analysis

Implementation is behind a seam. shotAnalysis/shotAnalysis.port.ts defines ShotVisualAnalyzer / ShotVisualFacts and exposes setShotVisualAnalyzer(analyzer) / getShotVisualAnalyzer(). The shipped ThinShotVisualAnalyzer is an explicit stand-in meant to be replaced with one call. Keep the seam.

Model settings (thinShotAnalyzer.ts:63-73):

OpenAIService.generateTextWithCompletions({
  systemPrompt: SHOT_VISUAL_FACTS_SYSTEM_PROMPT,
  userPrompt: this.buildUserPrompt(request),
  images: [shot.visual.url],
  modelId: OpenAIModelId.GPT_5_4,
  reasoningEffort: 'low',
  responseFormat: SHOT_VISUAL_FACTS_SCHEMA,
  prompt_cache_key: 'i2v-shot-visual-facts-v2',
  timeoutMs: 120000,
})
id = 'thin-a2-vision-v1', recorded in SequenceStageTiming.analyzerId.

One call, one shot, one picture. Not one call holding five pictures (that reads them against each other, which is Analysis B's job). Shots run in parallel, so the stage's wall clock is roughly the slowest shot, not the sum.

Two jobs, in order of importance:

  1. READ the picture into comparable fields — screenSide (left/center/right in the frame), pose (short, physical, no mood words), bodyOrientation, heldObjects (in a hand only), objects[].state + holder, depicts (start/middle/end/unclear). These are what Analysis B compares across shots, so a vague answer makes the comparison useless and a wrong one makes it lie. An omitted field means "not known" everywhere downstream — the prompt tells the model to omit rather than guess.
  2. Report picture-vs-own-text disagreements using the seven codes.

Behaviours to preserve:

  • No picture → no call. Returns textOnlyFacts(shot): depicts: 'unclear', cast/props from the shot text, confidence: 0, source: 'text_only'. Deliberately does not fill fields from the shot text — the value of these facts to the next stage is that they came from the picture.
  • One shot failing must not fail the sequence. A thrown call is logged and returns textOnlyFacts. The missing visual is already reported deterministically as SHOT_HAS_NO_VISUAL, so this must not add a second conflict.
  • characters[].name must be a name from the cast list. Analysis B lines rows up by name; "the woman in the dark coat" cannot be matched to anything. If a figure cannot be pinned to a name, the prompt says leave it out.
  • Compare against the frame the picture depicts. A picture of the opening pose will always "disagree" with the closing frame's text. That distinction is most of the false positives in this check.
  • Never judge the drawing. Not style, not line quality, not that it is a sketch, not that it lacks colour. A sequence mixes sketches and renders on purpose.
  • Nothing here is ever a blocker — one picture read on its own is not enough evidence to stop a generation.

Filters in shotAnalysisStage.ts (all deterministic, all after the model): - crossesAPositionBucket (263) — a SKETCH_POSITION_MISMATCH must cross a left/center/right boundary. A character a quarter of the way across the frame is frame-left; that was a real false positive. - crossesTwoShotSizeRungs (327) — one rung of difference is not a conflict, two is. - doesNotReopenASettledField (388) — §13. - oneFindingPerCode (438) — three copies of one pose problem is three banners for one problem. - In thinShotAnalyzer.keepReportableFindings: drop codes already raised or dismissed in T2I (suppressedCodesFor), then filterFindingsByAgreedScope (reuses the shipped T2I helper) for dimensions the user agreed the sketch does not govern.

Every drop is recorded as a DroppedFinding with a stable reason key: dismissed_in_t2i, already_open_in_t2i, outside_agreed_sketch_scope, etc.


12. Analysis B — cross_shot

Model settings (crossShotStage.ts:236-247): OpenAIModelId.GPT_5_4, reasoningEffort: 'low', prompt_cache_key: 'i2v-cross-shot-continuity-v4', timeoutMs: 120000, CROSS_SHOT_SCHEMA.

It sees every picture. All visuals attached to one call, in send order, with a manifest saying which attached picture is which shot. It was text-only until 2026-09-18; the pictures cost ~2,800 input tokens each, and the prompt cache returns ~95% of input.

THE RULE THAT DECIDES EVERYTHING — the fridge rule

Seedance invents the missing moment. "Walking to the fridge" then "holding the milk" needs no shot of the hand on the handle. That is not a conflict.

So whySeedanceCannotBridge is a required schema output, and a finding that cannot fill it with anything better than "the states are different" is dropped in code (filter ⑥). The schema also asks for bridgeableDifferences — what the model noticed and chose not to report — so "applied the rule" is distinguishable from "never saw it" in the diagnostics.

The 8 filters (header of crossShotStage.ts, lines 37-44)

# Filter Why
① unknown code the enum is strict, so an outside code means a regression, not a finding
② needs 3+ shots a flicker or a freeze is invisible in a pair (NEEDS_THREE_SHOTS, read off the catalog)
③ camera gate group E off across a camera change — a new angle legitimately swaps left/right (SCREEN_GEOMETRY_CODES × seam.cameraChanged)
④ wrong dimension the banner one-liner comes from the CODE, so a mislabelled code lies
⑤ severity clamp may lower, never raise
⑥ no justification the fridge rule
⑦ wrong evidence wardrobe read off a blank storyboard mannequin is read off nothing
⑧ no identity a look conflict about a person the picture cannot point at is dropped (IDENTITY_GATED_CODES = {PERSON_LOOK_CHANGES})

KNOWN_FIELD_PATHS (line 107) is the whitelist for fieldPaths — a path the UI cannot resolve is a card with a dead link.

The prompt has 17 sections; the ones you must not weaken: THE RULE THAT DECIDES EVERYTHING, THE PICTURES (the charter that forbids sketch-vs-render findings), WHO IS WHO IN THE PICTURE, THE COVERAGE TEST, THE LOCATION LADDER, NEVER FLAG THESE, SCREEN GEOMETRY IS GATED.

De-duplication against the reference checks is NOT this stage's job. PERSON_LOOK_CHANGES can be produced by both; collapsing them belongs in the orchestrator, which sees both lists (adjudicateLookCandidates, service.ts:647). The deterministic stage can only see that two files go into the request; Analysis B has looked at both. B's verdict decides.

The stage is told what the reference checks found (referenceConflicts is passed in) precisely so it does not re-report them. That costs no latency because those checks already ran.


13. Reference checks — references, deterministic

referenceChecks.ts, runReferenceChecks(timeline, assets). No model, ~0ms, the cheapest and most certain conflicts in the system. They answer "what will the request actually send":

  • no location reference at all (NO_LOCATION_REFERENCE_SENT)
  • a selected shot with no image and no sketch (SHOT_HAS_NO_VISUAL)
  • the same person entering under two asset ids (SAME_PERSON_LISTED_TWICE)
  • two different pictures of one person across shots (the free half of PERSON_LOOK_CHANGES, via collectLookCandidates)
  • one thing under several assets (SAME_OBJECT_SEVERAL_ASSETS)
  • a placeholder reference label (REFERENCE_LABEL_IS_PLACEHOLDER, isPlaceholderLabel)

The reference payload is keyed assetId:variantId, so two shots picking different variants of one asset send two pictures. For a character that is one person with two faces (a conflict). For a location it is the intended workflow (two angles on one room) — do not flag it.

variantSwapEffects(uses, assetId, targetVariantId) builds the set_shot_asset_variant effects that make several shots agree on one variant.


14. Resolution — the effects union

type SequenceResolutionEffect =
  | { kind: 'edit_shot_field';       shotId; path: SequenceEditableField; oldValue: string|null; newValue: string }
  | { kind: 'set_shot_camera_field'; shotId; field: SequenceCameraField;  value: string; oldValue: string|null }
  | { kind: 'set_shot_asset_variant';shotId; field: 'charactersVisible'; assetId; variantId: string|null; oldVariantId: string|null }
  | { kind: 'prompt_note';           shotId?; note: string }

Closed union on purpose: an option that cannot name its effect must not exist.

SEQUENCE_EDITABLE_FIELDS = ['description','actingInstructions','initialFrameInstructions',
  'middleFrameInstructions','finalFrameInstructions','shotSizeDetails','cameraAngleDetails','cameraMovementDetails']
SEQUENCE_CAMERA_FIELDS   = ['shotSize','cameraAngle','cameraMovement']

That boundary is not a matter of taste: a conflict may only propose a change the user could have made by hand in the shot properties panel. The editable list is the free-text half of the panel's own contract; the camera list is the dropdown half minus the reference lists (which set_shot_asset_variant owns).

ORDER MATTERS — the loudest feedback on v1

The check said "The picture seems much closer than an extreme wide shot. Consider updating Shot size" and offered one button — "Keep the picture and tell the video model to follow it" — which changed no field. Three of five findings on that sequence read the same way.

Rule: when a metadata field is one honest end of the disagreement, the option that MOVES THAT FIELD comes first, and prompt_note comes last. A prompt_note leaves the shot saying one thing and the video saying another, forever. It is the answer only when the picture is the sole thing that could change ("the drawing shows two carts and the shot needs one" — only a new drawing fixes that).

shotFieldFix.ts — the validator both model stages share

buildFieldFixEffect(shot, proposal) → {effect} or {rejected, detail}. A model proposal is untrusted in three ways, and each rejection is a button deliberately not rendered:

  1. field not in the allowed list → 'fix names a field we may not set/edit'
  2. dropdown value not a member of its enum → 'fix value is not one of the dropdown options'
  3. new value equals what is already there → 'fix would change nothing' — a button that visibly does nothing is worse than no button: the user clicks it, watches the conflict stay, and stops believing the screen.

CAMERA_FIELD_VALUES = Object.values(ShotSize | CameraAngleV2 | CameraMovement) from src/shared/cameraInstructions.ts. It is exported so the prompt can list the legal values to the model — one source, so what the model is shown and what the validator enforces cannot drift.

normalizeCameraValue accepts "medium_wide", "Medium Wide", "MEDIUM-WIDE" → MEDIUM_WIDE. Anything still not a member is rejected, never coerced.

oldValue is read from the SHOT DOCUMENT, never from the timeline. The timeline trims text (text() in sequenceTimeline.ts) and the staleness guard compares raw strings — a field with a trailing newline would 409 for no reason the user could discover.

Applying — and the 409

applyShotEdit (497), applyCameraFieldChange (534), applyVariantSwap (584). Each refuses with InvalidOperationError(..., 409) when the field no longer holds oldValue:

"<field> on this shot has changed since the check ran. Run the check again to get up-to-date options."

(throws at lines 509, 559, 596.)

These writes deliberately do NOT run the shot cascade. shotSize, cameraAngle, cameraMovement, description and actingInstructions are all SHOT_CASCADE_TRIGGER_FIELDS: changed through the properties panel they fire an LLM (shotCascadeUpdate.service.ts) that may rewrite up to 21 fields, including the three frame-instruction texts. Firing that from a conflict resolution would rewrite the text the check just read and the user just approved, and restampInputHash would then stamp the record against words nobody agreed to.

The honest cost: after setting shotSize, shotSizeDetails may still describe the old framing. That is why the *Details fields are editable.

The edited suggestion — overrideValue

The suggested wording is a starting point, not a verdict: the panel shows it in an editable box, and Update sends back whatever the user left in it (Isa, 2026-09-28). resolveConflict takes that as overrideValue and withEditedSuggestion (service.ts, immediately above applyShotEdit) swaps it into newValue on the option's single edit_shot_field effect.

Four things are deliberate:

  1. It is applied here, not through the ordinary shot route. Same reason as the paragraph above — description and actingInstructions are cascade triggers, so saving the edit the normal way fires the cascade LLM, which can rewrite the very text the check read, and the next fix button then 409s.
  2. Only when the option rewrites exactly one text field. Zero → 'This option does not rewrite any text…'; two or more → '…more than one field, so an edited wording cannot be matched to one of them'; both 400, and nothing is written. The frontend already only offers the box in the one-effect case (editableTextEffect in sequenceReviewPopover.tsx, PR #845), so a 400 here means the two sides have drifted.
  3. oldValue is untouched, so the 409 staleness guard still compares against what the check actually read. An edit does not become a way to overwrite somebody else's change.
  4. The option keeps the AI's wording; the resolution records the user's. conflict.resolutionOptions[].effects is the offer, conflict.resolution.effects is the answer — comparing the two is how anyone tells later that the user rewrote it. The prompt generator reads the resolution, so the video hears the user's sentence, not ours.

An edit on "Other" is a 400 ('"Other" has no suggested wording to edit…'): there is no suggestion there, the note is the resolution. Refusing rather than ignoring, because silently dropping typed text is the exact bug this path exists to prevent.

The chosen version — overrideVariantId

Same idea as the edited suggestion, for the one option that repoints a picture instead of rewriting text.

buildVariantSwapOptions (referenceChecks.ts) offers at most MAX_VARIANT_OPTIONS = 3 versions, and only versions already pinned on the selected shots — the ones the shots disagree about. A character with six saved looks therefore has looks the panel cannot reach, and "he should wear the coat in all three shots" is sometimes one of those. Without this the only way out is "Other", which writes a sentence and leaves both pictures in the request.

resolveConflict takes overrideVariantId and reaimVariantSwap (service.ts, immediately below withEditedSuggestion) rebuilds the effects.

Five things are deliberate:

  1. The client sends the id and nothing else; the writes are rebuilt here. Two things have to be true and neither can be trusted from a request body: the version must exist on the asset (assets.hasVariant, new in sequenceAssets.ts — an id that is not there resolves to the asset's own picture via SharedHelpers.getSelectedAssetImage, so a typo would quietly send a look nobody chose), and every effect's oldVariantId must be what is on the shot now, because applyVariantSwap only moves the entries that still hold it.
  2. Every use in the row moves, including shots the offered option left alone. The option skipped any shot already on the version it offered; a different version means those have to move too. This is the whole reason the effects are not copied off the option.
  3. A shot that lists the character twice has both copies moved. Two entries is how the product represents two copies of one person in one frame (PEA-2982); moving one still sends two looks.
  4. Refused, with nothing written, in four cases (all 400): a version not on the character ('…not one of this character's saved versions any more…'), an option that swaps no character version ('This option does not swap a character version…' — including an option that also rewrites text, whose wording was written for the version it offered), every shot in the row already on that version ('…nothing to change'), and "Other" ('"Other" changes no reference…').
  5. The option keeps the version the check offered; the resolution records the one applied. Same shape as an edited suggestion — comparing resolutionOptions[].effects with resolution.effects is how anyone tells later that the user picked their own.

One difference from every other write on this route, stated plainly: this path does not 409 on a stale click, it absorbs it. Because the effects are rebuilt from the live shots, a version a collaborator pinned a second ago is moved too. The offered options keep their refusal, because there the user agreed to a specific before → after that no longer holds; here the user named a version, and that is still the version they want.

The "Other" path

OTHER_OPTION_ID = 'other' (service.ts:125) — reserved, never offered as an option. When resolveConflict finds no matching option:

const written = note?.trim();
if (!written) throw new SharedErrors.InvalidOperationError('Choosing "Other" needs a sentence saying what should happen instead', 400);
conflict.resolution = { optionId: OTHER_OPTION_ID, at: new Date(), byUserId: userId, effects: [], note: written };
delete conflict.dismissed;
return { validation: await this.persistAction(...), updatedShots: [] };

So: nothing is edited, any prior dismissal is removed, the conflict counts as answered (the gate opens), the input hash is not re-stamped (so the cache survives and the user pays for no re-run), and the sentence reaches the prompt generator verbatim (§16). It goes to the prompt generator LLM, not straight to Seedance.

Rationale for not rewriting a field: we cannot tell which field the user meant, and rewriting the wrong one is not recoverable.


15. The three carry-overs

Applied after the merge, in this exact order (service.ts:236-238). All three enforced in code, not asked of a model — "same input, same decision" cannot depend on the model having read paragraph 11 correctly on this run.

1. applySettledFieldCarryOver (settledFields.ts) — DELETES. A field the user settled in the sketch conversation. sketchAnalysisGuru.ts now writes:

interface SettledShotField { field: string; value: string; source: 'sketch_analysis'; at: Date; sessionId?: string }
stored as StoryVideoShot.settledFields?.

value is what makes it falsifiable: liveSettledFields(shot) only counts a record whose field still holds that value. If somebody later changed the field, the conversation was overruled and the check must fire again. Drop reason: SETTLED_FIELD_DROP_REASON = 'field_settled_in_sketch_chat'.

This exists because of a real failure: the sketch chat asked which wins between the drawing's straight-on framing and the shot's low angle, the user picked the drawing, and I2V then offered LOW_ANGLE back as a fix.

It runs first because it deletes: a conflict the user already closed should never reach the T2I pass to be stamped, nor the stored record to be re-shown.

2. applyT2ICarryOver (t2iCarryOver.ts) — STAMPS, never deletes. When the same conflict was already marked "Not a conflict" in the T2I flow for one of these shots, the conflict is kept but pre-dismissed, and stamped with carriedOverFromT2I: {shotId, issueId} plus the original dismissed record. The user sees it was already handled instead of it vanishing silently.

3. carryForwardUserActions (service.ts:870) — RE-APPLIES the user's own prior answers. Keyed code::sorted(shotIds). Four guards: - never onto a non-dismissible conflict - never over an action taken this run - a resolution is only re-applied if answerStillHolds(resolution, shotsById) (service.ts:929) — i.e. the effects' oldValues still describe the shots - otherwise the conflict comes back open

This is what stops editing shot 4 from silently re-opening the decision made about shot 2.


16. Identity roster

identityRoster.ts. Why it exists: on a real sequence (Sentinels Ep3 scene 8, "Knocking Signals", two runs on 2026-09-22) Analysis B reported a wardrobe conflict about "the woman in the dark coat" and pointed at two different women in shots 1 and 2. A false blocker, and the answer was already on the shot — nobody passed it down.

Built from two sources:

  1. Mannequin colour (src/scenePreflight/characterMannequinColors.ts). Preflight paints each character as a featureless mannequin in a fixed colour, the same colour for that person in every shot of the scene. Claimed only for imageSource === GENERATED_SKETCH — a user-drawn sketch has no mannequins and no palette, so claiming "Shuki is the green one" about it invents evidence.
  2. sketchGuru's character mappings — the latest COMPLETED analysis, even of a different picture, because identity survives a redraw.

API: buildSequenceIdentityRoster(scene, timeline, shotsById) → Map<shotId, ShotIdentityEntry[]>; pictureHandleFor(name, cast, entries); describeIdentityRoster(entries) → the prompt block. The same block, the same words, goes into both stages (thinShotAnalyzer.ts:300).

Entries can be not_visible or background_or_extra — those are people the model must not describe or report about.


17. sketchType — the sketch classifier

Requested by Davide. One word recorded per drawing, telling the video prompt generator how the drawing was made, because that is what says how much of the picture can be trusted.

SketchType enum in src/shared/media.types.ts, plus sketchType?: SketchType on ImageMedia.

THE TAXONOMY IS ABOUT PROVENANCE, NOT APPEARANCE. The names come from Davide and they mislead:

Value What it actually means
cocoB-real-env Preflight drew it — Generate Sketch: location pick → variant pick → blocking → character placement. Blank flat-coloured mannequins, no hair, no clothing, no face.
cocoB-fake Built outside the platform by the India creative team: script into a chat model, then shot by shot with stand-in figures that carry the character (hair, collar, cuffs, belt, a cross for a face).
hand-drawn A thumbnail somebody drew — on paper or on a computer. The tool is irrelevant; a clean digital drawing is still this.

"real-env" does not mean the background is photographic. All three kinds routinely show the shot's actual location drawn in the show's own art style. The v0.7.0 prompt asked about the background and therefore labelled every stylised show cocoB-fake. Caught in the product on a preflight sketch that had been downloaded and re-uploaded: same picture, two different answers.

The 0.8.0 prompt is a three-step decision on the figures, with an explicit NOT-evidence list (limb contours, jaw contour, the outline, fingers, being posed) because the live model cited "interior arm/body contour lines" on blank mannequins and answered cocoB-fake.

Three rungs — each runs only if the one above had no answer

src/shared/sketchType.ts (64 lines): selectedSketch, storedSketchType, deterministicSketchType, isSketchType.

  1. storedSketchType(sketch) — already on the sketch. Wins forever. A drawing does not change after it is made. Consequence: a wrong stored answer is permanent until the field is deleted, and bumping the validator version does not re-label it.
  2. deterministicSketchType(sketch) — requires imageParams.generationRefs.environmentRefs.length and imageSource === GENERATED_SKETCH. Free and exact: preflight knows it drew it (sketchGeneration.service.ts labels its own output).
  3. Analysis A — the model, with the guards below.

Guards in thinShotAnalyzer.readSketchType (144): - ONLY FOR A DRAWING — shot.visual.kind !== 'sketch' → {}. A finished frame's "type" must never be written; readVisual prefers the render, and the value travels into the video prompt where nothing downstream can catch it. - NO EVIDENCE, NO cocoB-fake — the model must fill sketchTypeEvidence naming a concrete feature on a named figure ("red figure: ponytail, cuffs, belt"). Empty → the label is dropped, not downgraded, logged as a warn, and the next check asks again. "Same picture, same answer" cannot depend on the model having obeyed.

A TypeScript interface field alone ships nothing. OpenAI structured outputs are strict: true + additionalProperties: false with an explicit required array, so an unlisted key is REJECTED, not ignored. Optionality is a null union in both type and enum.

Link Where, for sketchType
① the prompt asks shotVisualFacts.prompt.ts §sketchType + §sketchTypeEvidence
② the strict schema allows SHOT_VISUAL_FACTS_SCHEMA in sequenceStageSchemas.ts
③ the analyzer maps parsed.x thinShotAnalyzer.readSketchType
④ the caller stores it shotAnalysisStage.persistSketchTypes (157-182)

persistSketchTypes skips non-sketch and already-stored shots, then writes deterministicSketchType(sketch) ?? (isSketchType(modelAnswer) ? modelAnswer : undefined) — the generation record beats the model — via StoryVideosModel.updateSketchType(storyId, sceneId, shotId, sketchImageUrl, sketchType) (model lines 5903-5913, matched by sketches.imageURL with arrayFilters). A failed write is a logger.warn only: it must never cost the user their check.

Read path is separate: toSketchType(shot) in sequenceFlow/sequenceFlow.utils.ts = storedSketchType(sketch) ?? deterministicSketchType(sketch) off selectedSketch(shot). It reads the selected sketch even on a rendered shot, because every selected shot's sketch goes into the storyboard collage.

Status at handover: sketchType is carried into Seedance2ShotBreakdown and no prompt template branches on it yet. So a wrong label costs nothing today. That changes the moment someone reads it.


18. The cascade into the video prompt

The last link, and it is invisible if you forget it. Without it a resolution click looks like it worked and changes nothing.

conflicts[].resolution.effects (prompt_note)  +  resolution.note
        ↓  promptNotes.collectResolvedConflictDirectives()
SequenceResolvedConflictDirective { shot?: string; decision: string }
        ↓  sequenceFlow.utils.buildInputBundleFromStory()
Seedance2InputBundle.resolved_conflicts?: SequenceResolvedConflictDirective[]
        ↓  buildSequenceUserPrompt.ts:378-381
JSON.stringify(...) into the user prompt  +  getResolvedConflictsAddendum()

promptNotes.ts (124 lines): - one directive per prompt_note effect, plus one for resolution.note - the note directive is labelled with a shot only when conflict.shotIds.length === 1 — otherwise it would attribute a sequence-wide decision to one shot - dedupes, sorts sequence-wide directives first then by play order - loadSequenceValidationForSelection fails open and logs loudly: a Mongo problem must not block prompt generation

Two rules:

  • The addendum is injected only when bundle.resolved_conflicts?.length. It is necessary because §3 of generateSequenceSeedancePrompt.template.ts already decides storyboard-vs-text on its own and would silently overrule the user. The addendum says a user decision outranks the normal priority rules.
  • edit_shot_field and set_shot_camera_field produce NO directive. The change is already in shot_breakdown; a directive would say it twice.
  • Shot labels must use the same numbering as the prompt headline — Shot ${shotLabel(shot, index + 1)}. A different scheme made the model renumber every shot in its output.

buildInputBundleFromStory loads the stored validation in the same Promise.all as its other reads and appends resolved_conflicts only when non-empty.


19. Timeline shapes

sequenceTimeline.ts, buildSequenceTimeline(shots, options).

TimelineShot: shotId, position (1-based send order), sceneShotNumber? (the number on the card), sceneShotIndex? (diagnostic only), isMasterShot, description, acting, frames{start,middle,end}, startState/endState {text, source: FrameTextSource}, characters/props/environment? as TimelineSubject {name, assetId?, variantId?}, camera{movement,movementDetails,angle,angleDetails,size,sizeDetails}, beatEmotion?, timing?, effect?, visual: TimelineVisual, dialogueLines.

Shot numbering, and this is a trap. userShotNumber(shot) = shot.sceneShotNumber ?? shot.position. - position is send order and is only what the model is told. - sceneShotNumber is the card number — 1-based position among the scene's live (non-deleted) shots, supplied by the caller via buildSceneShotNumbers(sceneShots), because only the caller holds the whole scene. - shot.shotIndex is unusable. Measured 2026-09-23 over 8,385 live shots in 1,293 scenes: 78% are not a number, and of the scenes that do store it 151 are 1-based and 68 are 0-based. - Anything a user reads uses userShotNumber. A message built from position contradicts its own chip in the banner.

TimelineVisual: kind: 'image'|'sketch'|'none', url?, connectedTo?: 'start'|'middle'|'end' (sketch only), hasAnalysis. readVisual prefers the render over the sketch — a rendered shot hands Analysis A the frame, and sketchType therefore stays absent for that shot.

SequenceSeam (seam i joins shots[i] → shots[i+1]): charactersOnBothSides (the only ones a continuity check can be about), charactersOnlyBefore (→ PERSON_LEAVES_UNEXPLAINED), charactersOnlyAfter (→ PERSON_APPEARS_UNEXPLAINED), cameraChanged (precomputed because group E must not fire across a camera change), environmentChanged, shotsSkippedBetween.

PresenceRow: {subject, type, present: boolean[], flickers}. hasFlicker = present → absent → present, which is invisible in a pair and so lives on the timeline, not a seam.

SequenceTimeline.isSingleShot — a third of real "sequences" are one shot (20 of the 60 most recent jobs). Cross-shot checks must no-op silently on those, not report nothing-found.

subjectKey is assetId ?? name.toLowerCase() — charactersVisible titles are hand-typed and drift in casing.


20. Latency, cost, and two measured traps

Measured 2026-09-18, 3 runs per cell, current prompt, via scripts/i2v/_bShippedLatency.ts:

Shots Analysis B wall clock B output tokens Analysis A whole fan-out
2 24.8s 1,663 9.5s
3 44.2s 3,438 15.2s
4 30.6s 1,771 19.0s

Since A and B run in parallel, the user's wait ≈ Analysis B. Cost is ~$0.25/sequence and is not the constraint; the target window is 15–25s, 30–40s tolerated.

Trap 1 — four shots is faster than three. What this call costs is what it writes, at roughly 60–80 tokens/second. Shot count barely enters into it. Do not optimise shot count.

Trap 2 — DO NOT ask the model for less prose. A variant capping bridgeableDifferences ran 34–39% faster and found nothing in 4 of 6 runs, where the shipped prompt found a conflict in 5 of 6 (scripts/i2v/_bOutputTokenAB.ts). On this model the output tokens ARE the reasoning. Same reason reasoningEffort is 'low' and not 'none': 'none' produced 349 tokens and zero conflicts; 'low' produced 3,748 tokens and two blockers. If the check is too weak, turn reasoning effort UP ('medium'), never the prose down.

Trap 3 — not reproducible. Analysis B does not return the same answer twice on byte-identical input, and did not before vision either. Any test that asserts a specific finding from a live call is flaky by construction. This is also the reason restampInputHash exists (§7).

Vision is what made the charter in the prompt necessary, and the charter is what spends the time. Output fell 2,266 → 1,202 at 3 shots when the pictures were first bolted on, then the charter pushed it back above 2,266. That trade is honest and worth knowing before anyone "simplifies" the prompt.


21. Invariants — break any of these and the feature regresses

  1. A and B in one Promise.all. Chaining measured ~70s.
  2. restampInputHash after any resolution that rewrites a shot. Otherwise: infinite rounds of conflicts.
  3. canGenerate counts every open conflict, not just blockers.
  4. The infrastructure conflict never gates and is never dismissible; every other code is dismissible.
  5. whySeedanceCannotBridge required; unjustified findings dropped in code.
  6. Never flag sketch-vs-render style, roughness, missing wardrobe or missing faces.
  7. Analysis A emits the seven T2I codes — the carry-over matches on the code.
  8. All three carry-overs applied after the merge, settled-fields first (it deletes).
  9. oldValue from the shot document, never the timeline; stale write → 409.
  10. No shot cascade from a resolution write.
  11. A fix proposal is validated (allowed field / legal enum value / actually different) before an option is offered.
  12. Order of shotIds is part of the storage key — never sort.
  13. userShotNumber for anything a user reads; position only for the model.
  14. description and traces stripped by forSequencePanel(); message is the only user-facing sentence.
  15. peekSequenceValidation never calls a model and never writes.
  16. enableHeartBeat(res) on the validate route.
  17. Every write guarded — a Mongo failure costs a re-run, never the user's click.
  18. Severity may be lowered by a model, never raised.
  19. sketchType never recomputed once stored; never written for a rendered frame.
  20. A new LLM output field needs all four links (§17).

22. Data persistence — a known gap, and the fix to port

Confirmed, by reading the storage path. Per conflict, from Mongo today:

User action Reconstructible?
"Not a conflict" Yes — dismissed {at, byUserId, reason, note}
Clicked an option Yes, with before/after — resolution {optionId, at, byUserId, effects[], note?}, and every effect carries oldValue/newValue
Went back and edited the field by hand No — nothing ties the edit to the conflict

And every re-run $sets the whole record for that selectionKey; only the newest run survives, and a scene keeps at most 8 selections.

So the third case disappears twice: the manual edit changes the input hash → the next Generate Prompt re-runs → the conflict is no longer raised → the row is replaced with no trace it existed.

The loss is biased, which is the real problem. A dismissal changes no input, so it is carried forward and accumulates. A fix that worked erases its own evidence, because the fix removes the reason the conflict was raised. Any dismissal rate computed from this data is the share of what was left behind, not of what was ever raised.

projectActivities does not cover it: neither dismiss nor resolve writes an activity row (grepped both folders, zero hits). A manual shot edit writes EDIT_SHOT, but createStoryActivity carries only actor / project / scene / shot / job — no field name, no old value.

The pattern to port: PR #860 (f-t2idata, merge 3a3fec9f)

Not an ancestor of PR #847 — verified with git merge-base --is-ancestor. It solved exactly this for the T2I checker. Key commit 4230bf71 added src/models/preGenValidationRuns.model.ts (235 lines), +210 in preGenValidation.service.ts, +27 in types, +4 in selectedFrameAlignmentStage.ts, and two test files (runHistoryWrite.test.ts 166, runOutcomeClassifier.test.ts 263).

Its shape, to copy: - one immutable document per run, holding the conflict verbatim as the user received it, the shot's text fields at run time, and the findings the filters dropped - a back-patched outcome written by the following run: still_raised / dismissed / fix_applied / field_changed / vanished_unchanged - dismissals as separate event documents, so no conflict count double-counts them - nothing user-visible changes; every write guarded

That PR's message records the measured T2I damage: 409 validation records destroyed across 339 shots (748 writes, 169 shots written more than once).

Existing history cannot be backfilled — it is already gone. That is the argument for porting this before the creatives start, not after.


23. Repo conventions

  • Linter/formatter: biome, not eslint.
  • lefthook. pre-commit: biome-format-staged, biome-lint-staged, no-console-log. commit-msg: check-prompts-tested — requires PROMPTS_TESTED=true (or skip) in the commit message when python_eval/ or src/promptTemplates/ changed. pre-push: build (tsc).
  • tsconfig excludes tests — tsc will not catch a test referring to a deleted method. Only npm test will.
  • PRs target staging, never main. Branch off origin/staging.
  • Branch names use an f- / f_ prefix.
  • Never merge a PR and never toggle draft state without being asked.
  • Commits end with Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>; PR descriptions end with 🤖 Generated with [Claude Code](https://claude.com/claude-code).
  • Throwaway scripts live at scripts/i2v/_*.ts and stay untracked (out of the PR).
  • Mongo, read-only: mongosh --quiet "$(grep -m1 '^MONGODB_URI' .env | cut -d= -f2-)". DB renderboard, collection storyvideos (lowercase). Assets hydrate from projects at read time, they are not raw in storyvideos.
  • Frontend preview against a specific backend: <frontend-url>/?be=<backend-preview-url> (note /?be=, not /be?= — the wrong form fails silently). It is sessionStorage, so tab-scoped. curl cannot prove a route exists here, because auth runs before routing.

Running the tests:

npm test -- tests/workbench/sequenceValidationService.test.ts
npm test -- tests/workbench/crossShotStage.test.ts
npm test                     # the only thing that typechecks tests/
npm run build                # tsc, also the pre-push hook

Running a live check against real data (read-only, untracked, model spend is real):

npx ts-node scripts/i2v/_sketchTypeRules08.ts     # pattern: dotenv/config, Mongo read, direct analyzer call

Note that script calls ThinShotVisualAnalyzer directly and not ShotAnalysisStage.run, because the stage persists — and calling the analyzer directly also bypasses rung ①, which is the point when a wrong answer is already stored.


24. Deliberately NOT built — do not "fix" these

  • CHOSEN_VARIANT_NOT_SENT is in the catalog and intentionally unimplemented.
  • SKETCH_MOMENT_MISMATCH is excluded from Analysis A (§10).
  • Text-vs-text inside one shot — the T2I checker owns it.
  • No sequenceId on stored records; there is none at check time.
  • No backfill of old sketches' sketchType. Only new sketches from here on. Stored answers from older rule versions are left alone.
  • traces not persisted (KBs per run; they are in the logs).
  • No shot cascade from resolution writes (§14).
  • The backend does not enforce the gate — it reports it.
  • Stale Nicola references in the headers of thinShotAnalyzer.ts (lines 5-8), shotAnalysis.port.ts (3-8), shotAnalysisStage.ts (6) and t2iCarryOver.ts name a person who is no longer the plan. Harmless, unchanged, ignore them.

25. Frontend contract notes

Not the spec — Jeet owns that. These are the backend-imposed constraints:

  1. The frontend owns the Generate button. Read validation.canGenerate.
  2. Render message only. description is already stripped; do not surface shortMessage as the whole row (it is the same 26 titles for every shot).
  3. resolveSequenceConflict returns updatedShots and you must apply them. The dialog builds the reference payload from shot.charactersVisible itself, so without them it keeps sending the pictures the user just resolved away. This is a real difference from T2I, where the frontend applies proposedFix itself.
  4. Call peekSequenceValidation when the panel opens, not validateSequenceForGeneration. A miss on the check route is ~20s and real money for a panel the user only opened to look at. validation: null means "sit empty, wait for Generate Prompt".
  5. "Other" requires a note or you get a 400. Do not smuggle an edited suggestion through note — the note goes to the prompt generator as a directive, so the shot would end up saying one thing and the video prompt another. Send it as overrideValue. 5b. overrideValue is only accepted on an option with exactly one text effect (§14). SEQUENCE_PREGEN_CAPS.overrideValue in sequencePregen.types.ts can be flipped to true once this backend is on staging. 5c. overrideVariantId is only accepted on an option whose effects are all set_shot_asset_variant for ONE asset (§14) — in practice the keep-* options on PERSON_LOOK_CHANGES. Which character the picker is for is on the option itself: effects[0].assetId. The versions it lists are the character's own, from whatever the panel already has the asset from — we do not send a version list. null = Default, field omitted = apply the option as offered, and a version the character no longer has comes back as a 400 saying so. Needs its own cap flag in SEQUENCE_PREGEN_CAPS — this backend is on f-i2vconflict/PR #847, not yet on staging.
  6. A resolution can 409 ("… has changed since the check ran. Run the check again …"). Surface it as "re-run the check", not as a failure.
  7. resolutionOptions may be empty; suggestion then carries the prose.
  8. Conflicts span several shots — shotIds, not shotId — and the chip must show the card number (the frontend maps shotIds through the scene's live-shot order; conflictScopeLabel in sequenceConflictsAlert.tsx in PR #814 does this).
  9. Do not invent presentation. ISA (design) and Disha (product) own it.

26. First hour for a new agent on this code

  1. Read docs/i2v/SEQUENCE_CONFLICT_TAXONOMY.md.
  2. Read the file headers of sequenceValidation.service.ts, stages/crossShotStage.ts, shotFieldFix.ts, sequenceSelection.ts. They contain the reasoning, not just the what.
  3. Read both prompts end to end. They are the product.
  4. npm test -- tests/workbench/sequenceValidationService.test.ts to see the contract asserted.
  5. Run one live sequence in the product before changing anything.