Skip to content

Sequence conflict taxonomy

Status: FINAL — this is the taxonomy. The earlier drafts are superseded and are not in the repo, including the old C1…X1 categories draft; its three checks that this document was missing are folded in below, each marked from the categories pass. Last change, 2026-09-17: added PERSON_POSE_CHANGES_UNEXPLAINED (§3A) after a teammate's production example that all 22 previous codes would have missed, and §2.2 — the end/start seam, which is the comparison the rest of the checks are built on. Read it as: the complete list of what we warn the user about before a multi-shot Seedance generation, what we deliberately stay quiet about, and what is our bug rather than theirs. Grounded in: 40 production sequences · 127 shots · 87 transitions · two independent passes, zero overlap · Mongo read-only. Start with: §3 for the conflicts, §4 for what we refuse to flag, §7 for how a conflict turns into a one-click fix, §9 if you read the old C1…X1 draft and want to know where each row went.


What this is for

The user picks 2, 3, 4 or 5 shots and presses generate. We fuse them into one video. Before we spend the money, we look at the text and images we are about to send and ask one question:

Is there something wrong here that this user can go and fix?

   user picks N shots  ─────►  [ sanity layer ]  ─────►  Seedance 2.0 / 2.5
                                     │
                                     ├─ nothing to say          → generate
                                     ├─ minor / risky           → warn, let them generate anyway
                                     └─ blocker                 → warn hard, name the shot and the field

Everything in this document exists to fill that middle box. A conflict earns a place here only if the user can act on it. If they can't, it belongs in §6 (our bug, not theirs) or §4 (not a problem at all).


How this was built

  • Data: 40 production sequences, in two independent passes with zero overlap.
  • First 20, frozen 2026-09-16 — evidence/I2V_20_SEQUENCE_DISSECTION.md. Spread across five stories. 64 shots, 44 transitions.
  • Second 20, frozen 2026-09-17 — evidence/I2V_20_MORE_SEQUENCES.md. 63 shots, 43 transitions. 19 of these 20 are from one show, The Countdown: deeper, narrower.
  • Combined: 127 shots, 87 transitions. Lengths run from 2 to 5 shots; 31 of 40 have three or more.
  • What I read per shot: description, actingInstructions, initialFrameInstructions, middleFrameInstructions, finalFrameInstructions, charactersVisible, props, environment + variant, shotJob, timing, beatEmotion, camera fields, dialogue, the images and sketches actually attached, sketchAnalyses, and the existing preGenValidation output.
  • Also read: the raw request — videoParams.jobMetaData.sequence.{characters,locations,props,duration}. This is the only way to see what the model is really handed.
  • Read-only. No writes.
  • What I could not check: whether the finished videos actually look broken. Nobody has watched them side by side against the inputs. Every severity below is a judgement on the inputs, not a measured defect rate in the output.
  • Existing system: src/workbench/preGenValidation/issueCatalog.ts has 26 shipped conflict codes with shortMessage + defaultSeverity (blocker | risky | minor | info). All 26 look at one shot in isolation. There is no sequence-level code today. Everything below is new and is written to slot into that same shape.

1. The rule that decides everything

Seedance fills in gaps. That is the whole point of it. So a difference between two shots is only a problem if the model cannot invent what happened in between.

              a state is different from one shot to the next
                                  │
              ┌───────────────────┴───────────────────┐
              │                                       │
   Is it a CONTINUATION of motion               Is it an EVENT that nobody
   the shots already describe?                  described anywhere?
              │                                       │
              ▼                                       ▼
      NOT A CONFLICT                        Can the user fix it in the UI?
   Seedance invents the bridge.              ┌────────┴────────┐
   Say nothing.                          YES │                 │ NO
                                             ▼                 ▼
                                      TELL THE USER      OUR PROBLEM
                                      → §3 taxonomy      → §6, don't show it

The test, in one sentence: if you saw only the last frame of shot N and the first frame of shot N+1, could you get from one to the other with no new information? If yes, so can Seedance.

Shot N ends with Shot N+1 starts with Seedance Why
walking toward the fridge fridge open, milk in hand fills it one continuous action — the hand on the handle is implied
raising a glass glass at the lips fills it same motion, later moment
turning away seen from behind fills it same
two forearms locked across her throat her own fingertips at her throat, no forearms cannot the attacker had to release and leave. That is an event, not motion
standing at the dining table standing on a cliff cannot somebody has to travel
raft finished raft half-built cannot time went backwards

The four "fills it" rows are the reason the old taxonomy was over-flagging. A skipped in-between moment is not a teleport.


2. A sequence is a timeline, not a pair of shots

This was my main mistake in the earlier drafts. I checked shot 1 against 2, then 2 against 3, and so on. Three of the five problem shapes below are invisible that way.

Take each thing that persists — a person, an object, the place, the light, the action — and lay its declared state across all N shots:

SEL 12 · Tadej's Triumphant Homecoming Dinner · shots 6,7,8,9,10

              shot 6    shot 7    shot 8    shot 9    shot 10
  Tadej         ●         ●         ●         ●          ●
  Father        ●         ·         ·         ●          ●
  Mother        ·         ·         ●         ●          ●
                          └────┬────┘
                    Father is gone for two shots and comes back.
                    Checking 6→7, 7→8, 8→9 separately never shows you this shape.
SEL 18 · Revelation and Countdown · shots 2,9,11,12

              shot 2    shot 9    shot 11   shot 12
  Shuki         ●         ●         ·         ●     ← "nobody visible" in a shot whose own
  attacker      ·         ·         ·         ·        text is entirely about her throat
                                    ↑
              shot 11's text: "Two forearms ... locked behind her throat"
              shot 11's cast list: (none)
              the attacker is in no shot's cast list, in any of the 20 sequences he appears in

The five shapes a timeline can have:

Shape What it looks like Visible pairwise?
Contradiction two shots claim states that can't both be true yes
Flicker present → absent → present no, needs 3+
Backwards the state regresses to an earlier point in the story sometimes
Frozen the state should have moved across the sequence and never does no, needs 3+
Unexplained jump an event-type change with no information anywhere in the N shots yes

Measured: 9 of the 16 multi-shot sequences in the first pass have at least one character listed present → absent → present (SEL 3, 7, 8, 11, 12, 14, 16, 18, 19). Most of those are ordinary cutaways and must not be flagged. See §4.

2.1 Order of operations — this matters more than it looks

The second pass showed that a data bug in the cast list manufactures fake flicker. R24 reads Shuki || Shuki || (none) and R40 reads Shuki|… || Shuki || (none). Both look like a character vanishing. Both are really PERSON_IN_TEXT_NOT_IN_CAST (§3A) — the person is described in the frame, just not listed.

   ①  fix the inputs        PERSON_IN_TEXT_NOT_IN_CAST
                            CHOSEN_VARIANT_NOT_SENT, NO_LOCATION_REFERENCE_SENT
                            SAME_OBJECT_SEVERAL_ASSETS
                                     │
                                     ▼
   ②  then read the timeline    flicker · frozen · backwards · unexplained jump

Run group F and the cast check first. Only on a clean cast list does a present → absent → present pattern mean anything.

Also: do not read "the sequence's environment" or "the sequence's props" from the first shot. In the second pass the environment variant sat on the first sent shot in only 13 of 20 — on the middle shot in R38, the last sent shot in R1 and R43, on two shots in R42 and R44, and nowhere at all in R46.

2.2 What exactly gets compared — the end/start seam

Earlier drafts said "compare the shot text", which is too loose to implement. The right comparison is narrower and much cheaper: the state a shot leaves behind against the state the next shot opens with.

   shot N                                     shot N+1
   ┌───────────────────────────────┐          ┌───────────────────────────────┐
   │ initial → middle → FINAL      │─── ✂ ───▶│ INITIAL → middle → final      │
   └───────────────────────────────┘          └───────────────────────────────┘
                     └──────────── the seam ────────────┘
              finalFrameInstructions  vs  initialFrameInstructions
              + actingInstructions on both sides, which is what supplies
                or fails to supply the transition

Every shot carries initialFrameInstructions, middleFrameInstructions, finalFrameInstructions and actingInstructions, and they are populated in practice — see the PERSON_POSE_CHANGES_UNEXPLAINED table in §3A, where all four are full sentences. So the seam check needs no vision read at all for the text-stated cases, which is why it is the first thing to build.

The seam is pairwise; the five shapes above are not. Both are needed and they are not the same check:

reads catches
seam shot N end vs shot N+1 start contradiction, unexplained jump
timeline all N shots at once flicker, frozen, backwards

Skipped shots are not read. Decided 2026-09-17 — provisional, Miki confirms. The user may select shots 1, 2 and 4; the unselected shots are simply not part of the check. Reason: Seedance only receives the selected shots, so the sequence has to stand on its own. The cost of this choice is a possible false positive — if the transition that explains a jump lives in a shot the user skipped, the story is coherent and we complain anyway. Worth knowing that the case which drove this rule does not have that problem: the UV sequence skips idx 12 and 13, and neither contains UV.


3. The taxonomy

Every row is something the user can go and change. seen = observed in the 40 sequences.

A. People

| Code | What the user reads | Severity | Seen | The fix they perform | Dimensions — the exact things compared | |---|---|---|---|------| | PERSON_LEAVES_UNEXPLAINED | "Shot 12 doesn't say how the attacker left shot 11" | blocker | yes | edit shot 12's start frame to show the release, or insert a shot | entity_presence | | PERSON_POSE_CHANGES_UNEXPLAINED | "Shot 02 starts with UV face-down, but shot 01 left him on his side and nothing turned him over" | blocker | yes, verified | make the two frames agree, or add the action that moves them | pose · body_orientation · location · held_state | | PERSON_IN_TEXT_NOT_IN_CAST | "Shot 7 describes Shuki in frame but she isn't in its cast list" | blocker | yes, 6 of 20 | add the character to the shot so a reference image gets sent | entity_presence · in_frame_position | | PERSON_APPEARS_UNEXPLAINED | "Shot 17 starts with someone who wasn't there and doesn't enter" | risky | yes | add the entrance, or accept it | entity_presence | | SAME_PERSON_LISTED_TWICE | "James is listed twice in shot 8" | risky | yes | remove the duplicate, or add a second copy properly | entity_presence (duplicate in charactersVisible) | | PERSON_LOOK_CHANGES | "Shuki is wearing a different outfit in shot 3 than in shot 1" | blocker | no — see §5 | pick one look for the whole sequence | identity · fixed_design · wardrobe · hair · accessory · injury · dirt · wetness |

PERSON_LEAVES_UNEXPLAINED, the canonical case, verbatim from the data:

shot 11, final frame: "Two forearms remain locked across her neck from frame-left and frame-right, centered tightly under her jaw. Her right hand remains hooked over the forearm at frame-left, and her left hand remains pressed against the wrist at frame-right."

shot 12, initial frame: "Only her eyes, nose, mouth, upper cheeks, and part of her throat are visible... The fingertips of her right hand appear at the lower frame edge pressed against the left side of her throat."

No forearms. Her left hand is gone too. Shot 12's acting notes treat her as alone. Both shots are EXTREME_CLOSE_UP, EYE_LEVEL, STATIC, back to back, both tagged suffocating_panic — the two shots claim to be one continuous moment, and a whole body disappears inside it. Blocker.

PERSON_POSE_CHANGES_UNEXPLAINED — found by a teammate, not by this pass. Job 2209e12c-74db-49b7-9044-734916212170, story 09c09721, scene b18cf43f, Seedance 2.5. An unconscious character changes body orientation between two shots with nothing on screen able to move him.

shot 01 (5ba4a544, idx 11) shot 02 (81816fe7, idx 14)
charactersVisible UV Shuki, UV
finalFrameInstructions "UV lies on his side near the center within the spiral" "UV rests on his side across frame-right"
initialFrameInstructions "UV lies on his side near the center" "UV lies face-down across the middle of frame"
actingInstructions "Keep UV lying still on his side… Keep all other human bodies off-screen" "Shuki… turns UV onto his side"

Shot 01 ends with UV on his side, motionless, alone. Shot 02 opens with him face-down. He is unconscious, so he cannot move himself, and shot 01 explicitly holds every other body off-screen, so nobody turns him. Then shot 02's own action turns him back onto his side — which is only coherent if he were face-down to begin with. Verified in Mongo, including the two shots the selection skips: idx 12 and idx 13 sit between them and contain Nawaz, Sunil, Bituin — no UV, so the missing transition is not hiding there either.

Why the other 21 codes miss it. Group B gives objects a state check (OBJECT_STATE_GOES_BACKWARDS, on condition). Group A only covered who is present, who they are, and what they look like. A person's body had no state check at all — an asymmetry, not a decision. The nearest code, ACTION_GOES_BACKWARDS, would have said "shot 3 happens before shot 2": a real problem reported with the wrong explanation, which is worse than silence because the user goes and fixes the wrong thing.

Why the fridge rule does not excuse it. The fridge rule forgives a missing action that the model can invent. Here there is no action to invent: the only person who could turn UV is stated to be off-screen, and the subject is unconscious. Blocker.

SAME_PERSON_LISTED_TWICE: SEL 2 shot 8 sends charactersVisible = James|James.

PERSON_LOOK_CHANGES — from the categories pass, and it is the case that started this whole investigation: one character, two different outfits inside one video. It splits into a free half and an expensive half, and only the free half should ship first.

  • The free version exists after all. Corrected 2026-09-17 — I had this backwards, and it is now shipped. I previously wrote that sequence.characters[] is deduped so an assetId can never repeat, on the strength of uniqueNamedAssets() in sequenceFlow.utils.ts:101. That function builds the prompt's slot legend, not the image payload. The payload is keyed assetId:variantId, so two variants of one character send two pictures. Re-verified over the 120 most recent sequence jobs: 5 of 120 send one character asset under two different variants.

The damage is visible in the request's own prompt. Job 142d6d9c-1aa0-4f47-88c2-60e3546a048c, story b7ffd148 — Tadej's homecoming dinner, where one shot is set to "Tadej normal" and the rest inherit the base cycling-kit reference, at a dinner table:

@image3: Tadej.
@image4: Tadej.
...
Constraints: Lock identity to @image3-@image6

Two pictures of one man, the same label on both because the legend dedupes by name, and an instruction to lock his identity to both at once. This is the complaint that started the investigation, detectable with a string comparison and no model. Implemented in referenceChecks.ts as checkCharacterLookAgreement, blocker, dismissible.

Environments are exempt, and that is not an oversight — see §4. A location variant per shot is the intended workflow, and 15 of those 120 jobs deliberately send two named views of one place. Two pictures of a room are two angles on one room; two pictures of a person are two people.

Props are left out. Prop assets in the sampled shows carry images: [], so a prop variant sends no second picture. Not verified beyond that sample. - The expensive version is still needed. A look change written into the prose or visible only in the pictures — wardrobe, hair, injury, dirt, wetness, with no variant difference behind it — needs the frame text plus a vision read. It rides along with whatever else looks at pixels; never its own call.

PERSON_IN_TEXT_NOT_IN_CAST — the gate that keeps it precise. Flag it only when the frame text gives the person a position in frame. Not when they are merely spoken to, spoken by, or reached toward.

cast list frame text verdict
R21 shot 7 (none) "Shuki's eyes and brow fill the frame, centered slightly left" flag
R28 shot 15 (none) "James stands frame-left, framed waist-up" flag
R39 shot 16 (none) "Badri's hand is entering from frame-left edge" flag
R1 shot 15 Becca Malcolm speaks, no frame position anywhere stay quiet
R44 shot 0 (none) "No characters are visible." stay quiet
R45 shot 2 (none) "Two small human figures … dwarfed by the environment" stay quiet

Cheap second signal: every flagged case has shotJob ∈ {reaction_emotion, interaction_relationship, action_progression}; every quiet case has shotJob ∈ {cinematic_anchor, transition, orientation_context, symbolic}. A reaction shot with nobody in the cast is suspect by construction.

Why it's a blocker: charactersVisible is what attaches the character reference image. R21 shot 7 sends an extreme close-up of Shuki's eyes with no picture of Shuki, in a sequence whose first shot does send her reference. One video, her established face and then her face invented from prose. Worst case is R28 shot 15: James speaks a line, both characters are described with full frame geometry, and the shot sends no characters, no image and no sketch.

B. Objects

| Code | What the user reads | Severity | Seen | The fix | Dimensions — the exact things compared | |---|---|---|---|------| | OBJECT_STATE_GOES_BACKWARDS | "The raft is finished in shot 14 and half-built in shot 17" | blocker | no — see §5 | reorder the shots, or fix the frame text | condition | | OBJECT_APPEARS_IN_HAND | "Shot 6 starts with her holding a camera she didn't have" | risky | yes | describe where it came from | holder · presence | | OBJECT_NEVER_CHANGES | "All four shots describe the same untouched object" | minor | no | probably fine, informational | condition (across 3+ shots) |

OBJECT_APPEARS_IN_HAND, now observed: R24 (none) || Shuki's Camera || (none) · R27 Shuki's Camera || Shuki's Camera || Rope Coil · R34 Dry Timber Logs|Audio Recorder || Rope Coil || (none).

Severity note: prop assets in The Countdown all have images: [] — a prop contributes a title and nothing visual. So a prop appearing mid-sequence changes only text. Keep it risky only when the prop is the thing the action is about (the camera in a shot about filming); otherwise minor.

C. The place and the light

| Code | What the user reads | Severity | Seen | The fix | Dimensions — the exact things compared | |---|---|---|---|------| | PLACE_CHANGES_WITHOUT_TRAVEL | "Shot 13 is in the dining room and shot 14 is in the kitchen, with no travel" | blocker | not yet, but one click away | split into two sequences | location_identity | | ENVIRONMENT_TAG_CONTRADICTS_FRAME_TEXT | "Shot 4 is tagged Cliffside Camp but the frame is in the forest looking at it" | risky | yes | change the tag on that shot | location_identity (tag vs frame text) | | LIGHT_OR_TIME_CHANGES | "Shot 1 reads as daylight, shot 3 reads as night" | risky | no — see §5 | pick one, or split the sequence | time_of_day · weather · lighting_direction · lighting_temperature · atmosphere | | EFFECT_STATE_JUMPS | "The fire is roaring in shot 2 and almost out in shot 4, with nothing saying it died down" | risky | no — see §5 | say what changed, or pick one intensity | presence · intensity · source · phase |

PLACE_CHANGES_WITHOUT_TRAVEL is reachable and I was wrong to say otherwise. I previously claimed a sequence cannot cross locations by construction. Confirmed false: scene Warm Homecoming: Lasagna and Laughter (story Chairs, sceneId 08ce9d26-4633-4eb3-9828-ef27b2829206) has 15 shots — shots 0–13 in Dinning room, shot 14 in Kitchen. Two environments in one scene. No sequence has spanned both yet, but selecting shots 13 and 14 together is one click, and only one location image is sent.

ENVIRONMENT_TAG_CONTRADICTS_FRAME_TEXT — R44 shot 4. Tagged Cliffside Camp v=7c04739d; the frame text puts the camera in the forest looking at the camp across distance, with dark tree trunks bordering frame-left and frame-right. The sequence reads Deep Forest || Cliffside Camp || Deep Forest, which looks like a place jump and is not one. Keep it separate from the row above: the fix is to change the tag, not the shot list. The damage is that it pulls the wrong location reference.

EFFECT_STATE_JUMPS — from the categories pass. Fire, smoke, rain, blood: a continuing effect that resets or leaps between shots. It sits here rather than under objects because an effect is ambient world state, not something a character holds. It is the one check with almost no field support — there is no stored effect state anywhere, so it is prose plus pixels or nothing. Risky, never a blocker: Seedance smooths small intensity changes, and a genuine change is often what the beat wants.

D. The action arc

| Code | What the user reads | Severity | Seen | The fix | Dimensions — the exact things compared | |---|---|---|---|------| | ACTION_GOES_BACKWARDS | "Shot 3 happens before shot 2" | blocker | no | reorder | action_phase · pose | | ACTION_NEVER_ADVANCES | "All 4 shots hold the same pose — nothing happens" | risky | no | change the frame text so the beat progresses | action_phase (across 3+ shots) | | EMOTION_ROUND_TRIP | "The emotion goes menace → betrayal → menace" | minor | yes | informational only | beatEmotion (across 3+ shots) |

EMOTION_ROUND_TRIP: SEL 4 shots 13,14,15 send quiet_menace → manipulation, creeping betrayal → quiet_menace. Only visible across three shots. Low severity — it may well be deliberate.

E. Screen geometry

| Code | What the user reads | Severity | Seen | The fix | Dimensions — the exact things compared | |---|---|---|---|------| | SCREEN_SIDE_FLIPS | "The father is on the left in shot 1 and the right in shot 2" | risky | yes | swap the side in one frame description | screen_side · relative_position | | SUBJECT_SCALE_JUMPS | "The creature is taller than the doorway in shot 2 and shorter than it in shot 5" | risky | no — see §5 | fix the size wording, or redo the sketch | subject_environment · character_character · character_prop |

This one is dangerous and must be gated. Fire it only when the two shots are adjacent in the send order AND the camera didn't change (same shot size, same angle, same movement). The same left/right swap appears twice in the data with opposite verdicts:

  SEL 13 · shots 1→2 · adjacent · same size, same angle, same movement   → real problem
  SEL 16 · shots 14→17 · a different shot in between · camera moved      → normal editing, say nothing

Ungated, this code will flag competent work and the feature gets switched off.

SUBJECT_SCALE_JUMPS — from the categories pass. Same gate problem as the row above, only worse: camera distance explains almost every apparent size change, so this can only fire when two shots agree on shot size and still disagree about how big something is relative to something else — a character against a doorway, a creature against a person. It is measured against a second object in frame, never in centimetres. Risky, never a blocker.

F. What actually gets sent as a reference

New group, from the second pass. Everything here is deterministic — no LLM, no cost — and none of it is visible in the UI today. The user picks a variant, or names a prop, and the request quietly carries something else.

The sent bundle is { id, imageUrl, label, description? } per reference, with id = "<assetId>:<variantId>" (trailing colon = no variant chosen). label is what names the picture for the model.

| Code | What the user reads | Severity | Seen | The fix | Dimensions — the exact things compared | |---|---|---|---|------| | NO_LOCATION_REFERENCE_SENT | "This sequence sends no picture of Cliffside Camp" | blocker | 4 of 20 | attach the environment image to a shot | locations[] is empty | | SHOT_HAS_NO_VISUAL | "Shot 15 has no image and no sketch" | blocker | 2 of 63 shots | generate an image or attach a sketch | imageUrl and sketch both absent | | CHOSEN_VARIANT_NOT_SENT | "Shot 2's chosen version of Cliffside Camp isn't in the request" | risky | needs the frontend — not in the MVP, see below | pick one variant for the whole sequence | variantId in id = "<assetId>:<variantId>" | | PERSON_LOOK_CHANGES (free half) | "Shots 6 and 7 use two different pictures of Tadej" | blocker | 5 of 120 jobs | pick one version for the sequence | variantId per character, across shots | | SAME_OBJECT_SEVERAL_ASSETS | "Shots 2 and 3 use two different assets for the wooden blocks" | risky | yes | merge them, or pick one for the sequence | assetId | | REFERENCE_LABEL_IS_PLACEHOLDER | "The prop is sent to the model named Upload from device" | risky | yes | rename the asset | label |

NO_LOCATION_REFERENCE_SENT — 4 of 20. The shots declare a place with a variant; the request carries nothing:

shots declare sent
R33 Cliffside Camp v=eb5d6014 (NONE)
R35 Cliffside Camp v=3e5f81e4 (NONE)
R44 Deep Forest + Cliffside Camp v=7c04739d (NONE)
R45 Cliffside Camp v=52f1ada5 on both shots (NONE)

The place is invented from prose, in a show whose whole point is that the camp looks the same every time.

CHOSEN_VARIANT_NOT_SENT — kept in the taxonomy, deliberately NOT in the MVP. Revised 2026-09-17. The original evidence was R42: shots declare 28afa663 || - || 27504ad8 || - and the request carries 28afa663, so shot 2's choice looks dropped. Two things have to be said about it now.

It is not straightforwardly true that a second variant gets dropped. The payload is keyed assetId:variantId, and 15 of the 120 most recent sequence jobs send two entries for one location asset — two named variants, two pictures, nothing lost. So whatever happened in R42 is not the general rule.

And the wider measurement cannot settle it. Comparing what the shots declare today against what 60 recent jobs carried gives "a declared environment variant was not sent" in 51 of 60 — a number too high to be a bug and almost certainly the confound: shots are mutable, the jobs are historical, so this measures how much the metadata has been edited since generation, not what the request got wrong. Discard it.

The structural reason it cannot ship yet: validation runs on Generate Prompt, before any job exists, so there is no sent payload to compare against. The reference payload is assembled by the frontend. To check this properly the frontend would have to send its resolved reference list into the validation call. That is a small change and worth making — but until it exists, the backend would have to predict the frontend's resolution, and a warning built on a prediction is a false positive waiting to happen.

Not verified: whether the frontend drops an unset-variant entry when another shot pins one, or correctly resolves both to the same picture. The UV sequence shows the ambiguous shape (shot 11 pins caef0615, shot 14 pins nothing, one picture sent) and cannot distinguish the two readings.

SAME_OBJECT_SEVERAL_ASSETS — confirmed in the asset registry, distinct assetId per row:

Scene Titles for one object
Raft Building and Accusations Dry Logs · Dry Timber Logs
Sentinels Close the Net Hollow Wooden Blocks · Hollowed Wooden Blocks · Wooden Signal Blocks
Panic at the Cliffside Phone · Shuki's phone
Return to Cliffside and Accusations Supply Crate · Nishi's Supply Crate

Two ID shapes — UUIDs and slugified titles — so two creation paths register the same object twice. It reaches the video: R43 is one 4-shot sequence where shot 3 declares Wooden Signal Blocks and shot 2 declares Hollow Wooden Blocks.

REFERENCE_LABEL_IS_PLACEHOLDER — the same asset lost its name over five weeks:

Sent at id label
2026-08-20 09:12 ffe9fe18…: Hollow Wooden Blocks
2026-08-20 09:56 ffe9fe18…: Upload from device
2026-08-21 11:34 hollowed-wooden-blocks: Upload from device (real name only in description)
2026-08-26 → 2026-09-16 asset-1787779559822-0 Upload from device

Upload from device is the UI button text. The model receives a picture captioned with the name of the button someone clicked. Across the whole collection, ~14,000 references are asset-<timestamp>-<n> — ad-hoc uploads with no asset identity, which therefore cannot carry a variant or be matched across shots.


4. What is NOT a conflict

Half the value of a taxonomy is the list of things it refuses to complain about.

Situation Why we stay quiet Evidence
A skipped in-between moment on one continuous action Seedance invents it. Fridge handle, glass to lips, turning around. §1
A character absent for one shot, then back That is a cutaway. Standard coverage. 9 of 16 multi-shot sequences in pass 1 do it
A cast list collapsing from a wide to singles Establishing shot → coverage. Normal grammar. R35 8→2→1 · R42 6→2→1→1 · R33 5→2→2
Mixed visual styles inside one sequence Deliberate. The creative team mixes a master shot, a rendered 3D sketch and a thumbnail sketch on purpose, and Seedance handles it. Never flag a style break. team confirmation
A different environment variant per shot This is the intended workflow, not drift. The variants are literally named per shot. DATA Center Exterior variants named EP02 SC01 SH01, SH02, SH03A; Cliffside Camp has SC-11_Sh-15, Sc-11_Sh-17, Sc-12_Sh-08
A dialogue speaker who isn't in the cast list Off-screen speech is normal. R1 shot 15: Malcolm speaks, Becca is the one on screen, no frame position given for Malcolm
An empty cast list on an environment shot Correct, and the text usually says so outright. R44 shot 0 "No characters are visible." · R45 shot 2, unnamed figures for scale
Camera angle changing sharply shot to shot Coverage and style. R20 LOW → GROUND → BIRDS_EYE in 8 s · R43 TOP → EYE → LOW → TOP
Two shots describing the same room in different words Prose varies. The reference image doesn't. —
Left/right swap when the camera moved Reframing legitimately changes screen sides. SEL 16 vs SEL 13 above
A line of dialogue split across shots mid-sentence By design — portion.start/end slices one line across shots on purpose. 25 of 69 portions don't end on a sentence boundary
Shot size text disagreeing with the enum during a dolly or zoom The frame legitimately changes size during the shot. SEL 18 shot 2

5. Must be checked even though these 40 didn't show it

Absence in 40 sequences is not evidence of absence. Five checks stay in the taxonomy with zero hits: two that this data exposed as gaps, and three carried over from the earlier categories pass, which looked at the problem from the model's side rather than the database's.

LIGHT_OR_TIME_CHANGES — keep it, and it is more important than the hit rate suggests. I checked all 13 environment variants used by the first 20 sequences. timeOfDay is empty on every single one. There is no structured field anywhere that says whether a shot is day or night. So if a user does put a daylight shot next to a night shot, we currently cannot see it at all — the only signals are the prose and the reference image. A check with no data behind it is a gap to close, not a check to delete.

OBJECT_STATE_GOES_BACKWARDS — keep it. The scene where this is most likely is Countdown's raft building. Across both passes I sampled shots 0,1,2,3,4,5,7,8,10,12,13,14,15,17 of it but never a pair that puts a half-built raft next to a finished one. Absence here is a sampling artifact, not a property of the product.

PERSON_LOOK_CHANGES, EFFECT_STATE_JUMPS, SUBJECT_SCALE_JUMPS — the three from the categories pass. None of them appeared in these 40 sequences, and none of them could have: all three live mainly in the pictures, and I read fields, not pixels. Two of the 40 sequences send no image at all and seven carry no sketch analysis, so a text-only reading is blind to exactly this class of problem. Dropping them because a field-level pass missed them would repeat the mistake LIGHT_OR_TIME_CHANGES above already demonstrates. PERSON_LOOK_CHANGES in particular is the complaint that started this work — one character, two outfits — and it is the only one of the three with a free deterministic half.

Not verified for any of the three: how often they actually occur. The honest position is that we know they are possible and we know we cannot see them today. Rates need a vision pass over sent bundles, which is a separate piece of work.


6. Real problems that are not the user's to fix

These showed up in the data and are worth fixing, but a warning banner is the wrong place for them. The user cannot repair any of them by editing a shot.

What's wrong Where Where it should go
The video length never matches the shots. Σtiming vs duration disagree in 20 of 20 in pass 2. 15 send a video longer than the content (R45: 6.5 s of shots in a 15 s video — 8.5 s of dead time the model fills freely); 5 send it shorter (R1: 15 s of shots in an 8 s video). Zero exact matches. 20 of 20 hard gate or auto-fit at build time. Highest-frequency finding in the whole sample
Shots sent out of story order. R25 [2,6,5,7], R43 [3,5,7,2], R1 [15,18,11], SEL 15 [11,4,7,12]. Every "backwards" check fires on these and every alarm is ours, not the user's. 4 of 40 needs a decision, not a warning — see the open question below
beatEmotion is empty on every shot of SEL 5, 6, 7, 15, 20 — all of Not Good, Really Not Good and the Becca scenes. 5 of 40 any emotion-based check is blind for a whole show. Know that before relying on it
A stale environment variant. SEL 18 sent Cliffside Camp var=1d2731f5; no shot declares it. SEL 19, eight minutes earlier, sent the correct 9c6e1ad6. 2 sends of the same shots deterministic check, free, no LLM needed
preGenValidation is sparse and is not a gate. 0 of 20 sequences have a status on every shot; the usual shape is one shot with a status and the rest - (R43 - \|\| - \|\| - \|\| pass/5c). status: pass coexists with attached conflicts (R39 pass/5c, R43 pass/5c). Sequences rendered anyway with blocked on a shot: R20, R42. 20 of 20 the sequence check cannot read per-shot validation as a precondition. It must compute its own view. If a sequence status is wanted, inherit the worst status among the shots that have one
7 of 20 sequences have zero sketchAnalyses on every shot. 7 of 20 the ten sketch-vs-text codes in issueCatalog.ts run blind on those. Since 15 of 20 sequences have only one shot with a selected image, the sketch is the visual input for most shots — so this gap matters more than it looks

Open question, not mine to decide — the send order. Are out-of-order sends a user mistake worth warning about, or a UI affordance working as intended? It changes whether ACTION_GOES_BACKWARDS can ever fire. Needs Miki or Disha.

Removed from this list: the middle frame

The earlier draft put middleFrameInstructions here, on the grounds that it is never sent to Seedance so a user's edit to it cannot affect the video. That changes with Seedance 2.5, which will send it. So:

  • middle-frame checks are live, not dead weight — MIDDLE_FRAME_MISSING is a real gap once 2.5 ships. Observed: R1 shot 18 and R44 shot 6 have MID: (empty) while their sibling shots have it.
  • the open question I raised for Miki about whether to send the field is already answered. Drop it.

7. How a conflict turns into a fix

This is the part that makes the taxonomy worth building, and it constrains what a conflict is allowed to be.

   conflict raised
        │
        ▼
   2–3 concrete options offered  ──►  user clicks one  ──►  cascade writes the metadata + text
        │                                                        │
        │                                                        ▼
        └─ we never hand-edit the video prompt          the next generate picks it up naturally

Three rules follow, and every row in §3 has to satisfy them:

  1. We do not touch the video prompt to paper over a conflict. The fix changes the underlying shot data.
  2. Every conflict needs 2–3 options the user can click, not a description of what's wrong. If a finding has no clickable option, it belongs in §6.
  3. The one sanctioned prompt override: text beats image. When the text is right and the visual is wrong, and the user confirms the text, we instruct the model to prefer the text over the reference. The worked example: the text says the hand is at her neck, the sketch shows it at her shoulder. If the user picks "neck", the model must be told to override the sketch. This is the only case where the conflict layer writes into the prompt.

Build notes for the layer itself:

  • The apply-fix endpoint is net-new and must be written generically, so the T2I (single-image) pre-gen flow can reuse it. ValidationConflict.proposedFix is already stored and returned today, but no route consumes it — that route is the missing piece.
  • For persisting a sequence draft before it has an ID: there is no sequenceId until generateSequenceVideo runs, so reuse the T2I pre-gen persistence pattern as the starting point rather than inventing a second one.

8. What the current 26-code catalog is missing

  1. There is no sequence-level code at all. All 26 codes in issueCatalog.ts look at a single shot. Everything in §3 is new surface area, not a rewording of what exists.
  2. Nothing looks at the request that actually leaves the building. Group F is entirely invisible to the current catalog: labels, dropped variants, duplicate assets, missing location references. All deterministic, all free, all currently unchecked.
  3. The old cross-shot list mixed four different things together — problems between shots, problems inside one shot, exemptions, and reference-bundle bugs. That is why severity could never be assigned: the same label covered a blocker and a non-issue.
  4. The biggest bucket in the earlier 193-selection pass, N8a at 39.4%, was two unrelated bugs wearing one label: a stale variant ID in the request (deterministic, free to check, ours to fix) and two shots describing a place in different prose (needs an LLM, and usually fine). One number, two fixes, no owner.
  5. Exemptions were coded as conflict types. N2n and N7c aren't things that go wrong; they're reasons to stay quiet. Listing them as categories inflated the catalogue and hid that one check produces both.
  6. "0 of 212, don't build the day/night gate" was wrong. See §5.

9. Where the old C1…X1 codes went

The archived categories draft (archive/I2V_CROSS_SHOT_CATEGORIES.md §10) listed nine codes. If you read that version, this is where each row landed. Nothing was lost silently, and one row was dropped on purpose.

Old Now
C1 world state — time, weather, light LIGHT_OR_TIME_CHANGES
C2 character appearance — outfit, hair, injury PERSON_LOOK_CHANGES
C3 prop state OBJECT_STATE_GOES_BACKWARDS · OBJECT_APPEARS_IN_HAND · OBJECT_NEVER_CHANGES
C4 environment PLACE_CHANGES_WITHOUT_TRAVEL · ENVIRONMENT_TAG_CONTRADICTS_FRAME_TEXT
C5 action seam ACTION_GOES_BACKWARDS · ACTION_NEVER_ADVANCES
C6 continuing effects — fire, smoke, blood EFFECT_STATE_JUMPS
C7 scale SUBJECT_SCALE_JUMPS
C8 style dropped. Mixed sketch and render styles are deliberate and Seedance handles them — §4
X1 spatial continuity SCREEN_SIDE_FLIPS

Two things changed beyond the renaming. The old draft split each code into dimensions — 50 of them across nine codes — and none of those 50 was ever tied to a severity or to something the user could click; the codes above are the ones a user can act on, which is the whole test in §1. And the old draft had no group F at all: it never looked at the request that actually leaves the building, which is where the most frequent and cheapest findings turned out to be.


Appendix — the 40 sequences

Pass 1 — frozen 2026-09-16

SEL show / scene shots sent N
1 The Countdown / Revelation and Countdown 14,16 2
2 The Countdown / Raft Building and Accusations 0,5,8 3
3 The Countdown / Revelation and Countdown 1,8,10 3
4 The Countdown / Revelation and Countdown 13,14,15 3
5 Not Good, Really Not Good / Data Center Cart Rampage 4,5 2
6 Not Good, Really Not Good / Data Center Cart Rampage 1,2,3 3
7 Not Good, Really Not Good / Data Bot's Plea and The Promise 0,1,2,3,4 5
8 Chairs / Warm Homecoming: Lasagna and Laughter 0,2,3 3
9 Chairs / Warm Homecoming: Lasagna and Laughter 0,1,3 3
10 Not Good, Really Not Good / Stan Arrives and Enters 0,1 2
11 enlarged_conflicts / Homecoming Lasagna Night 8,9,12,14 4
12 Enlarged_dev / Tadej's Triumphant Homecoming Dinner 6,7,8,9,10 5
13 Enlarged_dev / Tadej's Triumphant Homecoming Dinner 0,1,2,3 4
14 The Countdown / Revelation and Countdown 14,19,20 3
15 Not Good, Really Not Good / Becca Meets Malcolm 11,4,7,12 4
16 The Countdown / Raft Building and Accusations 14,15,17 3
17 Not Good, Really Not Good / Stan Arrives and Enters 0,1,2 3
18 The Countdown / Revelation and Countdown 2,9,11,12 4
19 The Countdown / Revelation and Countdown 2,11,12 3
20 Not Good, Really Not Good / Becca Meets Malcolm 1,11 2

Pass 2 — frozen 2026-09-17, zero overlap with pass 1

Rank show / scene shots sent N
1 Not Good, Really Not Good / Becca Meets Malcolm 15,18,11 3
20 The Countdown / Ash Spirals and Shadows 2,4,5 3
21 The Countdown / Revelation and Countdown 2,7,9 3
22 The Countdown / Raft Building and Accusations 3,7,8 3
24 The Countdown / Revelation and Countdown 2,6,7 3
25 The Countdown / Revelation and Countdown 2,6,5,7 4
27 The Countdown / Revelation and Countdown 14,17,19 3
28 The Countdown / Revelation and Countdown 14,15,16 3
33 The Countdown / Raft Building and Accusations 3,10,12 3
34 The Countdown / Raft Building and Accusations 3,4,5 3
35 The Countdown / Raft Building and Accusations 0,1,2 3
38 The Countdown / Raft Building and Accusations 13,14,15 3
39 The Countdown / Panic at the Cliffside 10,16,17 3
40 The Countdown / Panic at the Cliffside 10,11,13 3
42 The Countdown / Panic at the Cliffside 0,1,2,3 4
43 The Countdown / Sentinels Close the Net 3,5,7,2 4
44 The Countdown / Sentinels Close the Net 0,4,6 3
45 The Countdown / Red Mist Retreats 1,2 2
46 The Countdown / Knocking Signals, Group Flees 0,1,2,4 4
50 The Countdown / Footprints and UV Found 16,17,18 3