Skip to content

Cross-shot conflicts, found in the requests we actually sent

The question: we invented a cross-shot taxonomy (C1–C8, X1) from first principles. Does it survive contact with the sequences the creative team really generated over these past weeks?

The answer, in one paragraph: the contradictions are real and they are in the shots — but nobody ever sees them, because an LLM launders them on the way to Seedance. Of 21 selections where a character is placed on both screen sides across the selected shots, the prompt writer silently picked a winner in every single one and never told the creator a choice was made. A clean prompt is not a clean sequence: the prompt reads fine and the video still shows something the creative team did not ask for. That silent resolution — the template literally says "Do not flag, do not ask" — is the bug this project exists to fix. On top of it, the reference bundle is also wrong: 39% of sequences put a character on screen with no picture of them, 24% send no picture of the place, and 15 of the 23 sequences that move between two locations got fewer environment pictures than they had locations.

For the field-by-field dissection of one real sequence, see I2V_TADEJ_SEQUENCE_DISSECTION.md — eight contradictions in three shots of Chairs / Warm Homecoming, six of them findable without an LLM.

Method: Mongo reads only. Unit of analysis = one distinct shot selection (sceneId + sorted shotIds), not one generation attempt. For each selection I read the request that was actually sent — scenes[].sequences[].videos[].videoParams.jobMetaData.sequence — which stores the final prompt string plus the exact reference bundle with a resolved imageUrl per entry, and compared that against the shot metadata.

Corpus: 212 distinct selections, 57 scenes, 26 stories, 475 generation attempts.


0. The scoreboard

Retired — measurable in shot fields, but resolved before the model sees it

The contradictions that ARE in the shots, measured on 193 selections

These are found by comparing the selected shots against each other — description, initial/middle/finalFrameInstructions, actingInstructions, charactersVisible, characters, props, environment, dialogue. Full evidence in §2b; one sequence dissected end-to-end in I2V_TADEJ_SEQUENCE_DISSECTION.md.

Nature What contradicts Rate Surfaced today?
N8a same place, different environment.variantId across shots environment.variantId 39.4% (76) no
N2 a character is standing in one place and seated in another, no shot narrates the change frame text ↔ frame text 13.5% (26) no
N8b the shots are in genuinely different places environment.title 10.4% (20) no
N4b a registered prop is named in the frame text and no reference was ever sent scene props registry ↔ shot props ↔ frame text 8.3% (16) no
N7a a character is fully in frame in the text but missing from charactersVisible frame text ↔ charactersVisible 7.8% (15) no
N1a a character is on both screen sides inside ONE shot's own start/mid/end frames frame text ↔ frame text 6.2% (12) no
N1b a character is on both screen sides across DIFFERENT shots — a true 180° break frame text ↔ frame text 5.2% (10) no — silently resolved
N7b only a body part is in frame (hand, fingertips, shoulder), no sheet frame text ↔ charactersVisible 4.7% (9) no
N7c the character is explicitly off-screen ("eyes locked off-screen at James") → NOT a conflict — 5.2% (10) correctly ignored
N2n posture changes and a shot narrates it → NOT a conflict — 2.1% (4) correctly ignored

N2 vs N2n, and N7a vs N7c, are the calibration working. 26 unnarrated posture changes are conflicts; 4 narrated ones are not — the narrated ones look like the Chairs sequence ("Tadej drops heavily into the chair") and need no intervention. Likewise a character named only as an eyeline target ("chin angled right toward Badri off-frame") is correctly absent from charactersVisible; a character described "framed waist-up" and absent from it is not.

Numbers I measured, then threw out. Recorded so nobody rebuilds them:

Claim Why it is dead
~~N3 contact target lands on different body parts, 6.7%~~ 0 real cases. With sentence context required, the detector finds none. Your neck-vs-shoulder example does not appear in the frame text — it is a text-vs-sketch conflict, which needs the picture read, not the metadata.
~~N4 a prop vanishes between shots, 34.2%~~ Noise. "Shoebox present in shots 8,13, absent from 4,10, reference attached: yes" — the camera moved. Replaced by N4b.
~~N7 named but not listed, 14.0%~~ Conflated three things. Split above into N7a / N7b / N7c.
~~N1 both screen sides, 10.9%~~ Conflated intra-shot with cross-shot. Split into N1a / N1b.

Still not split: N2 (13.5%) mixes intra-shot with cross-shot the same way N1 did — e.g. The Countdown / Raft Building [13,14,15] has Sunil "stands beside UV at screen-right" and "sits on the same log … at screen-right" both inside shot 14. The cross-shot subset is smaller than 13.5% and I have not measured it.

The evidence — a clean cross-shot 180° break

  The Countdown / Raft Building and Accusations  [3, 4, 5]     att 2 / ok 2

  shot 3  start  "James stands in the FRAME-LEFT third near the raft materials,
                  body angled three-quarter right, right hand holding the Audio Recorder"
  shot 3  mid    "James stands in the FRAME-LEFT third…"      ← consistent
  shot 3  end    "James stands in the FRAME-LEFT third…"      ← consistent

  shot 5  start  "James stands FRAME-RIGHT near the right third, visible from
                  mid-thigh up, facing left toward Nishi."
  shot 5  mid    "James remains FRAME-RIGHT…"                 ← consistent
  shot 5  end    "James stands at FRAME-RIGHT edge…"          ← consistent

           Each shot is internally solid. Across the cut James jumps
           the axis. Nothing narrates a move. Rendered twice, no flag.

And the worst one, where both characters swap at once — a full axis flip:

  The Countdown / Raft Building and Accusations  [14, 17]      att 1 / ok 1

  shot 14   Shuki  SCREEN-LEFT  (all three frames)
            Sunil  SCREEN-RIGHT (all three frames)
  shot 17   Shuki  SCREEN-RIGHT
            Sunil  SCREEN-LEFT
                   └─ the two of them trade places across one cut

  PROMPT DECIDED: Shuki=LEFT, and no placement for Sunil at all

The old codes, re-judged

Old code Verdict
X1 SPATIAL_CONTINUITY_BREAK Confirmed real — N1b, 5.2%. Keep it. The reason it looked absent is that the prompt writer erases it before anyone can see it (§2). Add N1a (6.2%): one shot's own start/mid/end frames put the character on both sides — a shot cannot cut, so this one has no innocent reading.
C5 ACTION_SEAM_BREAK Split it. The useful half is N2 (unnarrated posture change, 13.5%). "Seam text missing" as a flat check is noise.
C4 ENVIRONMENT_BREAK Confirmed, and split. N8a variant drift 39.4% (mostly harmless) vs N8b real place change 10.4% (and 15 of 23 got fewer pictures than places → REF-3).
C2 CHARACTER_APPEARANCE_BREAK Reframe. Not "a character looks different" — it is N7a (in frame, absent from charactersVisible, 7.8%) and REF-1 (no sheet sent, 39.2%). Duplicate sheets are a designed feature, not a defect.
C3 PROP_STATE_BREAK Reframe as N4b (8.3%): a registered prop is named in the frame text and no reference was ever sent. The props[] array emptying between shots is not evidence — the array is not the object, and that check produces 34.2% noise.
C1 / C6 world & effect state Retire. No field stores them.
C7 SCALE_BREAK Retire. Wide → close-up is film grammar.
C8 STYLE_BREAK Retire. You confirmed the team mixes render + sketch deliberately and it works.
— NEW: actingInstructions performs a line that is not in dialogue — see T3 in the Tadej dissection. Deterministic: count lines, match the quoted word.
— NEW: characters carries the SCENE cast where charactersVisible carries the shot cast — see T2. Any check reading characters is wrong on every shot.
— NEW: an object's screen side flips across a cut (the chair in T1) — the single conflict in the Tadej sequence that would have changed the video.
your neck-vs-shoulder example Not findable in the metadata. 0 cases in the frame text (my earlier 6.7% was a regex artifact). That conflict is text vs sketch — it needs the picture read, not the fields compared. It belongs in the taxonomy; it cannot be detected by the checks in this document.
day/night inside a sequence 0 of 212 as a flat check. But see T4 in the Tadej dissection: the real form is night-vs-unspecified with a window visible in both, which a day/night check misses entirely.

Real — measured on the request that was actually sent (n = 212)

New code What it is Rate Severity
REF-1 MISSING_CHARACTER_SHEET character has a valid assetId on a selected shot; no sheet bound in the request 39.2% (83) BLOCKER when they speak or carry the shot
REF-2 NO_ENVIRONMENT_SHEET no picture of the place at all 24.1% (51) MEDIUM
REF-4 UNLABELLED_REFERENCES prompt tags fewer images than were attached 23.1% (49) MEDIUM
STR-1 SHOT_BLOCK_MISMATCH prompt describes a different number of shots than were selected 19.8% (42) MEDIUM
REF-3 PLACES_FLATTENED sequence moves between N places, gets fewer than N pictures 15 of 23 location-changing sequences BLOCKER
STR-3 NO_SHOT_STRUCTURE prompt has no per-shot blocks at all 11.3% (24) MEDIUM
REF-1b ZERO_CHARACTER_SHEETS characters on screen, characters: [] 4.2% (9) BLOCKER
REF-5 NO_TAG_BLOCK images attached, prompt never says what they are 3.8% (8) MEDIUM
REF-1c SPEAKER_NO_SHEET a character speaks with no reference picture 1.9% (4) BLOCKER
STR-2 DUPLICATE_SHOT_LABELS two blocks both labelled Shot 01 1.4% (3) MEDIUM
TXT-1 WRITER_INTRODUCED_FLIP prompt contradicts screen direction where the raw shots did not 0.5% (1) MEDIUM

1. Two methodological errors, and the corrected method

I got this wrong twice before arriving at the table above. Both errors are easy to repeat, so they are recorded here rather than quietly fixed.

Error 1 — the corpus was ~19× inflated by cloned test projects

  db.storyvideos.find({'scenes.sequences': {$exists: true}})
     ↳ 8,840 "sequences"   ← what the previous version of this document reported
     ↳ 160 story documents
     ↳ BUT sceneId 8c74019e ("Inside the Party Bubble") appears in 30 of them,
       under 30 different projectIds, ALL titled "malfunction"
     ↳ 1,095 sceneIds are duplicated this way

De-duplicating by sceneId (newest updatedAt wins): 475 attempts, 212 distinct selections, 57 scenes, 26 stories. Every percentage in the previous version was weighted by clones of test projects and should be discarded.

Error 2 — I measured fields the model never receives

I compared finalFrameInstructions of shot N against initialFrameInstructions of shot N+1. Those are inputs to a prompt-writing LLM, not the prompt. See §2. This is the more important of the two errors, because it invalidates the whole shape of the old taxonomy, not just its numbers.


2. The decisive finding: the contradictions get laundered, not resolved

There is already an LLM sitting between the shots and Seedance: src/promptTemplates/workbenchV2/generateSequenceSeedancePrompt.template.ts (SEEDANCE2_SYSTEM_PROMPT_v20260803), called from src/workbench/sequenceFlow/buildSequenceUserPrompt.ts.

It takes the selected shots as a shot_breakdown bundle and emits one prompt with per-shot Camera: / Starting Context: / Action: / Dialogue: / Performance: blocks. And that system prompt says, verbatim:

Conflict resolution: if the shot breakdown contradicts the storyboard, the storyboard wins when a storyboard slot exists. When no storyboard is provided, trust master_shot and asset sheets over contradictory shot_breakdown prose for spatial composition. Do not flag, do not ask.

Measured against real data — and read the arrow carefully, because it is not good news:

  RAW SHOT FIELDS                     THE PROMPT THAT WAS SENT
  ───────────────                     ────────────────────────
  21 selections place a character
  on BOTH screen sides across the      ──────▶   0 still contradict
  selected shots                                        ▲
                                          the writer PICKED A SIDE, 21 times,
                                          and told nobody. Example:

    The Countdown / Raft Building [14,17]
      shot 14  "Shuki stands SCREEN-LEFT at mid-depth, gripping one end
                of a log at chest height with both hands"
      shot 17  "Shuki stands behind UV at SCREEN-RIGHT, slightly out of
                focus, upright, both arms at her sides"
      PROMPT DECIDED: screen-LEFT        ← shot 17 lost. Nobody asked.

    The Countdown / Revelation [14,16]
      shot 14  "Shuki occupies SCREEN-RIGHT from upper torso up"
      shot 16  "Shuki occupies FRAME-RIGHT, head and upper shoulders
                visible, seated and turned three-quarter left"
      shot 14  "Shuki sits at SCREEN-LEFT on the rock in the foreground"
      PROMPT DECIDED: LEFT in both blocks   ← 6 attempts, 4 renders

  and 7 of the 21 are worse — the prompt dropped the character's placement
  entirely rather than choosing, so Seedance places them freely:
      "PROMPT DECIDED: (no placement for Sunil in the prompt)"

A clean prompt is not a clean sequence. The prompt has no contradiction because the writer is instructed to erase it. The creative team's intent is still ambiguous in the shots; an LLM resolved it by coin-flip and the video came back with a composition nobody signed off on. That is the failure mode, and it is invisible today at every layer: the shot fields still disagree, the prompt looks correct, the render succeeds, and no conflict is ever surfaced.

Consequence for the taxonomy. The old codes are not wrong about where to look — the shot fields are exactly right. They are wrong about when: the check has to run before the prompt writer, on the shots, and its output has to be a question for the creator, not a silent fix.

One extra case, in the other direction. In Arrival / James and Shuki Clash [4,5,6] the writer introduced a screen-direction flip that was not in the raw shots (James screen-right in block 04, screen-left in block 06). So the writer both hides real conflicts and manufactures new ones.

Worked proof — the Chairs sequence

The clearest case, because it is the one I previously reported as a defect. Selection [0, 2, 3] in Chairs / Warm Homecoming; shot 1 (the walk) not selected. In the raw shot fields, shot 0 ends with Tadej standing in the doorway and shot 3 begins with an empty chair and then has him seated. I called that a teleport. Here is what the request actually said:

Shot 03: 4.5s Action: Father's head tilts up two degrees to meet Tadej's gaze at screen-left. […] Tadej drops heavily into the chair at screen-left, weight falling into the seat, shoulders sagging forward, legs splaying under the table. He looks at Father with tired amusement […]

The sit-down is written out. The geometry is consistent across all three shots as well — Father screen-right, eyeline directed screen-left toward the doorway (shot 02), Tadej ending screen-left (shot 03). There is no conflict in this sequence. A cut bridging shot is a gap the writer fills, not a contradiction — which is the principle the rest of this document is calibrated on.


3. What actually reaches the model — the real conflicts

3.1 REF-1 — a character is on screen with no picture of them (39.2%)

The sharp form of the check: a character has a valid UUID assetId on one of the selected shots, and that assetId is not in the request's characters[]. So a resolvable reference exists in the data and simply was not attached.

Worked example — Arrival / Red Mist Attacks Ferry, selection [11, 13, 17, 19]:

  request payload:  characters: []              ← zero character sheets
  request tags:     @image1: Storyboard — visual and composition anchor.
                    @image2: Master shot — spatial-axis anchor.
                    (that is the entire tag block)

  but the prompt itself contains:
     Shot 11  Dialogue:  Badri: "like it has a will of its own."
     Shot 13  Dialogue:  Shuki: "(Loud whisper) Don't!"
     Shot 17  "Detail on UV's fingertip and the tendril…"

  and the shots' own metadata says:
     shot 13  charactersVisible: [ UV / 678d2f84-040… ]   ← valid, resolvable assetId
     shot 17  charactersVisible: [ UV / 678d2f84-040… ]
     shot 19  charactersVisible: [ UV / 678d2f84-040… ]
                                        ▲
                    UV has a real character sheet. It was not attached.
                    The model invents UV's hand three times and has no
                    reason to make it the same hand twice.

Rendered successfully, twice. Nothing flagged it.

Why this is the BLOCKER class. Everything else on this list, the model can paper over. It cannot invent an identity it was never shown and keep it consistent across four shots. This is the one case where the failure is guaranteed, not probabilistic.

The sub-cases, ranked by how unambiguous they are:

Rate Severity Why
REF-1c character speaks with no sheet — Granny, Stan Padnick, Badri, Shuki 1.9% (4) BLOCKER a talking face with no reference is invented at full size on screen
REF-1b characters: [] with characters on screen 4.2% (9) BLOCKER nothing anchors any identity
REF-1 valid assetId not bound 39.2% (83) BLOCKER → SLIGHT depends on whether the character is the subject or a passing element

Not verified — and this is the biggest open question in the document: whether the creator chooses which sheets to attach in the UI. If they do, part of 39.2% is deliberate (an insert of a fingertip may not need a full character sheet) and the right product response is a warning, not a fix. If the system chooses, it is a bug. Everything about REF-1's product shape depends on this answer.

3.2 REF-3 — the sequence moves between places and gets one picture (BLOCKER)

This is the mirror image of C4, and more damaging. C4 assumed the risk was too many pictures of one place. The real risk is too few pictures for several places.

  23 selections have shots tagged with DIFFERENT place titles.
  How many environment sheets did the request carry?

    0 sheets  ██████                       6    ← no picture of either place
    1 sheet   █████████                    9    ← one room for a two-room sequence
    2 sheets  ███████                      7    ← correct
    3 sheets  █                            1    ← correct

  15 of 23 (65%) got fewer pictures than they had places.

Worked example — Arrival / Arrival on Island, selection [3, 4], the second most-generated selection in the corpus (7 attempts / 6 renders, plus a separate run of 3 / 3):

  shot 3   environment = "Island Coast"
  shot 4   environment = "Cliffside Camp"
  request  locations: []          ← zero environment sheets for either

Two different beaches, one cut between them, no picture of either. The model invents both and has no reason to make them look like the same island.

More of the same: - malfunction / Gridlock and Goodbye [4,5] — LearnSpace + Medspace → Party Bubble, 0 sheets (4 attempts, 4 renders) - Arrival / Arrival on Island [14,15,16,17] — Cliffside Camp → Sentinel Jungle (Ancient), 0 sheets - malfunction / Bubble Car Breakdown [0,1,2,3,4] — Stan's House → Stan's House & Street View → Millionaire Town, 1 sheet

This also settles the point Gerard raised on comments_to_claude.md §0.6: the claim that "shot 1 is a kitchen and shot 2 is a street cannot happen by construction" is false. It happens in 23 of 212 selections, and most of the time the model is not given a picture of the second place.

3.3 REF-2 — no picture of the place at all (24.1%)

51 of 212 selections carry locations: [], independent of whether the place changes. A quarter of all sequences are generated with no environment reference. Combined with the fact that 82 of 212 (38.7%) have no storyboard either, a meaningful slice of requests reaches Seedance with no spatial anchor of any kind.

3.4 REF-4 / REF-5 — images attached that the prompt never explains

The system prompt calls the opening image-tag block "non-negotiable" and requires one line per attached image with no skipped slot numbers. Measured:

  • REF-4, 23.1% (49): fewer @imageN tag lines than references attached.
  • REF-5, 3.8% (8): no tag block at all, while images were attached.

REF-5 turned out to be creators writing their own prompt by hand and bypassing the generator. The stored prompts begin:

  "Camera: Static master wide shot, eye-level. No zoom, no pan, no cuts…"
  "Use Image-1 only for storyboard action layout…"      ← a different tag convention entirely
  "Use Image-1 as the storyboard reference…"
  "create a video sequence that includes al…"           ← 126 chars, 5 shots selected, 0 renders
  "The futuristic Party Bubble enters smoot…"           ← 11 attempts, 10 renders

malfunction / Assignment Day: Tim's Entrance [11,12,14] carries 3 character sheets, 1 location, 3 master shots and a storyboard — and the prompt never mentions a single image. The model receives eight pictures and no statement of which one is Tim, which is Stan, or which is the room. It rendered anyway.

This matters for the design of any gate: a validator that assumes the templated prompt format will fail to parse ~11–15% of real requests.

3.5 STR-1 / STR-2 / STR-3 — the prompt does not describe the shots that were selected

Rate What
STR-1 19.8% (42) shot-block count ≠ selected-shot count. malfunction / Assignment Day [11,12,14]: 1 block for 4 selected shots. Arrival / UV's Nightmare Vision [0,1,2,3]: 3 blocks for 4. malfunction / Inside the Party Bubble [1,2]: 3 blocks for 2.
STR-2 1.4% (3) duplicate labels — block labels come out as [1, 1, 2]
STR-3 11.3% (24) no Shot NN: blocks at all

Worked example of STR-2 — test-4 / Dawn Warning: The Loop [0,2]. Two shots selected. The prompt contains three blocks, two of them both labelled Shot 01:

  Shot 01: 2.5s   Meera foreground, Aarav at the console, over-shoulder
  Shot 01: 6s     camera glides through the exterior doorway into the control room   ← same label
  Shot 02: 5s     Aarav motionless in the chair, dolly out

Generated 5 times, 5 renders. This is also what produced a false positive in my own first pass: I read it as "Aarav appears in shot 2 with no narration" when the real defect was the duplicated block.


4. Calibrating severity

The calibration question: a missing transition can be a blocker, a medium miss, or something Seedance absorbs without anyone noticing. The discriminator is what the model has to invent, and the Chairs sequence supplies the signal.

  SLIGHT   — the prompt narrates the change. The model interpolates.
             "Tadej DROPS HEAVILY INTO THE CHAIR at screen-left, weight falling
              into the seat, legs splaying under the table"
             A bridging shot was cut and it does not matter.

  MEDIUM   — the change is large and the model must guess how it happened, but
             every element involved has a reference picture.
             A prop changes hands off-screen. A character crosses the room.
             Wrong in the details, right in substance.

  BLOCKER  — the model must invent something it was never shown, and no amount
             of text fixes it.
             · a character on screen with no character sheet          (REF-1)
             · a second location with no environment sheet            (REF-3)
             · a state change with no cause anywhere in the selection:
                 The Countdown / Raft Building [3,10,12]
                   shot 10  "both BARE PALMS angle toward camera"
                   shot 11  props: [Spiral Ash Pattern]   ← NOT SELECTED
                   shot 12  "palms facing up with ASH SPIRALS clearly visible"
                 Not interpolation. Invention of a cause.

So severity is a product of three things, per seam:

  severity  =  size of the state change
            ×  does any shot's text acknowledge it     ← the Chairs "drops heavily" test
            ×  does a reference picture exist for it   ← REF-1 / REF-3

Only the third factor is currently checkable with confidence, because the second requires parsing the generated prompt and 11% of prompts have no per-shot structure at all (STR-3).

Retracted: I measured "element appears in shot N+1, absent in shot N, arrival not narrated" at 25.9%. Hand-checking showed it was dominated by the duplicate-shot-block defect (STR-2), not by real unnarrated arrivals. That number is withdrawn. The check is still the right idea; STR-1 and STR-2 have to be fixed before it can be measured.


5. Honest negatives

Checked Result
Day/night contradiction inside a sequence 0 of 212. It exists at scene level (The Countdown shot 9 "night setting" vs shot 17 "muted daylight") but never inside one generated sequence. Do not build this gate.
Do any of these checks predict creator dissatisfaction? No. Baseline is 1.87 successful renders per selection. Every check sits within ±0.6 of baseline. STR-2 reaches 3.67 but n=3. No detector in this document is validated against an outcome.
Are generation failures caused by continuity? No. 178 failures: 103 content-moderation flags ("flagged as sensitive, E005"), 18 copyright rejections, 23 Byteplus API errors, plus S3 timeouts and one bad API key. Zero relate to cross-shot continuity.
Is isSelected a quality label? No. True on all 212, and up to 5 videos per selection. It does not mean "the creator chose this one".
isFavorite Present on 2 selections. Too rare to use.
Cross-scene shot references 0. Selections are always scene-local.
Duplicate character sheets (old C2) Not a defect. The builder deliberately labels them "(sheet 1 of 2)" — buildAssetOccurrenceLabels() in buildSequenceUserPrompt.ts. Tim Padnick legitimately has 3.
The 24× retry loop (test-4 / Aarav Alone [0,1,2,8]) Not dissatisfaction. 24 attempts, 0 renders — a technical failure loop, not a creator re-rolling.

6. What to change in the system

In priority order, based on the measurements above.

  1. Move the gate from the shot fields to the request bundle. The bundle is assembled in src/workbench/sequenceFlow/ and is fully deterministic — no LLM is needed for REF-1, REF-2, REF-3, REF-4, REF-5, STR-1, STR-2 or STR-3. That is eight of the eleven real checks for the cost of a set comparison.
  2. Answer the open question in §3.1 first. If the creator picks the reference sheets, REF-1 is a warning ("UV is on screen in 3 of these shots and has no reference — attach it?"). If the system picks them, it is a bug. This decides REF-1's entire product shape, and REF-1 is the biggest number here.
  3. REF-3 is the most under-appreciated finding. 65% of location-changing sequences get fewer environment pictures than they have locations, and 6 of them get none. It is a one-line check: distinct environment.title across the selected shots vs locations[].length.
  4. Fix STR-1 / STR-2 before building any text-level check. 19.8% of prompts do not describe the selected shots, and no continuity check can be trusted on a prompt whose shot blocks are miscounted or duplicated.
  5. Do not build C1, C6, C7, C8, or a day/night check. Nothing in the schema supports the first two; the third and fourth are normal craft; the fifth has zero occurrences.
  6. Keep exactly one text-level check: TXT-1. The prompt writer introduced a screen-direction contradiction that was not in the source data (Arrival / James and Shuki Clash [4,5,6]). Validating the writer's own output is a narrower and far more defensible job than second-guessing the shot fields.
  7. Any validator must tolerate hand-written prompts. ~11% have no per-shot structure and 3.8% have no tag block, because creators bypass the generator.

7. Not verified, and where the method stops

  • No output video was watched. Every finding here is a property of the request. Whether Seedance produced a visible glitch is unknown. The videos are in the data (videoURL on every completed attempt) and this is the obvious next study — it is the only thing that turns these rates into real severities.
  • No detector is validated against an outcome (§5). Read every rate as "how often the request has this property", never as "how often the creator got a bad video".
  • The severity tiers in §4 are reasoned, not measured. The three-factor rule is a proposal.
  • Whether reference attachment is a creator choice is unknown (§3.1). Biggest single caveat.
  • Screen-direction detection is text-based. It requires a placement verb (stands, sits, occupies, …) followed by screen-left / screen-right with no eyeline word in between. An earlier blacklist version over-reported by 2.4× because "eyes still fixed frame-right" reads as a body position. The whitelist version is used throughout, and it will under-report prompts that express position without those verbs.
  • REF-1 at 39.2% is the broadest number here. The tight, unambiguous subsets are REF-1c (a speaking character with no sheet, 1.9%) and REF-1b (characters: [], 4.2%).
  • Corpus concentration: 164 of 212 selections come from three shows — The Countdown (66), Arrival (58), malfunction (40). test-4 and Enlarged_dev are test shows and are included. A single show's habits can move any rate in this document.
  • storyboard is often a manual upload. In the Chairs request it is label: "Upload from device". The prompt gives the storyboard authority over "framing, staging, blocking, screen direction, and shot-to-shot composition" — so in 130 of 212 selections, composition is anchored on an image the pipeline did not generate and cannot inspect.