Cross-shot conflicts, found in the requests we actually sent¶
The question: we invented a cross-shot taxonomy (C1–C8, X1) from first principles. Does it survive contact with the sequences the creative team really generated over these past weeks?
The answer, in one paragraph: the contradictions are real and they are in the shots — but nobody ever sees them, because an LLM launders them on the way to Seedance. Of 21 selections where a character is placed on both screen sides across the selected shots, the prompt writer silently picked a winner in every single one and never told the creator a choice was made. A clean prompt is not a clean sequence: the prompt reads fine and the video still shows something the creative team did not ask for. That silent resolution — the template literally says "Do not flag, do not ask" — is the bug this project exists to fix. On top of it, the reference bundle is also wrong: 39% of sequences put a character on screen with no picture of them, 24% send no picture of the place, and 15 of the 23 sequences that move between two locations got fewer environment pictures than they had locations.
For the field-by-field dissection of one real sequence, see
I2V_TADEJ_SEQUENCE_DISSECTION.md — eight contradictions in
three shots of Chairs / Warm Homecoming, six of them findable without an LLM.
Method: Mongo reads only. Unit of analysis = one distinct shot selection
(sceneId + sorted shotIds), not one generation attempt. For each selection I read the request
that was actually sent — scenes[].sequences[].videos[].videoParams.jobMetaData.sequence — which
stores the final prompt string plus the exact reference bundle with a resolved imageUrl per
entry, and compared that against the shot metadata.
Corpus: 212 distinct selections, 57 scenes, 26 stories, 475 generation attempts.
0. The scoreboard¶
Retired — measurable in shot fields, but resolved before the model sees it¶
The contradictions that ARE in the shots, measured on 193 selections¶
These are found by comparing the selected shots against each other — description,
initial/middle/finalFrameInstructions, actingInstructions, charactersVisible, characters,
props, environment, dialogue. Full evidence in §2b; one sequence dissected end-to-end in
I2V_TADEJ_SEQUENCE_DISSECTION.md.
| Nature | What contradicts | Rate | Surfaced today? |
|---|---|---|---|
N8a same place, different environment.variantId across shots |
environment.variantId |
39.4% (76) | no |
| N2 a character is standing in one place and seated in another, no shot narrates the change | frame text ↔ frame text | 13.5% (26) | no |
| N8b the shots are in genuinely different places | environment.title |
10.4% (20) | no |
| N4b a registered prop is named in the frame text and no reference was ever sent | scene props registry ↔ shot props ↔ frame text |
8.3% (16) | no |
N7a a character is fully in frame in the text but missing from charactersVisible |
frame text ↔ charactersVisible |
7.8% (15) | no |
| N1a a character is on both screen sides inside ONE shot's own start/mid/end frames | frame text ↔ frame text | 6.2% (12) | no |
| N1b a character is on both screen sides across DIFFERENT shots — a true 180° break | frame text ↔ frame text | 5.2% (10) | no — silently resolved |
| N7b only a body part is in frame (hand, fingertips, shoulder), no sheet | frame text ↔ charactersVisible |
4.7% (9) | no |
| N7c the character is explicitly off-screen ("eyes locked off-screen at James") → NOT a conflict | — | 5.2% (10) | correctly ignored |
| N2n posture changes and a shot narrates it → NOT a conflict | — | 2.1% (4) | correctly ignored |
N2 vs N2n, and N7a vs N7c, are the calibration working. 26 unnarrated posture changes are conflicts;
4 narrated ones are not — the narrated ones look like the Chairs sequence ("Tadej drops heavily
into the chair") and need no intervention. Likewise a character named only as an eyeline target
("chin angled right toward Badri off-frame") is correctly absent from charactersVisible; a
character described "framed waist-up" and absent from it is not.
Numbers I measured, then threw out. Recorded so nobody rebuilds them:
| Claim | Why it is dead |
|---|---|
| ~~N3 contact target lands on different body parts, 6.7%~~ | 0 real cases. With sentence context required, the detector finds none. Your neck-vs-shoulder example does not appear in the frame text — it is a text-vs-sketch conflict, which needs the picture read, not the metadata. |
| ~~N4 a prop vanishes between shots, 34.2%~~ | Noise. "Shoebox present in shots 8,13, absent from 4,10, reference attached: yes" — the camera moved. Replaced by N4b. |
| ~~N7 named but not listed, 14.0%~~ | Conflated three things. Split above into N7a / N7b / N7c. |
| ~~N1 both screen sides, 10.9%~~ | Conflated intra-shot with cross-shot. Split into N1a / N1b. |
Still not split: N2 (13.5%) mixes intra-shot with cross-shot the same way N1 did — e.g.
The Countdown / Raft Building [13,14,15] has Sunil "stands beside UV at screen-right" and "sits on
the same log … at screen-right" both inside shot 14. The cross-shot subset is smaller than 13.5%
and I have not measured it.
The evidence — a clean cross-shot 180° break¶
The Countdown / Raft Building and Accusations [3, 4, 5] att 2 / ok 2
shot 3 start "James stands in the FRAME-LEFT third near the raft materials,
body angled three-quarter right, right hand holding the Audio Recorder"
shot 3 mid "James stands in the FRAME-LEFT third…" ← consistent
shot 3 end "James stands in the FRAME-LEFT third…" ← consistent
shot 5 start "James stands FRAME-RIGHT near the right third, visible from
mid-thigh up, facing left toward Nishi."
shot 5 mid "James remains FRAME-RIGHT…" ← consistent
shot 5 end "James stands at FRAME-RIGHT edge…" ← consistent
Each shot is internally solid. Across the cut James jumps
the axis. Nothing narrates a move. Rendered twice, no flag.
And the worst one, where both characters swap at once — a full axis flip:
The Countdown / Raft Building and Accusations [14, 17] att 1 / ok 1
shot 14 Shuki SCREEN-LEFT (all three frames)
Sunil SCREEN-RIGHT (all three frames)
shot 17 Shuki SCREEN-RIGHT
Sunil SCREEN-LEFT
└─ the two of them trade places across one cut
PROMPT DECIDED: Shuki=LEFT, and no placement for Sunil at all
The old codes, re-judged¶
| Old code | Verdict |
|---|---|
X1 SPATIAL_CONTINUITY_BREAK |
Confirmed real — N1b, 5.2%. Keep it. The reason it looked absent is that the prompt writer erases it before anyone can see it (§2). Add N1a (6.2%): one shot's own start/mid/end frames put the character on both sides — a shot cannot cut, so this one has no innocent reading. |
C5 ACTION_SEAM_BREAK |
Split it. The useful half is N2 (unnarrated posture change, 13.5%). "Seam text missing" as a flat check is noise. |
C4 ENVIRONMENT_BREAK |
Confirmed, and split. N8a variant drift 39.4% (mostly harmless) vs N8b real place change 10.4% (and 15 of 23 got fewer pictures than places → REF-3). |
C2 CHARACTER_APPEARANCE_BREAK |
Reframe. Not "a character looks different" — it is N7a (in frame, absent from charactersVisible, 7.8%) and REF-1 (no sheet sent, 39.2%). Duplicate sheets are a designed feature, not a defect. |
C3 PROP_STATE_BREAK |
Reframe as N4b (8.3%): a registered prop is named in the frame text and no reference was ever sent. The props[] array emptying between shots is not evidence — the array is not the object, and that check produces 34.2% noise. |
| C1 / C6 world & effect state | Retire. No field stores them. |
C7 SCALE_BREAK |
Retire. Wide → close-up is film grammar. |
C8 STYLE_BREAK |
Retire. You confirmed the team mixes render + sketch deliberately and it works. |
| — | NEW: actingInstructions performs a line that is not in dialogue — see T3 in the Tadej dissection. Deterministic: count lines, match the quoted word. |
| — | NEW: characters carries the SCENE cast where charactersVisible carries the shot cast — see T2. Any check reading characters is wrong on every shot. |
| — | NEW: an object's screen side flips across a cut (the chair in T1) — the single conflict in the Tadej sequence that would have changed the video. |
| your neck-vs-shoulder example | Not findable in the metadata. 0 cases in the frame text (my earlier 6.7% was a regex artifact). That conflict is text vs sketch — it needs the picture read, not the fields compared. It belongs in the taxonomy; it cannot be detected by the checks in this document. |
| day/night inside a sequence | 0 of 212 as a flat check. But see T4 in the Tadej dissection: the real form is night-vs-unspecified with a window visible in both, which a day/night check misses entirely. |
Real — measured on the request that was actually sent (n = 212)¶
| New code | What it is | Rate | Severity |
|---|---|---|---|
REF-1 MISSING_CHARACTER_SHEET |
character has a valid assetId on a selected shot; no sheet bound in the request |
39.2% (83) | BLOCKER when they speak or carry the shot |
REF-2 NO_ENVIRONMENT_SHEET |
no picture of the place at all | 24.1% (51) | MEDIUM |
REF-4 UNLABELLED_REFERENCES |
prompt tags fewer images than were attached | 23.1% (49) | MEDIUM |
STR-1 SHOT_BLOCK_MISMATCH |
prompt describes a different number of shots than were selected | 19.8% (42) | MEDIUM |
REF-3 PLACES_FLATTENED |
sequence moves between N places, gets fewer than N pictures | 15 of 23 location-changing sequences | BLOCKER |
STR-3 NO_SHOT_STRUCTURE |
prompt has no per-shot blocks at all | 11.3% (24) | MEDIUM |
REF-1b ZERO_CHARACTER_SHEETS |
characters on screen, characters: [] |
4.2% (9) | BLOCKER |
REF-5 NO_TAG_BLOCK |
images attached, prompt never says what they are | 3.8% (8) | MEDIUM |
REF-1c SPEAKER_NO_SHEET |
a character speaks with no reference picture | 1.9% (4) | BLOCKER |
STR-2 DUPLICATE_SHOT_LABELS |
two blocks both labelled Shot 01 |
1.4% (3) | MEDIUM |
TXT-1 WRITER_INTRODUCED_FLIP |
prompt contradicts screen direction where the raw shots did not | 0.5% (1) | MEDIUM |
1. Two methodological errors, and the corrected method¶
I got this wrong twice before arriving at the table above. Both errors are easy to repeat, so they are recorded here rather than quietly fixed.
Error 1 — the corpus was ~19× inflated by cloned test projects¶
db.storyvideos.find({'scenes.sequences': {$exists: true}})
↳ 8,840 "sequences" ← what the previous version of this document reported
↳ 160 story documents
↳ BUT sceneId 8c74019e ("Inside the Party Bubble") appears in 30 of them,
under 30 different projectIds, ALL titled "malfunction"
↳ 1,095 sceneIds are duplicated this way
De-duplicating by sceneId (newest updatedAt wins): 475 attempts, 212 distinct selections,
57 scenes, 26 stories. Every percentage in the previous version was weighted by clones of test
projects and should be discarded.
Error 2 — I measured fields the model never receives¶
I compared finalFrameInstructions of shot N against initialFrameInstructions of shot N+1. Those
are inputs to a prompt-writing LLM, not the prompt. See §2. This is the more important of the two
errors, because it invalidates the whole shape of the old taxonomy, not just its numbers.
2. The decisive finding: the contradictions get laundered, not resolved¶
There is already an LLM sitting between the shots and Seedance:
src/promptTemplates/workbenchV2/generateSequenceSeedancePrompt.template.ts
(SEEDANCE2_SYSTEM_PROMPT_v20260803), called from
src/workbench/sequenceFlow/buildSequenceUserPrompt.ts.
It takes the selected shots as a shot_breakdown bundle and emits one prompt with per-shot
Camera: / Starting Context: / Action: / Dialogue: / Performance: blocks. And that system
prompt says, verbatim:
Conflict resolution: if the shot breakdown contradicts the storyboard, the storyboard wins when a storyboard slot exists. When no storyboard is provided, trust
master_shotand asset sheets over contradictoryshot_breakdownprose for spatial composition. Do not flag, do not ask.
Measured against real data — and read the arrow carefully, because it is not good news:
RAW SHOT FIELDS THE PROMPT THAT WAS SENT
─────────────── ────────────────────────
21 selections place a character
on BOTH screen sides across the ──────▶ 0 still contradict
selected shots ▲
the writer PICKED A SIDE, 21 times,
and told nobody. Example:
The Countdown / Raft Building [14,17]
shot 14 "Shuki stands SCREEN-LEFT at mid-depth, gripping one end
of a log at chest height with both hands"
shot 17 "Shuki stands behind UV at SCREEN-RIGHT, slightly out of
focus, upright, both arms at her sides"
PROMPT DECIDED: screen-LEFT ← shot 17 lost. Nobody asked.
The Countdown / Revelation [14,16]
shot 14 "Shuki occupies SCREEN-RIGHT from upper torso up"
shot 16 "Shuki occupies FRAME-RIGHT, head and upper shoulders
visible, seated and turned three-quarter left"
shot 14 "Shuki sits at SCREEN-LEFT on the rock in the foreground"
PROMPT DECIDED: LEFT in both blocks ← 6 attempts, 4 renders
and 7 of the 21 are worse — the prompt dropped the character's placement
entirely rather than choosing, so Seedance places them freely:
"PROMPT DECIDED: (no placement for Sunil in the prompt)"
A clean prompt is not a clean sequence. The prompt has no contradiction because the writer is instructed to erase it. The creative team's intent is still ambiguous in the shots; an LLM resolved it by coin-flip and the video came back with a composition nobody signed off on. That is the failure mode, and it is invisible today at every layer: the shot fields still disagree, the prompt looks correct, the render succeeds, and no conflict is ever surfaced.
Consequence for the taxonomy. The old codes are not wrong about where to look — the shot fields are exactly right. They are wrong about when: the check has to run before the prompt writer, on the shots, and its output has to be a question for the creator, not a silent fix.
One extra case, in the other direction. In Arrival / James and Shuki Clash [4,5,6] the writer
introduced a screen-direction flip that was not in the raw shots (James screen-right in block
04, screen-left in block 06). So the writer both hides real conflicts and manufactures new ones.
Worked proof — the Chairs sequence¶
The clearest case, because it is the one I previously reported as a defect. Selection [0, 2, 3]
in Chairs / Warm Homecoming; shot 1 (the walk) not selected. In the raw shot fields, shot 0 ends
with Tadej standing in the doorway and shot 3 begins with an empty chair and then has him seated.
I called that a teleport. Here is what the request actually said:
Shot 03: 4.5s Action: Father's head tilts up two degrees to meet Tadej's gaze at screen-left. […] Tadej drops heavily into the chair at screen-left, weight falling into the seat, shoulders sagging forward, legs splaying under the table. He looks at Father with tired amusement […]
The sit-down is written out. The geometry is consistent across all three shots as well — Father screen-right, eyeline directed screen-left toward the doorway (shot 02), Tadej ending screen-left (shot 03). There is no conflict in this sequence. A cut bridging shot is a gap the writer fills, not a contradiction — which is the principle the rest of this document is calibrated on.
3. What actually reaches the model — the real conflicts¶
3.1 REF-1 — a character is on screen with no picture of them (39.2%)¶
The sharp form of the check: a character has a valid UUID assetId on one of the selected
shots, and that assetId is not in the request's characters[]. So a resolvable reference
exists in the data and simply was not attached.
Worked example — Arrival / Red Mist Attacks Ferry, selection [11, 13, 17, 19]:
request payload: characters: [] ← zero character sheets
request tags: @image1: Storyboard — visual and composition anchor.
@image2: Master shot — spatial-axis anchor.
(that is the entire tag block)
but the prompt itself contains:
Shot 11 Dialogue: Badri: "like it has a will of its own."
Shot 13 Dialogue: Shuki: "(Loud whisper) Don't!"
Shot 17 "Detail on UV's fingertip and the tendril…"
and the shots' own metadata says:
shot 13 charactersVisible: [ UV / 678d2f84-040… ] ← valid, resolvable assetId
shot 17 charactersVisible: [ UV / 678d2f84-040… ]
shot 19 charactersVisible: [ UV / 678d2f84-040… ]
▲
UV has a real character sheet. It was not attached.
The model invents UV's hand three times and has no
reason to make it the same hand twice.
Rendered successfully, twice. Nothing flagged it.
Why this is the BLOCKER class. Everything else on this list, the model can paper over. It cannot invent an identity it was never shown and keep it consistent across four shots. This is the one case where the failure is guaranteed, not probabilistic.
The sub-cases, ranked by how unambiguous they are:
| Rate | Severity | Why | |
|---|---|---|---|
REF-1c character speaks with no sheet — Granny, Stan Padnick, Badri, Shuki |
1.9% (4) | BLOCKER | a talking face with no reference is invented at full size on screen |
REF-1b characters: [] with characters on screen |
4.2% (9) | BLOCKER | nothing anchors any identity |
| REF-1 valid assetId not bound | 39.2% (83) | BLOCKER → SLIGHT | depends on whether the character is the subject or a passing element |
Not verified — and this is the biggest open question in the document: whether the creator chooses which sheets to attach in the UI. If they do, part of 39.2% is deliberate (an insert of a fingertip may not need a full character sheet) and the right product response is a warning, not a fix. If the system chooses, it is a bug. Everything about REF-1's product shape depends on this answer.
3.2 REF-3 — the sequence moves between places and gets one picture (BLOCKER)¶
This is the mirror image of C4, and more damaging. C4 assumed the risk was too many pictures of one place. The real risk is too few pictures for several places.
23 selections have shots tagged with DIFFERENT place titles.
How many environment sheets did the request carry?
0 sheets ██████ 6 ← no picture of either place
1 sheet █████████ 9 ← one room for a two-room sequence
2 sheets ███████ 7 ← correct
3 sheets █ 1 ← correct
15 of 23 (65%) got fewer pictures than they had places.
Worked example — Arrival / Arrival on Island, selection [3, 4], the second most-generated
selection in the corpus (7 attempts / 6 renders, plus a separate run of 3 / 3):
shot 3 environment = "Island Coast"
shot 4 environment = "Cliffside Camp"
request locations: [] ← zero environment sheets for either
Two different beaches, one cut between them, no picture of either. The model invents both and has no reason to make them look like the same island.
More of the same:
- malfunction / Gridlock and Goodbye [4,5] — LearnSpace + Medspace → Party Bubble, 0 sheets (4 attempts, 4 renders)
- Arrival / Arrival on Island [14,15,16,17] — Cliffside Camp → Sentinel Jungle (Ancient), 0 sheets
- malfunction / Bubble Car Breakdown [0,1,2,3,4] — Stan's House → Stan's House & Street View → Millionaire Town, 1 sheet
This also settles the point Gerard raised on comments_to_claude.md §0.6: the claim that "shot 1 is
a kitchen and shot 2 is a street cannot happen by construction" is false. It happens in 23 of 212
selections, and most of the time the model is not given a picture of the second place.
3.3 REF-2 — no picture of the place at all (24.1%)¶
51 of 212 selections carry locations: [], independent of whether the place changes. A quarter of
all sequences are generated with no environment reference. Combined with the fact that 82 of 212
(38.7%) have no storyboard either, a meaningful slice of requests reaches Seedance with no spatial
anchor of any kind.
3.4 REF-4 / REF-5 — images attached that the prompt never explains¶
The system prompt calls the opening image-tag block "non-negotiable" and requires one line per attached image with no skipped slot numbers. Measured:
- REF-4, 23.1% (49): fewer
@imageNtag lines than references attached. - REF-5, 3.8% (8): no tag block at all, while images were attached.
REF-5 turned out to be creators writing their own prompt by hand and bypassing the generator. The stored prompts begin:
"Camera: Static master wide shot, eye-level. No zoom, no pan, no cuts…"
"Use Image-1 only for storyboard action layout…" ← a different tag convention entirely
"Use Image-1 as the storyboard reference…"
"create a video sequence that includes al…" ← 126 chars, 5 shots selected, 0 renders
"The futuristic Party Bubble enters smoot…" ← 11 attempts, 10 renders
malfunction / Assignment Day: Tim's Entrance [11,12,14] carries 3 character sheets, 1
location, 3 master shots and a storyboard — and the prompt never mentions a single image. The
model receives eight pictures and no statement of which one is Tim, which is Stan, or which is the
room. It rendered anyway.
This matters for the design of any gate: a validator that assumes the templated prompt format will fail to parse ~11–15% of real requests.
3.5 STR-1 / STR-2 / STR-3 — the prompt does not describe the shots that were selected¶
| Rate | What | |
|---|---|---|
| STR-1 | 19.8% (42) | shot-block count ≠ selected-shot count. malfunction / Assignment Day [11,12,14]: 1 block for 4 selected shots. Arrival / UV's Nightmare Vision [0,1,2,3]: 3 blocks for 4. malfunction / Inside the Party Bubble [1,2]: 3 blocks for 2. |
| STR-2 | 1.4% (3) | duplicate labels — block labels come out as [1, 1, 2] |
| STR-3 | 11.3% (24) | no Shot NN: blocks at all |
Worked example of STR-2 — test-4 / Dawn Warning: The Loop [0,2]. Two shots selected. The
prompt contains three blocks, two of them both labelled Shot 01:
Shot 01: 2.5s Meera foreground, Aarav at the console, over-shoulder
Shot 01: 6s camera glides through the exterior doorway into the control room ← same label
Shot 02: 5s Aarav motionless in the chair, dolly out
Generated 5 times, 5 renders. This is also what produced a false positive in my own first pass: I read it as "Aarav appears in shot 2 with no narration" when the real defect was the duplicated block.
4. Calibrating severity¶
The calibration question: a missing transition can be a blocker, a medium miss, or something Seedance absorbs without anyone noticing. The discriminator is what the model has to invent, and the Chairs sequence supplies the signal.
SLIGHT — the prompt narrates the change. The model interpolates.
"Tadej DROPS HEAVILY INTO THE CHAIR at screen-left, weight falling
into the seat, legs splaying under the table"
A bridging shot was cut and it does not matter.
MEDIUM — the change is large and the model must guess how it happened, but
every element involved has a reference picture.
A prop changes hands off-screen. A character crosses the room.
Wrong in the details, right in substance.
BLOCKER — the model must invent something it was never shown, and no amount
of text fixes it.
· a character on screen with no character sheet (REF-1)
· a second location with no environment sheet (REF-3)
· a state change with no cause anywhere in the selection:
The Countdown / Raft Building [3,10,12]
shot 10 "both BARE PALMS angle toward camera"
shot 11 props: [Spiral Ash Pattern] ← NOT SELECTED
shot 12 "palms facing up with ASH SPIRALS clearly visible"
Not interpolation. Invention of a cause.
So severity is a product of three things, per seam:
severity = size of the state change
× does any shot's text acknowledge it ← the Chairs "drops heavily" test
× does a reference picture exist for it ← REF-1 / REF-3
Only the third factor is currently checkable with confidence, because the second requires parsing the generated prompt and 11% of prompts have no per-shot structure at all (STR-3).
Retracted: I measured "element appears in shot N+1, absent in shot N, arrival not narrated" at 25.9%. Hand-checking showed it was dominated by the duplicate-shot-block defect (STR-2), not by real unnarrated arrivals. That number is withdrawn. The check is still the right idea; STR-1 and STR-2 have to be fixed before it can be measured.
5. Honest negatives¶
| Checked | Result |
|---|---|
| Day/night contradiction inside a sequence | 0 of 212. It exists at scene level (The Countdown shot 9 "night setting" vs shot 17 "muted daylight") but never inside one generated sequence. Do not build this gate. |
| Do any of these checks predict creator dissatisfaction? | No. Baseline is 1.87 successful renders per selection. Every check sits within ±0.6 of baseline. STR-2 reaches 3.67 but n=3. No detector in this document is validated against an outcome. |
| Are generation failures caused by continuity? | No. 178 failures: 103 content-moderation flags ("flagged as sensitive, E005"), 18 copyright rejections, 23 Byteplus API errors, plus S3 timeouts and one bad API key. Zero relate to cross-shot continuity. |
Is isSelected a quality label? |
No. True on all 212, and up to 5 videos per selection. It does not mean "the creator chose this one". |
isFavorite |
Present on 2 selections. Too rare to use. |
| Cross-scene shot references | 0. Selections are always scene-local. |
| Duplicate character sheets (old C2) | Not a defect. The builder deliberately labels them "(sheet 1 of 2)" — buildAssetOccurrenceLabels() in buildSequenceUserPrompt.ts. Tim Padnick legitimately has 3. |
The 24× retry loop (test-4 / Aarav Alone [0,1,2,8]) |
Not dissatisfaction. 24 attempts, 0 renders — a technical failure loop, not a creator re-rolling. |
6. What to change in the system¶
In priority order, based on the measurements above.
- Move the gate from the shot fields to the request bundle. The bundle is assembled in
src/workbench/sequenceFlow/and is fully deterministic — no LLM is needed for REF-1, REF-2, REF-3, REF-4, REF-5, STR-1, STR-2 or STR-3. That is eight of the eleven real checks for the cost of a set comparison. - Answer the open question in §3.1 first. If the creator picks the reference sheets, REF-1 is a warning ("UV is on screen in 3 of these shots and has no reference — attach it?"). If the system picks them, it is a bug. This decides REF-1's entire product shape, and REF-1 is the biggest number here.
- REF-3 is the most under-appreciated finding. 65% of location-changing sequences get fewer
environment pictures than they have locations, and 6 of them get none. It is a one-line check:
distinct
environment.titleacross the selected shots vslocations[].length. - Fix STR-1 / STR-2 before building any text-level check. 19.8% of prompts do not describe the selected shots, and no continuity check can be trusted on a prompt whose shot blocks are miscounted or duplicated.
- Do not build C1, C6, C7, C8, or a day/night check. Nothing in the schema supports the first two; the third and fourth are normal craft; the fifth has zero occurrences.
- Keep exactly one text-level check: TXT-1. The prompt writer introduced a screen-direction
contradiction that was not in the source data (
Arrival / James and Shuki Clash [4,5,6]). Validating the writer's own output is a narrower and far more defensible job than second-guessing the shot fields. - Any validator must tolerate hand-written prompts. ~11% have no per-shot structure and 3.8% have no tag block, because creators bypass the generator.
7. Not verified, and where the method stops¶
- No output video was watched. Every finding here is a property of the request. Whether
Seedance produced a visible glitch is unknown. The videos are in the data (
videoURLon every completed attempt) and this is the obvious next study — it is the only thing that turns these rates into real severities. - No detector is validated against an outcome (§5). Read every rate as "how often the request has this property", never as "how often the creator got a bad video".
- The severity tiers in §4 are reasoned, not measured. The three-factor rule is a proposal.
- Whether reference attachment is a creator choice is unknown (§3.1). Biggest single caveat.
- Screen-direction detection is text-based. It requires a placement verb (
stands,sits,occupies, …) followed byscreen-left/screen-rightwith no eyeline word in between. An earlier blacklist version over-reported by 2.4× because"eyes still fixed frame-right"reads as a body position. The whitelist version is used throughout, and it will under-report prompts that express position without those verbs. - REF-1 at 39.2% is the broadest number here. The tight, unambiguous subsets are REF-1c (a
speaking character with no sheet, 1.9%) and REF-1b (
characters: [], 4.2%). - Corpus concentration: 164 of 212 selections come from three shows —
The Countdown(66),Arrival(58),malfunction(40).test-4andEnlarged_devare test shows and are included. A single show's habits can move any rate in this document. storyboardis often a manual upload. In the Chairs request it islabel: "Upload from device". The prompt gives the storyboard authority over "framing, staging, blocking, screen direction, and shot-to-shot composition" — so in 130 of 212 selections, composition is anchored on an image the pipeline did not generate and cannot inspect.