20 more sequences — second inspection pass¶
A second, independent read of 20 production sequences, chosen so that none of them appear in the first 20. Purpose: find patterns the first pass missed, and test the rules the taxonomy already claims.
Method¶
- Selection: all multi-shot sequence sends ranked by
videos[].createdAtdescending, then ranks 1, 20, 21, 22, 24, 25, 27, 28, 33, 34, 35, 38, 39, 40, 42, 43, 44, 45, 46, 50 taken. Six scenes in this batch had never been sampled before: Ash Spirals and Shadows, Panic at the Cliffside, Sentinels Close the Net, Red Mist Retreats, Knocking Signals Group Flees, Footprints and UV Found. - Zero overlap with the first 20, checked shot set by shot set.
- Size: 20 sequences, 63 shots, 43 shot-to-shot transitions. Lengths: 1 × two shots, 15 × three, 4 × four.
- Read per shot:
description,actingInstructions,initialFrameInstructions,middleFrameInstructions,finalFrameInstructions,shotSize(+details),cameraAngle(+details),cameraMovement(+details),charactersVisible,props,environment+ variant,shotJob,timing, dialogue, attached images and sketches,sketchAnalyses, andpreGenValidation. - Also read: the raw reference bundle actually sent —
videoParams.jobMetaData.sequence.{characters,locations,props,duration,selectedShots}. - Read-only. No writes. Scripts in
/tmp/i2v/, combined dump/tmp/i2v/all20.out(2,625 lines).
Limitation, stated up front¶
19 of the 20 are from one show, The Countdown. The first 20 spanned five stories. So this pass is deeper but narrower: it is strong evidence about how one team works, weaker evidence about the product as a whole. Nobody watched the output videos here either — every judgement below is about the inputs.
Finding 1 — the discriminator for "person in the text but not in the cast list"¶
This was already in the taxonomy as PERSON_IN_TEXT_NOT_IN_CAST with one example. It is now the
most frequent real defect in the sample: 6 shots across 6 of the 20 sequences (30%).
More importantly, this pass found the rule that separates a real miss from correct behaviour, and it is not "is a character name in the text".
charactersVisible |
Frame text | Verdict | |
|---|---|---|---|
| R21 shot 7 | (none) |
"Shuki's eyes and brow fill the frame, centered slightly left" | defect |
| R28 shot 15 | (none) |
"James stands frame-left, framed waist-up, leaning in toward Shuki" | defect |
| R39 shot 16 | (none) |
"Badri's hand is just entering from frame-left edge" | defect |
| R40 shot 13 | (none) |
"Shuki's eyes are centered symmetrically in frame" | defect |
| R1 shot 15 | Becca |
Malcolm speaks a line, no frame position given, "conference room wall fills the background" | correct |
| R44 shot 0 | (none) |
"No characters are visible." | correct |
| R45 shot 2 | (none) |
"Two small human figures … dwarfed by the environment" — nobody named | correct |
The rule: flag it only when the frame text gives the person a position in frame. Words like fills the frame, frame-left, centered, entering from the edge, framed waist-up. A person who is merely spoken to, spoken by, or reached toward is off-screen and legitimate.
Also true and useful as a cheap second signal: every defect above has
shotJob ∈ {reaction_emotion, interaction_relationship, action_progression}, and every correct case has
shotJob ∈ {cinematic_anchor, transition, orientation_context, symbolic}. A shot whose job is reaction
with nobody in the cast is suspect by construction.
Why it matters: charactersVisible is what puts a character reference image in the request. R21 shot 7
sends an extreme close-up of Shuki's eyes with no picture of Shuki. Shot 2 of the same sequence does
send her reference. So one video contains her established face and then her face invented from prose.
Worst single case — R28 shot 15. James speaks a line, both characters are described with full frame geometry, and the shot sends: no characters, no image, no sketch. Pure text.
Consequence for ordering¶
This defect manufactures false "flicker" alarms. R24 reads Shuki || Shuki || (none) and R40 reads
Shuki|… || Shuki || (none) — both look like "character disappears", both are really this bug. So:
run PERSON_IN_TEXT_NOT_IN_CAST first
│
├─ found? → fix the cast list, THEN re-evaluate the timeline
└─ clean? → now a present→absent→present pattern means something
Finding 2 — the reference labels sent to the model degrade over time¶
New category, deterministic, free to check, and invisible until you read the raw request.
The sent bundle is { id, imageUrl, label, description? } per reference, where
id = "<assetId>:<variantId>" (trailing colon = no variant chosen). label is what names the image.
Same physical object, same scene (Sentinels Close the Net), every send in order:
| Sent at | id |
label |
|---|---|---|
| 2026-08-20 09:12 | ffe9fe18…: |
Hollow Wooden Blocks |
| 2026-08-20 09:56 | ffe9fe18…: |
Upload from device |
| 2026-08-21 09:42 | ffe9fe18…: |
Upload from device |
| 2026-08-21 11:34 | hollowed-wooden-blocks: |
Upload from device (real name only in description) |
| 2026-08-26 23:38 | asset-1787779559822-0 |
Upload from device |
| 2026-09-16 18:28 → 18:42 | asset-1787779559822-0 |
Upload from device |
Two separate things went wrong:
- The label became a UI string.
Upload from deviceis the button text. The model is handed a picture captioned with the name of the button someone clicked. R43's send (2026-08-25,sent=[3,5,7,2]) isSENT props: Upload from device— one prop, no name. - The object drifted across three different asset identities — a UUID asset, a slug asset, then an ad-hoc upload with no registry entry at all.
Across every sequence send in the collection, reference IDs break down as:
characters 47,654 uuid + trailing colon (no variant)
10,927 uuid:uuid (variant chosen)
6,295 asset-<timestamp>-<n> ← ad-hoc upload, no asset identity
24 slug, e.g. "Nishi:"
locations 10,196 uuid:uuid
3,656 uuid + trailing colon
3,952 asset-<timestamp>-<n> ← ad-hoc upload
props 8,250 uuid + trailing colon
2,998 uuid:uuid
3,659 asset-<timestamp>-<n> ← ad-hoc upload
23 slug, e.g. "hollowed-wooden-blocks:"
asset-<timestamp>-<n> references cannot have a variant and cannot be matched across shots, because they
are not assets. That is ~14,000 references.
Finding 3 — the same object exists as several assets in one scene¶
Direct query of scenes[].props[] in The Countdown, distinct assetId per row:
| Scene | Prop title | assetId |
|---|---|---|
| Raft Building and Accusations | Dry Logs |
aa1bea7d… |
| Raft Building and Accusations | Dry Timber Logs |
e7c88306… |
| Sentinels Close the Net | Hollow Wooden Blocks |
ffe9fe18… |
| Sentinels Close the Net | Hollowed Wooden Blocks |
hollowed-… |
| Sentinels Close the Net | Wooden Signal Blocks |
wooden-s… |
| Panic at the Cliffside | Phone |
3ddde812… |
| Panic at the Cliffside | Shuki's phone |
5382a522… |
| Return to Cliffside and Accusations | Supply Crate |
supply-c… |
| Return to Cliffside and Accusations | Nishi's Supply Crate |
fa1dcd45… |
Two ID shapes — UUIDs and slugified titles — which is the mechanism: two creation paths, so the same object gets registered twice under slightly different names.
It reaches the user's video. R43 is a 4-shot sequence where shot 3 declares Wooden Signal Blocks and
shot 2 declares Hollow Wooden Blocks. Two assets, one object, one sequence.
Also measured: every prop asset in this story has images: []. Props contribute a title and nothing
visual, unless one is attached ad-hoc as in Finding 2.
Finding 4 — one shot carries the visual, the rest carry a sketch or nothing¶
The strongest structural fact in the batch.
sequences where exactly ONE shot has a selected reference image ......... 15 of 20
shots with no image AND no sketch at all ................................ 2 of 63 (R28 sh15, R44 sh0)
sequences with zero sketchAnalyses on every shot ........................ 7 of 20
The near-universal shape:
shot A shot B shot C
image 12img(1sel) 0img(0sel) 0img(0sel)
sketch 2sk(1sel) 1sk(1sel) 1sk(1sel)
env variant v=9c6e1ad6 v=- v=-
props listed (none) (none)
preGenValid pass/0c - -
Two things follow.
The sketch is the visual input for most shots. So the ten sketch-vs-text codes already in
issueCatalog.ts are the main visual check at sequence level too — but 7 of 20 sequences have no sketch
analysis at all, so they run blind.
The shot that carries the metadata is not always the first one sent. The environment variant sits on
the first sent shot in 13 of 20, on the middle shot in R38, on the last sent shot in R1 and R43, on
two shots in R42 and R44, and nowhere at all in R46. Any check that reads "the sequence's environment"
from shot 1 will read - in 7 of 20 cases.
Finding 5 — no location reference sent, while the shots name one¶
4 of 20 send locations: (NONE) even though the shots declare an environment with a variant:
| shots declare | sent | |
|---|---|---|
| R33 | Cliffside Camp v=eb5d6014 |
(NONE) |
| R35 | Cliffside Camp v=3e5f81e4 |
(NONE) |
| R44 | Deep Forest + Cliffside Camp v=7c04739d |
(NONE) |
| R45 | Cliffside Camp v=52f1ada5 on both shots |
(NONE) |
The place is then invented from prose, in a show whose whole point is that the cliffside camp looks the same every time.
And R42 sends only one of two chosen variants: shots declare
28afa663 || - || 27504ad8 || -, the request carries 28afa663. Shot 2's choice is dropped silently.
Finding 6 — an environment tag that contradicts its own frame text¶
R44 shot 4. environment: Cliffside Camp v=7c04739d, but the frame text puts the camera in the forest
looking at the camp across distance, with dark tree trunks bordering frame-left and frame-right.
The sequence reads Deep Forest || Cliffside Camp || Deep Forest, which looks like a place jump and is
not one. It is a mislabelled tag on a single shot, and the damage is that it pulls the wrong location
reference.
This is worth separating from PLACE_CHANGES_WITHOUT_TRAVEL, because the fix is different: change the tag,
not the shot list.
Finding 7 — video length almost never matches the shots¶
Σtiming of the sent shots vs the duration in the request. 20 of 20 disagree.
video LONGER than the content ......... 15 of 20 (worst: R45, 6.5 s of shots in a 15 s video)
video SHORTER than the content ........ 5 of 20 (worst: R1, 15 s of shots in an 8 s video)
exact match ........................... 0 of 20
| Direction | Sequences |
|---|---|
| longer | R21 10/6.8 · R22 15/13.5 · R25 12/9.6 · R27 12/11.0 · R28 15/11.7 · R33 14/13.4 · R35 12/9.2 · R39 10/9.0 · R40 12/10.6 · R42 14/9.5 · R43 15/13.9 · R44 12/8.0 · R45 15/6.5 · R46 15/13.3 · R50 14/8.1 |
| shorter | R1 8/15 · R20 8/8.5 · R24 6/7.6 · R34 10/12.6 · R38 15/18.4 |
The first pass called this "dialogue longer than the video" and found it 3 times. It is broader than that and it is the rule, not the exception. Direction matters:
- shorter → the described action is cut off or rushed.
- longer → dead time the shots never describe, which the model fills however it likes.
Not the user's bug to fix by editing a shot, so it belongs with the engineering items — but it is the highest-frequency finding in the whole sample.
Finding 8 — patterns the first pass listed as "not seen", now seen¶
| Pattern | Evidence |
|---|---|
| prop appears mid-sequence | R24 (none) \|\| Shuki's Camera \|\| (none) · R27 Shuki's Camera \|\| Shuki's Camera \|\| Rope Coil · R34 Dry Timber Logs\|Audio Recorder \|\| Rope Coil \|\| (none) |
| person appears with no entrance | R39 Shuki\|Nishi\|James\|Nawaz\|Bituin \|\| (none) \|\| Shuki\|Badri\|Nishi\|Sunil — Badri and Sunil arrive in the last shot of a 10 s video. R40 the same shape |
| middle frame missing on one shot | R1 shot 18 MID: (empty) while shots 15 and 11 have it. R44 shot 6 the same |
| shots sent out of story order | R25 [2,6,5,7] · R43 [3,5,7,2] · R1 [15,18,11] — 3 of 20 |
Note on the prop rows: because prop assets carry no image (Finding 3), a prop appearing mid-sequence changes only the text. Low severity on its own; it matters when the prop is the thing the action is about.
Finding 9 — per-shot validation is not a gate, and cannot be built on¶
| Sequences where every shot has a validation status | 0 of 20 |
| Typical shape | one shot has a status, the rest are - (R43 - \|\| - \|\| - \|\| pass/5c) |
| Sequences with no status on any shot | R44, R46, R50 |
status: pass with conflicts attached |
R39 pass/5c · R43 pass/5c · R35 pass/3c · R42 pass/3c |
Rendered anyway with a blocked status |
R20 (blocked/1c on shot 5) · R42 (blocked/1c on shot 2) |
So preGenValidation today is advisory, sparse, and per-shot. The sequence check cannot read it as a
precondition — it has to compute its own view, and if a sequence-level status is wanted it should
inherit the worst status among the shots it actually has.
Finding 10 — things that look wrong and are not¶
Negative evidence, to keep the detector quiet.
| Looks like | Actually | Evidence |
|---|---|---|
| cast list collapsing shot to shot | normal coverage: establishing wide → singles | R35 8→2→1 · R42 6→2→1→1 · R33 5→2→2 |
| camera angle whiplash | coverage and style, deliberate | R20 LOW → GROUND → BIRDS_EYE in 8 s · R43 TOP → EYE → LOW → TOP |
| dialogue speaker not in the cast list | off-screen speech, correct | R1 shot 15, Malcolm speaks, Becca is on screen |
| empty cast on an environment shot | correct, and the text says so | R44 shot 0 "No characters are visible." · R45 shot 2 unnamed figures for scale |
| two locations named in one sequence | one mislabelled tag, not a jump | R44, Finding 6 |
What this pass changes in the taxonomy¶
| Change | Reason |
|---|---|
PERSON_IN_TEXT_NOT_IN_CAST gets the in-frame-position rule and must run before any timeline check |
Finding 1 |
New group: what actually gets sent as a reference — REFERENCE_LABEL_IS_PLACEHOLDER, SAME_OBJECT_SEVERAL_ASSETS, NO_LOCATION_REFERENCE_SENT, CHOSEN_VARIANT_NOT_SENT |
Findings 2, 3, 5 |
New: ENVIRONMENT_TAG_CONTRADICTS_FRAME_TEXT, separate from PLACE_CHANGES_WITHOUT_TRAVEL |
Finding 6 |
New: SHOT_HAS_NO_VISUAL |
Finding 4 |
OBJECT_APPEARS_IN_HAND and PERSON_APPEARS_UNEXPLAINED move from "not seen" to observed |
Finding 8 |
| Duration mismatch re-stated as 20 of 20, both directions | Finding 7 |
| Do not read the sequence's environment or props from the first shot | Finding 4 |
Do not treat preGenValidation as a precondition |
Finding 9 |
| Do not flag collapsing cast lists, angle changes, or off-screen speakers | Finding 10 |