Skip to content

20 more sequences — second inspection pass

A second, independent read of 20 production sequences, chosen so that none of them appear in the first 20. Purpose: find patterns the first pass missed, and test the rules the taxonomy already claims.


Method

  • Selection: all multi-shot sequence sends ranked by videos[].createdAt descending, then ranks 1, 20, 21, 22, 24, 25, 27, 28, 33, 34, 35, 38, 39, 40, 42, 43, 44, 45, 46, 50 taken. Six scenes in this batch had never been sampled before: Ash Spirals and Shadows, Panic at the Cliffside, Sentinels Close the Net, Red Mist Retreats, Knocking Signals Group Flees, Footprints and UV Found.
  • Zero overlap with the first 20, checked shot set by shot set.
  • Size: 20 sequences, 63 shots, 43 shot-to-shot transitions. Lengths: 1 × two shots, 15 × three, 4 × four.
  • Read per shot: description, actingInstructions, initialFrameInstructions, middleFrameInstructions, finalFrameInstructions, shotSize(+details), cameraAngle(+details), cameraMovement(+details), charactersVisible, props, environment + variant, shotJob, timing, dialogue, attached images and sketches, sketchAnalyses, and preGenValidation.
  • Also read: the raw reference bundle actually sent — videoParams.jobMetaData.sequence.{characters,locations,props,duration,selectedShots}.
  • Read-only. No writes. Scripts in /tmp/i2v/, combined dump /tmp/i2v/all20.out (2,625 lines).

Limitation, stated up front

19 of the 20 are from one show, The Countdown. The first 20 spanned five stories. So this pass is deeper but narrower: it is strong evidence about how one team works, weaker evidence about the product as a whole. Nobody watched the output videos here either — every judgement below is about the inputs.


Finding 1 — the discriminator for "person in the text but not in the cast list"

This was already in the taxonomy as PERSON_IN_TEXT_NOT_IN_CAST with one example. It is now the most frequent real defect in the sample: 6 shots across 6 of the 20 sequences (30%).

More importantly, this pass found the rule that separates a real miss from correct behaviour, and it is not "is a character name in the text".

charactersVisible Frame text Verdict
R21 shot 7 (none) "Shuki's eyes and brow fill the frame, centered slightly left" defect
R28 shot 15 (none) "James stands frame-left, framed waist-up, leaning in toward Shuki" defect
R39 shot 16 (none) "Badri's hand is just entering from frame-left edge" defect
R40 shot 13 (none) "Shuki's eyes are centered symmetrically in frame" defect
R1 shot 15 Becca Malcolm speaks a line, no frame position given, "conference room wall fills the background" correct
R44 shot 0 (none) "No characters are visible." correct
R45 shot 2 (none) "Two small human figures … dwarfed by the environment" — nobody named correct

The rule: flag it only when the frame text gives the person a position in frame. Words like fills the frame, frame-left, centered, entering from the edge, framed waist-up. A person who is merely spoken to, spoken by, or reached toward is off-screen and legitimate.

Also true and useful as a cheap second signal: every defect above has shotJob ∈ {reaction_emotion, interaction_relationship, action_progression}, and every correct case has shotJob ∈ {cinematic_anchor, transition, orientation_context, symbolic}. A shot whose job is reaction with nobody in the cast is suspect by construction.

Why it matters: charactersVisible is what puts a character reference image in the request. R21 shot 7 sends an extreme close-up of Shuki's eyes with no picture of Shuki. Shot 2 of the same sequence does send her reference. So one video contains her established face and then her face invented from prose.

Worst single case — R28 shot 15. James speaks a line, both characters are described with full frame geometry, and the shot sends: no characters, no image, no sketch. Pure text.

Consequence for ordering

This defect manufactures false "flicker" alarms. R24 reads Shuki || Shuki || (none) and R40 reads Shuki|… || Shuki || (none) — both look like "character disappears", both are really this bug. So:

  run PERSON_IN_TEXT_NOT_IN_CAST first
        │
        ├─ found?  →  fix the cast list, THEN re-evaluate the timeline
        └─ clean?  →  now a present→absent→present pattern means something

Finding 2 — the reference labels sent to the model degrade over time

New category, deterministic, free to check, and invisible until you read the raw request.

The sent bundle is { id, imageUrl, label, description? } per reference, where id = "<assetId>:<variantId>" (trailing colon = no variant chosen). label is what names the image.

Same physical object, same scene (Sentinels Close the Net), every send in order:

Sent at id label
2026-08-20 09:12 ffe9fe18…: Hollow Wooden Blocks
2026-08-20 09:56 ffe9fe18…: Upload from device
2026-08-21 09:42 ffe9fe18…: Upload from device
2026-08-21 11:34 hollowed-wooden-blocks: Upload from device (real name only in description)
2026-08-26 23:38 asset-1787779559822-0 Upload from device
2026-09-16 18:28 → 18:42 asset-1787779559822-0 Upload from device

Two separate things went wrong:

  1. The label became a UI string. Upload from device is the button text. The model is handed a picture captioned with the name of the button someone clicked. R43's send (2026-08-25, sent=[3,5,7,2]) is SENT props: Upload from device — one prop, no name.
  2. The object drifted across three different asset identities — a UUID asset, a slug asset, then an ad-hoc upload with no registry entry at all.

Across every sequence send in the collection, reference IDs break down as:

  characters   47,654  uuid + trailing colon (no variant)
               10,927  uuid:uuid (variant chosen)
                6,295  asset-<timestamp>-<n>   ← ad-hoc upload, no asset identity
                   24  slug, e.g. "Nishi:"
  locations    10,196  uuid:uuid
                3,656  uuid + trailing colon
                3,952  asset-<timestamp>-<n>   ← ad-hoc upload
  props         8,250  uuid + trailing colon
                2,998  uuid:uuid
                3,659  asset-<timestamp>-<n>   ← ad-hoc upload
                   23  slug, e.g. "hollowed-wooden-blocks:"

asset-<timestamp>-<n> references cannot have a variant and cannot be matched across shots, because they are not assets. That is ~14,000 references.


Finding 3 — the same object exists as several assets in one scene

Direct query of scenes[].props[] in The Countdown, distinct assetId per row:

Scene Prop title assetId
Raft Building and Accusations Dry Logs aa1bea7d…
Raft Building and Accusations Dry Timber Logs e7c88306…
Sentinels Close the Net Hollow Wooden Blocks ffe9fe18…
Sentinels Close the Net Hollowed Wooden Blocks hollowed-…
Sentinels Close the Net Wooden Signal Blocks wooden-s…
Panic at the Cliffside Phone 3ddde812…
Panic at the Cliffside Shuki's phone 5382a522…
Return to Cliffside and Accusations Supply Crate supply-c…
Return to Cliffside and Accusations Nishi's Supply Crate fa1dcd45…

Two ID shapes — UUIDs and slugified titles — which is the mechanism: two creation paths, so the same object gets registered twice under slightly different names.

It reaches the user's video. R43 is a 4-shot sequence where shot 3 declares Wooden Signal Blocks and shot 2 declares Hollow Wooden Blocks. Two assets, one object, one sequence.

Also measured: every prop asset in this story has images: []. Props contribute a title and nothing visual, unless one is attached ad-hoc as in Finding 2.


Finding 4 — one shot carries the visual, the rest carry a sketch or nothing

The strongest structural fact in the batch.

  sequences where exactly ONE shot has a selected reference image ......... 15 of 20
  shots with no image AND no sketch at all ................................  2 of 63   (R28 sh15, R44 sh0)
  sequences with zero sketchAnalyses on every shot ........................  7 of 20

The near-universal shape:

                   shot A          shot B          shot C
  image         12img(1sel)      0img(0sel)      0img(0sel)
  sketch         2sk(1sel)        1sk(1sel)       1sk(1sel)
  env variant    v=9c6e1ad6         v=-             v=-
  props          listed            (none)          (none)
  preGenValid    pass/0c             -               -

Two things follow.

The sketch is the visual input for most shots. So the ten sketch-vs-text codes already in issueCatalog.ts are the main visual check at sequence level too — but 7 of 20 sequences have no sketch analysis at all, so they run blind.

The shot that carries the metadata is not always the first one sent. The environment variant sits on the first sent shot in 13 of 20, on the middle shot in R38, on the last sent shot in R1 and R43, on two shots in R42 and R44, and nowhere at all in R46. Any check that reads "the sequence's environment" from shot 1 will read - in 7 of 20 cases.


Finding 5 — no location reference sent, while the shots name one

4 of 20 send locations: (NONE) even though the shots declare an environment with a variant:

shots declare sent
R33 Cliffside Camp v=eb5d6014 (NONE)
R35 Cliffside Camp v=3e5f81e4 (NONE)
R44 Deep Forest + Cliffside Camp v=7c04739d (NONE)
R45 Cliffside Camp v=52f1ada5 on both shots (NONE)

The place is then invented from prose, in a show whose whole point is that the cliffside camp looks the same every time.

And R42 sends only one of two chosen variants: shots declare 28afa663 || - || 27504ad8 || -, the request carries 28afa663. Shot 2's choice is dropped silently.


Finding 6 — an environment tag that contradicts its own frame text

R44 shot 4. environment: Cliffside Camp v=7c04739d, but the frame text puts the camera in the forest looking at the camp across distance, with dark tree trunks bordering frame-left and frame-right.

The sequence reads Deep Forest || Cliffside Camp || Deep Forest, which looks like a place jump and is not one. It is a mislabelled tag on a single shot, and the damage is that it pulls the wrong location reference.

This is worth separating from PLACE_CHANGES_WITHOUT_TRAVEL, because the fix is different: change the tag, not the shot list.


Finding 7 — video length almost never matches the shots

Σtiming of the sent shots vs the duration in the request. 20 of 20 disagree.

  video LONGER than the content ......... 15 of 20   (worst: R45, 6.5 s of shots in a 15 s video)
  video SHORTER than the content ........  5 of 20   (worst: R1, 15 s of shots in an 8 s video)
  exact match ...........................  0 of 20
Direction Sequences
longer R21 10/6.8 · R22 15/13.5 · R25 12/9.6 · R27 12/11.0 · R28 15/11.7 · R33 14/13.4 · R35 12/9.2 · R39 10/9.0 · R40 12/10.6 · R42 14/9.5 · R43 15/13.9 · R44 12/8.0 · R45 15/6.5 · R46 15/13.3 · R50 14/8.1
shorter R1 8/15 · R20 8/8.5 · R24 6/7.6 · R34 10/12.6 · R38 15/18.4

The first pass called this "dialogue longer than the video" and found it 3 times. It is broader than that and it is the rule, not the exception. Direction matters:

  • shorter → the described action is cut off or rushed.
  • longer → dead time the shots never describe, which the model fills however it likes.

Not the user's bug to fix by editing a shot, so it belongs with the engineering items — but it is the highest-frequency finding in the whole sample.


Finding 8 — patterns the first pass listed as "not seen", now seen

Pattern Evidence
prop appears mid-sequence R24 (none) \|\| Shuki's Camera \|\| (none) · R27 Shuki's Camera \|\| Shuki's Camera \|\| Rope Coil · R34 Dry Timber Logs\|Audio Recorder \|\| Rope Coil \|\| (none)
person appears with no entrance R39 Shuki\|Nishi\|James\|Nawaz\|Bituin \|\| (none) \|\| Shuki\|Badri\|Nishi\|Sunil — Badri and Sunil arrive in the last shot of a 10 s video. R40 the same shape
middle frame missing on one shot R1 shot 18 MID: (empty) while shots 15 and 11 have it. R44 shot 6 the same
shots sent out of story order R25 [2,6,5,7] · R43 [3,5,7,2] · R1 [15,18,11] — 3 of 20

Note on the prop rows: because prop assets carry no image (Finding 3), a prop appearing mid-sequence changes only the text. Low severity on its own; it matters when the prop is the thing the action is about.


Finding 9 — per-shot validation is not a gate, and cannot be built on

Sequences where every shot has a validation status 0 of 20
Typical shape one shot has a status, the rest are - (R43 - \|\| - \|\| - \|\| pass/5c)
Sequences with no status on any shot R44, R46, R50
status: pass with conflicts attached R39 pass/5c · R43 pass/5c · R35 pass/3c · R42 pass/3c
Rendered anyway with a blocked status R20 (blocked/1c on shot 5) · R42 (blocked/1c on shot 2)

So preGenValidation today is advisory, sparse, and per-shot. The sequence check cannot read it as a precondition — it has to compute its own view, and if a sequence-level status is wanted it should inherit the worst status among the shots it actually has.


Finding 10 — things that look wrong and are not

Negative evidence, to keep the detector quiet.

Looks like Actually Evidence
cast list collapsing shot to shot normal coverage: establishing wide → singles R35 8→2→1 · R42 6→2→1→1 · R33 5→2→2
camera angle whiplash coverage and style, deliberate R20 LOW → GROUND → BIRDS_EYE in 8 s · R43 TOP → EYE → LOW → TOP
dialogue speaker not in the cast list off-screen speech, correct R1 shot 15, Malcolm speaks, Becca is on screen
empty cast on an environment shot correct, and the text says so R44 shot 0 "No characters are visible." · R45 shot 2 unnamed figures for scale
two locations named in one sequence one mislabelled tag, not a jump R44, Finding 6

What this pass changes in the taxonomy

Change Reason
PERSON_IN_TEXT_NOT_IN_CAST gets the in-frame-position rule and must run before any timeline check Finding 1
New group: what actually gets sent as a reference — REFERENCE_LABEL_IS_PLACEHOLDER, SAME_OBJECT_SEVERAL_ASSETS, NO_LOCATION_REFERENCE_SENT, CHOSEN_VARIANT_NOT_SENT Findings 2, 3, 5
New: ENVIRONMENT_TAG_CONTRADICTS_FRAME_TEXT, separate from PLACE_CHANGES_WITHOUT_TRAVEL Finding 6
New: SHOT_HAS_NO_VISUAL Finding 4
OBJECT_APPEARS_IN_HAND and PERSON_APPEARS_UNEXPLAINED move from "not seen" to observed Finding 8
Duration mismatch re-stated as 20 of 20, both directions Finding 7
Do not read the sequence's environment or props from the first shot Finding 4
Do not treat preGenValidation as a precondition Finding 9
Do not flag collapsing cast lists, angle changes, or off-screen speakers Finding 10