Skip to content

AIDA shot-write map — what writes shot text, and what breaks

Investigation for the AIDA update (shot breakdown, modify & auto-complete shot detail routines). The investigation below was read from code on branch f-t2ibenchmarking-system (PR #677). File:line given for every claim. §8 records what this branch actually changed — it is based on origin/staging (b7415cca), not on #677.

TL;DR — three defects, in severity order

  1. middleFrameInstructions is unwritable. Zero code outside src/workbench/preGenValidation/ ever writes it. It is absent from every allowlist and every prompt. Measured empty on 16/16 shots.
  2. The AI chat path (modify_shot) bypasses cascade entirely. Editing a shot through the scene guru changes one field and leaves the rest stale. Editing the same shot through the REST API reconciles it. Same user-visible action, two different behaviours.
  3. The three frames are treated as independent text boxes, not as one temporal unit. No block asserts that start → middle → end progresses coherently. Two prompts actively encourage degenerate output (see §4).

The reconciliation spec Miki wants already exists, fully written, in src/promptTemplates/preGenValidation/temporalFieldEditor.prompt.ts. It is used only to detect.


1. Every block that writes shot text

# Block Entry point Writes via Frame fields it can write
1 Shot breakdown generateShotsFromPlan.template.ts scene guru → updateSceneFields initial, final
2 Auto-complete shotAutocomplete.service.ts upsertShot (new shots only) initial, final
3 Cascade shotCascadeUpdate.service.ts upsertShot (existing shots) initial, final
4 SketchAIDA sketchAnalysisGuru.ts:449 updateShotFields initial, final
5 modify_shot (AI chat) workbenchSceneGuru.ts:319 updateSceneFields initial, final
6 merge_shots mergeShotsPrompt.template.ts scene guru initial only
7 enhanceShotsInScene enhanceShots.template.ts modifyShotsInScene initial only

middleFrameInstructions appears in none of them.

Verification:

grep -rn "middleFrameInstructions" src/agenticGuru src/promptTemplates src/workbench --include="*.ts" \
  | grep -v preGenValidation
→ (no output)

The allowlists that would have to change:

Allowlist Location Has middle?
ALLOWED_KEYS (autocomplete) shotAutocomplete.service.ts:7-23 no
ALLOWED_CASCADE_OUTPUT_KEYS shotCascadeUpdate.service.ts:47-68 no
SHOT_SNAPSHOT_KEYS shotCascadeUpdate.service.ts:70-95 no
SHOT_CASCADE_TRIGGER_FIELDS shotCascadeUpdate.service.ts:9-18 no
ALLOWED_METADATA_FIELDS (SketchAIDA) sketchAnalysisTools.ts:1-13 no
allowedFields (modify_shot) workbenchSceneGuru.ts:325-357 no

Note the model's own comment is wrong — storyVideos.model.ts:307:

middleFrameInstructions?: string; // … Filled by pre-gen validation or shot breakdown.

Shot breakdown does not fill it. That comment describes intent, not behaviour.

2. Two write paths, two behaviours — the cascade bypass

REST path (PUT shot from the UI form): workbench.controller.ts:67 → WorkbenchService.upsertShot → branches at workbench.service.ts:747-760:

  • shot exists → ShotCascadeUpdateService.mergeWithCascadeIfNeeded (reconciles other fields)
  • shot is new → ShotAutocompleteService.autocomplete (fills blanks)

AI chat path (modify_shot tool): workbenchSceneGuru.ts:319 modifyShot() → mutates this.sceneData.shots in memory → persistSceneShotsAndDialoguesToDb() (:2645) → StoryVideosModel.updateSceneFields.

ShotCascadeUpdateService is never imported anywhere under src/agenticGuru/:

grep -rn "ShotCascadeUpdate" src/agenticGuru/  → (no output)

Consequence. "Make shot 4 a close-up" typed in chat sets shotSize and nothing else. shotSizeDetails still describes a wide, both frame instructions still describe wide composition, charactersVisible still lists four people who don't fit a close-up. shotSize is the first entry in SHOT_CASCADE_TRIGGER_FIELDS (:10) — the REST path would have reconciled all of it.

This is a bigger deal than the missing middle frame, because chat is the primary editing surface and it silently produces the exact incoherence the pre-gen checker later flags as a conflict.

3. What "reconciled" has to mean

The three frames are three samples of one continuous action. Constraints:

  1. Same characters in all three, unless an entry/exit is explicitly stated.
  2. Same environment in all three.
  3. Framing consistent with cameraMovement — STATIC ⇒ identical framing; DOLLY_IN ⇒ initial and final are the start and end of that move, middle sits between.
  4. The action progresses. Three restatements of one instant is a failure; three unrelated instants is a worse failure.
  5. Each frame is independently drawable as a still — no "turns and then reaches" inside one frame.
  6. All three agree with description and actingInstructions.

These are already specified, in prose, in temporalFieldEditor.prompt.ts — including the self-containment rule (:18), the static-shot exemption (:69) and the definition of the middle as the peak, not the midpoint (:71).

4. Two prompts actively work against reconciliation

shotAutocomplete.template.ts:40 tells the model to write:

finalFrameInstructions … For STATIC camera with no subject movement, write "Same as initial frame"

That is forbidden by temporalFieldEditor.prompt.ts:18 ("Never reference another frame … A reader seeing ONLY that one frame description must have everything needed to draw it"). If the sketch depicts the end and finalFrameInstructions is the literal string "Same as initial frame", Stage 3 has nothing to compare against. shotCascadeUpdate.template.ts:154 has the same pattern ("Same framing, but…").

generateShotsFromPlan.template.ts:361 carries the identical instruction at shot-creation time, so shots are born this way.

This is why fixing allowlists alone is not enough — the prompts have to change too.

One shared reconciliation routine that every write path calls. Not per-block rules (7 copies will drift), and not the checker writing back (it runs at generate time — far too late; and it is deliberately read-only, only $setting its own namespaced preGenValidation field).

Cascade is the right home: it already exists for "one field changed, fix the others", and its trigger list is already exactly the set of fields that should force a frame rewrite (shotSize, cameraAngle, cameraMovement, environment, charactersVisible, props, description, actingInstructions).

Ordered work:

  1. Add middleFrameInstructions to all 6 allowlists in §1. Nothing else can have any effect until this lands — the field is silently dropped on write today.
  2. Route the AI chat path through cascade. modifyShot() should call mergeWithCascadeIfNeeded per shot instead of writing sceneData.shots directly. Closes the REST-vs-chat divergence. Watch latency: cascade is one GPT-5 call at 60s timeout (shotCascadeUpdate.service.ts timeoutMs: 60_000), and modify_shot is a batch tool, so this needs to be concurrent per shot, not sequential.
  3. Remove the "Same as initial frame" instruction from shotAutocomplete.template.ts:40, shotCascadeUpdate.template.ts:154 and generateShotsFromPlan.template.ts:361. Replace with the self-containment rule from temporalFieldEditor.prompt.ts:18.
  4. Extract the frame-reconciliation rules into one shared prompt fragment imported by shot breakdown, autocomplete and cascade — sourced from temporalFieldEditor.prompt.ts so detection and generation cannot drift.
  5. Treat the three frames as one unit in cascade — never rewrite initial without re-deriving middle and final.
  6. SketchAIDA writes the moment the sketch actually depicts. Stage 2 already computes this (sketchConnectedTo, storyVideos.model.ts:308). Today SketchAIDA can only write initial/final, so a sketch of the peak has nowhere to go.

6. Expected payoff

frameCandidates (what the checker derives) and the creator's own three fields converge. That kills the "derived-frame trap" in the frontend handoff — where a conflict quotes frame text the creator never wrote, they search for it, don't find it, and report the checker as broken. Conflict counts on new episodes should drop, because the inputs were coherent before validation ran.

7. Open — needs Miki

  • Does "auto-complete shot detail routines" mean shotAutocomplete.service.ts? Name and function match; not confirmed. Changes which files get touched.
  • ~~Is the cascade bypass (§2) in scope, or a separate PR?~~ Decided: in scope, in this PR. See §8.4 for the rationale, the three verified safety properties and the one-block revert path.
  • enhanceShotsInScene and merge_shots write initial only. In scope, or leave them?

8. What this branch shipped

Branch gerard/aida-frame-reconciliation, based on origin/staging (b7415cca). Not based on #677.

8.1 One shared prompt fragment (new)

src/promptTemplates/shared/frameReconciliation.ts — the single canonical statement of the frame contract. Three exports:

Export Used for
FRAME_FIELD_CONTRACT Miki's five field definitions, verbatim
FRAME_RECONCILIATION_RULES the full ## THE THREE FRAMES block (6 rules + GOOD/BAD examples)
FRAME_RECONCILIATION_RULES_SHORT one-paragraph variant, for prompts that are already long

The five definitions it encodes:

  • description — what happens over the shot (concise)
  • initialFrameInstructions — a single drawable start frame
  • middleFrameInstructions — a single drawable middle/key frame
  • finalFrameInstructions — a single drawable end frame
  • actingInstructions — the temporal performance connecting those states

The rule that does most of the work is #2, fully self-contained: a frame field may never reference another frame. "same as initial frame", "same framing but…", "unchanged from middle", "as before" are all forbidden. Reason: each frame's text is sent to the image model on its own, so a cross-reference resolves to nothing. Repetition across the three fields is required, not sloppy.

Rule #3 fixes the other misreading: middle is the peak, not the geometric midpoint. For a punch it is the instant of impact, even if that lands at 80% of the shot's duration.

Why one file rather than per-prompt rules: §5 already argued it — 7 copies drift. This is the counterpart of preGenValidation/temporalFieldEditor.prompt.ts, which states the same contract on the detection side. When #677 merges, temporalFieldEditor.prompt.ts should be rewritten to import this fragment so detection and generation cannot diverge.

8.2 middleFrameInstructions is now writable

Added to all six allowlists from §1 except one, deliberately:

Allowlist File Added?
ALLOWED_KEYS shotAutocomplete.service.ts yes
ALLOWED_CASCADE_OUTPUT_KEYS shotCascadeUpdate.service.ts yes
SHOT_SNAPSHOT_KEYS shotCascadeUpdate.service.ts yes
ALLOWED_METADATA_FIELDS sketchAnalysisTools.ts yes
allowedFields (modify_shot) workbenchSceneGuru.ts yes
SHOT_ACTIVITY_AUDIT_FIELDS storyVideos.model.ts yes
SHOT_CASCADE_TRIGGER_FIELDS shotCascadeUpdate.service.ts no — see below

Why the trigger list is deliberately untouched. Frame fields are cascade outputs, not triggers. initialFrameInstructions and finalFrameInstructions were never triggers either. If middleFrameInstructions became a trigger, a creator hand-editing one frame would fire an LLM call that overwrites the other two — the opposite of what they asked for.

The wrong model comment at storyVideos.model.ts:307 from §1 is gone; the field now sits inline with its two siblings and the comment states the peak-not-midpoint rule.

8.3 Prompts rewritten (6 files)

Every prompt that writes frame text now names all three fields and carries the reconciliation rules.

File Change
workbenchV2/shotAutocomplete.template.ts two anti-pattern spots removed (field def + a RULES bullet); rules injected via {{FRAME_RECONCILIATION_RULES}}; JSON example rewritten with three fully-written frames
workbenchV2/shotCascadeUpdate.template.ts 6 cascade rule blocks (shotSize, cameraMovement, environment, props, charactersVisible, actingInstructions) now say "the three frame fields (all three, together)"; "Same framing, but…" example replaced; added "the three frame fields move as a UNIT"
workbenchV2/generateShotsFromPlan.template.ts "Same as initial frame" removed; _SHORT rules interpolated
workbenchV2/generateBeatPackage.template.ts all three fields + three fully-written example frames
agenticGuruSystemPrompts/workbenchSceneGuruPromptv20260803.ts FIELD QUALITY RULES + shotSize/cameraMovement/prop cascade blocks
agenticGuruSystemPrompts/sketchAnalysisGuruPrompt.ts middle field def + THE THREE FRAMES para: write the frame the sketch actually depicts

Only the live guru prompt version was edited. workbenchGuruPrompts.index.ts:14 wires workbenchSceneGuruPromptv20260803 and nothing else; the older workbenchSceneGuruPromptv2026* files are frozen history and were left alone on purpose.

This step matters more than it looks: the guru system prompts also enumerate the writable fields, so allowlist edits without prompt edits would have been dead code — the model would never emit the field.

Rewriting each prompt's JSON example was not cosmetic either. An example that violates the rule teaches the violation, and few-shot examples beat prose instructions in practice.

8.4 The chat cascade bypass is fixed (§2)

workbenchSceneGuru.ts now imports ShotCascadeUpdateService — the first reference to it anywhere under src/agenticGuru/. In modifyShot():

  1. per shot, getChangedFieldsFromIncoming + shouldRunCascade decide whether repair is needed;
  2. qualifying shots are queued in cascadePending (skipped entirely when testMode);
  3. before persistSceneShotsAndDialoguesToDb(), all queued shots run through mergeWithCascadeIfNeeded inside a single Promise.all.

Concurrent, not sequential — modify_shot is a batch tool, so N sequential 60s-timeout LLM calls would be unacceptable. Measured cost of one cascade call: 5.1s and 7.0s on two runs of measureCascadeBypass.ts — not the 60s the timeoutMs: 60_000 constant suggests. Two samples, so treat it as single-digit seconds, not a precise figure. Measured staleness it removes: 6 fields left inconsistent by the chat path.

mergeWithCascadeIfNeeded is read-only (reads scene + assets, calls the LLM, returns the shot), so the write still happens in exactly one place: persistSceneShotsAndDialoguesToDb.

Decision: kept in this PR rather than split out. §7 flagged it as an open question; the answer is in scope. Reason: the frame work alone would have left the same incoherence reachable through chat, which is the primary editing surface. Shipping a frame contract that the main editing path ignores would fix the symptom on one route and leave it on the other.

Why this is safe to ship in the same PR — three properties, each verified in code, not assumed:

  1. A cascade failure cannot lose the user's edit. mergeWithCascadeIfNeeded sets result = merged (the user's literal edit) at shotCascadeUpdate.service.ts:237, retries twice, and its catch at :273-282 only logs — :285 returns result either way. The log line says it outright: "persisting user merge only". A failed cascade degrades to exactly the old chat behaviour; it never throws into the chat handler.
  2. Cascade cannot overwrite the field the user just set. sanitizeCascadePatch:219 puts changedFields into skip, so the edited field is dropped from the patch before it is applied. This is the invariant the pre-existing shot_cascade_*_no_overwrite tests already cover.
  3. No new write path. Cascade returns a shot; the single existing write (persistSceneShotsAndDialoguesToDb) is untouched.

What changes for the user. Typing "make shot 4 a close-up" in chat now also reconciles shotSizeDetails, description, actingInstructions, the three frame fields and charactersVisible, instead of setting shotSize and leaving six fields describing a wide. Same action through the shot panel already did this.

Cost. One extra LLM call per edited shot, only when the edit touches a trigger field (SHOT_CASCADE_TRIGGER_FIELDS), and only outside testMode. Measured 5.1s and 7.0s (two runs), concurrent across shots, so a 3-shot batch costs one call's latency rather than three.

Residual risk, stated plainly: chat edits to trigger fields are now slower by one LLM round-trip. That is the intended trade — it is what the REST path has always paid for coherence. If the latency turns out to be unacceptable in practice, the revert is small and local: drop the cascadePending block at workbenchSceneGuru.ts:460-475 and the queueing at :421-428. Nothing else depends on it.

8.5 Tooling

  • scripts/aida/checkFramePrompts.ts — guards the two silent failure modes: an unsubstituted {{PLACEHOLDER}} shipping to the model, and the "same as initial frame" anti-pattern creeping back. Asserts each built prompt contains THE THREE FRAMES, middleFrameInstructions and FULLY SELF-CONTAINED, with no leftover placeholders. Both prompts PASS.
  • scripts/aida/measureCascadeBypass.ts — read-only chat-vs-panel field diff, the source of the 5.1s / 7.0s latency and 6-stale-fields numbers above.
  • scripts/aida/testFrameOutput.ts — read-only; runs the real auto-complete on a real shot with the three frame fields blanked, prints the output and fails on an empty or cross-referencing frame.

Also in this branch, unrelated to frames: python_eval/tests/test_definitions/__init__.py imported .scene_context_drift, which was never committed (the import landed in 4b0c88f3; the module was written in a separate repo copy that got deleted). That broke from tests.test_definitions import TESTS on staging and main — the whole g-eval suite, not just those tests. Commented out with a TODO(miki); 453 test definitions import again. No CI workflow runs python_eval (.github/workflows/test.yml runs vitest only), which is why it shipped.

8.6 Verification

Check Result
npx tsc --noEmit clean
npm run build clean
npx tsx scripts/aida/checkFramePrompts.ts PASS on all 6 frame-writing prompts + clean anti-pattern sweep
grep -rn "Same as initial frame" src/ only inside frameReconciliation.ts, where it is quoted as forbidden
npx biome check on changed files clean (repo-wide 134 errors are pre-existing import ordering in untouched files)
npx vitest run (full suite) 949 passed / 4 failed; the 4 pass 86/86 in isolation (resource contention, no frame/cascade code in either file)
g-eval --collection shot_autocomplete 23/23 pass (200s) — 17 pre-existing + 6 added here
g-eval --collection shot_cascade 22/22 pass (123s) — 18 pre-existing + 4 added here
Real auto-complete output, 2 scenes three self-contained frames each, zero cross-references (below)

LLM output actually inspected

scripts/aida/testFrameOutput.ts blanks a real shot's three frame fields and runs the real ShotAutocompleteService.autocomplete(). Read-only — nothing persists.

  • Scene fd95fea2, "Stan Under the Flavinator" (STATIC two-shot, 10.3s). Before: middle and final empty; initial read "Stan and Becca face the Flavinator, then pivot slightly toward each other" — temporal language describing two instants, not a drawable frame. After: three frames, each repeating camera, both characters' positions and the background in full, progressing look-away → eyes meet → held contact.
  • Scene 61365420, "Eagle's Flight" (STEADICAM, 7.2s). Three progressing frames; stalk tops climb the frame as altitude drops; initial ≠ final.

8.6.1 The pre-existing tests could not have caught this work

Checked before writing any new ones: grep -rn "middleFrameInstructions" python_eval/ returned zero hits across all 453 test definitions. The main change on this branch had no test coverage at all.

Worse, the two existing frame tests pass while the anti-pattern is present:

  • test_frame_instructions_present (shot_autocomplete.py) asserts only len(initial.strip()) > 10. "Same as initial frame" is 21 characters — it passes.
  • test_frame_instructions_differ_for_movement asserts initial != final. "Same as initial frame" is a different string from the initial frame's text — it passes.

So 35/35 green proved the old behaviour was intact. It could not prove the new behaviour worked.

8.6.2 Ten tests added, and the bug they found

Shared helpers in python_eval/tests/test_definitions/utils/frame_contract.py (one copy, used by both collections — same argument as the shared prompt fragment).

Test Kind Asserts
shot_autocomplete_three_frames_present deterministic all three frames ≥40 chars
shot_autocomplete_frames_self_contained deterministic no cross-reference; zero-cost tripwire on the exact strings removed from the prompts
shot_autocomplete_frames_distinct_moments deterministic no two frames byte-identical
shot_autocomplete_frames_independently_drawable G-Eval, 0–10 rubric each frame drawable with the other two hidden — catches paraphrases the regex cannot
shot_autocomplete_middle_frame_is_peak G-Eval, 0–10 rubric middle is the peak, one instant, distinct from both neighbours
shot_autocomplete_frames_temporal_progression G-Eval, 0–10 rubric cast/location stable, framing matches cameraMovement, action advances in order
shot_cascade_three_frames_present deterministic cascaded shot carries all three
shot_cascade_frames_self_contained deterministic tripwire for "Same framing, but…"
shot_cascade_frames_move_as_unit deterministic a trigger edit patches all three or none — never two
shot_cascade_frames_coherent_with_edit G-Eval, 0–10 rubric the rewritten frames describe the NEW value, not the old one

The four G-Eval metrics use deepeval's rubric=[Rubric(score_range=…, expected_outcome=…)] (0–10, non-overlapping, full coverage; the runner passes at score >= 0.5, i.e. 6/10). Bands: 0–2 broken, 3–5 fails the contract, 6–8 passes, 9–10 exemplary. strict_mode is deliberately off — it forces binary 0/1 and would discard the grading. Nothing else in python_eval used rubric before this.

Observed scores on the current prompts: independently-drawable 0.84, middle-is-peak 0.70, temporal-progression 0.81, cascade-coherence 0.80. The 0.70 is the rubric working — the judge found the middle frame drawable and distinct but not landing on the strongest instant, which is the "passes, not exemplary" band.

The deterministic checks were falsified before being trusted: injecting "Same as initial frame", "Same framing, but…", "As described above", an empty middle, and three identical frames — all five caught, correct output passes.

shot_autocomplete_three_frames_present failed on first run and found a real bug in this branch: src/testDebug/testDebug.service.ts hand-copies three allowlists rather than importing them, and I had updated only the production ones. Three copies were still missing middleFrameInstructions:

Line Copy of Effect
:1072 SHOT_SNAPSHOT_KEYS the cascade prompt never saw the existing middle frame
:1137 ALLOWED_CASCADE_OUTPUT_KEYS cascade's middle frame stripped from the patch
:1211 ALLOWED_KEYS (autocomplete) autocomplete's middle frame stripped — measured middle=0 chars

The model was emitting the field correctly; the allowlist threw it away. Only the testDebug route is affected, so production writes were never broken — but it means the g-eval harness would have reported the feature as working while measuring a path that silently dropped it. Fixed, and a sweep now confirms every allowlist in src/ that names initialFrameInstructions also names middle.

The g-eval runs are the stronger evidence, because shot_autocomplete_frame_instructions, shot_autocomplete_frames_differ_movement, shot_autocomplete_frame_instructions_spatial, shot_autocomplete_spatial_specificity, shot_cascade_frame_instructions_differ, shot_cascade_environment_coherence and shot_cascade_prop_coherence assert exactly the contract this branch adds, and they are judged by an LLM against the output rather than by grepping the prompt.

Run them with:

ENABLE_TEST_ROUTES=true npm run dev            # port 8081 is hardcoded (src/index.ts:19)
cd python_eval && ./venv/bin/python test_runner.py --collection shot_autocomplete -v
cd python_eval && ./venv/bin/python test_runner.py --collection shot_cascade -v
Note: no --fixture. Both collections are standalone (test_runner.py:3538) with their own input fixtures; passing -f sdf3 routes them down the generic path and they fail on a null template_id. python_eval needs Python 3.10+ — the system python3 is 3.9 and cannot even build the venv.

Not verified: the read side. Nothing consumes middleFrameInstructions yet — generateSequenceSeedancePrompt.template.ts:50-51 documents only initial/final and sequenceFlow.utils.ts:81-82 does not put it in the bundle. The field is now written correctly and is visible to the pre-gen checker, but it does not yet reach video generation. That is a follow-up.

8.7 Out of scope for this PR

  • sketchConnectedTo (§5 step 6) — that field is #677's Stage 2 output; SketchAIDA writing the peak depends on it, so it lands with #677.
  • enhanceShotsInScene and merge_shots — still initial-only, pending Miki's answer in §7.
  • copyShot and moveShotsToScene adapt two fields and leave the frames behind. Both spread ...sourceShot and then overwrite with { description, actingInstructions } from adaptShotForScene.template.ts, whose AdaptedShot schema contains only description, actingDirection, environmentNote and adaptationRationale — no frame fields (workbenchSceneGuru.ts:1250 and :1122).

So the three frames survive the copy verbatim while description and actingInstructions are rewritten for the new scene, and moveShotsToScene additionally replaces environment with the target scene's. The frames can end up describing the old location.

This is not a middle-frame problem — the spread copies all three frames equally. It is the same gap for initial and final, and it predates this branch. What makes it worth naming now: description, actingInstructions and environment are all in SHOT_CASCADE_TRIGGER_FIELDS, so this is precisely the change that the cascade exists to reconcile — and neither of these two paths calls it. Adding frame fields to AdaptedShot would be the wrong fix (a second prompt writing frames, competing with the cascade); the right fix is to route both paths through mergeWithCascadeIfNeeded the way modifyShot now does.

Left out of this PR deliberately: it is two more write paths and one more LLM call per copied shot, on tools nobody reported a problem with. Raised by the review bot on #682, framed there as a middle-frame issue; the framing is wrong but the underlying gap is real.