Skip to content

Analysis A — the variants, and the "keep it always aligned" question

Status: design options. Nothing implemented, Mongo read only.

Read it as: two questions on one page. (1) Analysis A looks a lot like T2I pre-gen — what else could it be, and which version is cheaper? (2) Miki asked Nicola for one function that can run anywhere in the backend and check that a picture still agrees with its own metadata. Does Analysis A cover that, does T2I pre-gen, or neither?


0. The three short answers

1. T2I pre-gen makes two LLM calls because it has one picture and one shot — not because it skips checks. It does skip two things, and both look like accidents rather than decisions: stages/sketchMomentStage.ts and stages/finalReviewStage.ts are written and tested, and nothing imports either of them.

2. Analysis A can be built five ways. They differ on one axis only: how many pictures and how many shots is a single prompt allowed to hold? Fewer calls is not the same as faster, because the calls run in parallel — §5 has the arithmetic.

3. Nothing in the product does Nicola's job today, but two thirds of it is already in the repo. There is a generic post-generation checker — src/assetGen/ qualityAnalysis/, a harness plus a plug-in registry — that sends a finished image plus its references to a judge and scores it against an intent built from the shot's own fields. It fires on one event: an asset-gen job settling. Uploads, tweaks, tone-mapping and text edits are invisible to it. What is missing is a media-keyed alignment record that any trigger can write and every checker can read.


1. T2I pre-gen, checked against the code

Your diagram is accurate. Two corrections and two additions:

Claim Verdict Evidence
One route, POST /workbench/validateShotForGeneration Confirmed workbench.router.ts:75
Only 2 LLM calls, both at the same time Confirmed preGenValidation.service.ts:536-562, one Promise.all over TextTemporalStage.run and SelectedFrameAlignmentStage.run
Both GPT-5 Almost — the model is OpenAIModelId.GPT_5_4 preGenValidation.service.ts:540, :558
Cache: same hash and not forced, return the stored answer, 0 calls Confirmed :426 — stored.inputHash === inputHash && stored.status !== 'validation_failed'
Hard pre-check: never-analysed sketch is blocked, canGenerate=false, stop Confirmed :475-507. It is also the only thing that sets canGenerate=false — :998
Which moment the sketch is: no LLM, connectedTo or assume start at confidence 0.34 Confirmed :509-523, and the constant at :387
Merge + judge caps severity, collapses shared fixes, re-applies dismissals, labels fixTarget Confirmed :878-890 (BLOCKABLE_SKETCH_CODES, 6 codes), :796 (collapseSharedFixes), :754-763, :934
Saves to shot.preGenValidation, never touches the shot's own fields Confirmed :1046-1060
Addition: Stage 4, the final review, does not run at all Confirmed :574 passes null where the final-review result goes; stages/finalReviewStage.ts exists (123 lines) and nothing imports it
Addition: neither does the sketch-moment vision stage Confirmed stages/sketchMomentStage.ts (117 lines), no importer. :512 builds the result by hand instead
Addition: timings[] is computed and then thrown away Confirmed :1047-1058 — persistOnShot never copies it, so there is no stored duration for any stage anywhere

So the shape it actually runs is:

  collect ──► hash ──► cache hit? ──► return, 0 calls
                │ miss
                ▼
        sketch analysed? ──no──► BLOCKED, stop
                │ yes
                ▼
        which moment?  (plain code, no LLM)
                │
        ┌───────┴────────┐
        ▼                ▼
   STAGE 1 text     STAGE 3 vision      ← the only two calls, in parallel
        └───────┬────────┘
                ▼
        merge + judge (plain code)
                ▼
        write shot.preGenValidation      ← the shot's own fields are never touched

Written but never wired: STAGE 2 vision (which moment), STAGE 4 final review.


2. Why two calls there, and more here

Your read was that T2I is cheaper because it skips steps in the middle. Half right. Two dead stage files are genuinely skipped work. The rest of the gap is not skipping — it is that T2I has less to look at.

T2I pre-gen has one shot and one picture, so exactly two questions exist for it: does the shot contradict itself, and does the picture match the shot. Two calls, one each. There is no third question available.

Thing to check T2I pre-gen I2V Analysis A
Shot's text fields against each other yes — Stage 1 yes (A2)
The picture against the shot's text yes — Stage 3 yes (A3)
Camera movement, timing, dialogue duration never looks. Movement and TIMING are context-only since validator 1.17.1, and dialogue was dropped in 1.1.0 new (A4). Nothing upstream has ever read these
The picture is a generated image, not the sketch that was judged no — it judges the sketch, then the image is generated afterwards must. This is the residual risk on SHOT 1 in tab 3
Several pictures in one check no — one sketch yes, one per shot
Two shots against each other no — the unit is one shot the whole of Analysis B
The @image1..N reference bundle as a bundle partly — reference findings on one shot yes (B3)

Both extra per-shot questions (video fields, generated-image-vs-sketch) read the same two inputs A2 and A3 already read: the shot's fields, and the picture. That is the opening every variant below exploits.


3. The system nobody put on the whiteboard

src/assetGen/qualityAnalysis/ — 14 files, ~4,600 lines, a README, its own tests. This matters because it is already the generic shape Nicola was asked to build, for one trigger.

  a job settles (runJob | pollJob | webhook | updateJobAsFailed)
        │        4 call sites: assetGen.service.ts:310, :374, :547,
        │                      webhooks.service.ts:169
        ▼
  QualityAnalysisHarness.scheduleOnJobSettled     fire and forget
        │
        │  QUALITY_ANALYSIS_ENABLED !== 'true' ──► return, nothing happens
        │  no useCaseId stamped on the job     ──► return, nothing happens
        ▼
  resolve the use case from a registry          registry.ts, 15 lines
        │                                       today one entry: SHOT_IMAGE_V1
        ▼
  claim it (one atomic write to PROCESSING, so it is never judged twice)
        ▼
  useCase.buildContext(job)
        │   intent  ← built from the SHOT'S OWN FIELDS, not from the model prompt:
        │             description, acting, the frame the sketch depicts, camera,
        │             characters visible, environment, props
        │             (shotImage.useCase.ts:158-195)
        │   references ← every image the model received, each given a NAMED ROLE
        │             (:88-130) so identity is judged only against character sheets
        ▼
  one LLM judge call — gpt-5.4, medium         (or N images at once, batch path)
        ▼
  verdict: 1-5 per criterion, critical criteria cap the overall
        ▼
  pass ──► ACCEPTED
  fail ──► REGENERATE {prompt} | EDIT_PREVIOUS {instruction} | ESCALATE
             │
             └── may auto-apply and re-enter the loop, up to maxAttempts

What that already is: a trigger-agnostic-ish harness, a plug-in registry so each kind of asset says what to compare, a per-item claim so nothing double-runs, realtime events, and a comparison of a finished picture against the metadata it came from.

What it is not:

Gap Evidence
One trigger only — a job settling. An upload creates no job at all, so it never fires uploads go straight to StoryVideosModel.addShotImage(... USER_UPLOAD ...); the 4 schedule sites are all job sinks
Even some jobs are invisible. The reject-and-regenerate batch is unstamped workbench.service.ts:3529, a TODO(quality-analysis follow-up) says so in full sentences
So are edit batches, and so is every imageEdit job README "Future work"; imageEdit.service.ts:61, :85 call createJob with no jobMetaData
Output is a score plus a next job, not a conflict the user can read and dismiss qualityAnalysis.types.ts, NextAction = ACCEPT / REGENERATE / EDIT_PREVIOUS / ESCALATE
No reusable read of the picture. Every judge call re-sends the pixels; nothing is cached per media no media-keyed store anywhere in the folder
Off by default sharedConstants.ts:45-46, and QUALITY_ANALYSIS_ENABLED is not in the local .env. Not verified: whether it is on in staging or prod

4. Three systems that all compare a picture to its own words

T2I pre-gen QC harness (SHOT_IMAGE_V1) I2V Analysis A (proposed)
Trigger the user clicks the check an asset-gen job settles the user opens the sequence and clicks
Unit one shot one job (or one batch of slots) one sequence of shots
Picture it judges the sketch the generated image whatever the shot will actually send: sketch or image
Sees the shot's text? yes yes — rebuilt as intent yes
Sees other shots? no no yes, that is Analysis B
LLM calls 2, parallel 1 (or 1 per batch) see §5
Output conflicts, each with a fixTarget the UI routes on score 1-5 + a follow-up job conflicts, same type as T2I
Stored on shot.preGenValidation job.qualityAnalysis the sequence draft (§4.4)
Cache key inputHash over the shot + VALIDATOR_VERSION none — re-judges on every call per-stage hashes + mediaId for the frame read
Blocks anything? only a missing sketch analysis no, it proposes product decision, open

Read the row "cache key" as the real difference. T2I keys on the shot, the harness keys on nothing, and only the I2V proposal keys anything on the picture. That is the row Nicola's function has to fix, and §6 is about that.


5. Five ways to build Analysis A

One axis: how much is one prompt allowed to hold?

  MORE CALLS, narrower prompts                      FEWER CALLS, wider prompts
  ──────────────────────────────────────────────────────────────────────────►

  A-1              A-0                A-2                A-3
  T2I twin         as proposed        one call/shot      one call, whole sequence
  2 calls/shot     3 calls/shot       1 call/shot        1 call, N pictures
  + video stage    + cross-shot       + cross-shot       everything in it

                          A-4  always aligned
                          0 per-shot calls at click — they already ran
                          on the event that changed something

A-4 is not further right on that axis; it is off it. It changes when the work happens, not how much a prompt holds. It can be combined with any of the others, and it is where Nicola's function lands.

A-0 · as proposed (the baseline)

Five per-shot stages, one purpose each. A0 deterministic (no LLM), A1 the frame read (pixels only, fired when the media is attached, cached on the media), A2 text vs text, A3 the frame read vs the text, A4 the video-only fields. Then B0-B3 across shots.

  shot 1   A0 ─ A1* ─ A2 ─ A3 ─ A4 ┐
  shot 2   A0 ─ A1* ─ A2 ─ A3 ─ A4 ├─► B0 ─ B1 ─ B2 ─ B3 ─ merge ─ final review
  shot 3   A0 ─ A1* ─ A2 ─ A3 ─ A4 ┘
  shot 4   A0 ─ A1* ─ A2 ─ A3 ─ A4 ┘        * A1 is paid at attach time, not at click

Good: every stage has one job, so a bad finding is traceable to one prompt. Per stage caching means a text edit re-runs A2/A3 and leaves A1 and A4 alone. Bad: 3 calls per shot. On four shots that is 12 at click plus 4 across, and the per-shot prompts overlap heavily — A2, A3 and A4 all receive the same shot fields.

A-1 · the T2I twin

Call the shipped thing. For each shot in the sequence, run validateShotForGeneration as it exists today, then add one new stage for the video fields, then run Analysis B on top.

  shot 1 ──► POST /validateShotForGeneration  (2 calls, its own cache) ─┐
  shot 2 ──► POST /validateShotForGeneration  (2 calls)                 ├─► B0..B3
  shot 3 ──► POST /validateShotForGeneration  (2 calls)                 │
  shot 4 ──► POST /validateShotForGeneration  (2 calls)                 ┘
             + A4 video stage per shot (new, 1 call each)

Good: by far the least new code, and it inherits 22 versions of prompt tuning, the issue catalogue, the conflict language, the dismissal logic and Isa's wording rules for free. Demo-ready soonest. Bad, and it is serious: its Stage 3 judges the sketch. It has no idea a generated image exists, so on a shot that already has an image it checks the wrong picture, which is precisely the residual risk on SHOT 1. It also has no frame read to hand to Analysis B, so B2 would have to re-read every picture itself. And its cache key is one hash for the whole shot, so it cannot express "the text check is still valid, the picture check is not".

A-2 · one call per shot — the recommendation

Merge A2, A3 and A4 into a single multimodal call per shot. Legitimate because all three read the same two things: this shot's fields, and this shot's picture.

  shot 1   A0 ─ A1* ─┐
  shot 2   A0 ─ A1* ─┤   one call per shot:
  shot 3   A0 ─ A1* ─┤   text-vs-text + picture-vs-text + video fields
  shot 4   A0 ─ A1* ─┘   ────────────────────────────────► B0 ─ B1 ─ B2 ─ B3 ─ review

Good: halves the per-shot calls, and the prompt is better informed than three narrow ones — the same model that sees the picture also sees the timing, so "she crosses the room in 2 seconds" can be weighed against a picture of her standing still, which no split version can do. Bad: one prompt with three jobs is where quality regressions hide, and you lose per-stage cache granularity — a text edit re-runs the visual half too. That is cheap, because the expensive part (the frame read) is cached separately on the media.

A-3 · one call for the whole sequence

One prompt: all four pictures, all four shot texts, per-shot and cross-shot findings in one answer.

  A0 all shots ─ A1* all shots ─► ONE CALL (4 pictures + 4 shot texts + 50 dimensions) ─► merge

Good: one call. Cheapest possible on tokens-per-call count. Bad, and this is the part worth saying out loud to Miki: it is probably not faster. The calls in A-2 run in parallel, so four shots is one wave, not four calls' worth of waiting. A-3 replaces four parallel medium calls with one long serial call carrying four images. On a 15-25 second budget that trades the thing we have (parallelism) for the thing we do not need (a lower call count — ~$0.25 a sequence is already fine). It also has no partial cache at all: edit one word in shot 3 and the entire sequence is re-judged. And one parse failure loses everything.

Worth listing so the axis is complete. Not worth shipping.

A-4 · always aligned (off the axis)

Nothing per-shot runs at click. Every event that can break the alignment enqueues the per-shot check when it happens, and the click reads what is already stored.

  image finished generating ─┐
  image uploaded            ─┤
  sketch uploaded           ─┤      enqueue    ┌───────────────────────┐
  crop / magic select       ─┼───────────────► │  per-shot alignment   │
  tone-mapping correction   ─┤                 │  (A1 + A2 + A3 + A4)  │
  shot text edited          ─┘                 └───────────┬───────────┘
                                                           │ writes
                                                           ▼
                                            stored alignment for that shot
                                                           │
   user clicks generate video ─────────────────────────────┤
                                                           ▼
                              anything stale? run just that. then B0..B3.

Good: the click costs only the cross-shot half, which is the part that can never be pre-computed anyway (no cached result describes a pair of shots). Everything else was paid while the user was doing something else. This is the only variant that comfortably fits a 15-25 second ceiling on a long sequence. Bad: the most moving parts. Needs a queue, needs every mutation site to fire an event, and spends calls on shots the user may never put in a sequence. A text edit per keystroke would be pathological, so it needs debouncing, or a "mark stale now, compute on demand" split.

Side by side, four-shot sequence, nothing cached

A-1 T2I twin A-0 as proposed A-2 one/shot A-3 one call A-4 always aligned
Per-shot calls at click 12 12 4 — 0
Cross-shot calls at click 4 4 4 included 4
Total at click 16 16 8 1 4
Parallel waves at click 3 3 2 1 1
Judges the generated image? no — the sketch yes yes yes yes
Gives B2 a frame read to reuse? no yes yes n/a yes
Survives a one-word text edit cheaply? no yes mostly no yes
New code to write least most medium least-but-riskiest most
Where a bad finding is traceable to a shipped prompt one narrow prompt one wide prompt one huge prompt one wide prompt

Latency is a shape here, not a measurement. No per-stage duration is stored anywhere: preGenValidation computes timings[] and persistOnShot never writes it (:1047-1058), and there is no duration field in analyticsLogs either. Waves are the only honest proxy until someone measures it.

Recommendation: A-2 for the per-shot half, with A-4's trigger for the frame read only (which the proposal already assumes — A1 fires when media is attached). That is 8 calls in 2 waves on a cold four-shot sequence, it judges the right picture, and it hands B2 a cached read. Then, if the sequence check turns out too slow in practice, move A-2's per-shot call to the A-4 trigger as a second step — the prompt does not change, only who calls it.


6. Nicola's function: one aligner, many triggers

The ask, as you described it: one thing that can run at any point in the backend and confirm that a visual piece — sketch or image — still agrees with its own metadata.

Does Analysis A cover it? No, and neither does T2I pre-gen. Both are gates in front of a button. They answer "is this shot ready to generate right now", they run when a user clicks, and they store their answer where their own caller looks. Neither can be called from an upload handler.

But Analysis A already contains the hard half of it. §4.5 of the proposal splits the work in a way that is exactly what a trigger-agnostic function needs:

  ┌──────────────────────────────────────────────────────────────────┐
  │  HALF 1 — READ THE PICTURE          pixels only, never the text  │
  │  key: mediaId + analyserVersion + schemaVersion                  │
  │  a text edit CANNOT stale it. Written once per picture, ever.    │
  └───────────────────────────┬──────────────────────────────────────┘
                              │ structured facts, not pixels
  ┌───────────────────────────▼──────────────────────────────────────┐
  │  HALF 2 — COMPARE                   the read vs the metadata     │
  │  key: mediaId + metadataHash + checkerVersion                    │
  │  cheap, text-vs-text, no vision call. Re-run on any text edit.   │
  └──────────────────────────────────────────────────────────────────┘

That split is the whole trick. Half 1 is expensive and never needs redoing. Half 2 is cheap and needs redoing constantly. Today's systems fuse them, which is why none of them can be called from anywhere.

Today

  generate image ──► QC harness ──► score 1-5      ──► job.qualityAnalysis
  upload image   ──► (nothing)
  upload sketch  ──► (nothing until the user opens sketch chat)
  crop / tweak   ──► (nothing)
  tone-mapping   ──► (nothing)
  edit shot text ──► (nothing)
  click T2I gate ──► pre-gen    ──► conflicts      ──► shot.preGenValidation
  click I2V gate ──► Analysis A ──► conflicts      ──► the sequence draft

  three checkers · three triggers · three output shapes · three stores
  the same picture is re-read by each one, every time

After

                       ┌───────────────────────────────────────┐
  generate image ─────►│                                       │
  upload image   ─────►│   alignMedia({ mediaId, subject })     │
  upload sketch  ─────►│                                       │
  crop / tweak   ─────►│   half 1: read the picture  (cached)  │──► media.imageAnalysis
  tone-mapping   ─────►│   half 2: compare to metadata         │──► media.alignment
  edit shot text ─────►│                                       │
  on demand      ─────►└───────────────────┬───────────────────┘
                                           │  one stored answer
                    ┌──────────────────────┼──────────────────────┐
                    ▼                      ▼                      ▼
              T2I pre-gen            I2V Analysis A          QC harness
              skips its own          skips A1 and A3         judges quality only;
              vision stage           for read shots          stops re-deriving
                                                             what the aligner knows

Same three product surfaces. One thing looking at pictures.

The trigger table — what fires today, what would have to

Event Creates an asset-gen job? QC harness fires? Anything checks alignment?
Generate shot image, initial batch yes, stamped yes, if the flag is on the QC judge
Regenerate after reject yes, unstamped no nothing. workbench.service.ts:3529
Edit batch yes, unstamped no nothing. README, Future work
imageEdit (crop, upscale, edit-image guru) yes, no jobMetaData at all no nothing. imageEdit.service.ts:61, :85
Upload an image no job no nothing
Upload a sketch no job no only when the user opens sketch chat
Tone-mapping correction writes a new shot image directly no nothing. projects.service.ts:1107, :1187
Magic Select / tweak apply mints a new mediaId no nothing
Edit any shot text field no no nothing until the user clicks a gate

Nine ways a shot can fall out of agreement with itself. One of them is watched today, and only behind a flag that is off by default.

What the generic function needs that nothing has yet

  1. A subject, not a job. The input is a media id plus what this picture is supposed to be. A job id is one way to get there, not the only one.
  2. Two keys, not one. mediaId + analyserVersion for the read, and mediaId + metadataHash + checkerVersion for the comparison. Both halves must be able to say "still valid" independently.
  3. Its own store, on the media. Not on the job (uploads have none), not on the shot (the read belongs to the picture and outlives the shot's text). The proposal already puts the read on ImageMedia.imageAnalysis, §4.4 — the comparison goes next to it.
  4. A stale marker that is free to set. Any text edit marks it stale in one write, with no LLM call. Whoever needs a fresh answer pays for it — the gate at click, or a background worker, or both.
  5. One output shape. A conflict list, as T2I already defines it. The QC harness's 1-5 score is a different question (how good is this picture) and should stay its own thing; it can consume the alignment findings rather than re-deriving them.
  6. The existing registry pattern is the right one. registerQualityAnalysisUseCase already solves "each kind of asset says what to compare" in 15 lines. Nicola's function should reuse that idea, not invent a second plug-in system.

Where this leaves Analysis A

Unchanged, and simpler. If the aligner exists, Analysis A is only the part that no event can pre-compute: the cross-shot half. No stored per-shot result can ever describe a pair of shots, so B0-B3 will always run at click. Everything per-shot becomes a cache read.

That is worth saying to Miki in one line: Nicola's function and Analysis B are the same project seen from two ends. Build the aligner, and Analysis A mostly stops existing as separate code.


7. What I would put in front of Miki

  1. T2I pre-gen is 2 calls because its unit is one shot with one picture, and it is judging the sketch, not the image that gets sent to video. Reusing it as-is for I2V would check the wrong picture on any shot that already has an image.
  2. For Analysis A, pick A-2: one multimodal call per shot, then the cross-shot pass. 8 calls in 2 waves on a cold four-shot sequence, against 16 in 3 waves for the baseline. The one-call-for-everything version is cheaper on paper and probably slower in practice — we have parallelism and we do not need a lower call count.
  3. Nicola's generic aligner is real and it is not any of the three current systems — but the harness in src/assetGen/qualityAnalysis/ is already two thirds of the machinery, and the frame-read/compare split in §4.5 is the other third. Nine events can break a picture's agreement with its metadata; exactly one is watched today, behind a flag that defaults to off.

8. Not verified

  • Whether QUALITY_ANALYSIS_ENABLED is on in staging or production. Only checked the local .env, where it is absent, and the default is off.
  • Any latency number in §5. There is no stored duration for any LLM stage anywhere in Mongo, so the "waves" column is reasoning about parallelism, not measurement. This is the first thing to measure before choosing between A-0 and A-2.
  • Whether the two dead stage files (sketchMomentStage.ts, finalReviewStage.ts) were deliberately parked or dropped in a merge. Worth one question to whoever wrote them before assuming either.
  • Whether the QC judge's findings are good. shotImage.useCase.ts:153 carries its own TODO: validate this properly ... no measured comparison on real shot images.
  • The A-3 attention claim (one prompt with four pictures degrades). That is the standard result for multi-image prompts, but I have not measured it on this taxonomy, and it is the sort of thing Miki will rightly ask for evidence on.