Skip to content

To Yashasewi Singh

The aim of this document is to share with Yash the main fundamentals and characteristics of the I2V pregeneration conflict generation system.

A more detailed document aimed to be shared with a coding agent (Claude, GPT…) will be also forwarded: i2v_agent_handover_yash.md.

Management ideally wants the system to be in staging and able to be tested by the Creative team (at least two creatives) by Wednesday 30th of September.

The first version of the system was implemented under the supervision of Miquel Àngel Farré and revisited by Disha, Isa and Jeet. It can be found here:

BACKEND: https://github.com/Scenarix/studio-backend/pull/847

FRONTEND: As far as I know, Jeet shared all of the details regarding how the frontend should look like for this version.

In case it helps, to be able to test the backend and get some visuals, I did implement a mock up myself that captures the PRD basics that Disha requested: https://github.com/Scenarix/studio-frontend/pull/814 — it can be used as inspiration or naive example.


What I'm asking you for

The code is written, reviewed and running. So this is not a spec to build from — it's a review and a trip to staging.

Concretely: read PR #847, land it, and get the frontend side to staging with it. If something looks wrong to you, tell me and I'll fix it — I'd rather you ask than rewrite. The list further down is what has to still be true after the merge, so you know what to check and what not to take out.


BACKEND DETAILS

The content of this section is covered in more detail within the PR or the document to be shared with the coding agent. But I wanted to enumerate the main must-have's in human readable wording. Let's dive into it:

- I2V conflict: When the user clicks Generate Prompt in the I2V pregen screen it will trigger a system of conflict detection that works with two objectives:

  1. Find conflicts within a shot individually (Analysis A, VLM): conflicts between the text inputs and the visual input of a shot, analysed individually against itself. So all the text inputs (description, acting, init, mid, end frame…) of the shot against the visual input (image, sketch) it carries. (E.g. Shot 8 text vs visual.)

  2. Find conflicts across all the shots of a sequence (Analysis B, VLM): since I2V gets fed with a sequence of shots, let's analyse the conflicts that sequence might carry. In this stage a VLM looks into all the conflicts that (examples) Shot 1, Shot 2, Shot 3 might have against each other (e.g. Shot 1 is daylight, Shot 2 is night, Shot 3 is raining).

All the conflicts raised by Individual Analysis (Analysis A) and Cross shot or Sequence Analysis (Analysis B) are raised to the user once the computation has finished.

There is also a third check running alongside them, with no model at all: it looks at what the request will actually send as pictures — no picture of the location, a selected shot with no image and no sketch, the same character listed twice, two different pictures of the same person. Cheapest and most certain conflicts we have. Don't drop them because they're not AI.

IMPORTANT: with the conflicts raised the user can then:

  • Mark them as Not a conflict
  • Resolve them by doing 1 click !!!!! The system is prepared in a way that for every conflict 2–3 alternatives to solve the conflict are provided to the user. (E.g. the daylight vs night vs rainy day case — three alternatives are proposed to the user)
  • It's daylight
  • It's dark
  • It's rainy
  • Other:

Whatever option is selected by the user it will change the metadata of the shots that option names.

One precision, because getting it wrong is a bad bug: an option changes only the shots it names, never all of them. And some options change nothing at all — that's not a bug. The picture already cost the user money, so "keep the picture and tell the video model to follow it" is a legitimate way out: nothing is redrawn, nothing typed is rewritten, a sentence goes to the prompt instead. "Other" works the same way — we can't tell which field the user meant, and rewriting the wrong one isn't recoverable.


THE 5 THINGS THAT MUST STILL BE TRUE AFTER THIS LANDS

If one of these breaks, the creatives stop trusting the screen in the first hour, and we don't get a second chance.

1. The gate: no prompt until every conflict has been answered

Disha's rule. The prompt is generated automatically only when there are no conflicts, or the user has answered every single one — "Not a conflict" or "Resolve". Not "no blockers". Every conflict, whatever its colour, is a question, because the answer is what changes the prompt.

One exception: if our own check breaks (model times out, API down) we raise a conflict saying "we could not check this sequence". That one reports, never blocks. Our failure must not stop a user generating their video.

The backend reports whether generation is allowed; the frontend owns the button. Same as the single-shot flow.

2. One round of questions, not a loop

Generate Prompt → conflicts → answer them → press again → prompt. Two presses. Always.

The trap: some resolutions rewrite a shot. If the system then sees the shot changed and re-runs, the new run produces different conflicts (this model doesn't return the same answer twice on the same input). The user answers those, more shots get rewritten, and they never reach a prompt.

So the system recognises its own writes. We keep a fingerprint of everything the checks read, and when our own resolution rewrites a shot we re-take the fingerprint at the same time. Only somebody else's edit invalidates the answer. Don't remove that — it looks redundant and it's the whole reason the flow terminates.

Same fingerprint makes the screen cheap: re-opening a selection returns the stored answer with the user's dismissals on it, for free. There's a read-only endpoint just for that, because opening a panel to look at it must not cost 20 seconds and real money.

3. Nothing the user already answered comes back

Three sources, all three have to hold:

  • The single-shot flow. Marked "Not a conflict" in the T2I screen → must not appear here. Hard requirement.
  • The sketch chat. If the sketch conversation settled "the drawing reads straight-on, the shot says low angle" in favour of the drawing, we must not offer low angle back as a fix. This one bit me on a real shot. The sketch compiler now records which field was settled and to what value — the value matters, because if someone later changes that field by hand the old decision is dead and the check should fire again.
  • Their own earlier answers here. Editing shot 4 must not re-open the decision about shot 2.

All three are enforced in code, not by asking the model nicely.

4. If the message names a field, an option must move that field

Loudest feedback on the first version. The check said "The picture seems much closer than an extreme wide shot. Consider updating Shot size" — and offered one button, "Keep the picture and tell the video model to follow it", which changed no field. Three of five findings on that sequence read the same way.

So when a picture disagrees with a metadata field and the field is the easy end to move, the fix moves the field. Telling the video model to follow the picture is the LAST option.

And a model's proposal is untrusted: before an option is shown we check in code that the field is one we may write, that a dropdown value really belongs to that dropdown, and that the new value differs from what's there. A button that visibly does nothing is worse than no button.

5. The wait has to stay in the 30–35 second range

Cost is not our problem (~$0.25 a sequence is fine). Wall clock is, because the user is sitting on the button.

The two AI analyses run at the same time, never one after the other. Chained — read all the pictures first, then hand the facts to the cross-shot check — it measured around 70 seconds. In parallel the wait is just the slower of the two: measured on real scenes the cross-shot call is 25–44s and the per-shot fan-out is 9–19s, so the cross-shot call is the wait and the other one is free.

Two measured facts that save a wasted day:

  • Four shots is faster than three. What this costs is how much the model writes, not how many shots there are.
  • Asking the model to be brief makes it stop finding things. A version told to keep its reasoning short ran ~35% faster and found nothing in 4 of 6 runs; the shipped one found something in 5 of 6. On this model the output tokens are the reasoning. To make the check stronger, turn reasoning effort up, not the prose down.

Also: the request needs a heartbeat, or the proxy closes the socket before the answer comes back.


Smaller things that are easy to miss, and every one was a real bug

  • Tell the VLM who is who in the picture. On a real sequence it reported a wardrobe conflict about "the woman in the dark coat" — and meant two different women in shots 1 and 2. False, and it came out as a blocker. The answer was already on the shot: our generated sketches paint each character as a mannequin in a fixed colour for the whole scene, and the sketch chat records a label per character. Only claim the colour for sketches we generated — a user's own drawing has no mannequins, so that would invent evidence.

  • Number the shots the way the user sees them. "Shot 13" if the card says 13, not "Shot 3" because it was third in the selection.

  • Two checks finding the same problem is ONE row. The free reference check spots two pictures of one person by the ids; the AI spots it by the prose. Same problem to the user — merge them.

  • The conflict has two texts and only one is for the user. We store a diagnostic paragraph (asset ids, compared fields, the model's justification) for Mongo and the evals; the user reads a different plain sentence. Strip the diagnostic one, and the raw model output, on the way out of the API.

  • Refuse a stale fix instead of applying it. If the shot changed after the check ran, applying the old option would destroy the user's newer words. Return a 409 and say re-run the check.

  • Order is part of the identity. 8→11→14 and 14→11→8 are different sequences — half the checks are about what happens between one shot and the next. Never sort the selection.

  • Screen-side conflicts are gated on the camera. "She was on the left, now she's on the right" is correct when the angle changed. Same for shot size: one step apart is not a conflict, two is.

  • The resolution is applied on the backend, not the frontend. Real difference from T2I, where the frontend applies the fix itself. Here one click can rewrite text, set a dropdown, repoint a reference and record a decision for the prompt — so the backend does it and hands back every shot it rewrote. The frontend needs those shots, or it keeps sending the pictures the user just resolved away.

  • A resolved decision has to reach the video prompt. Last link, invisible if you break it. "Keep the picture, follow it" has to travel into the prompt the video model receives, and the prompt needs a paragraph saying a user decision outranks the normal priority rules — otherwise our own prompt decides storyboard-vs-text by itself and silently overrules the user. The click would look like it worked and change nothing.


Data persistence

I know there is an overwrite. But I don't want data to be lost. For the T2I pregen conflict system we created a new MongoDB collection preGenValidationRuns. Worth implementing it in I2V? Maybe for the first weeks of the product, to keep track of data.

My answer after checking the code: yes, it's the same bug, so the same fix ports over. What is recoverable from Mongo today, per conflict:

What the user did Recoverable?
Marked it "Not a conflict" Yes — who, when, the reason, their note
Clicked one of the options Yes, with before and after — the option, who, when, every field changed with old and new value
Went back and changed the field by hand No — nothing ties the edit to the conflict

And every re-run replaces the whole record for that selection. A scene keeps at most 8 selections.

So the third case disappears twice: the manual edit changes the fingerprint, the next Generate Prompt re-runs, the conflict is no longer raised, and the row is gone with no trace it existed.

The loss is biased, which is the real problem. A dismissal changes no input, gets carried forward, and piles up. A fix that worked erases its own evidence, because it removes the reason the conflict was raised. Any dismissal rate from this data is the share of what was left behind, not of what was raised. On T2I that had already destroyed 409 records across 339 shots before we noticed. projectActivities doesn't save us — dismiss and resolve write no row at all, and a manual shot edit writes one with no field name and no old value.

What I'd like: one immutable document per run in its own collection — the conflict verbatim as the user received it, the shot's text at run time, the findings our filters dropped, and (back-patched by the next run) what became of each conflict: still raised / dismissed / fix applied / field changed by hand / vanished on its own. Dismissals as separate events. Nothing user-visible changes. Copy the shape from preGenValidationRuns, it's already in staging and already reviewed.

Past history can't be backfilled — it's gone. That's the argument for doing it before the creatives start.


Sketch classifier

Davide requests the system, let's make sure it's created and with the appropriate data fields.

What it is: one word on each drawing saying how the drawing was made, because that's what tells the video prompt generator how much of the picture it can trust. The names are Davide's and they mislead, but the product already uses them:

  • cocoB-real-env — preflight drew it, through the Generate Sketch flow. Blank flat-coloured mannequins: no hair, no clothing, no face.
  • cocoB-fake — built outside the platform by the creative team. Stand-ins that carry the character: a ponytail, a collar, cuffs, a belt, a cross where the face is.
  • hand-drawn — a thumbnail somebody drew. On paper or on a computer; the tool is irrelevant.

The thing we got wrong first: it is not a question about the background. All three kinds usually show the shot's real location drawn in the show's art style, so "real-env" does not mean photographic. The difference is in the figures. Our first version asked about the background, so every stylised show came back cocoB-fake — I caught it on a preflight sketch that had been downloaded and re-uploaded: same picture, two different answers.

Two rules worth keeping:

  • Make the model name the feature it saw, on a named figure — "red figure: ponytail, cuffs, belt". No named feature, no cocoB-fake. That field is what let me debug a wrong answer in one run.
  • Write the answer at generation time for what we draw ourselves. Preflight knows it drew it, so it labels its own work and the AI is never asked. Free and exact.

Decided once, never recomputed — a drawing doesn't change after it's made. Which also means a wrong stored answer is permanent until somebody deletes the field. If a shot is holding an answer from older rules, bumping the version does not re-label it.


How we'll know it works before Wednesday

These are the checks that actually caught bugs.

  1. Run a real sequence end to end in the product. Not a unit test — unit tests only prove what text was sent to the model, never what it answered.
  2. Press Generate Prompt twice on a clean sequence. Second press instant and free, still showing the first press's dismissals.
  3. Answer every conflict, then press again. You must get a prompt. A fresh round of conflicts instead means #2 above is broken.
  4. Dismiss something in the single-shot screen, then run the sequence check. It must not come back.
  5. Resolve one conflict with "keep the picture", then read the generated prompt. The user's sentence has to be in it. If it isn't, the click did nothing.
  6. Pick a sequence mixing a rendered frame with sketches. Zero conflicts about the sketches being rough.
  7. Take one shot out of the selection and re-run. Different sequence, so it re-runs — but the decisions about the shots that stayed must still be there.
  8. Watch the clock. A 3-shot check over 45 seconds means something got chained that shouldn't be.

What I need from you in the meeting

  1. Do we port the run-history collection now or after the creatives start? I want now, before there's history worth losing. Your call on effort.
  2. Which frontend are you starting from — our mock up, or Jeet's spec? I want to know so I don't send you in two directions.
  3. Who owns the sketch classifier field going to the video prompt? Nothing downstream reads it yet, so a wrong label costs nothing today. That changes the day someone uses it.
  4. Drawings already in the database with an answer from the old rules — my position is leave them, only new sketches from here on.
  5. Anything above you think we should cut for Wednesday — say it now, not in staging.