Sketch Generation — Frontend Integration Guide¶
Real-time Events (Socket)¶
All events are published via JaduSpinePublisher.publishToProject on the project channel. We are sending multiple scene updated / shot updated events to refresh the screen
sceneUpdated¶
| When | Message | Key payload fields | UI action |
|---|---|---|---|
| NER + variant selection completes | "Asset linking and variant selection complete" | Full fresh scene with updated charactersVisible, environment.variantId |
Refresh shot cards (characters, variant thumbnails) |
| Blocking starts | "Scene blocking started" | scene.preflight.blocking.status = 'computing' |
Show progress indicator on scene header |
| Blocking completes | "Scene blocking computed" | scene.preflight.blocking.status = 'ready' |
Refresh shot data, enable sketch generation |
| Blocking fails | "Scene blocking failed: ..." | scene.preflight.blocking.status = 'failed' |
Show error, allow retry |
| Re-trigger completes | "Blocking refreshed for edited shots" | Updated scene with refreshed plan + suggestedVariant on dirty shots |
Show variant suggestions on affected shots |
shotUpdated¶
| When | Message | Key payload fields |
|---|---|---|
| NER links written to one shot | — | { shot, storyId, sceneId, shotId } |
| Sketch generation succeeds | "Sketch generation completed" | { shot, storyId, sceneId, shotId, prerequisiteReport } |
| Sketch generation fails | "Sketch generation failed: ..." | Same shape, status 500 |
Pipeline Lifecycle (post-confirm_shots)¶
After the user confirms shots, the backend runs a fire-and-forget pipeline:
NER → Variant Selection → Blocking → Screen Projection ∥ Variant Validation
Timing & dirty-shot handling¶
- A
pipelineStartedAttimestamp is recorded at pipeline start - Before writing NER/variant results to a shot, the pipeline checks
shot.updatedAt > pipelineStartedAt— if true, that shot is dirty and writes are skipped - On completion, any dirty shots trigger a re-trigger for their blocking group
Re-trigger flow (automatic, per dirty setup)¶
- refreshSetupBlocking — re-run blocking positions for the group (lighter LLM call, group membership stays fixed)
- Screen Projection — re-run for that setup (positions may have changed)
- Variant Validation — re-run for all shots in the setup:
- Dirty shots →
suggestedVariantstored (user must accept/dismiss) - Clean shots → auto-swap
variantIdif a better plate is found
Other triggers¶
| Trigger | What happens |
|---|---|
insert_shot |
extendBlockingPlan — fits into existing setup or creates new one |
User accepts suggestedVariant |
Writes variantId, clears suggestion, re-runs screen projection for that setup |
User dismisses suggestedVariant |
Clears suggestion, nothing else |
Suggested Variants¶
Per-shot suggestions (suggestedVariant)¶
When variant validation finds a better environment variant for a shot (one that covers more anchor objects), it stores a suggestion:
shot.environment: {
title: string;
assetId: string;
variantId?: string;
suggestedVariant?: {
variantId: string;
label: string;
reason: string; // e.g. "anchor objects [stove, shelf] not visible in current plate"
suggestedAt: Date;
};
}
Endpoints¶
| Endpoint | Action |
|---|---|
POST /api/workbench/acceptSuggestedVariant |
{ storyId, sceneId, shotId } — applies variant, clears suggestion, refreshes screen projection |
POST /api/workbench/dismissSuggestedVariant |
{ storyId, sceneId, shotId } — clears suggestion |
Frontend behavior (Discuss with Isa. This is something we can implement later. If you implement it later add a linear ticket to keep track please)¶
- When
shot.environment.suggestedVariantis present, show an inline card on the shot: - Current plate thumbnail vs suggested plate thumbnail
reasontext explaining why the suggestion was made- Accept / Dismiss buttons
- A new suggestion overwrites the previous one (check
suggestedAt)
Missing / Suggested variants (scene-level)¶
Stored at scene.preflight.variantSelection.missingVariants:
{
environmentAssetId: string;
environmentTitle: string;
suggestedLabel: string; // e.g. "front left lounge — Stan seated in Party Bubble"
guidance: string; // "Needs to show: Lounge Seating, Disco Ball..."
usedByShotIds: string[];
usedByShotDescriptions: string[];
}
These represent variants that don't exist yet but would be needed. Frontend should show them as actionable cards prompting the user to create/upload the variant.
Pre-generation Validation (Prerequisite Gate)¶
This was mainly part of testing. You might want to use the issues field as part of pregen.
POST /api/workbench/generateSketch¶
Before generating, the service checks prerequisites. If not met, throws 422 with:
{
"error": "Sketch prerequisites not met:\n...",
"prerequisiteReport": {
"variant": "ok | missing",
"environmentAnalysis": "ok | missing",
"blocking": "ok | missing | computing",
"screenProjection": "ok | missing",
"characterAnalysis": { "asset-id-1": "ok", "asset-id-2": "missing" },
"ready": false,
"issues": ["blocking: missing — run computeSceneBlocking", "Character 'Alviss' has no analyzed reference image"]
}
}
Hard gates (always block): variant, blocking, screenProjection
Soft gates (block by default, override with allowMissingPrerequisites: true): environmentAnalysis, characterAnalysis
ENVIRONMENT_VARIANT_NOT_USABLE (422)¶
Fires after the LLM assesses the variant image. Only when the image is fundamentally wrong (different location):
{
"error": "ENVIRONMENT_VARIANT_NOT_USABLE",
"shotId": "...",
"setupId": "...",
"environmentAssetId": "...",
"variantId": "...",
"reasons": ["The plate shows a road between crop fields, not the required sidewalk..."]
}
Placement Overrides (user-adjusted previz → generation)¶
The user can preview character/prop positions and crop, adjust them, then generate with those adjustments applied.
Flow¶
- Preview:
POST /api/workbench/planSketchwith{ storyId, sceneId, shotId }— returns center points for characters/props, crop region, and prompt text - User adjusts: Frontend makes center markers draggable and crop resizable. Track whether anything moved from original positions.
- Generate: If nothing moved, call
generateSketchnormally. If anything moved: - Call
planSketchagain withoverrides— backend re-reads current shot properties and rewrites the prompt to match the new layout - Call
generateSketchwithprecomputedPlan: { prompt, cropRegion }from the override response
POST /api/workbench/planSketch (with overrides)¶
Request:
{
"storyId": "...",
"sceneId": "...",
"shotId": "...",
"overrides": {
"characterPlacements": [
{ "name": "Stan Padnick", "centerX": 0.4, "centerY": 0.6 }
],
"propPlacements": [
{ "name": "Party Bubble", "centerX": 0.55, "centerY": 0.7 }
],
"cropRegion": { "x": 0.1, "y": 0.2, "width": 0.7, "height": 0.6 }
}
}
All three override fields are optional — only send what the user actually changed.
Response: Same shape as without overrides:
{
"characterPlacements": [{ "name": "Stan Padnick", "color": "green", "centerX": 0.36, "centerY": 0.625 }],
"propPlacements": [{ "name": "Party Bubble", "centerX": 0.6, "centerY": 0.7 }],
"cropRegion": { "x": 0.1, "y": 0.2, "width": 0.7, "height": 0.6 },
"prompt": "reconciled prompt text...",
"reasoning": "...",
"plateUsable": true
}
The response includes full bboxes (x, y, width, height) from the LLM plus computed centerX/centerY. The center point is what the frontend should display and let users drag. The full bbox is available if needed for context but is not pixel-accurate to the final generation.
POST /api/workbench/generateSketch (with precomputedPlan)¶
Request:
{
"storyId": "...",
"sceneId": "...",
"shotId": "...",
"precomputedPlan": {
"prompt": "the reconciled prompt from planSketch override response",
"cropRegion": { "x": 0.1, "y": 0.2, "width": 0.7, "height": 0.6 }
}
}
When precomputedPlan is present, the backend skips the prompt-writer LLM entirely and uses the provided prompt + crop directly for image generation.
Priority chain (how the backend reconciles overrides)¶
When overrides are present, the prompt-writer LLM follows this priority: 1. Visual overrides (user-dragged center points) — highest authority 2. Current shot properties (description, action, characters) — re-read from DB at call time 3. Blocking data (positions, anchors, screen order) — advisory context only, ignored where it contradicts #1 or #2
Coordinates¶
- All coordinates are normalized 0–1 on the full plate
characterPlacements/propPlacementsnames must match exactly what the initialplanSketchreturnedpropPlacementsonly includes props that need to be DRAWN (not already visible in the plate)- Overrides use
centerX/centerY(point), crop usesx, y, width, height(rectangle)
Pre-gen Preview (experimental)¶
POST /api/workbench/getShotSketchContext¶
Returns the exact inputs that feed the prompt-writer LLM. No LLM call.
Request: { storyId, sceneId, shotId }
Response:
- shotContext — full text block sent to LLM (blocking positions, anchors, zones, objects with bboxes, shot metadata, scale guidance)
- characterRefs — resolved character images + mannequin colors + height
- propRefs — resolved prop images + identity descriptions
- colorAssignments — e.g. "Stan = green mannequin, Alviss = orange mannequin (holographic/translucent)"
- plateUrl — the resolved plate image URL
- envRef — environment asset reference
- prerequisiteReport — same structure as above
POST /api/workbench/planSketch¶
Runs the prompt-writer LLM and returns character/prop placements without generating images. Accepts optional overrides to reconcile user-adjusted positions (see Placement Overrides section above).
Request: { storyId, sceneId, shotId, frame?: 'start' | 'middle' | 'end', overrides?: { characterPlacements?, propPlacements?, cropRegion? } }
Response:
{
"characterPlacements": [
{ "name": "Stan Padnick", "color": "green", "centerX": 0.36, "centerY": 0.625 }
],
"propPlacements": [
{ "name": "Party Bubble", "centerX": 0.6, "centerY": 0.7 }
],
"cropRegion": { "x": 0.0, "y": 0.1, "width": 0.8, "height": 0.7 },
"prompt": "the complete prompt for the image model...",
"plateUsable": true,
"unusableReasons": [],
"reasoning": "brief explanation of choices"
}
Frontend overlay:
- Draw center-point markers (dot/crosshair + label) for each character placement — colored by color field
- Draw center-point markers (dot/crosshair + label) for each prop placement — use a distinct style (e.g. dashed outline)
- Draw a red border rectangle for cropRegion to show framing
- Character/prop markers should be draggable — track dirty state by comparing centerX/centerY to original values
- Crop rectangle should be draggable/resizable
- Display reasoning and prompt below the preview
Reference implementation: Branch f-miki-e2e-test has a working frontend implementation of the blocking visualization and edit UI that can serve as a starting point.
Blocking — Key Concepts¶
Blocking groups a scene's shots into setups — stable staging states where characters don't fundamentally change position.
- Setup: A group of consecutive shots sharing the same character positions + environment variant
- Character positions: Room-space descriptions anchored to named objects in the plate (e.g. "seated on the lounge bench, relation: seated_at")
- presenceType:
physical(solid mannequin),depicted_live(holographic/translucent — video calls),hologram - screenOrder: Left-to-right arrangement of characters on screen (computed by Screen Projection for multi-character setups)
Blocking is computed once after confirm_shots. The user owns it from that point — no automatic staleness invalidation. Extensions happen via insert_shot → extendBlockingPlan.
Debug / Admin Endpoints¶
| Endpoint | Purpose |
|---|---|
POST /getShotSketchContext |
Full LLM input preview for one shot |
POST /validateShotForGeneration |
Prerequisite check without generating |
POST /planSketch |
LLM plan + placements without image gen |
POST /selectSceneVariants |
Run variant selection, see recommendations |
POST /computeSceneBlocking |
Trigger blocking computation |
POST /acceptSuggestedVariant |
Accept a per-shot variant suggestion |
POST /dismissSuggestedVariant |
Dismiss a per-shot variant suggestion |