Skip to content

Sketch Generation — Frontend Integration Guide

Real-time Events (Socket)

All events are published via JaduSpinePublisher.publishToProject on the project channel. We are sending multiple scene updated / shot updated events to refresh the screen

sceneUpdated

When Message Key payload fields UI action
NER + variant selection completes "Asset linking and variant selection complete" Full fresh scene with updated charactersVisible, environment.variantId Refresh shot cards (characters, variant thumbnails)
Blocking starts "Scene blocking started" scene.preflight.blocking.status = 'computing' Show progress indicator on scene header
Blocking completes "Scene blocking computed" scene.preflight.blocking.status = 'ready' Refresh shot data, enable sketch generation
Blocking fails "Scene blocking failed: ..." scene.preflight.blocking.status = 'failed' Show error, allow retry
Re-trigger completes "Blocking refreshed for edited shots" Updated scene with refreshed plan + suggestedVariant on dirty shots Show variant suggestions on affected shots

shotUpdated

When Message Key payload fields
NER links written to one shot — { shot, storyId, sceneId, shotId }
Sketch generation succeeds "Sketch generation completed" { shot, storyId, sceneId, shotId, prerequisiteReport }
Sketch generation fails "Sketch generation failed: ..." Same shape, status 500

Pipeline Lifecycle (post-confirm_shots)

After the user confirms shots, the backend runs a fire-and-forget pipeline:

NER → Variant Selection → Blocking → Screen Projection ∥ Variant Validation

Timing & dirty-shot handling

  • A pipelineStartedAt timestamp is recorded at pipeline start
  • Before writing NER/variant results to a shot, the pipeline checks shot.updatedAt > pipelineStartedAt — if true, that shot is dirty and writes are skipped
  • On completion, any dirty shots trigger a re-trigger for their blocking group

Re-trigger flow (automatic, per dirty setup)

  1. refreshSetupBlocking — re-run blocking positions for the group (lighter LLM call, group membership stays fixed)
  2. Screen Projection — re-run for that setup (positions may have changed)
  3. Variant Validation — re-run for all shots in the setup:
  4. Dirty shots → suggestedVariant stored (user must accept/dismiss)
  5. Clean shots → auto-swap variantId if a better plate is found

Other triggers

Trigger What happens
insert_shot extendBlockingPlan — fits into existing setup or creates new one
User accepts suggestedVariant Writes variantId, clears suggestion, re-runs screen projection for that setup
User dismisses suggestedVariant Clears suggestion, nothing else

Suggested Variants

Per-shot suggestions (suggestedVariant)

When variant validation finds a better environment variant for a shot (one that covers more anchor objects), it stores a suggestion:

shot.environment: {
  title: string;
  assetId: string;
  variantId?: string;
  suggestedVariant?: {
    variantId: string;
    label: string;
    reason: string;       // e.g. "anchor objects [stove, shelf] not visible in current plate"
    suggestedAt: Date;
  };
}

Endpoints

Endpoint Action
POST /api/workbench/acceptSuggestedVariant { storyId, sceneId, shotId } — applies variant, clears suggestion, refreshes screen projection
POST /api/workbench/dismissSuggestedVariant { storyId, sceneId, shotId } — clears suggestion

Frontend behavior (Discuss with Isa. This is something we can implement later. If you implement it later add a linear ticket to keep track please)

  • When shot.environment.suggestedVariant is present, show an inline card on the shot:
  • Current plate thumbnail vs suggested plate thumbnail
  • reason text explaining why the suggestion was made
  • Accept / Dismiss buttons
  • A new suggestion overwrites the previous one (check suggestedAt)

Missing / Suggested variants (scene-level)

Stored at scene.preflight.variantSelection.missingVariants:

{
  environmentAssetId: string;
  environmentTitle: string;
  suggestedLabel: string;        // e.g. "front left lounge — Stan seated in Party Bubble"
  guidance: string;              // "Needs to show: Lounge Seating, Disco Ball..."
  usedByShotIds: string[];
  usedByShotDescriptions: string[];
}

These represent variants that don't exist yet but would be needed. Frontend should show them as actionable cards prompting the user to create/upload the variant.


Pre-generation Validation (Prerequisite Gate)

This was mainly part of testing. You might want to use the issues field as part of pregen.

POST /api/workbench/generateSketch

Before generating, the service checks prerequisites. If not met, throws 422 with:

{
  "error": "Sketch prerequisites not met:\n...",
  "prerequisiteReport": {
    "variant": "ok | missing",
    "environmentAnalysis": "ok | missing",
    "blocking": "ok | missing | computing",
    "screenProjection": "ok | missing",
    "characterAnalysis": { "asset-id-1": "ok", "asset-id-2": "missing" },
    "ready": false,
    "issues": ["blocking: missing — run computeSceneBlocking", "Character 'Alviss' has no analyzed reference image"]
  }
}

Hard gates (always block): variant, blocking, screenProjection Soft gates (block by default, override with allowMissingPrerequisites: true): environmentAnalysis, characterAnalysis

ENVIRONMENT_VARIANT_NOT_USABLE (422)

Fires after the LLM assesses the variant image. Only when the image is fundamentally wrong (different location):

{
  "error": "ENVIRONMENT_VARIANT_NOT_USABLE",
  "shotId": "...",
  "setupId": "...",
  "environmentAssetId": "...",
  "variantId": "...",
  "reasons": ["The plate shows a road between crop fields, not the required sidewalk..."]
}

Placement Overrides (user-adjusted previz → generation)

The user can preview character/prop positions and crop, adjust them, then generate with those adjustments applied.

Flow

  1. Preview: POST /api/workbench/planSketch with { storyId, sceneId, shotId } — returns center points for characters/props, crop region, and prompt text
  2. User adjusts: Frontend makes center markers draggable and crop resizable. Track whether anything moved from original positions.
  3. Generate: If nothing moved, call generateSketch normally. If anything moved:
  4. Call planSketch again with overrides — backend re-reads current shot properties and rewrites the prompt to match the new layout
  5. Call generateSketch with precomputedPlan: { prompt, cropRegion } from the override response

POST /api/workbench/planSketch (with overrides)

Request:

{
  "storyId": "...",
  "sceneId": "...",
  "shotId": "...",
  "overrides": {
    "characterPlacements": [
      { "name": "Stan Padnick", "centerX": 0.4, "centerY": 0.6 }
    ],
    "propPlacements": [
      { "name": "Party Bubble", "centerX": 0.55, "centerY": 0.7 }
    ],
    "cropRegion": { "x": 0.1, "y": 0.2, "width": 0.7, "height": 0.6 }
  }
}

All three override fields are optional — only send what the user actually changed.

Response: Same shape as without overrides:

{
  "characterPlacements": [{ "name": "Stan Padnick", "color": "green", "centerX": 0.36, "centerY": 0.625 }],
  "propPlacements": [{ "name": "Party Bubble", "centerX": 0.6, "centerY": 0.7 }],
  "cropRegion": { "x": 0.1, "y": 0.2, "width": 0.7, "height": 0.6 },
  "prompt": "reconciled prompt text...",
  "reasoning": "...",
  "plateUsable": true
}

The response includes full bboxes (x, y, width, height) from the LLM plus computed centerX/centerY. The center point is what the frontend should display and let users drag. The full bbox is available if needed for context but is not pixel-accurate to the final generation.

POST /api/workbench/generateSketch (with precomputedPlan)

Request:

{
  "storyId": "...",
  "sceneId": "...",
  "shotId": "...",
  "precomputedPlan": {
    "prompt": "the reconciled prompt from planSketch override response",
    "cropRegion": { "x": 0.1, "y": 0.2, "width": 0.7, "height": 0.6 }
  }
}

When precomputedPlan is present, the backend skips the prompt-writer LLM entirely and uses the provided prompt + crop directly for image generation.

Priority chain (how the backend reconciles overrides)

When overrides are present, the prompt-writer LLM follows this priority: 1. Visual overrides (user-dragged center points) — highest authority 2. Current shot properties (description, action, characters) — re-read from DB at call time 3. Blocking data (positions, anchors, screen order) — advisory context only, ignored where it contradicts #1 or #2

Coordinates

  • All coordinates are normalized 0–1 on the full plate
  • characterPlacements / propPlacements names must match exactly what the initial planSketch returned
  • propPlacements only includes props that need to be DRAWN (not already visible in the plate)
  • Overrides use centerX/centerY (point), crop uses x, y, width, height (rectangle)

Pre-gen Preview (experimental)

POST /api/workbench/getShotSketchContext

Returns the exact inputs that feed the prompt-writer LLM. No LLM call.

Request: { storyId, sceneId, shotId }

Response: - shotContext — full text block sent to LLM (blocking positions, anchors, zones, objects with bboxes, shot metadata, scale guidance) - characterRefs — resolved character images + mannequin colors + height - propRefs — resolved prop images + identity descriptions - colorAssignments — e.g. "Stan = green mannequin, Alviss = orange mannequin (holographic/translucent)" - plateUrl — the resolved plate image URL - envRef — environment asset reference - prerequisiteReport — same structure as above

POST /api/workbench/planSketch

Runs the prompt-writer LLM and returns character/prop placements without generating images. Accepts optional overrides to reconcile user-adjusted positions (see Placement Overrides section above).

Request: { storyId, sceneId, shotId, frame?: 'start' | 'middle' | 'end', overrides?: { characterPlacements?, propPlacements?, cropRegion? } }

Response:

{
  "characterPlacements": [
    { "name": "Stan Padnick", "color": "green", "centerX": 0.36, "centerY": 0.625 }
  ],
  "propPlacements": [
    { "name": "Party Bubble", "centerX": 0.6, "centerY": 0.7 }
  ],
  "cropRegion": { "x": 0.0, "y": 0.1, "width": 0.8, "height": 0.7 },
  "prompt": "the complete prompt for the image model...",
  "plateUsable": true,
  "unusableReasons": [],
  "reasoning": "brief explanation of choices"
}

Frontend overlay: - Draw center-point markers (dot/crosshair + label) for each character placement — colored by color field - Draw center-point markers (dot/crosshair + label) for each prop placement — use a distinct style (e.g. dashed outline) - Draw a red border rectangle for cropRegion to show framing - Character/prop markers should be draggable — track dirty state by comparing centerX/centerY to original values - Crop rectangle should be draggable/resizable - Display reasoning and prompt below the preview

Reference implementation: Branch f-miki-e2e-test has a working frontend implementation of the blocking visualization and edit UI that can serve as a starting point.


Blocking — Key Concepts

Blocking groups a scene's shots into setups — stable staging states where characters don't fundamentally change position.

  • Setup: A group of consecutive shots sharing the same character positions + environment variant
  • Character positions: Room-space descriptions anchored to named objects in the plate (e.g. "seated on the lounge bench, relation: seated_at")
  • presenceType: physical (solid mannequin), depicted_live (holographic/translucent — video calls), hologram
  • screenOrder: Left-to-right arrangement of characters on screen (computed by Screen Projection for multi-character setups)

Blocking is computed once after confirm_shots. The user owns it from that point — no automatic staleness invalidation. Extensions happen via insert_shot → extendBlockingPlan.


Debug / Admin Endpoints

Endpoint Purpose
POST /getShotSketchContext Full LLM input preview for one shot
POST /validateShotForGeneration Prerequisite check without generating
POST /planSketch LLM plan + placements without image gen
POST /selectSceneVariants Run variant selection, see recommendations
POST /computeSceneBlocking Trigger blocking computation
POST /acceptSuggestedVariant Accept a per-shot variant suggestion
POST /dismissSuggestedVariant Dismiss a per-shot variant suggestion