Skip to content

AssetGen Quality Analysis

After an asset-gen job finishes, this module has an LLM look at the output, score it 1–5 against what the job was meant to produce, and decide what to do next: accept it, regenerate with a better prompt, edit the image, or hand it to a human. When a fix is chosen it can run automatically as a new job, and that job is scored again, up to a small number of attempts.

The generic machinery (the harness) handles scheduling, the scale, the judge call, the attempt chain, credits, events and de-duplication. Each kind of asset plugs in as a use case that says what to judge, what the judge should see, and which fixes are allowed. Today there is one use case, workbench shot images.

A failed or skipped analysis never flips the generated asset to failed. Unfamiliar terms are defined in the Glossary at the end.


Enablement

Flag Effect
QUALITY_ANALYSIS_ENABLED=true Master switch for the harness. Off → no evaluate, no stamp.

Defaults off. Opt a use case in by stamping jobMetaData.qualityAnalysisLineage.useCaseId at launch and leaving enabled: true on the use case object. Flip enabled: false in code to take a use case out without an env var.

The harness only runs stamped jobs. Everything else that tunes behaviour (pass threshold, allowed actions, auto-apply, judge model) is set per use case; see Configuration. The old Python /scoreAsset / /rewritePrompt path is unchanged and deprecated.


Flow

job settles (runJob / pollJob / webhook / updateJobAsFailed)
        │
        ▼
QualityAnalysisHarness.scheduleOnJobSettled  (fire-and-forget)
        │  resolve use case from jobMetaData.qualityAnalysisLineage.useCaseId
        │
        ├── batch root slot (jobBatchId, attempt 1, useCase.batch.enabled) ──► batch path, see below
        │
        │  (single job, COMPLETED only)
        │  claim: qualityCheckStatus → PROCESSING (one atomic write; a second caller gets null and stops)
        ▼
useCase.evaluate(job)  OR  LLM judge (gpt-5.4 / medium)
        │  overall capped by lowest CRITICAL criterion
        │  score wins over a contradicting ACCEPT recommendation
        ▼
verdict.pass ──► followUp.status = ACCEPTED
        │
        ├── attempt >= maxAttempts ──► EXHAUSTED
        │
        └── PROPOSED { action, requiresReview }
                │
                ├── shouldAutoApply ──► applyFollowUp({ triggeredBy: 'AUTO' })
                └── else wait for POST /applyFollowUp
                                    │
                                    ▼
                        child job (attempt + 1) → re-enters harness on completion

Manual POST /evaluate re-runs a COMPLETED check, appends to verdictHistory, and does not auto-apply. Terminal follow-up statuses (APPLIED, DISMISSED, EXHAUSTED, ACCEPTED) are left in place; only PROPOSED may be replaced.

Batch path (comparative judge)

Shot images are generated N at a time (generateShotImageInitialBatch, 2–4 slots sharing jobMetaData.jobBatchId). When the use case sets batch.enabled, the harness judges the slots together instead of one at a time. The trigger is generic: nothing in the workbench calls the harness for batches, every job sink already does.

Assumption: all slots in a batch were generated from the same prompt and inputs. The harness calls buildContext once, on the first completed candidate, and sends that single brief (intent, generation prompt, references) with all N images. This holds for createJobInitialBatch, which fans one config out N times. A launch path whose slots differ in prompt or references must not be routed through the comparative judge, otherwise N−1 slots are scored against the wrong brief.

any root slot settles (COMPLETED | FAILED | CANCELLED)
        │  list siblings by jobBatchId, keep root slots (attempt 1, no parentJobId)
        │  any sibling non-terminal ──► return; the last one to settle triggers
        │  candidates = COMPLETED siblings with a result
        │    0 → nothing; 1 → single-slot path (nothing to compare)
        │  claim the batch: assetGenJobBatches.qualityAnalysis.qualityCheckStatus → PROCESSING (atomic, dedupes concurrent finishers)
        │  each candidate: qualityCheckStatus → PROCESSING, STARTED event with { batch: { batchId, slot } }
        ▼
one LLM call: candidates as Image 1..N (slot order) + shared references
        │  per candidate: criteria, overall (critical cap), issues, own recommended action
        │  ranking (best first), sharedIssues (flaws in EVERY candidate), requiresReview
        │  re-read candidates (an apply/dismiss during the call must not be overwritten)
        ▼
per slot: normal Verdict + qualityAnalysis.batch = { batchId, slot, rank, isBest }
batch document: qualityAnalysis = { qualityCheckStatus: COMPLETED, summary, summaryHistory, judgedJobIds }
        │
        ├── pin rule: any slot APPLIED / ACCEPTED / EXHAUSTED / user-DISMISSED ──► REEVALUATED on every slot, no follow-up change
        ├── ≥1 slot passes ──► passing slots ACCEPTED; failing slots DISMISSED by 'BATCH' (own action kept)
        ├── none pass, attempt < maxAttempts ──► best slot PROPOSED; others DISMISSED by 'BATCH'
        └── none pass, attempt ≥ maxAttempts ──► best slot EXHAUSTED; others DISMISSED by 'BATCH'
        │
        ▼
BATCH_JUDGED event (once), then per-slot phase events with { batch: pointer }
major sharedIssues ──► logger.warn QualityAnalysisBatchSharedFlaw (Sentry issue: the prompt is at fault, not one generation)
  • At most one harness proposal per batch. Auto-apply on the best slot follows the normal rules; shot images leave allowAutoApplyInBatch off, so it stays a proposal.
  • Override. A batch-dismissed slot keeps its own followUp.action. POST /applyFollowUp on it moves DISMISSED(BATCH) → APPLIED and creates a child from that slot; the child is judged via the single-slot path. A slot a person dismissed cannot be applied. The best slot's pending PROPOSED is not touched by an override.
  • Pin rule. Once anyone has decided on a slot, re-evaluation only refreshes verdicts (and verdictHistory on every slot) and never opens a new proposal anywhere in the batch.
  • Slot numbers are 1-based positions in the batch document's slotJobIds, so a failed slot keeps its number and the judge's Image k for k ≤ N is candidate k in that order. Batches created before the document existed fall back to createdAt then jobId.
  • Failure. Anything that throws after the batch claim (marking slots PROCESSING, the judge, persisting a slot) marks the batch document qualityCheckStatus: FAILED with error and releases the claim, so a manual evaluate can retry. Slots not yet persisted get qualityCheckStatus: FAILED and a FAILED event; slots already persisted with a verdict keep their COMPLETED state and follow-up (an ACCEPTED slot flipped to FAILED would pin the batch and make the retry a no-op). Use-case hooks (onAccepted, onExhausted, onBatchJudged) are logged and swallowed, never fail the batch. A claim left PROCESSING by a process crash still blocks manual re-evaluation with "already running" (same limitation as the single path; no lease yet).
  • Edit batches (dispatchShotImageEditBatch) and reject-and-regenerate batches (rejectAndRegenerateShotImageBatch) are not stamped and never enter the harness; each carries a TODO(quality-analysis follow-up) describing what stamping them needs. Restart-from-sketch batches reuse dispatchShotImageEditBatch but pass their stamp via slotJobMetaData, so they are judged like initial batches.

Files

File Role
qualityAnalysis.harness.ts Settle hook, claim the check (per job or per batch), evaluate, persist, decide, publish.
qualityAnalysis.batch.ts Pure batch helpers: isBatchRootSlot, sortBatchSlots, slotVerdictsFromBatchOutput, hasUserDecidedFollowUp (pin rule), resolveBatchOutcome, majorSharedIssues.
qualityAnalysis.followUp.ts resolveAction, shouldAutoApply, applyFollowUp (also on batch-dismissed slots), dismissFollowUp, default buildNextJob.
qualityAnalysis.judge.ts LLM judge (OpenAIService.generateTextWithCompletions, JSON schema, one retry) and the comparative runLlmBatchJudge. Prompts live in src/promptTemplates/qualityAnalysis/qualityAnalysisJudge.template.ts and qualityAnalysisBatchJudge.template.ts.
src/models/assetGenJobBatch.model.ts Generic assetGenJobBatches collection (asset-gen layer). The harness only writes its qualityAnalysis subdocument.
qualityAnalysis.scale.ts 1–5 anchors, critical-criterion cap, score fallback action.
qualityAnalysis.lineage.ts Attempt / root / parent stamps. loadRootJob for original generation inputs.
qualityAnalysis.registry.ts Use cases register at module load.
qualityAnalysis.events.ts JaduSpine events (project channel when projectLinks.projectId exists).
qualityAnalysis.controller.ts / .validator.ts HTTP endpoints.
useCases/shotImage.useCase.ts First LLM use case.
scripts/qualityAnalysisBatchEval.ts Offline runner for the comparative judge on an existing batch (see Manual runs).

Data

Two places on the job:

jobMetaData.qualityAnalysisLineage (type QualityAnalysisLineage) — set at launch, copied onto children.

{
  useCaseId: string;
  rootJobId: string;       // attempt-1 job; filled in createJob / createJobInitial
  parentJobId?: string;
  attempt: number;         // 1-based; only applyFollowUp increments
  promptSource: 'GENERATED' | 'USER_EDITED';
  context?: Record<string, unknown>; // use-case data the judge needs later (shot image: generationRefs)
}

job.qualityAnalysis (type QualityAnalysisData in sharedTypes.ts) — harness result. LLM use cases write score as 1–5. The harness never writes rewrittenPrompt: a regenerate prompt lives on followUp.action.prompt, next to instruction for edits and reason for escalations. The deprecated Python path writes 0–10 score, rewrittenPrompt, transformedScore and actionableFeedback.

{
  qualityCheckStatus?: 'processing' | 'completed' | 'failed';
  score?: number;
  reasoning?: string;
  verdict?: Verdict;
  verdictHistory?: Verdict[];
  followUp?: {
    status: 'ACCEPTED' | 'PROPOSED' | 'APPLIED' | 'DISMISSED' | 'EXHAUSTED';
    action?: NextAction;
    applied?: NextAction;
    appliedJobId?: string;
    triggeredBy?: 'AUTO' | 'BATCH' | string;   // 'BATCH' = parked by the batch judge, still applicable by a person
    requiresReview: boolean;
    reason?: string;
  };
  batch?: QualityAnalysisBatchPointer; // { batchId, slot, rank, isBest } — written by the batch judge only
}

NextAction is one of: ACCEPT, REGENERATE { prompt }, EDIT_PREVIOUS { instruction }, ESCALATE { reason }.

Batch document (assetGenJobBatches, type AssetGenJobBatchType in sharedTypes.ts) — one per batch, created by createJobInitialBatch and owned by the asset-gen layer. The harness writes only qualityAnalysis:

{
  batchId, userId, slotJobIds: string[], batchSize, createdVia?, storyLinks?, projectLinks?, createdAt, updatedAt,
  qualityAnalysis?: BatchQualityAnalysisData; // = {
    useCaseId: string;
    qualityCheckStatus: 'processing' | 'completed' | 'failed';   // the batch-level claim, same field name as on the job
    judgedJobIds: string[];                          // completed slots sent to the judge
    summary?: BatchVerdictSummary;                   // { ranking: [{ jobId, slot, rank, overall, pass }], bestJobId, anyPassed,
    summaryHistory?: BatchVerdictSummary[];          //   sharedIssues, requiresReview, reasoning, evaluator, evaluatedAt }
    error?: string;
  };
}

The best slot is a recommendation only (isBest on the slot, summary.bestJobId on the document). shot.images[].isSelected and shot.batches[].selectedAssetGenJobId stay user-driven.


Scale and action resolution

Pass threshold is useCase.passThreshold, default 4 (DEFAULT_PASS_THRESHOLD). pass is overall >= passThreshold. Critical criteria cap overall: overall cannot exceed the lowest critical score. A criterion the judge marked unjudgeable (occluded, too small, no reference to compare) never caps.

Each criterion score carries observed (what the judge saw for that criterion) next to reason (why that score).

Score fallback when the judge's recommendation is missing, not allowed, or is ACCEPT on a failing score. The bands are derived from the pass threshold, so raising or lowering it moves all of them; there is no separate edit or regenerate threshold.

Overall Action
≥ threshold ACCEPT
threshold − 1 EDIT_PREVIOUS if every major issue is localized and edit is allowed; else REGENERATE
below that REGENERATE
Final attempt below threshold ESCALATE if allowed

With the default threshold that is 4–5 / 3 / 1–2.

REGENERATE with an empty prompt copies the root job prompt. If that is also empty, the harness escalates instead of inventing a prompt.


Root resolution

EDIT_PREVIOUS children have an i2i config whose image input is the previous output. Grading and a later REGENERATE must not treat that as the original generation.

  • buildContext (shot image) and empty-prompt fill load root prompt + reference images via loadRootJob.
  • Default REGENERATE clones the root modelConfig and replaces the prompt.
  • EDIT_PREVIOUS still edits the immediate parent output.

Auto-apply

All of these must hold:

  • reviewPromptBeforeRun is false
  • verdict does not requiresReview
  • promptSource is not USER_EDITED
  • attempt <= autoApplyAttempts
  • not an API job (createdVia: API — created via /apiViaKey/assetGen, not Studio). Auto-apply would spawn extra charged generations the API client did not request and has no Studio UI to review.
  • not a batch job, unless allowAutoApplyInBatch
  • if the parent was auto-applied, this attempt's overall must be higher than the parent's

Shot image: { autoApplyAttempts: 1, reviewPromptBeforeRun: false, maxAttempts: 3 } — attempt 1 may auto-regenerate/edit; attempt 2 waits.

Follow-up generations are charged as normal asset-gen jobs. Judge scoring is logged, not charged. The verdict stores no log ids; the judge's analytics log entry (QUALITY_ANALYSIS_JUDGE) carries jobId (single path) or batchId (comparative path) plus the job's user, project and story links in its context, so it is found from the log side.


Endpoints

Mounted under /assetGen/qualityAnalysis. Owner or ACCESS_ALL_JOBS.

Route Body Behaviour
POST /evaluate { jobId } Manual (re-)evaluation. On a batch root slot the whole batch is re-judged: the response is the requested slot, siblings arrive via events, every slot's verdictHistory grows. Fails while siblings are still running.
POST /applyFollowUp { jobId, override? } Moves PROPOSED → APPLIED in one atomic write (a concurrent apply gets the same child back), creates the child. Also accepts a slot DISMISSED by 'BATCH' (moves it to APPLIED from its own action); a user-dismissed slot is rejected. Override may change type, prompt / instruction, and modelConfigId (a server-side config id from the use case allowlist — never a full modelConfig). Changing prompt/instruction text or picking a model stamps USER_EDITED.
POST /dismissFollowUp { jobId } DISMISSED. No child, no attempt increment. PROPOSED only.
GET /lineage/:rootJobId — Jobs sharing rootJobId, sorted by attempt.
POST /humanScore { jobId, humanScore: 1–5 } Calibration signal on qualityAnalysis.humanScore.

Generic, on the asset-gen router (not QC-specific): GET /assetGen/batch/:batchId returns { batch, jobs } — the batch document and its slot jobs in slot order. Owner of the first slot or ACCESS_ALL_JOBS.

POST /scoreAsset and POST /rewritePrompt are deprecated and untouched by the harness. They remain the Python AssetGen-page path.


Realtime

Published to the project channel when projectLinks.projectId exists, otherwise the user's STUDIO channel.

One event: qualityAnalysisUpdated. Discriminate on data.phase:

phase When Extra
STARTED Judge running { job }
PROPOSED Follow-up waiting { job }
APPLIED Child created { job, childJob }
ACCEPTED Passed { job }
EXHAUSTED Max attempts, still failing { job }
DISMISSED Human dismissed the proposal, or the batch judge parked the slot (job.qualityAnalysis.followUp.triggeredBy === 'BATCH') { job }
REEVALUATED Verdict refreshed, follow-up unchanged { job }
FAILED Judge threw { job, error } (status 500)
BATCH_JUDGED Once per comparative batch run, before the per-slot phases { job, batch } — job is the first judged slot (kept so data.job readers keep working), batch is the batch document with qualityAnalysis.summary

On the batch path every per-slot event (STARTED through FAILED) also carries data.batch = { batchId, slot, rank?, isBest? } (rank and isBest are absent on STARTED / FAILED).

The harness does not emit assetGenQualityCheckResponse. That event is only the deprecated Python /scoreAsset path.

Shot-image follow-ups run through WorkbenchService.runShotImageGenerationAsync, which commits shot media and publishes shotUpdated (including on failure). Do not use runJobAsync for that use case — a failed child would leave the tile spinning.


Use cases

Register at module load (registerQualityAnalysisUseCase). The harness side-effect-imports shotImage.useCase.ts; add a new import there (or next to it) so the use case is loaded.

Configuration

Everything that tunes the loop lives on the use case object (QualityAnalysisUseCase in qualityAnalysis.types.ts). There are no env vars per use case and nothing is read from the job except the launch stamp.

Field Default Effect
id — Matches jobMetaData.qualityAnalysisLineage.useCaseId.
enabled true false takes the use case out of the harness without an env var.
criteria — { key, label, critical?, anchors? }. critical caps overall. anchors are the 1 / 3 / 5 descriptions shown to the judge.
judgeRules none Extra rules appended to the judge system prompt.
batch.enabled false Judge the root slots of a jobBatchId batch together (see Batch path). Requires the launch path to go through createJobInitialBatch and assumes every slot shares the same prompt, references and intent.
batch.judgeRules none Ranking / shared-flaw rules appended after judgeRules in the comparative prompt.
passThreshold 4 Pass line and the origin of the action bands (see Scale and action resolution).
allowedActions — Subset of ACCEPT, REGENERATE, EDIT_PREVIOUS, ESCALATE. A recommendation outside it falls back to the score bands.
followUpPolicy.autoApplyAttempts — Highest attempt number that may auto-apply. 1 means only attempt 1.
followUpPolicy.reviewPromptBeforeRun — true disables auto-apply entirely.
followUpPolicy.maxAttempts — Chain length. The final attempt escalates instead of proposing.
followUpPolicy.allowAutoApplyInBatch false Let batch slots auto-apply.
promptInputId 'prompt' Input id read when a REGENERATE needs the root prompt.
judge.modelId / judge.reasoningEffort gpt-5.4 / medium Judge model.
buildContext(job) — Intent, named reference images, notes and generation prompt for the judge. Optional extraCriteria / extraJudgeRules apply to that job only.
evaluate(job) none Replace the LLM judge with your own verdict.
buildEditModelConfig / buildRegenerateModelConfig / buildNextJob see Adding a use case How follow-up jobs are built.
runFollowUpJob and the on* hooks none Own run path and side effects. onBatchJudged(candidates, batch) fires once per comparative run after the slots and the batch document are persisted.

Per job, the only input is the launch stamp: initAttemptOneMeta(useCaseId, context?) on jobMetaData.qualityAnalysisLineage. context carries use-case data the judge needs later (shot image stores its generationRefs there) and is copied onto children. Everything else about a chain (rootJobId, parentJobId, attempt, promptSource) is managed by the harness.

SHOT_IMAGE_V1

Stamped by workbench shot-image launch and by the tweak chat's restart-from-sketch (single and batch, shotImageEditGuru.restartFromSketch) when QUALITY_ANALYSIS_ENABLED is on and shotImageUseCase.enabled is true. Restart also puts the user's typed feedback on jobMetaData.userFeedback, the field pre-gen review already uses.

  • Criteria: character identity vs. named character references (critical; face and defining head gear decide recognition, costume only confirms it; a different face or missing signature item is a 2, a missing costume piece a 3), art style vs. the character / environment references and the art-style text in the generation prompt (critical; different medium is a 1), composition vs. sketch (placement, relative scale and pose), camera framing vs. sketch and shot size, artifacts (critical; the judge counts hands, arms, legs and faces and attributes each to a character before scoring), prompt adherence limited to acting, props and scene details that no other criterion covers. Composition / camera anchors use the pass / minor / major definitions from the T2I composition-fidelity evaluation (same left/right and foreground/background arrangement, relative scale, camera distance and angle). Anchors may define any subset of the 1–5 rungs; missing rungs fall back to the generic scale.
  • Judge: gpt-5.4 / medium unless judge is overridden. judgeRules tell the judge the sketch is a layout guide only (never judge style or identity against it), that sketches may be renders of the scene with green / blue / orange mannequins that look like an output (a mannequin means it is the sketch; re-check the image labels before scoring), that a sketch unrelated to the intent makes composition unjudgeable, and that identity is judged against the named character references. Both judge templates also require issue severity to agree with the criterion score (a major issue means that criterion is 3 or lower). This is prompt-only for now; enforcing it in code after parsing is a TODO.
  • Judge context (buildContext):
  • intent is built from shot fields: description, acting, the frame the sketch depicts (sketches[].connectedTo → initial / middle / final frame instructions; all frames listed when the moment is unrecorded), camera (shot size, angle, details), characters visible, environment, props. A trailing Requested change: line carries the root job's jobMetaData.userFeedback when set.
  • Only when that feedback is set, the context also carries extraCriteria / extraJudgeRules with the requested change criterion (critical; no precedence over the sketch, so a change away from the sketch shows as a high requested change next to a lower composition) and its rule. The harness merges them into the use case with withContextExtras for the prompt and the critical cap. Jobs without feedback are judged with the use case's criteria and rules unchanged.
  • generationPrompt is the root job prompt, shown under its own heading. Its Image [0] indices refer to the image model's inputs, not the judge's labels.
  • references are the root job's image inputs with a named role. Roles come from jobMetaData.qualityAnalysisLineage.context.generationRefs when stamped (see below), else from matching URLs against the shot's sketches (raw and annotated) and the story's character / environment / prop images (including variants). An unresolved first image on a shot with sketches is treated as the sketch, because generateFromSketch always sends the sketch first. Anything else is reference (role unknown).
  • withShotImageGenerationRefs stamps a compact copy of generateFromSketch's generationRefs into jobMetaData.qualityAnalysisLineage.context on the batch launch path and in processSketchToShotInputs, so the judge names references as they were at generation time.
  • onFollowUpCreated writes a PROCESSING placeholder via addShotImage.
  • runFollowUpJob calls runShotImageGenerationAsync (dynamic import to avoid a cycle with workbench.service).
  • buildEditModelConfig defaults to i2i-gpt-image-2. A human apply may pass override.modelConfigId: shot-image-gen (REGENERATE only), i2i-gpt-image-2, or google-i2i-nanoBananaPro.
  • batch.enabled is on for the initial batch (generateShotImageInitialBatch, count 2–4). batch.judgeRules rank sketch composition and named character references above rendering polish and put any critical failure last. No onBatchJudged: the ranking is already on each slot and on the batch document, and selection stays with the user.

Adding a use case

  1. Create useCases/<name>.useCase.ts implementing QualityAnalysisUseCase.
  2. Register at module load. Side-effect-import it from qualityAnalysis.harness.ts.
  3. Stamp initAttemptOneMeta(YOUR_USE_CASE_ID) on jobs at launch. Gate with QUALITY_ANALYSIS_ENABLED and useCase.enabled (default true). No per-use-case env vars.
  4. If the use case has its own run path (commit media, publish an event, mark FAILED on the tile), implement runFollowUpJob. Otherwise the default is AssetGenJobService.runJobAsync.
  5. If EDIT_PREVIOUS is allowed, implement buildEditModelConfig. If a human may pick a model on REGENERATE, implement buildRegenerateModelConfig. Do not import another use case's model registry from followUp.ts.
  6. Override buildNextJob only when the default (clone root config / edit parent output) is wrong.

Judge cost is not charged. Follow-up jobs are.


Tests

tests/assetGen/qualityAnalysis/ — scale, follow-up policy, lineage/root resolution, harness claim and re-eval (single and batch), batch helpers (qualityAnalysis.batch.test.ts), judge parse retry (single and comparative), shot-image context. tests/assetGen/assetGenJobBatch.model.test.ts covers the batch document and its atomic claim on in-memory Mongo.

npx vitest run tests/assetGen/qualityAnalysis tests/assetGen/assetGenJobBatch.model.test.ts

Manual runs

scripts/qualityAnalysisBatchEval.ts runs the comparative judge on an existing batch from the shell, using MONGODB_URI and OPENAI_API_KEY from .env. Useful for tuning batch.judgeRules and checking a verdict without regenerating.

npx ts-node scripts/qualityAnalysisBatchEval.ts --batch <jobBatchId>          # or --job <any slot jobId>
npx ts-node scripts/qualityAnalysisBatchEval.ts --batch <jobBatchId> --json   # raw judge output + derived summary
npx ts-node scripts/qualityAnalysisBatchEval.ts --job <jobId> --write         # go through the harness, persist + events

Read-only by default: it prints the brief the judge sees (intent, generation prompt, image labels), the ranking, shared issues, per-slot criteria with observed/reason, each slot's own recommended action, and the follow-up resolveBatchOutcome would set. Only the OpenAI log entry is written. Slots that predate the lineage stamp are judged with --useCase (default SHOT_IMAGE_V1); --write refuses them because the harness needs the stamp.


Future work

Raised in review; deferred to follow-up PRs. Each has a matching TODO(quality-analysis follow-up) comment in the code.

Item Where Idea
Few-shot calibration qualityAnalysis.scale.ts LLM judges compress scores toward 3–4. Add optional calibrationExamples (image, human score, reason) per use case, appended to the judge call as labelled anchor images. Source them from POST /humanScore.
Stamp edit / regenerate batches workbench.service.ts (dispatchShotImageEditBatch, rejectAndRegenerateShotImageBatch) Only the initial batch is judged comparatively. Edit batches need an intent built from the edit instruction and reference roles for the edited source; regenerate batches need the user's rejection feedback folded into the intent.
Auto-select the best slot storyVideos.model.ts (BatchAsset) The batch judge only recommends (isBest, summary.bestJobId). Flip isSelected / selectedAssetGenJobId from the ranking once it has been validated against human picks.
Learning from lineage qualityAnalysis.harness.ts (verdictHistory) verdictHistory and the parent → applied prompt → child chain already record which rewrites improved scores. Mine that offline and feed it back into the judge or rewrite prompt.
Adversarial prompt distillation useCases/shotImage.useCase.ts (buildShotIntent) Run a heavier model once per prompt version to distill acceptance assertions and an adversarial checklist, and decide how to split QC into specialized judge calls. QC consumes the stored assertions. First step: pass assertions via context.notes.

Glossary

Term Meaning
Harness The generic loop in this folder: claim the check, run the judge, store the verdict, decide the follow-up, publish events.
Use case One kind of asset plugged into the harness (e.g. SHOT_IMAGE_V1). Supplies criteria, judge context and allowed actions. See Configuration.
Judge The LLM call that scores the output. Returns a score per criterion, an overall score, issues and a recommended action.
Verdict The stored result of one judge run: overall, pass, per-criterion scores, issues, recommendation, reasoning.
Criterion One thing the judge scores (identity, composition, artifacts, …). Critical criteria cap the overall score. Unjudgeable means the judge could not assess it from the image; it never caps.
Anchors The 1 / 3 / 5 descriptions for a criterion that tell the judge what each score looks like.
Pass threshold Overall score at or above which the output is accepted. Default 4. The edit / regenerate bands are derived from it.
Recommendation vs. action The judge recommends an action; the harness resolves the action actually taken, falling back to the score bands when the recommendation is missing, not allowed, or ACCEPT on a failing score.
Follow-up The action taken after a failing verdict: REGENERATE { prompt }, EDIT_PREVIOUS { instruction } or ESCALATE { reason }. followUp.status tracks whether it is proposed, applied, dismissed, accepted or exhausted.
Localized An issue confined to a region an image edit can fix (a hand, a prop). All major issues localized → EDIT_PREVIOUS; otherwise → REGENERATE.
requiresReview The judge asked for a human to read the proposed prompt before it runs. Blocks auto-apply.
Auto-apply The harness runs the follow-up itself instead of waiting for a human. Gated by the rules under Auto-apply.
Stamp Writing jobMetaData.qualityAnalysisLineage onto a job when it is created. Presence of useCaseId is what opts the job into the harness.
Lineage / chain The sequence of jobs produced by follow-ups. Root is the attempt-1 job with the original inputs, parent is the job a follow-up was applied to, attempt is 1-based.
Claim The single atomic write that moves qualityCheckStatus to processing. A second concurrent caller gets null and stops, so a job is never judged twice at once. The batch path claims qualityAnalysis.qualityCheckStatus on the batch document instead.
Batch root slot An attempt-1 job with jobMetaData.jobBatchId and no parentJobId whose use case has batch.enabled. Judged together with its siblings. Follow-up children inherit jobBatchId but are never root slots.
Batch document The assetGenJobBatches record for one batch: slot order (slotJobIds), core links, and the batch-level QC state (qualityAnalysis.qualityCheckStatus, summary). Generic and owned by the asset-gen layer.
Slot / rank / best slot is the 1-based position in slotJobIds; rank is the judge's order (1 = best); isBest marks the top-ranked candidate. A recommendation only, never a selection.
Shared flaw An issue the judge saw in every candidate of a batch. Points at the prompt or the references, not at one generation. A major one opens a Sentry issue (QualityAnalysisBatchSharedFlaw).
Pin rule Once any slot of a batch holds a person's decision (APPLIED, user DISMISSED) or a settled state (ACCEPTED, EXHAUSTED), re-evaluation refreshes verdicts but never opens a new proposal in that batch.
Override Applying a slot the batch judge parked (DISMISSED by 'BATCH'). Uses that slot's own followUp.action; the best slot's proposal is left as is.
Intent The text the judge is told the output must show. For shot images it is built from shot fields, not the raw model prompt.
Generation prompt The prompt the image model was given. Shown to the judge separately from the intent because its Image [0] indices refer to the model's inputs, not the judge's image labels.
References The images the image model received (sketch, character sheet, environment, prop), each shown to the judge with a named role.
i2i Image-to-image: a model config whose input is an existing image. EDIT_PREVIOUS children are i2i jobs on the previous output.
JaduSpine The realtime publish/subscribe layer used for the qualityAnalysisUpdated event.
Human score A 1–5 score a person gives via POST /humanScore, stored next to the judge's score to measure how well the judge agrees with people.