AssetGen Quality Analysis¶
After an asset-gen job finishes, this module has an LLM look at the output, score it 1–5 against what the job was meant to produce, and decide what to do next: accept it, regenerate with a better prompt, edit the image, or hand it to a human. When a fix is chosen it can run automatically as a new job, and that job is scored again, up to a small number of attempts.
The generic machinery (the harness) handles scheduling, the scale, the judge call, the attempt chain, credits, events and de-duplication. Each kind of asset plugs in as a use case that says what to judge, what the judge should see, and which fixes are allowed. Today there is one use case, workbench shot images.
A failed or skipped analysis never flips the generated asset to failed. Unfamiliar terms are defined in the Glossary at the end.
Enablement¶
| Flag | Effect |
|---|---|
QUALITY_ANALYSIS_ENABLED=true |
Master switch for the harness. Off → no evaluate, no stamp. |
Defaults off. Opt a use case in by stamping jobMetaData.qualityAnalysisLineage.useCaseId at launch and leaving enabled: true on the use case object. Flip enabled: false in code to take a use case out without an env var.
The harness only runs stamped jobs. Everything else that tunes behaviour (pass threshold, allowed actions, auto-apply, judge model) is set per use case; see Configuration. The old Python /scoreAsset / /rewritePrompt path is unchanged and deprecated.
Flow¶
job settles (runJob / pollJob / webhook / updateJobAsFailed)
│
▼
QualityAnalysisHarness.scheduleOnJobSettled (fire-and-forget)
│ resolve use case from jobMetaData.qualityAnalysisLineage.useCaseId
│
├── batch root slot (jobBatchId, attempt 1, useCase.batch.enabled) ──► batch path, see below
│
│ (single job, COMPLETED only)
│ claim: qualityCheckStatus → PROCESSING (one atomic write; a second caller gets null and stops)
▼
useCase.evaluate(job) OR LLM judge (gpt-5.4 / medium)
│ overall capped by lowest CRITICAL criterion
│ score wins over a contradicting ACCEPT recommendation
▼
verdict.pass ──► followUp.status = ACCEPTED
│
├── attempt >= maxAttempts ──► EXHAUSTED
│
└── PROPOSED { action, requiresReview }
│
├── shouldAutoApply ──► applyFollowUp({ triggeredBy: 'AUTO' })
└── else wait for POST /applyFollowUp
│
▼
child job (attempt + 1) → re-enters harness on completion
Manual POST /evaluate re-runs a COMPLETED check, appends to verdictHistory, and does not auto-apply. Terminal follow-up statuses (APPLIED, DISMISSED, EXHAUSTED, ACCEPTED) are left in place; only PROPOSED may be replaced.
Batch path (comparative judge)¶
Shot images are generated N at a time (generateShotImageInitialBatch, 2–4 slots sharing jobMetaData.jobBatchId). When the use case sets batch.enabled, the harness judges the slots together instead of one at a time. The trigger is generic: nothing in the workbench calls the harness for batches, every job sink already does.
Assumption: all slots in a batch were generated from the same prompt and inputs. The harness calls buildContext once, on the first completed candidate, and sends that single brief (intent, generation prompt, references) with all N images. This holds for createJobInitialBatch, which fans one config out N times. A launch path whose slots differ in prompt or references must not be routed through the comparative judge, otherwise N−1 slots are scored against the wrong brief.
any root slot settles (COMPLETED | FAILED | CANCELLED)
│ list siblings by jobBatchId, keep root slots (attempt 1, no parentJobId)
│ any sibling non-terminal ──► return; the last one to settle triggers
│ candidates = COMPLETED siblings with a result
│ 0 → nothing; 1 → single-slot path (nothing to compare)
│ claim the batch: assetGenJobBatches.qualityAnalysis.qualityCheckStatus → PROCESSING (atomic, dedupes concurrent finishers)
│ each candidate: qualityCheckStatus → PROCESSING, STARTED event with { batch: { batchId, slot } }
▼
one LLM call: candidates as Image 1..N (slot order) + shared references
│ per candidate: criteria, overall (critical cap), issues, own recommended action
│ ranking (best first), sharedIssues (flaws in EVERY candidate), requiresReview
│ re-read candidates (an apply/dismiss during the call must not be overwritten)
▼
per slot: normal Verdict + qualityAnalysis.batch = { batchId, slot, rank, isBest }
batch document: qualityAnalysis = { qualityCheckStatus: COMPLETED, summary, summaryHistory, judgedJobIds }
│
├── pin rule: any slot APPLIED / ACCEPTED / EXHAUSTED / user-DISMISSED ──► REEVALUATED on every slot, no follow-up change
├── ≥1 slot passes ──► passing slots ACCEPTED; failing slots DISMISSED by 'BATCH' (own action kept)
├── none pass, attempt < maxAttempts ──► best slot PROPOSED; others DISMISSED by 'BATCH'
└── none pass, attempt ≥ maxAttempts ──► best slot EXHAUSTED; others DISMISSED by 'BATCH'
│
▼
BATCH_JUDGED event (once), then per-slot phase events with { batch: pointer }
major sharedIssues ──► logger.warn QualityAnalysisBatchSharedFlaw (Sentry issue: the prompt is at fault, not one generation)
- At most one harness proposal per batch. Auto-apply on the best slot follows the normal rules; shot images leave
allowAutoApplyInBatchoff, so it stays a proposal. - Override. A batch-dismissed slot keeps its own
followUp.action.POST /applyFollowUpon it movesDISMISSED(BATCH) → APPLIEDand creates a child from that slot; the child is judged via the single-slot path. A slot a person dismissed cannot be applied. The best slot's pendingPROPOSEDis not touched by an override. - Pin rule. Once anyone has decided on a slot, re-evaluation only refreshes verdicts (and
verdictHistoryon every slot) and never opens a new proposal anywhere in the batch. - Slot numbers are 1-based positions in the batch document's
slotJobIds, so a failed slot keeps its number and the judge'sImage kfork ≤ Nis candidatekin that order. Batches created before the document existed fall back tocreatedAtthenjobId. - Failure. Anything that throws after the batch claim (marking slots PROCESSING, the judge, persisting a slot) marks the batch document
qualityCheckStatus: FAILEDwitherrorand releases the claim, so a manual evaluate can retry. Slots not yet persisted getqualityCheckStatus: FAILEDand aFAILEDevent; slots already persisted with a verdict keep their COMPLETED state and follow-up (an ACCEPTED slot flipped to FAILED would pin the batch and make the retry a no-op). Use-case hooks (onAccepted,onExhausted,onBatchJudged) are logged and swallowed, never fail the batch. A claim leftPROCESSINGby a process crash still blocks manual re-evaluation with "already running" (same limitation as the single path; no lease yet). - Edit batches (
dispatchShotImageEditBatch) and reject-and-regenerate batches (rejectAndRegenerateShotImageBatch) are not stamped and never enter the harness; each carries aTODO(quality-analysis follow-up)describing what stamping them needs. Restart-from-sketch batches reusedispatchShotImageEditBatchbut pass their stamp viaslotJobMetaData, so they are judged like initial batches.
Files¶
| File | Role |
|---|---|
qualityAnalysis.harness.ts |
Settle hook, claim the check (per job or per batch), evaluate, persist, decide, publish. |
qualityAnalysis.batch.ts |
Pure batch helpers: isBatchRootSlot, sortBatchSlots, slotVerdictsFromBatchOutput, hasUserDecidedFollowUp (pin rule), resolveBatchOutcome, majorSharedIssues. |
qualityAnalysis.followUp.ts |
resolveAction, shouldAutoApply, applyFollowUp (also on batch-dismissed slots), dismissFollowUp, default buildNextJob. |
qualityAnalysis.judge.ts |
LLM judge (OpenAIService.generateTextWithCompletions, JSON schema, one retry) and the comparative runLlmBatchJudge. Prompts live in src/promptTemplates/qualityAnalysis/qualityAnalysisJudge.template.ts and qualityAnalysisBatchJudge.template.ts. |
src/models/assetGenJobBatch.model.ts |
Generic assetGenJobBatches collection (asset-gen layer). The harness only writes its qualityAnalysis subdocument. |
qualityAnalysis.scale.ts |
1–5 anchors, critical-criterion cap, score fallback action. |
qualityAnalysis.lineage.ts |
Attempt / root / parent stamps. loadRootJob for original generation inputs. |
qualityAnalysis.registry.ts |
Use cases register at module load. |
qualityAnalysis.events.ts |
JaduSpine events (project channel when projectLinks.projectId exists). |
qualityAnalysis.controller.ts / .validator.ts |
HTTP endpoints. |
useCases/shotImage.useCase.ts |
First LLM use case. |
scripts/qualityAnalysisBatchEval.ts |
Offline runner for the comparative judge on an existing batch (see Manual runs). |
Data¶
Two places on the job:
jobMetaData.qualityAnalysisLineage (type QualityAnalysisLineage) — set at launch, copied onto children.
{
useCaseId: string;
rootJobId: string; // attempt-1 job; filled in createJob / createJobInitial
parentJobId?: string;
attempt: number; // 1-based; only applyFollowUp increments
promptSource: 'GENERATED' | 'USER_EDITED';
context?: Record<string, unknown>; // use-case data the judge needs later (shot image: generationRefs)
}
job.qualityAnalysis (type QualityAnalysisData in sharedTypes.ts) — harness result. LLM use cases write score as 1–5. The harness never writes rewrittenPrompt: a regenerate prompt lives on followUp.action.prompt, next to instruction for edits and reason for escalations. The deprecated Python path writes 0–10 score, rewrittenPrompt, transformedScore and actionableFeedback.
{
qualityCheckStatus?: 'processing' | 'completed' | 'failed';
score?: number;
reasoning?: string;
verdict?: Verdict;
verdictHistory?: Verdict[];
followUp?: {
status: 'ACCEPTED' | 'PROPOSED' | 'APPLIED' | 'DISMISSED' | 'EXHAUSTED';
action?: NextAction;
applied?: NextAction;
appliedJobId?: string;
triggeredBy?: 'AUTO' | 'BATCH' | string; // 'BATCH' = parked by the batch judge, still applicable by a person
requiresReview: boolean;
reason?: string;
};
batch?: QualityAnalysisBatchPointer; // { batchId, slot, rank, isBest } — written by the batch judge only
}
NextAction is one of: ACCEPT, REGENERATE { prompt }, EDIT_PREVIOUS { instruction }, ESCALATE { reason }.
Batch document (assetGenJobBatches, type AssetGenJobBatchType in sharedTypes.ts) — one per batch, created by createJobInitialBatch and owned by the asset-gen layer. The harness writes only qualityAnalysis:
{
batchId, userId, slotJobIds: string[], batchSize, createdVia?, storyLinks?, projectLinks?, createdAt, updatedAt,
qualityAnalysis?: BatchQualityAnalysisData; // = {
useCaseId: string;
qualityCheckStatus: 'processing' | 'completed' | 'failed'; // the batch-level claim, same field name as on the job
judgedJobIds: string[]; // completed slots sent to the judge
summary?: BatchVerdictSummary; // { ranking: [{ jobId, slot, rank, overall, pass }], bestJobId, anyPassed,
summaryHistory?: BatchVerdictSummary[]; // sharedIssues, requiresReview, reasoning, evaluator, evaluatedAt }
error?: string;
};
}
The best slot is a recommendation only (isBest on the slot, summary.bestJobId on the document). shot.images[].isSelected and shot.batches[].selectedAssetGenJobId stay user-driven.
Scale and action resolution¶
Pass threshold is useCase.passThreshold, default 4 (DEFAULT_PASS_THRESHOLD). pass is overall >= passThreshold. Critical criteria cap overall: overall cannot exceed the lowest critical score. A criterion the judge marked unjudgeable (occluded, too small, no reference to compare) never caps.
Each criterion score carries observed (what the judge saw for that criterion) next to reason (why that score).
Score fallback when the judge's recommendation is missing, not allowed, or is ACCEPT on a failing score. The bands are derived from the pass threshold, so raising or lowering it moves all of them; there is no separate edit or regenerate threshold.
| Overall | Action |
|---|---|
| ≥ threshold | ACCEPT |
| threshold − 1 | EDIT_PREVIOUS if every major issue is localized and edit is allowed; else REGENERATE |
| below that | REGENERATE |
| Final attempt below threshold | ESCALATE if allowed |
With the default threshold that is 4–5 / 3 / 1–2.
REGENERATE with an empty prompt copies the root job prompt. If that is also empty, the harness escalates instead of inventing a prompt.
Root resolution¶
EDIT_PREVIOUS children have an i2i config whose image input is the previous output. Grading and a later REGENERATE must not treat that as the original generation.
buildContext(shot image) and empty-prompt fill load root prompt + reference images vialoadRootJob.- Default
REGENERATEclones the rootmodelConfigand replaces the prompt. EDIT_PREVIOUSstill edits the immediate parent output.
Auto-apply¶
All of these must hold:
reviewPromptBeforeRunis false- verdict does not
requiresReview promptSourceis notUSER_EDITEDattempt <= autoApplyAttempts- not an API job (
createdVia: API— created via/apiViaKey/assetGen, not Studio). Auto-apply would spawn extra charged generations the API client did not request and has no Studio UI to review. - not a batch job, unless
allowAutoApplyInBatch - if the parent was auto-applied, this attempt's overall must be higher than the parent's
Shot image: { autoApplyAttempts: 1, reviewPromptBeforeRun: false, maxAttempts: 3 } — attempt 1 may auto-regenerate/edit; attempt 2 waits.
Follow-up generations are charged as normal asset-gen jobs. Judge scoring is logged, not charged. The verdict stores no log ids; the judge's analytics log entry (QUALITY_ANALYSIS_JUDGE) carries jobId (single path) or batchId (comparative path) plus the job's user, project and story links in its context, so it is found from the log side.
Endpoints¶
Mounted under /assetGen/qualityAnalysis. Owner or ACCESS_ALL_JOBS.
| Route | Body | Behaviour |
|---|---|---|
POST /evaluate |
{ jobId } |
Manual (re-)evaluation. On a batch root slot the whole batch is re-judged: the response is the requested slot, siblings arrive via events, every slot's verdictHistory grows. Fails while siblings are still running. |
POST /applyFollowUp |
{ jobId, override? } |
Moves PROPOSED → APPLIED in one atomic write (a concurrent apply gets the same child back), creates the child. Also accepts a slot DISMISSED by 'BATCH' (moves it to APPLIED from its own action); a user-dismissed slot is rejected. Override may change type, prompt / instruction, and modelConfigId (a server-side config id from the use case allowlist — never a full modelConfig). Changing prompt/instruction text or picking a model stamps USER_EDITED. |
POST /dismissFollowUp |
{ jobId } |
DISMISSED. No child, no attempt increment. PROPOSED only. |
GET /lineage/:rootJobId |
— | Jobs sharing rootJobId, sorted by attempt. |
POST /humanScore |
{ jobId, humanScore: 1–5 } |
Calibration signal on qualityAnalysis.humanScore. |
Generic, on the asset-gen router (not QC-specific): GET /assetGen/batch/:batchId returns { batch, jobs } — the batch document and its slot jobs in slot order. Owner of the first slot or ACCESS_ALL_JOBS.
POST /scoreAsset and POST /rewritePrompt are deprecated and untouched by the harness. They remain the Python AssetGen-page path.
Realtime¶
Published to the project channel when projectLinks.projectId exists, otherwise the user's STUDIO channel.
One event: qualityAnalysisUpdated. Discriminate on data.phase:
phase |
When | Extra |
|---|---|---|
STARTED |
Judge running | { job } |
PROPOSED |
Follow-up waiting | { job } |
APPLIED |
Child created | { job, childJob } |
ACCEPTED |
Passed | { job } |
EXHAUSTED |
Max attempts, still failing | { job } |
DISMISSED |
Human dismissed the proposal, or the batch judge parked the slot (job.qualityAnalysis.followUp.triggeredBy === 'BATCH') |
{ job } |
REEVALUATED |
Verdict refreshed, follow-up unchanged | { job } |
FAILED |
Judge threw | { job, error } (status 500) |
BATCH_JUDGED |
Once per comparative batch run, before the per-slot phases | { job, batch } — job is the first judged slot (kept so data.job readers keep working), batch is the batch document with qualityAnalysis.summary |
On the batch path every per-slot event (STARTED through FAILED) also carries data.batch = { batchId, slot, rank?, isBest? } (rank and isBest are absent on STARTED / FAILED).
The harness does not emit assetGenQualityCheckResponse. That event is only the deprecated Python /scoreAsset path.
Shot-image follow-ups run through WorkbenchService.runShotImageGenerationAsync, which commits shot media and publishes shotUpdated (including on failure). Do not use runJobAsync for that use case — a failed child would leave the tile spinning.
Use cases¶
Register at module load (registerQualityAnalysisUseCase). The harness side-effect-imports shotImage.useCase.ts; add a new import there (or next to it) so the use case is loaded.
Configuration¶
Everything that tunes the loop lives on the use case object (QualityAnalysisUseCase in qualityAnalysis.types.ts). There are no env vars per use case and nothing is read from the job except the launch stamp.
| Field | Default | Effect |
|---|---|---|
id |
— | Matches jobMetaData.qualityAnalysisLineage.useCaseId. |
enabled |
true |
false takes the use case out of the harness without an env var. |
criteria |
— | { key, label, critical?, anchors? }. critical caps overall. anchors are the 1 / 3 / 5 descriptions shown to the judge. |
judgeRules |
none | Extra rules appended to the judge system prompt. |
batch.enabled |
false |
Judge the root slots of a jobBatchId batch together (see Batch path). Requires the launch path to go through createJobInitialBatch and assumes every slot shares the same prompt, references and intent. |
batch.judgeRules |
none | Ranking / shared-flaw rules appended after judgeRules in the comparative prompt. |
passThreshold |
4 |
Pass line and the origin of the action bands (see Scale and action resolution). |
allowedActions |
— | Subset of ACCEPT, REGENERATE, EDIT_PREVIOUS, ESCALATE. A recommendation outside it falls back to the score bands. |
followUpPolicy.autoApplyAttempts |
— | Highest attempt number that may auto-apply. 1 means only attempt 1. |
followUpPolicy.reviewPromptBeforeRun |
— | true disables auto-apply entirely. |
followUpPolicy.maxAttempts |
— | Chain length. The final attempt escalates instead of proposing. |
followUpPolicy.allowAutoApplyInBatch |
false |
Let batch slots auto-apply. |
promptInputId |
'prompt' |
Input id read when a REGENERATE needs the root prompt. |
judge.modelId / judge.reasoningEffort |
gpt-5.4 / medium |
Judge model. |
buildContext(job) |
— | Intent, named reference images, notes and generation prompt for the judge. Optional extraCriteria / extraJudgeRules apply to that job only. |
evaluate(job) |
none | Replace the LLM judge with your own verdict. |
buildEditModelConfig / buildRegenerateModelConfig / buildNextJob |
see Adding a use case | How follow-up jobs are built. |
runFollowUpJob and the on* hooks |
none | Own run path and side effects. onBatchJudged(candidates, batch) fires once per comparative run after the slots and the batch document are persisted. |
Per job, the only input is the launch stamp: initAttemptOneMeta(useCaseId, context?) on jobMetaData.qualityAnalysisLineage. context carries use-case data the judge needs later (shot image stores its generationRefs there) and is copied onto children. Everything else about a chain (rootJobId, parentJobId, attempt, promptSource) is managed by the harness.
SHOT_IMAGE_V1¶
Stamped by workbench shot-image launch and by the tweak chat's restart-from-sketch (single and batch, shotImageEditGuru.restartFromSketch) when QUALITY_ANALYSIS_ENABLED is on and shotImageUseCase.enabled is true. Restart also puts the user's typed feedback on jobMetaData.userFeedback, the field pre-gen review already uses.
- Criteria: character identity vs. named character references (critical; face and defining head gear decide recognition, costume only confirms it; a different face or missing signature item is a 2, a missing costume piece a 3), art style vs. the character / environment references and the art-style text in the generation prompt (critical; different medium is a 1), composition vs. sketch (placement, relative scale and pose), camera framing vs. sketch and shot size, artifacts (critical; the judge counts hands, arms, legs and faces and attributes each to a character before scoring), prompt adherence limited to acting, props and scene details that no other criterion covers. Composition / camera anchors use the pass / minor / major definitions from the T2I composition-fidelity evaluation (same left/right and foreground/background arrangement, relative scale, camera distance and angle). Anchors may define any subset of the 1–5 rungs; missing rungs fall back to the generic scale.
- Judge:
gpt-5.4/ medium unlessjudgeis overridden.judgeRulestell the judge the sketch is a layout guide only (never judge style or identity against it), that sketches may be renders of the scene with green / blue / orange mannequins that look like an output (a mannequin means it is the sketch; re-check the image labels before scoring), that a sketch unrelated to the intent makes composition unjudgeable, and that identity is judged against the named character references. Both judge templates also require issue severity to agree with the criterion score (amajorissue means that criterion is 3 or lower). This is prompt-only for now; enforcing it in code after parsing is a TODO. - Judge context (
buildContext): intentis built from shot fields: description, acting, the frame the sketch depicts (sketches[].connectedTo→ initial / middle / final frame instructions; all frames listed when the moment is unrecorded), camera (shot size, angle, details), characters visible, environment, props. A trailingRequested change:line carries the root job'sjobMetaData.userFeedbackwhen set.- Only when that feedback is set, the context also carries
extraCriteria/extraJudgeRuleswith the requested change criterion (critical; no precedence over the sketch, so a change away from the sketch shows as a high requested change next to a lower composition) and its rule. The harness merges them into the use case withwithContextExtrasfor the prompt and the critical cap. Jobs without feedback are judged with the use case's criteria and rules unchanged. generationPromptis the root job prompt, shown under its own heading. ItsImage [0]indices refer to the image model's inputs, not the judge's labels.referencesare the root job's image inputs with a named role. Roles come fromjobMetaData.qualityAnalysisLineage.context.generationRefswhen stamped (see below), else from matching URLs against the shot's sketches (raw and annotated) and the story's character / environment / prop images (including variants). An unresolved first image on a shot with sketches is treated as the sketch, becausegenerateFromSketchalways sends the sketch first. Anything else isreference (role unknown).withShotImageGenerationRefsstamps a compact copy ofgenerateFromSketch'sgenerationRefsintojobMetaData.qualityAnalysisLineage.contexton the batch launch path and inprocessSketchToShotInputs, so the judge names references as they were at generation time.onFollowUpCreatedwrites a PROCESSING placeholder viaaddShotImage.runFollowUpJobcallsrunShotImageGenerationAsync(dynamic import to avoid a cycle withworkbench.service).buildEditModelConfigdefaults toi2i-gpt-image-2. A human apply may passoverride.modelConfigId:shot-image-gen(REGENERATE only),i2i-gpt-image-2, orgoogle-i2i-nanoBananaPro.batch.enabledis on for the initial batch (generateShotImageInitialBatch, count 2–4).batch.judgeRulesrank sketch composition and named character references above rendering polish and put any critical failure last. NoonBatchJudged: the ranking is already on each slot and on the batch document, and selection stays with the user.
Adding a use case¶
- Create
useCases/<name>.useCase.tsimplementingQualityAnalysisUseCase. - Register at module load. Side-effect-import it from
qualityAnalysis.harness.ts. - Stamp
initAttemptOneMeta(YOUR_USE_CASE_ID)on jobs at launch. Gate withQUALITY_ANALYSIS_ENABLEDanduseCase.enabled(default true). No per-use-case env vars. - If the use case has its own run path (commit media, publish an event, mark FAILED on the tile), implement
runFollowUpJob. Otherwise the default isAssetGenJobService.runJobAsync. - If
EDIT_PREVIOUSis allowed, implementbuildEditModelConfig. If a human may pick a model on REGENERATE, implementbuildRegenerateModelConfig. Do not import another use case's model registry fromfollowUp.ts. - Override
buildNextJobonly when the default (clone root config / edit parent output) is wrong.
Judge cost is not charged. Follow-up jobs are.
Tests¶
tests/assetGen/qualityAnalysis/ — scale, follow-up policy, lineage/root resolution, harness claim and re-eval (single and batch), batch helpers (qualityAnalysis.batch.test.ts), judge parse retry (single and comparative), shot-image context. tests/assetGen/assetGenJobBatch.model.test.ts covers the batch document and its atomic claim on in-memory Mongo.
npx vitest run tests/assetGen/qualityAnalysis tests/assetGen/assetGenJobBatch.model.test.ts
Manual runs¶
scripts/qualityAnalysisBatchEval.ts runs the comparative judge on an existing batch from the shell, using MONGODB_URI and OPENAI_API_KEY from .env. Useful for tuning batch.judgeRules and checking a verdict without regenerating.
npx ts-node scripts/qualityAnalysisBatchEval.ts --batch <jobBatchId> # or --job <any slot jobId>
npx ts-node scripts/qualityAnalysisBatchEval.ts --batch <jobBatchId> --json # raw judge output + derived summary
npx ts-node scripts/qualityAnalysisBatchEval.ts --job <jobId> --write # go through the harness, persist + events
Read-only by default: it prints the brief the judge sees (intent, generation prompt, image labels), the ranking, shared issues, per-slot criteria with observed/reason, each slot's own recommended action, and the follow-up resolveBatchOutcome would set. Only the OpenAI log entry is written. Slots that predate the lineage stamp are judged with --useCase (default SHOT_IMAGE_V1); --write refuses them because the harness needs the stamp.
Future work¶
Raised in review; deferred to follow-up PRs. Each has a matching TODO(quality-analysis follow-up) comment in the code.
| Item | Where | Idea |
|---|---|---|
| Few-shot calibration | qualityAnalysis.scale.ts |
LLM judges compress scores toward 3–4. Add optional calibrationExamples (image, human score, reason) per use case, appended to the judge call as labelled anchor images. Source them from POST /humanScore. |
| Stamp edit / regenerate batches | workbench.service.ts (dispatchShotImageEditBatch, rejectAndRegenerateShotImageBatch) |
Only the initial batch is judged comparatively. Edit batches need an intent built from the edit instruction and reference roles for the edited source; regenerate batches need the user's rejection feedback folded into the intent. |
| Auto-select the best slot | storyVideos.model.ts (BatchAsset) |
The batch judge only recommends (isBest, summary.bestJobId). Flip isSelected / selectedAssetGenJobId from the ranking once it has been validated against human picks. |
| Learning from lineage | qualityAnalysis.harness.ts (verdictHistory) |
verdictHistory and the parent → applied prompt → child chain already record which rewrites improved scores. Mine that offline and feed it back into the judge or rewrite prompt. |
| Adversarial prompt distillation | useCases/shotImage.useCase.ts (buildShotIntent) |
Run a heavier model once per prompt version to distill acceptance assertions and an adversarial checklist, and decide how to split QC into specialized judge calls. QC consumes the stored assertions. First step: pass assertions via context.notes. |
Glossary¶
| Term | Meaning |
|---|---|
| Harness | The generic loop in this folder: claim the check, run the judge, store the verdict, decide the follow-up, publish events. |
| Use case | One kind of asset plugged into the harness (e.g. SHOT_IMAGE_V1). Supplies criteria, judge context and allowed actions. See Configuration. |
| Judge | The LLM call that scores the output. Returns a score per criterion, an overall score, issues and a recommended action. |
| Verdict | The stored result of one judge run: overall, pass, per-criterion scores, issues, recommendation, reasoning. |
| Criterion | One thing the judge scores (identity, composition, artifacts, …). Critical criteria cap the overall score. Unjudgeable means the judge could not assess it from the image; it never caps. |
| Anchors | The 1 / 3 / 5 descriptions for a criterion that tell the judge what each score looks like. |
| Pass threshold | Overall score at or above which the output is accepted. Default 4. The edit / regenerate bands are derived from it. |
| Recommendation vs. action | The judge recommends an action; the harness resolves the action actually taken, falling back to the score bands when the recommendation is missing, not allowed, or ACCEPT on a failing score. |
| Follow-up | The action taken after a failing verdict: REGENERATE { prompt }, EDIT_PREVIOUS { instruction } or ESCALATE { reason }. followUp.status tracks whether it is proposed, applied, dismissed, accepted or exhausted. |
| Localized | An issue confined to a region an image edit can fix (a hand, a prop). All major issues localized → EDIT_PREVIOUS; otherwise → REGENERATE. |
| requiresReview | The judge asked for a human to read the proposed prompt before it runs. Blocks auto-apply. |
| Auto-apply | The harness runs the follow-up itself instead of waiting for a human. Gated by the rules under Auto-apply. |
| Stamp | Writing jobMetaData.qualityAnalysisLineage onto a job when it is created. Presence of useCaseId is what opts the job into the harness. |
| Lineage / chain | The sequence of jobs produced by follow-ups. Root is the attempt-1 job with the original inputs, parent is the job a follow-up was applied to, attempt is 1-based. |
| Claim | The single atomic write that moves qualityCheckStatus to processing. A second concurrent caller gets null and stops, so a job is never judged twice at once. The batch path claims qualityAnalysis.qualityCheckStatus on the batch document instead. |
| Batch root slot | An attempt-1 job with jobMetaData.jobBatchId and no parentJobId whose use case has batch.enabled. Judged together with its siblings. Follow-up children inherit jobBatchId but are never root slots. |
| Batch document | The assetGenJobBatches record for one batch: slot order (slotJobIds), core links, and the batch-level QC state (qualityAnalysis.qualityCheckStatus, summary). Generic and owned by the asset-gen layer. |
| Slot / rank / best | slot is the 1-based position in slotJobIds; rank is the judge's order (1 = best); isBest marks the top-ranked candidate. A recommendation only, never a selection. |
| Shared flaw | An issue the judge saw in every candidate of a batch. Points at the prompt or the references, not at one generation. A major one opens a Sentry issue (QualityAnalysisBatchSharedFlaw). |
| Pin rule | Once any slot of a batch holds a person's decision (APPLIED, user DISMISSED) or a settled state (ACCEPTED, EXHAUSTED), re-evaluation refreshes verdicts but never opens a new proposal in that batch. |
| Override | Applying a slot the batch judge parked (DISMISSED by 'BATCH'). Uses that slot's own followUp.action; the best slot's proposal is left as is. |
| Intent | The text the judge is told the output must show. For shot images it is built from shot fields, not the raw model prompt. |
| Generation prompt | The prompt the image model was given. Shown to the judge separately from the intent because its Image [0] indices refer to the model's inputs, not the judge's image labels. |
| References | The images the image model received (sketch, character sheet, environment, prop), each shown to the judge with a named role. |
| i2i | Image-to-image: a model config whose input is an existing image. EDIT_PREVIOUS children are i2i jobs on the previous output. |
| JaduSpine | The realtime publish/subscribe layer used for the qualityAnalysisUpdated event. |
| Human score | A 1–5 score a person gives via POST /humanScore, stored next to the judge's score to measure how well the judge agrees with people. |