Skip to content

Shot-image QA offline evaluation

Read-only evaluation of the SHOT_IMAGE_V1 quality-analysis judge on real sketch-to-shot batches from the test DB. Nothing here writes to Mongo; the only side effect is the OpenAI call log entry each judge call creates.

file what
regression-set.json Committed. The batches (and one job per batch for the single judge) every prompt revision is judged on, with a note on why each was picked. Add cases here when a new failure class turns up.
snapshots/<date>-<label>.json Committed. Compact scores, issues, actions and reasoning per slot for one prompt revision, with the git commit it was judged at.
run-selection.sh [outDir] Judges the regression set into outDir (one dir per prompt revision, about 1 min per call, 2 calls per batch). Skips what already exists.
snapshot.ts write / diff Writes a snapshot from one or more out dirs (later dirs override earlier per batch); diffs two snapshots in the terminal.
buildReport.ts Renders index.html from out dirs: new-prompt tab, prod tab, diff tab with images. Needs the out dirs, so local only.
out*/ Gitignored. Full judge payloads (context, prompt, raw output) and readable .txt reports.

Evaluating a prompt change

zsh scripts/qa-eval/run-selection.sh scripts/qa-eval/out-v4            # judge everything with the working-tree prompts
npx ts-node scripts/qa-eval/snapshot.ts write --label v4-<what-changed> --dirs scripts/qa-eval/out-v4
npx ts-node scripts/qa-eval/snapshot.ts diff scripts/qa-eval/snapshots/<previous>.json scripts/qa-eval/snapshots/<new>.json
npx ts-node scripts/qa-eval/buildReport.ts --dir scripts/qa-eval/out --compare scripts/qa-eval/out-v4 --afterLabel "v4"

To re-judge only some batches after a follow-up tweak, run the two eval scripts by hand into another dir and layer it: snapshot.ts write --dirs out-v4,out-v5 and buildReport.ts --compare out-v4 --overlay out-v5.

Snapshots so far

snapshot prompts notes
2026-09-16-prod as deployed (commit cc43c812) 10 batches. Sketch mistaken for the output twice; major costume issue next to a passing identity score; no 5s.
2026-09-17-new-prompts working tree of 2026-09-17 15 batches. Sketch-vs-output rule, severity ↔ score rule, pose in composition, costume rungs in identity, narrowed prompt_adherence, new critical art_style, five-rung artifacts with limb count. 12 batches judged with the first revision, 3 re-judged after art_style / limb-count were added.

Open items recorded in regression-set.json notes: pose mismatch still scores composition 3 (13fa9c81), a different face with a matching hoodie still reads as the character (d54717e2).