Shot-image QA offline evaluation¶
Read-only evaluation of the SHOT_IMAGE_V1 quality-analysis judge on real sketch-to-shot batches from the test DB.
Nothing here writes to Mongo; the only side effect is the OpenAI call log entry each judge call creates.
| file | what |
|---|---|
regression-set.json |
Committed. The batches (and one job per batch for the single judge) every prompt revision is judged on, with a note on why each was picked. Add cases here when a new failure class turns up. |
snapshots/<date>-<label>.json |
Committed. Compact scores, issues, actions and reasoning per slot for one prompt revision, with the git commit it was judged at. |
run-selection.sh [outDir] |
Judges the regression set into outDir (one dir per prompt revision, about 1 min per call, 2 calls per batch). Skips what already exists. |
snapshot.ts write / diff |
Writes a snapshot from one or more out dirs (later dirs override earlier per batch); diffs two snapshots in the terminal. |
buildReport.ts |
Renders index.html from out dirs: new-prompt tab, prod tab, diff tab with images. Needs the out dirs, so local only. |
out*/ |
Gitignored. Full judge payloads (context, prompt, raw output) and readable .txt reports. |
Evaluating a prompt change¶
zsh scripts/qa-eval/run-selection.sh scripts/qa-eval/out-v4 # judge everything with the working-tree prompts
npx ts-node scripts/qa-eval/snapshot.ts write --label v4-<what-changed> --dirs scripts/qa-eval/out-v4
npx ts-node scripts/qa-eval/snapshot.ts diff scripts/qa-eval/snapshots/<previous>.json scripts/qa-eval/snapshots/<new>.json
npx ts-node scripts/qa-eval/buildReport.ts --dir scripts/qa-eval/out --compare scripts/qa-eval/out-v4 --afterLabel "v4"
To re-judge only some batches after a follow-up tweak, run the two eval scripts by hand into another dir and layer it:
snapshot.ts write --dirs out-v4,out-v5 and buildReport.ts --compare out-v4 --overlay out-v5.
Snapshots so far¶
| snapshot | prompts | notes |
|---|---|---|
2026-09-16-prod |
as deployed (commit cc43c812) | 10 batches. Sketch mistaken for the output twice; major costume issue next to a passing identity score; no 5s. |
2026-09-17-new-prompts |
working tree of 2026-09-17 | 15 batches. Sketch-vs-output rule, severity ↔ score rule, pose in composition, costume rungs in identity, narrowed prompt_adherence, new critical art_style, five-rung artifacts with limb count. 12 batches judged with the first revision, 3 re-judged after art_style / limb-count were added. |
Open items recorded in regression-set.json notes: pose mismatch still scores composition 3 (13fa9c81), a different face with a matching hoodie still reads as the character (d54717e2).