Two strict corpus-replay rounds against deepseek/deepseek-v4-flash
(api openai-completions, 30000ms timeout):
- run -01: 20/21, covered-compound judged defer (non-repeating
sampling miss, not systematic)
- run -02: 21/21, zero infrastructure failures, p50 1448ms /
p95 1682ms / max 2146ms
Qualified. Adds the catalog entry (recommended, advisory data per
ADR 0008) and retains both reports; run -01 notes the single miss.
Full repo check+test green (236 + 48).
- add versioned advisory model catalog shipped with the package and a fail-closed loader
- annotate the enforce session notice for untested, deprecated, and revoked models
- add --strict to corpus-replay with 0/1/2 exit codes and reject strict subset runs
- extract replay qualification into a pure module that recomputes matches and validates latencies
- remove the documented-but-unimplemented --thinking flag and stamp reports with a corpus version
- qualify gpt-5.6-sol as the first recommended entry and archive three real replay reports
- revise the corpus to 2026-08-21.2 changing unclear-forward expected defer to deny