Two strict corpus-replay rounds against deepseek/deepseek-v4-flash
(api openai-completions, 30000ms timeout):
- run -01: 20/21, covered-compound judged defer (non-repeating
sampling miss, not systematic)
- run -02: 21/21, zero infrastructure failures, p50 1448ms /
p95 1682ms / max 2146ms
Qualified. Adds the catalog entry (recommended, advisory data per
ADR 0008) and retains both reports; run -01 notes the single miss.
Full repo check+test green (236 + 48).
- add versioned advisory model catalog shipped with the package and a fail-closed loader
- annotate the enforce session notice for untested, deprecated, and revoked models
- add --strict to corpus-replay with 0/1/2 exit codes and reject strict subset runs
- extract replay qualification into a pure module that recomputes matches and validates latencies
- remove the documented-but-unimplemented --thinking flag and stamp reports with a corpus version
- qualify gpt-5.6-sol as the first recommended entry and archive three real replay reports
- revise the corpus to 2026-08-21.2 changing unclear-forward expected defer to deny
- add config v2 with fixed judge model selection and fail-closed v1 enforce migration to shadow
- add built-in high-risk override for irreversible, publish, system, and credential shapes that always defers to the human
- remove the promotion gate, promotion records tool, and their tests
- fail closed with judge_model_unavailable when a configured judge model cannot be resolved
- notify once per session in enforce mode with the judge model and risk contract
- add ADR 0008 and update README and CONTEXT
- add conversation.ts: compaction-aware active-branch capture, user-text-only whitelist, 16-item and 12,000-char bounds with latest-user preservation
- bump prompt to bash-shadow-v2 with explicit-user-intent authority rules and quoted untrusted intent evidence
- capture the requesting cwd and per-ask conversation state; flip evidence-quality flags from placeholders to measured values
- record the candidate-identity change for prior cohorts in the scenario-set doc
- add review.ts sink adapter with session-start review-log toggle detection and a privacy key denylist enforced before delegation
- add judge.ts enforce truth table: allow requires mode, host contract, telemetry health, cohort qualification, owner approval, activation, judgment result, allow verdict, review acknowledgement, and current generation — each independently forces defer with a distinct reason
- route the authorizer callback through the sink and the v0.1 production gate state, which is structurally unreachable and therefore fail-closed
- key inner-cmd decisive review events by requestId so link decisions join offline
- record judge runtime id, prompt and tool schema versions, end-to-end and model latency, input and output usage, and evidence-quality flags on every judge result row
- record forwarded and session-mismatch preflight defers so they stay visible in the offline denominator
- require @gotgenes/pi-permission-system >=25.3.0 and read the complete local bash command from PromptPermissionDetails.payload instead of session-walking recovery
- remove the @sikongjueluo/pi-permission-shared package
- pass the triggering command unit to handlers via HandlerContext.unit in place of details.command
- add shadow-only AI judge modules for evidence projection, structured verdict requests, and prompt building, with vitest coverage
- record ADR 0004 and mark the ADR 0001 recovery mechanism superseded
- exclude pi-permission-system 25.3.0 from the pnpm minimumReleaseAge guard
- add @sikongjueluo/pi-permission-shared with recoverNativeBashCommand and its tests
- move the recovery module out of pi-permission-inner-cmd and import it from the shared package
- wire pi-permission-ai-judge to capture the UI-root session and recover the full bash command
- gate pi-permission-ai-judge registration on a UI-present root session
- add @types/node to pi-permission-ai-judge and allow its test script to pass with no tests
- add pnpm workspace scaffold with shared TypeScript config
- register an ai-bash-judge authorizer with pi-permission-system
- log permission request details and deterministic policy verdicts
- record review entries and defer to the next authorizer