- stage-ify judgeAuthorize into auditEnrollment, runPreflightGates, prepareModelCall, and enforceAndEmit
- extract tallyJoinedRows and tallyAttributable from computeMetrics, removing three dead locals
- extract validateVerdictResponse from requestStructuredVerdict
- extract collectUserTexts and charBudgetStart from buildConversationEvidence
- table-drive corpus-replay parseArgs and split main into resolveReplayModel, selectCorpusCases, and replayCorpus
- split analyzer cli main into loadReviewEvents, withinWindow, loadAuditEnrolled, and printReport with a run-as-script guard
- add 81 tests covering parseEntry, extractBashCommandEvidence, validateVerdictResponse, forcedToolChoice, classifyGit dry-run paths, CLI arg/window/report rendering, and the infra-failure result path
- add @vitest/coverage-istanbul for exact per-function CRAP scoring via fallow health --coverage
- key inner-cmd decisive review events by requestId so link decisions join offline
- record judge runtime id, prompt and tool schema versions, end-to-end and model latency, input and output usage, and evidence-quality flags on every judge result row
- record forwarded and session-mismatch preflight defers so they stay visible in the offline denominator
- require @gotgenes/pi-permission-system >=25.3.0 and read the complete local bash command from PromptPermissionDetails.payload instead of session-walking recovery
- remove the @sikongjueluo/pi-permission-shared package
- pass the triggering command unit to handlers via HandlerContext.unit in place of details.command
- add shadow-only AI judge modules for evidence projection, structured verdict requests, and prompt building, with vitest coverage
- record ADR 0004 and mark the ADR 0001 recovery mechanism superseded
- exclude pi-permission-system 25.3.0 from the pnpm minimumReleaseAge guard
- accept GNU timeout durations without a unit suffix and with decimals (timeout 240 …)
- detect the wrapper on details.command and strip it from the full command so scaffolded inputs (cd … && timeout … | tail) unwrap
- re-evaluate the full de-wrapped compound so sibling commands cannot hide behind the wrapper allow
- defer fail-closed when the unit is not a unique substring of the full command
- amend ADR 0001 with the relaxed grammar and the scaffolded-command handling
- add handlers/xargs.ts mirroring env: claim xargs-leading commands and defer
- register xargsHandler so leading-xargs commands log and defer instead of falling through silently
- add CONTEXT.md xargs example and ADR 0003 (xargs args come from stdin, so even the AI judge cannot know them)
- replace the hardcoded timeout switch with an engine that iterates registered handlers
- extract the timeout logic into handlers/timeout.ts and add handlers/env.ts that defers env as non-transparent
- thread a partial-evidence bag so the engine exception log retains handler-derived values like innerCommand
- add CONTEXT.md with the transparent vs non-transparent wrapper glossary
- record ADR 0002: env always defers to the AI judge and is never unwrapped
- add @sikongjueluo/pi-permission-shared with recoverNativeBashCommand and its tests
- move the recovery module out of pi-permission-inner-cmd and import it from the shared package
- wire pi-permission-ai-judge to capture the UI-root session and recover the full bash command
- gate pi-permission-ai-judge registration on a UI-present root session
- add @types/node to pi-permission-ai-judge and allow its test script to pass with no tests
- walk entries in reverse and stop at the latest assistant message containing the id
- require the id to match exactly one block within that message rather than across the whole session
- resolve a cross-message id reuse to the latest call being authorized
- update ADR 0001 wording for the narrowed scope
- add a regression test for cross-message id reuse
- recover the full bash command from the session by tool-call id
- add recognizer for the strict timeout wrapper grammar
- add authorizer mapping inner allow/ask/deny and forwarding agent name
- defer fail-closed on session mismatch, nested wrappers, and errors
- add unit tests for recovery, recognizer, authorizer, and lifecycle
- document the decision in ADR 0001