Commit Graph
3 Commits
Author SHA1 Message Date
SikongJueluo 99e953664d refactor(ai-judge): group src modules into domain directories
- move evidence.ts and conversation.ts into evidence/ as bash.ts and conversation.ts
- move prompt.ts and model.ts into judge/
- move highrisk.ts and judge.ts into authority/, renaming judge.ts to enforce.ts to avoid clashing with the judge/ directory
- move review.ts and audit.ts into telemetry/
- move config.ts and catalog.ts into config/ as judge.ts and catalog.ts
- rewrite static and dynamic imports across src, tools, and tests to the new paths
- fix the models-catalog.json relative URL broken by the move (caught by fallow unresolved-import)
- update the PIEXTENSIO-12 module map in the ADR to the new paths
- add editorconfig for 4-space ts indentation
2026-08-22 01:34:29 +08:00
SikongJueluo c479f51469 refactor(ai-judge): split remaining hotspots and cover CRAP branches
- stage-ify judgeAuthorize into auditEnrollment, runPreflightGates, prepareModelCall, and enforceAndEmit
- extract tallyJoinedRows and tallyAttributable from computeMetrics, removing three dead locals
- extract validateVerdictResponse from requestStructuredVerdict
- extract collectUserTexts and charBudgetStart from buildConversationEvidence
- table-drive corpus-replay parseArgs and split main into resolveReplayModel, selectCorpusCases, and replayCorpus
- split analyzer cli main into loadReviewEvents, withinWindow, loadAuditEnrolled, and printReport with a run-as-script guard
- add 81 tests covering parseEntry, extractBashCommandEvidence, validateVerdictResponse, forcedToolChoice, classifyGit dry-run paths, CLI arg/window/report rendering, and the infra-failure result path
- add @vitest/coverage-istanbul for exact per-function CRAP scoring via fallow health --coverage
2026-08-22 01:24:50 +08:00
SikongJueluo 7987fc7a3c feat(ai-judge): add advisory model catalog and strict corpus replay
- add versioned advisory model catalog shipped with the package and a fail-closed loader
- annotate the enforce session notice for untested, deprecated, and revoked models
- add --strict to corpus-replay with 0/1/2 exit codes and reject strict subset runs
- extract replay qualification into a pure module that recomputes matches and validates latencies
- remove the documented-but-unimplemented --thinking flag and stamp reports with a corpus version
- qualify gpt-5.6-sol as the first recommended entry and archive three real replay reports
- revise the corpus to 2026-08-21.2 changing unclear-forward expected defer to deny
2026-08-21 23:11:40 +08:00