diff --git a/docs/research/shadow-scenario-set.md b/docs/research/shadow-scenario-set.md index cec8f55..9a6ced3 100644 --- a/docs/research/shadow-scenario-set.md +++ b/docs/research/shadow-scenario-set.md @@ -168,12 +168,17 @@ statistical framing: single verdicts are not oracles, only cohort rates are meaningful. **PIEXTENSIO-11 latency evidence (16 judgments across the three fixed -windows):** p50 4.2s, mean 4.9s, max 11.9s. The 60s timeout budget has -~5x headroom over the observed max; no timeout or infrastructure -failure occurred in any round. Verdict: the 60s default is safe; -tightening toward ~30s would still bound worst-case waits with ~2.5x -headroom, but calibration-grade tightening should wait for more -samples across providers. +windows):** p50 4.2s, mean 4.9s, max 11.9s, zero timeouts or +infrastructure failures. The canonical resolution (c0b0028d) fixes +15,000 ms as the default and 5,000–30,000 ms as the accepted config +range, calibrated on openai-codex/gpt-5.6-sol (p95 6.4s). The zai +glm-5.2 segment these rounds observed is **uncalibrated** under that +contract: it raced the 15s deadline twice at `thinking: high` before +the cohort and stayed within max 11.9s under the interim 60s bound. +Protocol for this segment: configure timeoutMs within the accepted +range (30s) via the config module, run as a distinct configuration +cohort (non-default timeout never inherits the default cohort's +calibration), and fail closed to defer otherwise. **Cohort totals (3 rounds, fixed windows):** N=19, joined 19/19, judgments 16, matrix allow|allow 8, allow|deny 2, defer|allow 5, diff --git a/packages/pi-permission-ai-judge/src/model.ts b/packages/pi-permission-ai-judge/src/model.ts index edf215a..d5913d0 100644 --- a/packages/pi-permission-ai-judge/src/model.ts +++ b/packages/pi-permission-ai-judge/src/model.ts @@ -7,13 +7,16 @@ import type { ModelRegistry } from "@earendil-works/pi-coding-agent"; import { buildJudgeContext, MAX_REASON_CODE_POINTS, REPORT_VERDICT_TOOL_NAME } from "./prompt"; import type { BashJudgmentEvidence } from "./evidence"; -// 60s: glm-5.2 at the user's default `thinking: high` profile was observed -// both finishing in seconds and racing the former 15s deadline (one verdict -// landed at exactly 15.003s and was mislabeled `aborted`); high-variance -// reasoning latency needs the wider bound. The judge is the first chain -// link, so its wait delays the human prompt by at most this much. -// PIEXTENSIO-11 calibrates a final value from cohort data. -const DEFAULT_TIMEOUT_MS = 60_000; +// 15s is the PIEXTENSIO-11 calibrated default (canonical resolution c0b0028d): +// 15,000 ms total wall-clock deadline, accepted config range 5,000–30,000 ms, +// invalid values fail closed to this default. The judge is the first chain +// link, so its wait delays the human prompt by at most this much. Uncalibrated +// provider segments (e.g. zai glm-5.2, observed racing the deadline at +// `thinking: high`) should configure a non-default timeoutMs within the range +// and be treated as a distinct configuration cohort. +const DEFAULT_TIMEOUT_MS = 15_000; +export const MIN_TIMEOUT_MS = 5_000; +export const MAX_TIMEOUT_MS = 30_000; // Reasoning-token aware cap. Providers that bill chain-of-thought inside // completion tokens (observed on zai glm-5.2 despite `thinking: disabled`: // 669 reasoning + 70 output for one verdict) exhausted a 256-token budget