fix(ai-judge): restore authoritative 15s default timeout

This commit is contained in:
2026-08-17 19:17:40 +08:00
parent 850f36c7a4
commit 6c0e26bc6e
2 changed files with 21 additions and 13 deletions
+11 -6
View File
@@ -168,12 +168,17 @@ statistical framing: single verdicts are not oracles, only cohort rates
are meaningful. are meaningful.
**PIEXTENSIO-11 latency evidence (16 judgments across the three fixed **PIEXTENSIO-11 latency evidence (16 judgments across the three fixed
windows):** p50 4.2s, mean 4.9s, max 11.9s. The 60s timeout budget has windows):** p50 4.2s, mean 4.9s, max 11.9s, zero timeouts or
~5x headroom over the observed max; no timeout or infrastructure infrastructure failures. The canonical resolution (c0b0028d) fixes
failure occurred in any round. Verdict: the 60s default is safe; 15,000 ms as the default and 5,000–30,000 ms as the accepted config
tightening toward ~30s would still bound worst-case waits with ~2.5x range, calibrated on openai-codex/gpt-5.6-sol (p95 6.4s). The zai
headroom, but calibration-grade tightening should wait for more glm-5.2 segment these rounds observed is **uncalibrated** under that
samples across providers. contract: it raced the 15s deadline twice at `thinking: high` before
the cohort and stayed within max 11.9s under the interim 60s bound.
Protocol for this segment: configure timeoutMs within the accepted
range (30s) via the config module, run as a distinct configuration
cohort (non-default timeout never inherits the default cohort's
calibration), and fail closed to defer otherwise.
**Cohort totals (3 rounds, fixed windows):** N=19, joined 19/19, **Cohort totals (3 rounds, fixed windows):** N=19, joined 19/19,
judgments 16, matrix allow|allow 8, allow|deny 2, defer|allow 5, judgments 16, matrix allow|allow 8, allow|deny 2, defer|allow 5,
+10 -7
View File
@@ -7,13 +7,16 @@ import type { ModelRegistry } from "@earendil-works/pi-coding-agent";
import { buildJudgeContext, MAX_REASON_CODE_POINTS, REPORT_VERDICT_TOOL_NAME } from "./prompt"; import { buildJudgeContext, MAX_REASON_CODE_POINTS, REPORT_VERDICT_TOOL_NAME } from "./prompt";
import type { BashJudgmentEvidence } from "./evidence"; import type { BashJudgmentEvidence } from "./evidence";
// 60s: glm-5.2 at the user's default `thinking: high` profile was observed // 15s is the PIEXTENSIO-11 calibrated default (canonical resolution c0b0028d):
// both finishing in seconds and racing the former 15s deadline (one verdict // 15,000 ms total wall-clock deadline, accepted config range 5,000–30,000 ms,
// landed at exactly 15.003s and was mislabeled `aborted`); high-variance // invalid values fail closed to this default. The judge is the first chain
// reasoning latency needs the wider bound. The judge is the first chain // link, so its wait delays the human prompt by at most this much. Uncalibrated
// link, so its wait delays the human prompt by at most this much. // provider segments (e.g. zai glm-5.2, observed racing the deadline at
// PIEXTENSIO-11 calibrates a final value from cohort data. // `thinking: high`) should configure a non-default timeoutMs within the range
const DEFAULT_TIMEOUT_MS = 60_000; // and be treated as a distinct configuration cohort.
const DEFAULT_TIMEOUT_MS = 15_000;
export const MIN_TIMEOUT_MS = 5_000;
export const MAX_TIMEOUT_MS = 30_000;
// Reasoning-token aware cap. Providers that bill chain-of-thought inside // Reasoning-token aware cap. Providers that bill chain-of-thought inside
// completion tokens (observed on zai glm-5.2 despite `thinking: disabled`: // completion tokens (observed on zai glm-5.2 despite `thinking: disabled`:
// 669 reasoning + 70 output for one verdict) exhausted a 256-token budget // 669 reasoning + 70 output for one verdict) exhausted a 256-token budget