fix(ai-judge): restore authoritative 15s default timeout

This commit is contained in:
2026-08-17 19:17:40 +08:00
parent 850f36c7a4
commit 6c0e26bc6e
2 changed files with 21 additions and 13 deletions
+11 -6
View File
@@ -168,12 +168,17 @@ statistical framing: single verdicts are not oracles, only cohort rates
are meaningful.
**PIEXTENSIO-11 latency evidence (16 judgments across the three fixed
windows):** p50 4.2s, mean 4.9s, max 11.9s. The 60s timeout budget has
~5x headroom over the observed max; no timeout or infrastructure
failure occurred in any round. Verdict: the 60s default is safe;
tightening toward ~30s would still bound worst-case waits with ~2.5x
headroom, but calibration-grade tightening should wait for more
samples across providers.
windows):** p50 4.2s, mean 4.9s, max 11.9s, zero timeouts or
infrastructure failures. The canonical resolution (c0b0028d) fixes
15,000 ms as the default and 5,000–30,000 ms as the accepted config
range, calibrated on openai-codex/gpt-5.6-sol (p95 6.4s). The zai
glm-5.2 segment these rounds observed is **uncalibrated** under that
contract: it raced the 15s deadline twice at `thinking: high` before
the cohort and stayed within max 11.9s under the interim 60s bound.
Protocol for this segment: configure timeoutMs within the accepted
range (30s) via the config module, run as a distinct configuration
cohort (non-default timeout never inherits the default cohort's
calibration), and fail closed to defer otherwise.
**Cohort totals (3 rounds, fixed windows):** N=19, joined 19/19,
judgments 16, matrix allow|allow 8, allow|deny 2, defer|allow 5,
+10 -7
View File
@@ -7,13 +7,16 @@ import type { ModelRegistry } from "@earendil-works/pi-coding-agent";
import { buildJudgeContext, MAX_REASON_CODE_POINTS, REPORT_VERDICT_TOOL_NAME } from "./prompt";
import type { BashJudgmentEvidence } from "./evidence";
// 60s: glm-5.2 at the user's default `thinking: high` profile was observed
// both finishing in seconds and racing the former 15s deadline (one verdict
// landed at exactly 15.003s and was mislabeled `aborted`); high-variance
// reasoning latency needs the wider bound. The judge is the first chain
// link, so its wait delays the human prompt by at most this much.
// PIEXTENSIO-11 calibrates a final value from cohort data.
const DEFAULT_TIMEOUT_MS = 60_000;
// 15s is the PIEXTENSIO-11 calibrated default (canonical resolution c0b0028d):
// 15,000 ms total wall-clock deadline, accepted config range 5,000–30,000 ms,
// invalid values fail closed to this default. The judge is the first chain
// link, so its wait delays the human prompt by at most this much. Uncalibrated
// provider segments (e.g. zai glm-5.2, observed racing the deadline at
// `thinking: high`) should configure a non-default timeoutMs within the range
// and be treated as a distinct configuration cohort.
const DEFAULT_TIMEOUT_MS = 15_000;
export const MIN_TIMEOUT_MS = 5_000;
export const MAX_TIMEOUT_MS = 30_000;
// Reasoning-token aware cap. Providers that bill chain-of-thought inside
// completion tokens (observed on zai glm-5.2 despite `thinking: disabled`:
// 669 reasoning + 70 output for one verdict) exhausted a 256-token budget