mirror of
https://github.com/SikongJueluo/pi-extensions.git
synced 2026-10-05 11:52:55 +08:00
fix(ai-judge): restore authoritative 15s default timeout
This commit is contained in:
@@ -168,12 +168,17 @@ statistical framing: single verdicts are not oracles, only cohort rates
|
||||
are meaningful.
|
||||
|
||||
**PIEXTENSIO-11 latency evidence (16 judgments across the three fixed
|
||||
windows):** p50 4.2s, mean 4.9s, max 11.9s. The 60s timeout budget has
|
||||
~5x headroom over the observed max; no timeout or infrastructure
|
||||
failure occurred in any round. Verdict: the 60s default is safe;
|
||||
tightening toward ~30s would still bound worst-case waits with ~2.5x
|
||||
headroom, but calibration-grade tightening should wait for more
|
||||
samples across providers.
|
||||
windows):** p50 4.2s, mean 4.9s, max 11.9s, zero timeouts or
|
||||
infrastructure failures. The canonical resolution (c0b0028d) fixes
|
||||
15,000 ms as the default and 5,000–30,000 ms as the accepted config
|
||||
range, calibrated on openai-codex/gpt-5.6-sol (p95 6.4s). The zai
|
||||
glm-5.2 segment these rounds observed is **uncalibrated** under that
|
||||
contract: it raced the 15s deadline twice at `thinking: high` before
|
||||
the cohort and stayed within max 11.9s under the interim 60s bound.
|
||||
Protocol for this segment: configure timeoutMs within the accepted
|
||||
range (30s) via the config module, run as a distinct configuration
|
||||
cohort (non-default timeout never inherits the default cohort's
|
||||
calibration), and fail closed to defer otherwise.
|
||||
|
||||
**Cohort totals (3 rounds, fixed windows):** N=19, joined 19/19,
|
||||
judgments 16, matrix allow|allow 8, allow|deny 2, defer|allow 5,
|
||||
|
||||
@@ -7,13 +7,16 @@ import type { ModelRegistry } from "@earendil-works/pi-coding-agent";
|
||||
import { buildJudgeContext, MAX_REASON_CODE_POINTS, REPORT_VERDICT_TOOL_NAME } from "./prompt";
|
||||
import type { BashJudgmentEvidence } from "./evidence";
|
||||
|
||||
// 60s: glm-5.2 at the user's default `thinking: high` profile was observed
|
||||
// both finishing in seconds and racing the former 15s deadline (one verdict
|
||||
// landed at exactly 15.003s and was mislabeled `aborted`); high-variance
|
||||
// reasoning latency needs the wider bound. The judge is the first chain
|
||||
// link, so its wait delays the human prompt by at most this much.
|
||||
// PIEXTENSIO-11 calibrates a final value from cohort data.
|
||||
const DEFAULT_TIMEOUT_MS = 60_000;
|
||||
// 15s is the PIEXTENSIO-11 calibrated default (canonical resolution c0b0028d):
|
||||
// 15,000 ms total wall-clock deadline, accepted config range 5,000–30,000 ms,
|
||||
// invalid values fail closed to this default. The judge is the first chain
|
||||
// link, so its wait delays the human prompt by at most this much. Uncalibrated
|
||||
// provider segments (e.g. zai glm-5.2, observed racing the deadline at
|
||||
// `thinking: high`) should configure a non-default timeoutMs within the range
|
||||
// and be treated as a distinct configuration cohort.
|
||||
const DEFAULT_TIMEOUT_MS = 15_000;
|
||||
export const MIN_TIMEOUT_MS = 5_000;
|
||||
export const MAX_TIMEOUT_MS = 30_000;
|
||||
// Reasoning-token aware cap. Providers that bill chain-of-thought inside
|
||||
// completion tokens (observed on zai glm-5.2 despite `thinking: disabled`:
|
||||
// 669 reasoning + 70 output for one verdict) exhausted a 256-token budget
|
||||
|
||||
Reference in New Issue
Block a user