mirror of
https://github.com/SikongJueluo/pi-extensions.git
synced 2026-10-05 11:52:55 +08:00
feat(ai-judge): rework enforce mode as user-assumed-risk contract
- add config v2 with fixed judge model selection and fail-closed v1 enforce migration to shadow - add built-in high-risk override for irreversible, publish, system, and credential shapes that always defers to the human - remove the promotion gate, promotion records tool, and their tests - fail closed with judge_model_unavailable when a configured judge model cannot be resolved - notify once per session in enforce mode with the judge model and risk contract - add ADR 0008 and update README and CONTEXT
This commit is contained in:
+20
-5
@@ -8,11 +8,16 @@ extensions that inspect and re-evaluate Bash commands before they are allowed.
|
|||||||
### Authority
|
### Authority
|
||||||
|
|
||||||
**Enforce authority**:
|
**Enforce authority**:
|
||||||
Allow-only delegation — when the Judge's verdict is `allow` and every gate in
|
Allow-only delegation under a user-assumed-risk contract (ADR 0008):
|
||||||
the Enforce truth table passes, the command runs without the human dialog. The
|
writing `mode: "enforce"` in config v2 consents to the selected judge model
|
||||||
Judge can never answer `deny` with authority; every uncertain case (defer,
|
approving ordinary operations on the user's behalf. When the Judge's verdict
|
||||||
deny, preflight, infrastructure failure) falls back to the human dialog.
|
is `allow` and every fail-closed health gate in the Enforce truth table
|
||||||
_Avoid_: AI takeover, auto-deny, full delegation
|
passes (audit, telemetry, result kind, review acknowledgement,
|
||||||
|
generation-current), the command runs without the human dialog. The Judge can
|
||||||
|
never answer `deny` with authority; every uncertain case (defer, deny,
|
||||||
|
preflight, infrastructure failure) falls back to the human dialog.
|
||||||
|
_Avoid_: model safety certification (dropped — see ADR 0008), AI takeover,
|
||||||
|
auto-deny, full delegation
|
||||||
|
|
||||||
**Audit log (Judge-owned)**:
|
**Audit log (Judge-owned)**:
|
||||||
The Enforce-era accountability record written by the Judge package itself —
|
The Enforce-era accountability record written by the Judge package itself —
|
||||||
@@ -21,6 +26,16 @@ record, self-checked for health; an unhealthy audit log refuses authority.
|
|||||||
_Avoid_: host contract (dropped — see ADR 0006), acknowledged write (upstream
|
_Avoid_: host contract (dropped — see ADR 0006), acknowledged write (upstream
|
||||||
sense)
|
sense)
|
||||||
|
|
||||||
|
**High-risk override**:
|
||||||
|
A narrow code-level preflight recognizing clear-cut command shapes in four
|
||||||
|
categories — data loss/history rewrite, publish/deploy/infrastructure
|
||||||
|
destruction, privilege escalation/system modification, direct credential
|
||||||
|
access — that always defers to the human dialog regardless of verdict. In
|
||||||
|
Enforce it skips the model call entirely; in Shadow the model is still called
|
||||||
|
for quality observation. Deliberately not a sandbox: only well-known explicit
|
||||||
|
shapes, no alias/script/variable analysis, no `alwaysPrompt` config.
|
||||||
|
_Avoid_: command sandbox, alwaysPrompt, opaque-command blanket defer
|
||||||
|
|
||||||
**Irreversibility boundary**:
|
**Irreversibility boundary**:
|
||||||
An allow verdict requires every operation's effects to be recoverable —
|
An allow verdict requires every operation's effects to be recoverable —
|
||||||
reversible, or reproducible from the repository or the evidence at hand.
|
reversible, or reproducible from the repository or the evidence at hand.
|
||||||
|
|||||||
@@ -0,0 +1,98 @@
|
|||||||
|
---
|
||||||
|
status: accepted
|
||||||
|
---
|
||||||
|
|
||||||
|
# Enforce as user-assumed risk — promotion gates leave the runtime
|
||||||
|
|
||||||
|
PIEXTENSIO-23. This ADR supersedes the **runtime effect** of the
|
||||||
|
per-segment promotion floor (PIEXTENSIO-10, as consumed by PIEXTENSIO-21)
|
||||||
|
while keeping every historical artifact immutable. Enforce stops being a
|
||||||
|
"model earned safety certification" mode and becomes what it actually is:
|
||||||
|
a convenience mode where the user explicitly assumes the risk of model
|
||||||
|
misjudgment to reduce manual review.
|
||||||
|
|
||||||
|
## What changes
|
||||||
|
|
||||||
|
1. **The three promotion record gates (`cohort_qualified`,
|
||||||
|
`owner_approval`, `activation`) are no longer Enforce authority
|
||||||
|
inputs.** The truth table keeps its fail-closed health gates — audit
|
||||||
|
health, telemetry, result kind, review acknowledgement,
|
||||||
|
generation-current — but a session with no promotion records can now
|
||||||
|
hold Enforce authority.
|
||||||
|
2. **Risk contract replaces certification.** Writing `mode: "enforce"`
|
||||||
|
in config v2 *is* the consent: the user authorizes the selected judge
|
||||||
|
model to approve ordinary operations on their behalf and accepts
|
||||||
|
missed/wrong judgments. The project makes no "model safety
|
||||||
|
certification" claim and does not attempt to cover all dangerous shell
|
||||||
|
behavior. AI `allow` can auto-run; AI `deny`/`defer` still fall back
|
||||||
|
to the human dialog (no auto-deny authority is introduced).
|
||||||
|
3. **Narrow built-in high-risk override (code-level, not a sandbox).**
|
||||||
|
A small deterministic classifier recognizes clear-cut command shapes in
|
||||||
|
four categories — data loss / history rewrite (ADR 0007's enumerated
|
||||||
|
shapes), publish/deploy/infrastructure destruction, privilege
|
||||||
|
escalation / system modification, and direct credential
|
||||||
|
read/output/delete/replace. In Shadow it still calls the model
|
||||||
|
(quality observation stays) and records the override; in Enforce it
|
||||||
|
skips the model and defers immediately with an explicit risk code.
|
||||||
|
It matches only well-known explicit shapes — no alias resolution,
|
||||||
|
script-content analysis, variable expansion, or "every opaque command
|
||||||
|
defers" rule. There is deliberately **no `alwaysPrompt` user
|
||||||
|
config**: that was rejected as re-inventing policy text-matching; if a
|
||||||
|
"this permission rule always forces a human" need ever emerges, it
|
||||||
|
belongs in the permission-system's own `force_prompt` /
|
||||||
|
delegation-exclusion semantics, not in the Judge.
|
||||||
|
4. **Config v2 with explicit migration.** `version: 2` adds
|
||||||
|
`model: {provider, id}` — a fixed judge model, resolved per ask; if it
|
||||||
|
cannot be resolved (unknown model, unauthenticated, unsupported API)
|
||||||
|
the ask defers with a visible infrastructure code, never silently
|
||||||
|
falling back to the session model. Unset `model` keeps following the
|
||||||
|
session model (legacy behavior). Legacy v1 `mode: "enforce"` meant
|
||||||
|
"exact identity certified" under the old governance; upgrading must
|
||||||
|
not silently widen that consent, so a config without `version: 2`
|
||||||
|
that sets `enforce` falls back to Shadow with a diagnostic requiring
|
||||||
|
one explicit migration.
|
||||||
|
5. **One session-start notification** in Enforce mode shows the actual
|
||||||
|
judge model and the risk contract; it does not repeat per ask.
|
||||||
|
|
||||||
|
## What is explicitly preserved
|
||||||
|
|
||||||
|
- **Historical evidence immutability.** `promotion-records.jsonl`, the
|
||||||
|
cohort declarations/reports under `docs/testing/`, and every archived
|
||||||
|
analyzer artifact stay byte-for-byte as they are. The runtime stops
|
||||||
|
*consuming* them; nobody deletes or rewrites them. The failed cohorts
|
||||||
|
are not reopened.
|
||||||
|
- **ADR 0006 self-sufficiency stands.** The Judge-owned audit log, its
|
||||||
|
sticky health gate, and the fail-closed truth-table discipline are
|
||||||
|
unchanged; only the three promotion inputs leave the table.
|
||||||
|
- **ADR 0007's boundary is unchanged in substance** — irreversible
|
||||||
|
operations always defer. v4 keeps it at the prompt level; this ADR
|
||||||
|
adds a deterministic preflight backstop for the enumerated shapes
|
||||||
|
(exactly the fallback 0007 anticipated: "deterministic preflight defer
|
||||||
|
for the enumerated irreversible shapes").
|
||||||
|
|
||||||
|
## Why
|
||||||
|
|
||||||
|
The certification framing asked a personal convenience tool to carry an
|
||||||
|
enterprise-grade evidentiary apparatus, and the apparatus could not even
|
||||||
|
carry itself: v4 exposed a protocol contradiction — the declaration froze
|
||||||
|
`earliest report asOf 22:05Z` while the report's actual `asOf` is
|
||||||
|
`21:15:06Z`, the PASS outcome was recorded at `21:20Z` (before the frozen
|
||||||
|
earliest report), and the `-04` declaration incorporates `-02`'s frozen
|
||||||
|
sections "by reference" although the current file no longer contains them.
|
||||||
|
These are recorded here as **legacy-mechanism limitations**, not rewritten
|
||||||
|
away. The honest conclusion is that the runtime gate derived no real
|
||||||
|
safety from a process this fragile, while the user-visible value of
|
||||||
|
Enforce (fewer dialogs on ordinary commands) never needed it. What
|
||||||
|
actually needs to stay certain — irreversible shapes reach a human,
|
||||||
|
auditability, fail-closed behavior on infra failures — is kept by the
|
||||||
|
narrow override and the retained health gates.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
- `promotion.ts`, its tests, and `tools/promotion-record.ts` are deleted;
|
||||||
|
disk records survive. Future model-quality evidence (advisory catalog,
|
||||||
|
replay runs) informs *recommendations* and documentation only, never
|
||||||
|
runtime authority — milestone 2 of PIEXTENSIO-23.
|
||||||
|
- Rollback is `mode: "shadow"`.
|
||||||
|
- Revisit trigger: multi-user / delegated deployments where "the config
|
||||||
|
author is the risk bearer" stops being true.
|
||||||
@@ -1,4 +1,14 @@
|
|||||||
|
|
||||||
|
> **Status note (2026-08-21):** superseded as a runtime governance
|
||||||
|
> mechanism by [ADR 0008](../adr/0008-enforce-user-assumed-risk.md)
|
||||||
|
> (PIEXTENSIO-23) — Enforce authority no longer consumes promotion records
|
||||||
|
> or cohort qualification. This document is preserved unchanged as
|
||||||
|
> historical audit material. It also carries a recorded protocol
|
||||||
|
> contradiction (frozen `earliest report asOf 22:05Z` vs the report's
|
||||||
|
> `asOf 21:15:06Z` and a PASS recorded at `21:20Z`; `-02` frozen sections
|
||||||
|
> incorporated by reference are no longer in this file) — noted in ADR 0008
|
||||||
|
> as a legacy-mechanism limitation, not rewritten.
|
||||||
|
|
||||||
## Outcome of `piextensio-22-v4-gpt56sol-20260820-03`: FAIL (activity floor), 87 enrollments
|
## Outcome of `piextensio-22-v4-gpt56sol-20260820-03`: FAIL (activity floor), 87 enrollments
|
||||||
|
|
||||||
Recorded `2026-08-20T18:50:00Z`, after the window closed at `13:20:00Z`.
|
Recorded `2026-08-20T18:50:00Z`, after the window closed at `13:20:00Z`.
|
||||||
|
|||||||
@@ -1,3 +1,10 @@
|
|||||||
|
> **Status note (2026-08-21):** superseded as a runtime governance
|
||||||
|
> mechanism by [ADR 0008](../adr/0008-enforce-user-assumed-risk.md)
|
||||||
|
> (PIEXTENSIO-23) — Enforce authority no longer consumes cohort
|
||||||
|
> qualification. This report is preserved unchanged as historical audit
|
||||||
|
> material; its results speak to observed model quality, not to any
|
||||||
|
> certification of safety.
|
||||||
|
|
||||||
# PIEXTENSIO-22 v4 Enforce promotion cohort report
|
# PIEXTENSIO-22 v4 Enforce promotion cohort report
|
||||||
|
|
||||||
## Decision
|
## Decision
|
||||||
|
|||||||
@@ -1,14 +1,14 @@
|
|||||||
# pi-permission-ai-judge
|
# pi-permission-ai-judge
|
||||||
|
|
||||||
[`pi`](https://github.com/earendil-works/pi-coding-agent) 的 Bash 权限 AI 判官——对每一条进入权限对话框的命令,用会话模型给出第二意见。
|
[`pi`](https://github.com/earendil-works/pi-coding-agent) 的 Bash 权限 AI 判官——对每一条进入权限对话框的命令,用 AI 模型给出第二意见。
|
||||||
|
|
||||||
- **影子模式**(默认):判官并行给出"允许 / 拒绝 / 交给人类"的判决,人类照常决策;判决与证据元数据全部落盘,可离线分析
|
- **影子模式**(默认):判官并行给出"允许 / 拒绝 / 交给人类"的判决,人类照常决策;判决与证据元数据全部落盘,可离线分析
|
||||||
- **强制模式**:仅允许委托——仅当判决为"允许"且全部治理门通过时跳过人类对话框;任何不确定(交人类 / 拒绝 / 超时 / 健康异常)一律回落对话框。不可逆操作(如 `git clean -xfd`)永远交给人类,详见 ADR 0007
|
- **强制模式**(config v2):**用户自担风险的便利模式**——判官判决"允许"且运行时健康门全部通过时跳过人类对话框;判官"拒绝/交人类/超时/健康异常"一律回落对话框,不引入自动拒绝。明确的高风险命令形状(不可逆删除、发布/基础设施销毁、提权/系统修改、直接读删凭据)永远交人类。详见 ADR 0008
|
||||||
|
|
||||||
## 安装
|
## 安装
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
pi install github.com/Sikongjueluo/pi-extensions
|
pi install github.com/SikongJueluo/pi-extensions
|
||||||
```
|
```
|
||||||
|
|
||||||
依赖项目启用 [`@gotgenes/pi-permission-system`](https://github.com/gotgenes/pi-permission-system) ≥ 25.4,并配置授权链(仅 UI 会话生效):
|
依赖项目启用 [`@gotgenes/pi-permission-system`](https://github.com/gotgenes/pi-permission-system) ≥ 25.4,并配置授权链(仅 UI 会话生效):
|
||||||
@@ -23,27 +23,30 @@ pi install github.com/Sikongjueluo/pi-extensions
|
|||||||
`~/.pi/agent/pi-permission-ai-judge.config.json`,会话启动时读取:
|
`~/.pi/agent/pi-permission-ai-judge.config.json`,会话启动时读取:
|
||||||
|
|
||||||
```json
|
```json
|
||||||
{ "mode": "shadow", "timeoutMs": 30000 }
|
{
|
||||||
|
"version": 2,
|
||||||
|
"mode": "enforce",
|
||||||
|
"model": { "provider": "openai-codex", "id": "gpt-5.6-sol" },
|
||||||
|
"timeoutMs": 30000
|
||||||
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
- `mode`:`shadow`(影子)| `enforce`(强制),默认影子;非法值一律回退影子
|
- `version`:配置协议版本。**旧 v1 配置(无 `version` 字段)里的 `mode: "enforce"` 不会静默获得新授权**——会回退 shadow 并提示迁移;写入 `"version": 2` 即完成一次显式迁移,这个动作本身就是风险同意
|
||||||
- `timeoutMs`:5000–30000,默认 15000。**非默认超时是独立的配置群组**,不继承其他群组的晋升记录
|
- `mode`:`shadow`(默认)| `enforce`(强制)。非法值一律回退 shadow
|
||||||
|
- `model`(可选,仅 v2):固定判官模型,不随会话模型切换。配置的模型不存在、无认证或 API 不支持时,该次询问记为基础设施失败并交人类,**绝不静默改用会话模型**;未配置时跟随当前会话模型
|
||||||
|
- `timeoutMs`:5000–30000,默认 15000
|
||||||
|
|
||||||
## 强制模式晋升(仅仓库所有者;普通用户无需任何操作)
|
强制模式每次会话启动时弹一次非阻塞通知,显示实际判官模型与风险契约;不逐次重复。
|
||||||
|
|
||||||
默认影子模式对使用者零门槛。强制模式意味着“判官说允许就跳过对话框”,因此启用前需要证据链:影子模式下收集一批真实流量(群组,如 110 条)证明零误放,然后所有者显式记录三次——群组合格、批准、激活(三个独立动作,精确绑定候选身份:模型 × 提示词版本 × 超时群组等 9 个字段)。任一缺失或身份漂移即自动关闸。
|
## 强制模式 = 风险契约(不是安全认证)
|
||||||
|
|
||||||
**使用者两种选择:**信任本仓库已归档的群组证据,直接拷贝 `promotion-records.jsonl` 并配置 `mode: "enforce"`;或换用自己的模型,重新收集群组后自行记录。所有者记录命令:
|
手写 `mode: "enforce"` 即表示:你授权所选判官模型代为批准普通操作,并**自行承担漏判/误判风险**(ADR 0008)。本包不承诺任何"模型已获安全认证",也不试图覆盖所有危险 shell 行为。保留的运行时保护:
|
||||||
|
|
||||||
```bash
|
- **健康门**(全部 fail-closed):审计日志健康、遥测健康、判决结果类型、审计回执、会话世代——任一失败即回退对话框
|
||||||
cd packages/pi-permission-ai-judge
|
- **内置高风险 override**(窄范围代码级规则,不是 sandbox):明确的数据丢失/历史重写(`git clean -xfd`、`git reset --hard`、`git push --force`、`rm -rf ~` 等)、发布/部署/基础设施销毁(`npm publish`、`terraform destroy` 等)、提权/系统修改(`sudo`、`mkfs`、`dd of=/dev/*`、`shutdown` 等)、直接凭据读取/输出/删除(`cat ~/.ssh/id_*`、`~/.aws/credentials`、`~/.gnupg` 等)。Enforce 命中时不调模型、立即交人类;Shadow 命中时仍调模型(保留质量观测)并记录 override,最终照常人类决策。不解析别名/脚本内容/变量展开,没有 `alwaysPrompt` 配置
|
||||||
npx tsx tools/promotion-record.ts --kind cohort_qualified \
|
- 回滚 = 配置切回 `shadow`
|
||||||
--provider openai-codex --model gpt-5.6-sol --api openai-codex-responses \
|
|
||||||
--timeout-cohort 30000 --basis "cohort id; report path"
|
|
||||||
# 再依次 --kind owner_approval、--kind activation(三个独立显式动作)
|
|
||||||
```
|
|
||||||
|
|
||||||
记录写入 `~/.pi/agent/extensions/pi-permission-ai-judge/promotion-records.jsonl`(只追加)。回滚 = 配置切回 `shadow`;记录保留作审计轨迹。
|
历史治理(v4 及更早的 promotion cohort、三重记录门)已被 ADR 0008 取代,运行时不再读取 `promotion-records.jsonl`;旧记录与 cohort 报告原样保留作审计资料,见 `docs/testing/`。
|
||||||
|
|
||||||
## 日志与工具
|
## 日志与工具
|
||||||
|
|
||||||
@@ -54,5 +57,6 @@ npx tsx tools/promotion-record.ts --kind cohort_qualified \
|
|||||||
|
|
||||||
## 文档
|
## 文档
|
||||||
|
|
||||||
- 治理与晋升底线:PIEXTENSIO-10;v3 失败与 v4 修复:PIEXTENSIO-19/20/22,见 `docs/testing/`
|
- 风险契约与治理变更:ADR 0008 / PIEXTENSIO-23
|
||||||
- ADR 0006(审计自持)、0007(不可逆边界):见 `docs/adr/`
|
- 审计自持:ADR 0006;不可逆边界:ADR 0007,见 `docs/adr/`
|
||||||
|
- 旧 cohort 证据(v4 前含晋升治理):`docs/testing/`
|
||||||
|
|||||||
@@ -4,21 +4,32 @@ import { DEFAULT_TIMEOUT_MS, MAX_TIMEOUT_MS, MIN_TIMEOUT_MS } from "./model";
|
|||||||
|
|
||||||
/**
|
/**
|
||||||
* Judge configuration (PIEXTENSIO-3 Config acceptance, PIEXTENSIO-11
|
* Judge configuration (PIEXTENSIO-3 Config acceptance, PIEXTENSIO-11
|
||||||
* configuration contract).
|
* configuration contract; version 2 — PIEXTENSIO-23 / ADR 0008).
|
||||||
*
|
*
|
||||||
* The single user-configurable field is `timeoutMs` because acceptable
|
* Version 2 adds `model: {provider, id}` — an explicit fixed judge model
|
||||||
* interactive wait varies by deployment. Mode is read but v0.1 authority is
|
* — and makes hand-written `mode: "enforce"` the user's risk consent. A
|
||||||
* mechanically fail-closed: `enforce` loads but the judge truth table
|
* legacy (version 1 / unversioned) `mode: "enforce"` meant "exact
|
||||||
* (PIEXTENSIO-12) can never produce real authority until the promotion
|
* identity certified" under the retired promotion governance; it must not
|
||||||
* gates exist upstream, so it always defers.
|
* silently gain the new, broader consent, so it fails closed to shadow
|
||||||
|
* with a diagnostic requiring one explicit migration to `version: 2`.
|
||||||
*
|
*
|
||||||
* Global config path only — no project/env override, matching the trusted
|
* Everything invalid fails closed to the documented defaults
|
||||||
* user-global Judge configuration of PIEXTENSIO-10.
|
* (PIEXTENSIO-11) and records a diagnostic. Global config path only — no
|
||||||
|
* project/env override, matching the trusted user-global Judge
|
||||||
|
* configuration.
|
||||||
*/
|
*/
|
||||||
|
|
||||||
export type JudgeMode = "shadow" | "enforce";
|
export type JudgeMode = "shadow" | "enforce";
|
||||||
|
|
||||||
|
/** A user-selected fixed judge model (config v2). */
|
||||||
|
export interface JudgeModelSelection {
|
||||||
|
readonly provider: string;
|
||||||
|
readonly id: string;
|
||||||
|
}
|
||||||
|
|
||||||
export interface EffectiveJudgeConfig {
|
export interface EffectiveJudgeConfig {
|
||||||
|
/** Config schema version the effective values were loaded under. */
|
||||||
|
readonly configVersion: 1 | 2;
|
||||||
readonly mode: JudgeMode;
|
readonly mode: JudgeMode;
|
||||||
readonly timeoutMs: number;
|
readonly timeoutMs: number;
|
||||||
/**
|
/**
|
||||||
@@ -27,6 +38,8 @@ export interface EffectiveJudgeConfig {
|
|||||||
* (PIEXTENSIO-11). `default` marks the calibrated default.
|
* (PIEXTENSIO-11). `default` marks the calibrated default.
|
||||||
*/
|
*/
|
||||||
readonly timeoutCohort: "default" | number;
|
readonly timeoutCohort: "default" | number;
|
||||||
|
/** Fixed judge model (v2 only); undefined follows the session model. */
|
||||||
|
readonly judgeModel: JudgeModelSelection | undefined;
|
||||||
/** Validation diagnostics for the loaded raw file, newest wins per key. */
|
/** Validation diagnostics for the loaded raw file, newest wins per key. */
|
||||||
readonly diagnostics: readonly ConfigDiagnostic[];
|
readonly diagnostics: readonly ConfigDiagnostic[];
|
||||||
}
|
}
|
||||||
@@ -46,17 +59,44 @@ export interface ConfigLoadDeps {
|
|||||||
|
|
||||||
const CONFIG_FILENAME = "pi-permission-ai-judge.config.json";
|
const CONFIG_FILENAME = "pi-permission-ai-judge.config.json";
|
||||||
const DEFAULT_CONFIG: EffectiveJudgeConfig = {
|
const DEFAULT_CONFIG: EffectiveJudgeConfig = {
|
||||||
|
configVersion: 1,
|
||||||
mode: "shadow",
|
mode: "shadow",
|
||||||
timeoutMs: DEFAULT_TIMEOUT_MS,
|
timeoutMs: DEFAULT_TIMEOUT_MS,
|
||||||
timeoutCohort: "default",
|
timeoutCohort: "default",
|
||||||
|
judgeModel: undefined,
|
||||||
diagnostics: [],
|
diagnostics: [],
|
||||||
};
|
};
|
||||||
|
|
||||||
|
function fileDefaults(problem: string): EffectiveJudgeConfig {
|
||||||
|
return {
|
||||||
|
...DEFAULT_CONFIG,
|
||||||
|
diagnostics: [{ key: "file", problem, fallback: "all defaults" }],
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
function isNonEmptyString(value: unknown): value is string {
|
||||||
|
return typeof value === "string" && value.trim().length > 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
function parseJudgeModel(value: unknown): JudgeModelSelection | undefined {
|
||||||
|
if (typeof value !== "object" || value === null || Array.isArray(value)) {
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
const record = value as Record<string, unknown>;
|
||||||
|
if (!isNonEmptyString(record.provider) || !isNonEmptyString(record.id)) {
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
return {
|
||||||
|
provider: record.provider.trim(),
|
||||||
|
id: record.id.trim(),
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Load and validate the global config. Missing file, malformed JSON,
|
* Load and validate the global config. Missing file, malformed JSON, or
|
||||||
* unknown mode, or out-of-range timeout all fail closed to the documented
|
* an unknown version resolve to the documented defaults with one
|
||||||
* defaults (PIEXTENSIO-11: "invalid values fail closed to the documented
|
* diagnostic. Version 1 / unversioned files keep v1 semantics except
|
||||||
* default rather than becoming unbounded") and record a diagnostic.
|
* that `enforce` fails closed to shadow pending explicit migration.
|
||||||
*/
|
*/
|
||||||
export function loadJudgeConfig(
|
export function loadJudgeConfig(
|
||||||
deps: ConfigLoadDeps,
|
deps: ConfigLoadDeps,
|
||||||
@@ -67,51 +107,49 @@ export function loadJudgeConfig(
|
|||||||
try {
|
try {
|
||||||
raw = read(path);
|
raw = read(path);
|
||||||
} catch (error) {
|
} catch (error) {
|
||||||
return {
|
return fileDefaults(
|
||||||
...DEFAULT_CONFIG,
|
`config not readable at ${path}: ${error instanceof Error ? error.message : String(error)}`,
|
||||||
diagnostics: [
|
);
|
||||||
{
|
|
||||||
key: "file",
|
|
||||||
problem: `config not readable at ${path}: ${error instanceof Error ? error.message : String(error)}`,
|
|
||||||
fallback: "all defaults",
|
|
||||||
},
|
|
||||||
],
|
|
||||||
};
|
|
||||||
}
|
}
|
||||||
|
|
||||||
let parsed: unknown;
|
let parsed: unknown;
|
||||||
try {
|
try {
|
||||||
parsed = JSON.parse(raw);
|
parsed = JSON.parse(raw);
|
||||||
} catch (error) {
|
} catch (error) {
|
||||||
return {
|
return fileDefaults(
|
||||||
...DEFAULT_CONFIG,
|
`malformed JSON: ${error instanceof Error ? error.message : String(error)}`,
|
||||||
diagnostics: [
|
);
|
||||||
{
|
|
||||||
key: "file",
|
|
||||||
problem: `malformed JSON: ${error instanceof Error ? error.message : String(error)}`,
|
|
||||||
fallback: "all defaults",
|
|
||||||
},
|
|
||||||
],
|
|
||||||
};
|
|
||||||
}
|
}
|
||||||
|
|
||||||
if (typeof parsed !== "object" || parsed === null || Array.isArray(parsed)) {
|
if (typeof parsed !== "object" || parsed === null || Array.isArray(parsed)) {
|
||||||
return {
|
return fileDefaults("top-level value is not an object");
|
||||||
...DEFAULT_CONFIG,
|
|
||||||
diagnostics: [
|
|
||||||
{
|
|
||||||
key: "file",
|
|
||||||
problem: "top-level value is not an object",
|
|
||||||
fallback: "all defaults",
|
|
||||||
},
|
|
||||||
],
|
|
||||||
};
|
|
||||||
}
|
}
|
||||||
|
|
||||||
const record = parsed as Record<string, unknown>;
|
const record = parsed as Record<string, unknown>;
|
||||||
const diagnostics: ConfigDiagnostic[] = [];
|
const diagnostics: ConfigDiagnostic[] = [];
|
||||||
|
|
||||||
// Mode: unknown or missing resolves to shadow (fail-closed).
|
// Version: unversioned files are legacy v1; anything but 1 or 2 fails
|
||||||
|
// closed to all defaults (the consent expression is uninterpretable).
|
||||||
|
let configVersion: 1 | 2 = 1;
|
||||||
|
if (record.version !== undefined) {
|
||||||
|
if (record.version === 1 || record.version === 2) {
|
||||||
|
configVersion = record.version;
|
||||||
|
} else {
|
||||||
|
return {
|
||||||
|
...DEFAULT_CONFIG,
|
||||||
|
diagnostics: [
|
||||||
|
{
|
||||||
|
key: "version",
|
||||||
|
problem: `unknown version ${JSON.stringify(record.version)}`,
|
||||||
|
fallback: "all defaults",
|
||||||
|
},
|
||||||
|
],
|
||||||
|
};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Mode: unknown or missing resolves to shadow (fail-closed). A v1
|
||||||
|
// enforce cannot silently inherit the v2 risk contract (ADR 0008).
|
||||||
let mode: JudgeMode = "shadow";
|
let mode: JudgeMode = "shadow";
|
||||||
if (record.mode !== undefined) {
|
if (record.mode !== undefined) {
|
||||||
if (record.mode === "shadow" || record.mode === "enforce") {
|
if (record.mode === "shadow" || record.mode === "enforce") {
|
||||||
@@ -124,6 +162,39 @@ export function loadJudgeConfig(
|
|||||||
});
|
});
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
if (configVersion === 1 && mode === "enforce") {
|
||||||
|
mode = "shadow";
|
||||||
|
diagnostics.push({
|
||||||
|
key: "mode",
|
||||||
|
problem:
|
||||||
|
"v1 enforce means certified identity under the retired promotion governance (ADR 0008); add \"version\": 2 to consent to the user-assumed-risk contract",
|
||||||
|
fallback: "shadow (v1 enforce requires explicit migration)",
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
// Judge model: v2-only. A malformed model poisons the consent file —
|
||||||
|
// fail the whole config closed to shadow rather than silently
|
||||||
|
// substituting the session model.
|
||||||
|
let judgeModel: JudgeModelSelection | undefined;
|
||||||
|
if (record.model !== undefined) {
|
||||||
|
if (configVersion === 2) {
|
||||||
|
judgeModel = parseJudgeModel(record.model);
|
||||||
|
if (judgeModel === undefined) {
|
||||||
|
mode = "shadow";
|
||||||
|
diagnostics.push({
|
||||||
|
key: "model",
|
||||||
|
problem: `invalid model ${JSON.stringify(record.model)} (expected {provider, id} with non-empty strings)`,
|
||||||
|
fallback: "shadow (no judge model)",
|
||||||
|
});
|
||||||
|
}
|
||||||
|
} else {
|
||||||
|
diagnostics.push({
|
||||||
|
key: "model",
|
||||||
|
problem: "model selection requires \"version\": 2",
|
||||||
|
fallback: "session model",
|
||||||
|
});
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
// Timeout: integers in [5_000, 30_000]; anything else falls back to the
|
// Timeout: integers in [5_000, 30_000]; anything else falls back to the
|
||||||
// documented 15,000 ms default. Boundary semantics: 4,999 and 30,001
|
// documented 15,000 ms default. Boundary semantics: 4,999 and 30,001
|
||||||
@@ -149,5 +220,12 @@ export function loadJudgeConfig(
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
return Object.freeze({ mode, timeoutMs, timeoutCohort, diagnostics });
|
return Object.freeze({
|
||||||
|
configVersion,
|
||||||
|
mode,
|
||||||
|
timeoutMs,
|
||||||
|
timeoutCohort,
|
||||||
|
judgeModel,
|
||||||
|
diagnostics,
|
||||||
|
});
|
||||||
}
|
}
|
||||||
@@ -0,0 +1,360 @@
|
|||||||
|
/**
|
||||||
|
* Narrow built-in high-risk override (ADR 0008; PIEXTENSIO-23).
|
||||||
|
*
|
||||||
|
* A deterministic preflight that recognizes clear-cut command shapes in
|
||||||
|
* four categories and always defers them to the human dialog:
|
||||||
|
*
|
||||||
|
* - `data_loss` / `history_rewrite`: the ADR 0007 enumerated irreversible
|
||||||
|
* shapes (git clean -xfd, git reset --hard, git checkout/restore .,
|
||||||
|
* git push --force, rm -rf against home/root targets);
|
||||||
|
* - `publish_deploy`: package publishing and infrastructure destruction
|
||||||
|
* (npm/pnpm/yarn/cargo publish, terraform destroy);
|
||||||
|
* - `system_modify`: privilege escalation and system-level modification
|
||||||
|
* (sudo, mkfs, dd onto devices, shutdown/reboot/halt/poweroff);
|
||||||
|
* - `credential_access`: direct read/output/delete/replace of well-known
|
||||||
|
* credential files (~/.ssh/id_*, authorized_keys, ~/.aws/credentials,
|
||||||
|
* ~/.netrc, ~/.gnupg, gcloud application-default credentials), including
|
||||||
|
* redirect writes (`>`, `>>`, `N>`) and `tee`.
|
||||||
|
*
|
||||||
|
* Scanning is quote-, escape-, and separator-aware (`;`, `&&`, `||`, `|`,
|
||||||
|
* `&`, newlines outside quotes split units), so ordinary compound inputs
|
||||||
|
* are neither bypassed (`echo ready & npm publish`) nor false-positived
|
||||||
|
* (`printf 'x; npm publish;'`). Deliberately NOT a sandbox: only
|
||||||
|
* well-known explicit command shapes are matched. No alias resolution,
|
||||||
|
* no script-content analysis, no variable expansion, no "every opaque
|
||||||
|
* command defers" rule, and no user-facing `alwaysPrompt` config.
|
||||||
|
* Behavior: in Enforce a hit skips the model and defers immediately with
|
||||||
|
* code `high_risk_override`; in Shadow the model is still called (quality
|
||||||
|
* observation) and the override is recorded.
|
||||||
|
*/
|
||||||
|
|
||||||
|
export type HighRiskCategory =
|
||||||
|
| "data_loss"
|
||||||
|
| "history_rewrite"
|
||||||
|
| "publish_deploy"
|
||||||
|
| "system_modify"
|
||||||
|
| "credential_access";
|
||||||
|
|
||||||
|
export interface HighRiskMatch {
|
||||||
|
readonly category: HighRiskCategory;
|
||||||
|
readonly rule: string;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Characters that separate command units when unquoted and unescaped. */
|
||||||
|
const UNIT_SEPARATORS = new Set([";", "|", "&", "\n"]);
|
||||||
|
/** Output-redirection operators, emitted as standalone tokens (`>`, `>>`, `2>`). */
|
||||||
|
const REDIRECT_OP = /^(?:\d+)?>{1,2}$/;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Scan a complete Bash input into units of tokens. Single- and
|
||||||
|
* double-quoted segments become one token (quotes stripped); backslash
|
||||||
|
* escapes are taken literally; separators inside quotes or escapes do
|
||||||
|
* not split. Collapsed operators (`&&`, `||`) produce empty units that
|
||||||
|
* are dropped. This is lexical hygiene only — no alias, variable, or
|
||||||
|
* substitution semantics.
|
||||||
|
*/
|
||||||
|
function scanCommand(command: string): string[][] {
|
||||||
|
const units: string[][] = [];
|
||||||
|
let tokens: string[] = [];
|
||||||
|
let current = "";
|
||||||
|
let quote: string | null = null;
|
||||||
|
|
||||||
|
const pushToken = (): void => {
|
||||||
|
if (current.length > 0) {
|
||||||
|
tokens.push(current);
|
||||||
|
current = "";
|
||||||
|
}
|
||||||
|
};
|
||||||
|
const pushUnit = (): void => {
|
||||||
|
pushToken();
|
||||||
|
if (tokens.length > 0) {
|
||||||
|
units.push(tokens);
|
||||||
|
tokens = [];
|
||||||
|
}
|
||||||
|
};
|
||||||
|
|
||||||
|
for (let i = 0; i < command.length; i += 1) {
|
||||||
|
const ch = command[i];
|
||||||
|
if (quote !== null) {
|
||||||
|
if (ch === quote) {
|
||||||
|
quote = null;
|
||||||
|
} else {
|
||||||
|
current += ch;
|
||||||
|
}
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (ch === '"' || ch === "'") {
|
||||||
|
pushToken();
|
||||||
|
quote = ch;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (ch === "\\") {
|
||||||
|
current += command[i + 1] ?? "";
|
||||||
|
i += 1;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (UNIT_SEPARATORS.has(ch)) {
|
||||||
|
pushUnit();
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (ch === ">") {
|
||||||
|
const fd = /^\d+$/.exec(current);
|
||||||
|
pushToken();
|
||||||
|
let op = fd === null ? ">" : `${fd[0]}>`;
|
||||||
|
if (command[i + 1] === ">") {
|
||||||
|
op += ">";
|
||||||
|
i += 1;
|
||||||
|
}
|
||||||
|
tokens.push(op);
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (/\s/.test(ch)) {
|
||||||
|
pushToken();
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
current += ch;
|
||||||
|
}
|
||||||
|
pushUnit();
|
||||||
|
return units;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Combined single-letter git flags, e.g. `-xfd` → has x, f, d. */
|
||||||
|
function shortFlagLetters(token: string): Set<string> {
|
||||||
|
const letters = new Set<string>();
|
||||||
|
if (!/^-[a-zA-Z]+$/.test(token)) {
|
||||||
|
return letters;
|
||||||
|
}
|
||||||
|
for (const letter of token.slice(1)) {
|
||||||
|
letters.add(letter);
|
||||||
|
}
|
||||||
|
return letters;
|
||||||
|
}
|
||||||
|
|
||||||
|
function hasGitFlag(args: string[], letter: string, long: string): boolean {
|
||||||
|
for (const arg of args) {
|
||||||
|
if (arg === `--${long}`) {
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (shortFlagLetters(arg).has(letter)) {
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
function classifyGit(args: string[]): HighRiskMatch | undefined {
|
||||||
|
const subcommand = args[0];
|
||||||
|
const rest = args.slice(1);
|
||||||
|
switch (subcommand) {
|
||||||
|
case "clean": {
|
||||||
|
// git clean with force + ignored-file removal (-f/-x/-d shapes).
|
||||||
|
// Dry-run (-n) never matches.
|
||||||
|
const hasDryRun = rest.some((arg) => arg === "-n" || shortFlagLetters(arg).has("n"));
|
||||||
|
const force = hasGitFlag(rest, "f", "force");
|
||||||
|
const ignored = hasGitFlag(rest, "x", "x") || rest.includes("-x");
|
||||||
|
if (force && ignored && !hasDryRun) {
|
||||||
|
return { category: "data_loss", rule: "git_clean_forced_x" };
|
||||||
|
}
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
case "reset":
|
||||||
|
if (rest.includes("--hard")) {
|
||||||
|
return { category: "data_loss", rule: "git_reset_hard" };
|
||||||
|
}
|
||||||
|
return undefined;
|
||||||
|
case "checkout":
|
||||||
|
case "restore": {
|
||||||
|
// Whole-tree discard: `git checkout -- .` / `git checkout .` /
|
||||||
|
// `git restore .` (optionally `--worktree`-style flags first).
|
||||||
|
const positional = rest.filter((arg) => !arg.startsWith("-"));
|
||||||
|
const dotIndex = positional.lastIndexOf(".");
|
||||||
|
if (dotIndex !== -1 && positional.slice(dotIndex).every((arg) => arg === ".")) {
|
||||||
|
return {
|
||||||
|
category: "data_loss",
|
||||||
|
rule: subcommand === "checkout" ? "git_checkout_dot" : "git_restore_dot",
|
||||||
|
};
|
||||||
|
}
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
case "push": {
|
||||||
|
const forced =
|
||||||
|
rest.includes("--force") ||
|
||||||
|
rest.includes("-f") ||
|
||||||
|
rest.includes("--force-with-lease") ||
|
||||||
|
rest.some((arg) => arg.startsWith("+") && arg.length > 1);
|
||||||
|
if (forced) {
|
||||||
|
return { category: "history_rewrite", rule: "git_push_force" };
|
||||||
|
}
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
default:
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function classifyRm(args: string[]): HighRiskMatch | undefined {
|
||||||
|
const recursive = args.some((arg) => /^-[a-zA-Z]*[rR][a-zA-Z]*$/.test(arg) || arg === "-r" || arg === "-R" || arg === "--recursive");
|
||||||
|
const forced = hasGitFlag(args, "f", "force");
|
||||||
|
if (!recursive || !forced) {
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
const targets = args.filter((arg) => !arg.startsWith("-") && !REDIRECT_OP.test(arg));
|
||||||
|
const dangerous = targets.some(
|
||||||
|
(target) =>
|
||||||
|
target === "/" ||
|
||||||
|
target === "~" ||
|
||||||
|
target === "$HOME" ||
|
||||||
|
target.startsWith("/*") ||
|
||||||
|
target.startsWith("~/") ||
|
||||||
|
target.startsWith("~") ||
|
||||||
|
target.startsWith("$HOME/"),
|
||||||
|
);
|
||||||
|
return dangerous
|
||||||
|
? { category: "data_loss", rule: "rm_rf_home_or_root" }
|
||||||
|
: undefined;
|
||||||
|
}
|
||||||
|
|
||||||
|
function classifyPublish(head: string, args: string[]): HighRiskMatch | undefined {
|
||||||
|
const sub = args[0];
|
||||||
|
switch (head) {
|
||||||
|
case "npm":
|
||||||
|
case "pnpm":
|
||||||
|
case "yarn":
|
||||||
|
case "cargo":
|
||||||
|
if (sub === "publish") {
|
||||||
|
return { category: "publish_deploy", rule: `${head}_publish` };
|
||||||
|
}
|
||||||
|
return undefined;
|
||||||
|
case "terraform":
|
||||||
|
if (sub === "destroy") {
|
||||||
|
return { category: "publish_deploy", rule: "terraform_destroy" };
|
||||||
|
}
|
||||||
|
return undefined;
|
||||||
|
default:
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function classifySystem(head: string, args: string[]): HighRiskMatch | undefined {
|
||||||
|
if (head === "sudo") {
|
||||||
|
return { category: "system_modify", rule: "sudo" };
|
||||||
|
}
|
||||||
|
if (head.startsWith("mkfs")) {
|
||||||
|
return { category: "system_modify", rule: "mkfs" };
|
||||||
|
}
|
||||||
|
if (head === "shutdown" || head === "reboot" || head === "halt" || head === "poweroff") {
|
||||||
|
return { category: "system_modify", rule: head };
|
||||||
|
}
|
||||||
|
if (head === "dd") {
|
||||||
|
if (args.some((arg) => /^of=\/dev\//.test(arg))) {
|
||||||
|
return { category: "system_modify", rule: "dd_to_device" };
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Read/alter verbs that touch a credential path directly. */
|
||||||
|
const CREDENTIAL_VERBS = new Set([
|
||||||
|
"cat",
|
||||||
|
"less",
|
||||||
|
"more",
|
||||||
|
"head",
|
||||||
|
"tail",
|
||||||
|
"bat",
|
||||||
|
"zcat",
|
||||||
|
"tee",
|
||||||
|
"cp",
|
||||||
|
"mv",
|
||||||
|
"rm",
|
||||||
|
"install",
|
||||||
|
"scp",
|
||||||
|
"rsync",
|
||||||
|
]);
|
||||||
|
|
||||||
|
const CREDENTIAL_PATH_PATTERNS: ReadonlyArray<{ pattern: RegExp; rule: string }> = [
|
||||||
|
{ pattern: /^~\/\.ssh\/id_./, rule: "ssh_private_key" },
|
||||||
|
{ pattern: /^~\/\.ssh\/authorized_keys/, rule: "ssh_authorized_keys" },
|
||||||
|
{ pattern: /^~\/\.aws\/credentials/, rule: "aws_credentials" },
|
||||||
|
{ pattern: /^~\/\.netrc/, rule: "netrc" },
|
||||||
|
{ pattern: /^~\/\.gnupg(\/|$)/, rule: "gnupg" },
|
||||||
|
{
|
||||||
|
pattern: /^~\/\.config\/gcloud\/application_default_credentials/,
|
||||||
|
rule: "gcloud_default_credentials",
|
||||||
|
},
|
||||||
|
{ pattern: /^\$HOME\/\.ssh\/id_./, rule: "ssh_private_key" },
|
||||||
|
{ pattern: /^\$HOME\/\.ssh\/authorized_keys/, rule: "ssh_authorized_keys" },
|
||||||
|
{ pattern: /^\$HOME\/\.aws\/credentials/, rule: "aws_credentials" },
|
||||||
|
{ pattern: /^\$HOME\/\.netrc/, rule: "netrc" },
|
||||||
|
{ pattern: /^\$HOME\/\.gnupg(\/|$)/, rule: "gnupg" },
|
||||||
|
];
|
||||||
|
|
||||||
|
function matchCredentialPath(
|
||||||
|
value: string,
|
||||||
|
): { rule: string } | undefined {
|
||||||
|
for (const { pattern, rule } of CREDENTIAL_PATH_PATTERNS) {
|
||||||
|
if (pattern.test(value)) {
|
||||||
|
return { rule };
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
|
||||||
|
function classifyCredential(args: string[]): HighRiskMatch | undefined {
|
||||||
|
for (const arg of args) {
|
||||||
|
if (arg.startsWith("-") || REDIRECT_OP.test(arg)) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
const match = matchCredentialPath(arg);
|
||||||
|
if (match !== undefined) {
|
||||||
|
return { category: "credential_access", rule: match.rule };
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Direct credential replacement via output redirection — any verb writing
|
||||||
|
* a credential path (`>`, `>>`, fd-prefixed). Checked per unit regardless
|
||||||
|
* of the leading verb, since `echo x > ~/.aws/credentials` replaces the
|
||||||
|
* credential whatever the writer is.
|
||||||
|
*/
|
||||||
|
function classifyCredentialRedirect(tokens: string[]): HighRiskMatch | undefined {
|
||||||
|
for (let i = 0; i < tokens.length; i += 1) {
|
||||||
|
if (!REDIRECT_OP.test(tokens[i])) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
const target = tokens[i + 1];
|
||||||
|
if (target === undefined) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
const match = matchCredentialPath(target);
|
||||||
|
if (match !== undefined) {
|
||||||
|
return {
|
||||||
|
category: "credential_access",
|
||||||
|
rule: `${match.rule}_redirect_write`,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Classify a complete Bash input against the built-in high-risk shapes.
|
||||||
|
* Any matching unit marks the whole input; the first match wins.
|
||||||
|
*/
|
||||||
|
export function classifyHighRisk(fullCommand: string): HighRiskMatch | undefined {
|
||||||
|
for (const tokens of scanCommand(fullCommand)) {
|
||||||
|
const [head, ...args] = tokens;
|
||||||
|
const match =
|
||||||
|
classifyCredentialRedirect(tokens) ??
|
||||||
|
(CREDENTIAL_VERBS.has(head) ? classifyCredential(args) : undefined) ??
|
||||||
|
classifySystem(head, args) ??
|
||||||
|
(head === "git" ? classifyGit(args) : undefined) ??
|
||||||
|
(head === "rm" ? classifyRm(args) : undefined) ??
|
||||||
|
classifyPublish(head, args);
|
||||||
|
if (match !== undefined) {
|
||||||
|
return match;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
@@ -24,16 +24,8 @@ import {
|
|||||||
conversationProbeFromSession,
|
conversationProbeFromSession,
|
||||||
type ConversationEvidence,
|
type ConversationEvidence,
|
||||||
} from "./conversation";
|
} from "./conversation";
|
||||||
import {
|
import { classifyHighRisk, type HighRiskMatch } from "./highrisk";
|
||||||
evaluateEnforceAuthority,
|
import { evaluateEnforceAuthority, type EnforceGateState } from "./judge";
|
||||||
type EnforceGateState,
|
|
||||||
} from "./judge";
|
|
||||||
import {
|
|
||||||
loadPromotionRecords,
|
|
||||||
resolvePromotionGates,
|
|
||||||
type CandidateIdentity,
|
|
||||||
type PromotionRecordsSnapshot,
|
|
||||||
} from "./promotion";
|
|
||||||
|
|
||||||
const LINK_NAME = "ai-bash-judge";
|
const LINK_NAME = "ai-bash-judge";
|
||||||
const REVIEW_SCHEMA_VERSION = 1;
|
const REVIEW_SCHEMA_VERSION = 1;
|
||||||
@@ -58,15 +50,6 @@ interface RootSession {
|
|||||||
readonly getCwd: () => string;
|
readonly getCwd: () => string;
|
||||||
/** Judge-owned audit log (ADR 0006); unhealthy refuses Enforce authority. */
|
/** Judge-owned audit log (ADR 0006); unhealthy refuses Enforce authority. */
|
||||||
readonly auditLog: AuditLog;
|
readonly auditLog: AuditLog;
|
||||||
/** Promotion-records snapshot loaded at session start (PIEXTENSIO-21);
|
|
||||||
* gates are resolved per ask against the live candidate identity. */
|
|
||||||
readonly promotionRecords: PromotionRecordsSnapshot;
|
|
||||||
/** Static candidate-identity fields (everything except the model
|
|
||||||
* segment, which is captured per ask). */
|
|
||||||
readonly identityBase: Omit<
|
|
||||||
CandidateIdentity,
|
|
||||||
"provider" | "model" | "api"
|
|
||||||
>;
|
|
||||||
}
|
}
|
||||||
|
|
||||||
const EMPTY_CONVERSATION: ConversationEvidence = {
|
const EMPTY_CONVERSATION: ConversationEvidence = {
|
||||||
@@ -122,9 +105,6 @@ function resultBase(
|
|||||||
requestId: details.requestId,
|
requestId: details.requestId,
|
||||||
judgeRuntimeId,
|
judgeRuntimeId,
|
||||||
mode: config.mode,
|
mode: config.mode,
|
||||||
// v0.1 fail-closed: `enforce` loads here but the truth table can
|
|
||||||
// never grant real authority, so the effective verdict below stays
|
|
||||||
// defer; the configured mode is recorded for audit.
|
|
||||||
timeoutCohort: config.timeoutCohort,
|
timeoutCohort: config.timeoutCohort,
|
||||||
origin:
|
origin:
|
||||||
details.forwarding !== undefined ||
|
details.forwarding !== undefined ||
|
||||||
@@ -137,7 +117,6 @@ function resultBase(
|
|||||||
};
|
};
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
/** Read the permission-system review-log toggle (default true when unset). */
|
/** Read the permission-system review-log toggle (default true when unset). */
|
||||||
function readPermissionReviewLogEnabled(): boolean {
|
function readPermissionReviewLogEnabled(): boolean {
|
||||||
try {
|
try {
|
||||||
@@ -157,6 +136,30 @@ function readPermissionReviewLogEnabled(): boolean {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Resolve the judge model for one ask (ADR 0008 / PIEXTENSIO-23).
|
||||||
|
*
|
||||||
|
* A configured v2 `model` is a fixed judge model resolved per ask against
|
||||||
|
* the live registry. Resolution failure (unknown model, no configured
|
||||||
|
* auth) is reported as an explicit failure — never a silent fallback to
|
||||||
|
* the session model. No configured model follows the session model
|
||||||
|
* (legacy behavior).
|
||||||
|
*/
|
||||||
|
function resolveJudgeModel(
|
||||||
|
config: EffectiveJudgeConfig,
|
||||||
|
sessionModel: Model<any> | undefined,
|
||||||
|
registry: ModelRegistry,
|
||||||
|
): { kind: "model"; model: Model<any> | undefined; source: "configured" | "session" } | { kind: "unavailable" } {
|
||||||
|
if (config.judgeModel === undefined) {
|
||||||
|
return { kind: "model", model: sessionModel, source: "session" };
|
||||||
|
}
|
||||||
|
const found = registry.find(config.judgeModel.provider, config.judgeModel.id);
|
||||||
|
if (found === undefined || !registry.hasConfiguredAuth(found)) {
|
||||||
|
return { kind: "unavailable" };
|
||||||
|
}
|
||||||
|
return { kind: "model", model: found, source: "configured" };
|
||||||
|
}
|
||||||
|
|
||||||
/** Register a Shadow-only structured-output judge for local native Bash asks. */
|
/** Register a Shadow-only structured-output judge for local native Bash asks. */
|
||||||
export default function permissionAiJudge(pi: ExtensionAPI): void {
|
export default function permissionAiJudge(pi: ExtensionAPI): void {
|
||||||
let root: RootSession | undefined;
|
let root: RootSession | undefined;
|
||||||
@@ -282,14 +285,69 @@ export default function permissionAiJudge(pi: ExtensionAPI): void {
|
|||||||
return { kind: "defer" };
|
return { kind: "defer" };
|
||||||
}
|
}
|
||||||
|
|
||||||
// Per-request model capture (PIEXTENSIO-3 cat.3): an
|
// Built-in high-risk override (ADR 0008): clear-cut
|
||||||
// in-request switch does not change an in-flight attempt;
|
// irreversible/system shapes always defer. In Enforce the
|
||||||
// a between-request switch affects the next attempt. The
|
// model is skipped entirely; in Shadow it still runs for
|
||||||
// probe reads the live current model here, not at start.
|
// quality observation and the override is recorded.
|
||||||
const availability = createModelAvailability(
|
const risk: HighRiskMatch | undefined = classifyHighRisk(
|
||||||
|
evidence.fullCommand,
|
||||||
|
);
|
||||||
|
if (risk !== undefined && captured.config.mode === "enforce") {
|
||||||
|
emitResult({
|
||||||
|
...resultBase(
|
||||||
|
captured.judgeRuntimeId,
|
||||||
|
details,
|
||||||
|
startedAt,
|
||||||
|
captured.config,
|
||||||
|
),
|
||||||
|
resultKind: "preflight_defer",
|
||||||
|
verdict: null,
|
||||||
|
effectiveVerdict: "defer",
|
||||||
|
modelCalled: false,
|
||||||
|
code: "high_risk_override",
|
||||||
|
riskCategory: risk.category,
|
||||||
|
riskRule: risk.rule,
|
||||||
|
evidenceQuality: evidenceQuality(true, EMPTY_CONVERSATION, captured.getCwd()),
|
||||||
|
});
|
||||||
|
return { kind: "defer" };
|
||||||
|
}
|
||||||
|
|
||||||
|
// Per-request judge-model resolution (PIEXTENSIO-3 cat.3
|
||||||
|
// for the session model; ADR 0008 for a configured fixed
|
||||||
|
// model). A configured model that cannot be resolved is
|
||||||
|
// an observable infrastructure failure — never a silent
|
||||||
|
// fallback to the session model.
|
||||||
|
const resolved = resolveJudgeModel(
|
||||||
|
captured.config,
|
||||||
captured.getModel(),
|
captured.getModel(),
|
||||||
captured.modelRegistry,
|
captured.modelRegistry,
|
||||||
);
|
);
|
||||||
|
if (resolved.kind === "unavailable") {
|
||||||
|
emitResult({
|
||||||
|
...resultBase(
|
||||||
|
captured.judgeRuntimeId,
|
||||||
|
details,
|
||||||
|
startedAt,
|
||||||
|
captured.config,
|
||||||
|
),
|
||||||
|
resultKind: "infrastructure_failure",
|
||||||
|
verdict: null,
|
||||||
|
effectiveVerdict: "defer",
|
||||||
|
modelCalled: false,
|
||||||
|
code: "judge_model_unavailable",
|
||||||
|
provider: captured.config.judgeModel?.provider ?? null,
|
||||||
|
model: captured.config.judgeModel?.id ?? null,
|
||||||
|
api: null,
|
||||||
|
riskOverride: risk ?? null,
|
||||||
|
evidenceQuality: evidenceQuality(true, EMPTY_CONVERSATION, captured.getCwd()),
|
||||||
|
});
|
||||||
|
return { kind: "defer" };
|
||||||
|
}
|
||||||
|
const modelSource = resolved.source;
|
||||||
|
const availability: ModelAvailability = createModelAvailability(
|
||||||
|
resolved.model,
|
||||||
|
captured.modelRegistry,
|
||||||
|
);
|
||||||
// Conversation evidence is captured at ask time from the
|
// Conversation evidence is captured at ask time from the
|
||||||
// live serving branch, not at session start: the newest
|
// live serving branch, not at session start: the newest
|
||||||
// user intent is the ask's intent.
|
// user intent is the ask's intent.
|
||||||
@@ -303,36 +361,19 @@ export default function permissionAiJudge(pi: ExtensionAPI): void {
|
|||||||
conversation,
|
conversation,
|
||||||
);
|
);
|
||||||
|
|
||||||
// Enforce truth table (PIEXTENSIO-3 cat.4 / M5): the
|
// Enforce truth table (PIEXTENSIO-3 cat.4 / M5; ADR 0008):
|
||||||
// promotion gates resolve per ask from the session-start
|
// the fail-closed runtime health gates — audit health,
|
||||||
// records snapshot against the live candidate identity
|
// telemetry, result kind, verdict, review
|
||||||
// (PIEXTENSIO-21) — the current model segment is part of
|
// acknowledgement, generation currency — are the single
|
||||||
// that identity, so a mid-session model switch cannot
|
// authority seam; the retired promotion gates are no
|
||||||
// inherit another segment's records, and identity drift
|
// longer inputs. reviewAcknowledged is true in the ADR
|
||||||
// after recording fails closed. The truth table is the
|
// 0006 sense: the Judge-owned audit write for this result
|
||||||
// single authority seam: the owner's records flip the
|
// happens before the authority return, and a failed write
|
||||||
// gate inputs, never the callback's return path.
|
// flips auditHealthy sticky-unhealthy, closing authority
|
||||||
// reviewAcknowledged is true in the ADR 0006 sense: the
|
// for every later ask.
|
||||||
// Judge-owned audit write for this result happens before
|
|
||||||
// the authority return, and a failed write flips
|
|
||||||
// auditHealthy sticky-unhealthy, closing authority for
|
|
||||||
// every later ask.
|
|
||||||
const liveIdentity: CandidateIdentity = {
|
|
||||||
...captured.identityBase,
|
|
||||||
provider: result.metadata?.provider ?? "",
|
|
||||||
model: result.metadata?.model ?? "",
|
|
||||||
api: result.metadata?.api ?? "",
|
|
||||||
};
|
|
||||||
const gates = resolvePromotionGates(
|
|
||||||
captured.promotionRecords,
|
|
||||||
liveIdentity,
|
|
||||||
);
|
|
||||||
const gateState: EnforceGateState = {
|
const gateState: EnforceGateState = {
|
||||||
auditHealthy: captured.auditLog.healthy(),
|
auditHealthy: captured.auditLog.healthy(),
|
||||||
telemetryHealth: sink.health(),
|
telemetryHealth: sink.health(),
|
||||||
cohortQualified: gates.cohortQualified,
|
|
||||||
ownerApprovalRecorded: gates.ownerApprovalRecorded,
|
|
||||||
activationRecorded: gates.activationRecorded,
|
|
||||||
resultKind:
|
resultKind:
|
||||||
result.kind === "judgment"
|
result.kind === "judgment"
|
||||||
? "judgment"
|
? "judgment"
|
||||||
@@ -363,9 +404,11 @@ export default function permissionAiJudge(pi: ExtensionAPI): void {
|
|||||||
authorityBlockedBy,
|
authorityBlockedBy,
|
||||||
modelCalled: true,
|
modelCalled: true,
|
||||||
code: null,
|
code: null,
|
||||||
|
modelSource,
|
||||||
provider: result.metadata.provider,
|
provider: result.metadata.provider,
|
||||||
model: result.metadata.model,
|
model: result.metadata.model,
|
||||||
api: result.metadata.api,
|
api: result.metadata.api,
|
||||||
|
riskOverride: risk ?? null,
|
||||||
// Log keys deliberately avoid the substring
|
// Log keys deliberately avoid the substring
|
||||||
// "token": permission-system masks any key matching
|
// "token": permission-system masks any key matching
|
||||||
// /token/i (structural key-name redaction), which
|
// /token/i (structural key-name redaction), which
|
||||||
@@ -390,9 +433,11 @@ export default function permissionAiJudge(pi: ExtensionAPI): void {
|
|||||||
authorityBlockedBy,
|
authorityBlockedBy,
|
||||||
modelCalled: result.modelCalled,
|
modelCalled: result.modelCalled,
|
||||||
code: result.code,
|
code: result.code,
|
||||||
|
modelSource,
|
||||||
provider: result.metadata?.provider ?? null,
|
provider: result.metadata?.provider ?? null,
|
||||||
model: result.metadata?.model ?? null,
|
model: result.metadata?.model ?? null,
|
||||||
api: result.metadata?.api ?? null,
|
api: result.metadata?.api ?? null,
|
||||||
|
riskOverride: risk ?? null,
|
||||||
inputUsage: result.inputTokens ?? null,
|
inputUsage: result.inputTokens ?? null,
|
||||||
outputUsage: result.outputTokens ?? null,
|
outputUsage: result.outputTokens ?? null,
|
||||||
modelLatencyMs: result.modelLatencyMs,
|
modelLatencyMs: result.modelLatencyMs,
|
||||||
@@ -426,26 +471,6 @@ export default function permissionAiJudge(pi: ExtensionAPI): void {
|
|||||||
|
|
||||||
const runtimeId = crypto.randomUUID();
|
const runtimeId = crypto.randomUUID();
|
||||||
const config = loadJudgeConfig({ agentDir: getAgentDir() });
|
const config = loadJudgeConfig({ agentDir: getAgentDir() });
|
||||||
// Static candidate-identity fields. The judge and permission-system
|
|
||||||
// versions must track package.json / the peer floor; the cohort
|
|
||||||
// declaration used exactly these values.
|
|
||||||
const identityBase = {
|
|
||||||
judge: "@sikongjueluo/pi-permission-ai-judge@0.0.1",
|
|
||||||
permissionSystem: "25.4.0",
|
|
||||||
promptVersion: PROMPT_VERSION,
|
|
||||||
toolSchemaVersion: TOOL_SCHEMA_VERSION,
|
|
||||||
reviewSchemaVersion: String(REVIEW_SCHEMA_VERSION),
|
|
||||||
timeoutCohort: config.timeoutCohort,
|
|
||||||
};
|
|
||||||
const promotionRecords = loadPromotionRecords({
|
|
||||||
agentDir: getAgentDir(),
|
|
||||||
});
|
|
||||||
if (promotionRecords.diagnostic !== null) {
|
|
||||||
ctx.ui.notify(
|
|
||||||
`ai-bash-judge promotion records: ${promotionRecords.diagnostic}; Enforce gates stay closed`,
|
|
||||||
"warning",
|
|
||||||
);
|
|
||||||
}
|
|
||||||
root = {
|
root = {
|
||||||
getSessionId: () => ctx.sessionManager.getSessionId(),
|
getSessionId: () => ctx.sessionManager.getSessionId(),
|
||||||
expectedSessionId: sessionId,
|
expectedSessionId: sessionId,
|
||||||
@@ -459,8 +484,6 @@ export default function permissionAiJudge(pi: ExtensionAPI): void {
|
|||||||
agentDir: getAgentDir(),
|
agentDir: getAgentDir(),
|
||||||
runtimeId,
|
runtimeId,
|
||||||
}),
|
}),
|
||||||
promotionRecords,
|
|
||||||
identityBase,
|
|
||||||
conversation: conversationProbeFromSession(ctx.sessionManager),
|
conversation: conversationProbeFromSession(ctx.sessionManager),
|
||||||
getCwd: () => ctx.sessionManager.getCwd(),
|
getCwd: () => ctx.sessionManager.getCwd(),
|
||||||
};
|
};
|
||||||
@@ -470,6 +493,25 @@ export default function permissionAiJudge(pi: ExtensionAPI): void {
|
|||||||
"warning",
|
"warning",
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
|
// One non-blocking session notice in Enforce mode: the risk
|
||||||
|
// contract and the effective judge model (ADR 0008). Not repeated
|
||||||
|
// per ask.
|
||||||
|
if (root.config.mode === "enforce") {
|
||||||
|
const configured = root.config.judgeModel;
|
||||||
|
const judgeModelDescription =
|
||||||
|
configured !== undefined
|
||||||
|
? `${configured.provider}/${configured.id} (configured)`
|
||||||
|
: (() => {
|
||||||
|
const sessionModel = ctx.model;
|
||||||
|
return sessionModel === undefined
|
||||||
|
? "the current session model (none resolved yet)"
|
||||||
|
: `${sessionModel.provider}/${sessionModel.id} (current session model)`;
|
||||||
|
})();
|
||||||
|
ctx.ui.notify(
|
||||||
|
`ai-bash-judge Enforce active: ${judgeModelDescription} judges Bash asks; allow skips the dialog — you accept the risk of model misjudgment (ADR 0008). High-risk shapes (irreversible, publish, system, credentials) always ask.`,
|
||||||
|
"info",
|
||||||
|
);
|
||||||
|
}
|
||||||
tryRegister();
|
tryRegister();
|
||||||
});
|
});
|
||||||
|
|
||||||
|
|||||||
@@ -1,13 +1,14 @@
|
|||||||
import type { TelemetryHealth } from "./review";
|
import type { TelemetryHealth } from "./review";
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Enforce truth table (PIEXTENSIO-3 cat.4 / M5).
|
* Enforce truth table (PIEXTENSIO-3 cat.4 / M5; ADR 0008 / PIEXTENSIO-23).
|
||||||
*
|
*
|
||||||
* `allow` authority requires every gate to hold **independently**; any one
|
* `allow` authority requires every gate to hold **independently**; any one
|
||||||
* false forces defer. Since PIEXTENSIO-21 the promotion gates read the
|
* false forces defer. Since ADR 0008 the promotion gates (cohort
|
||||||
* Judge-owned records file (exact-identity qualification, fail-closed);
|
* qualification, owner approval, activation) are no longer authority
|
||||||
* with no records, every mode mechanically defers — verified by the
|
* inputs: hand-written `mode: "enforce"` in config v2 is the user's risk
|
||||||
* truth-table tests toggling each condition.
|
* consent, and the remaining gates are the fail-closed runtime health
|
||||||
|
* checks — verified by the truth-table tests toggling each condition.
|
||||||
*/
|
*/
|
||||||
|
|
||||||
export type EnforceGateState = {
|
export type EnforceGateState = {
|
||||||
@@ -15,12 +16,6 @@ export type EnforceGateState = {
|
|||||||
readonly auditHealthy: boolean;
|
readonly auditHealthy: boolean;
|
||||||
/** Runtime telemetry healthy at decision time. */
|
/** Runtime telemetry healthy at decision time. */
|
||||||
readonly telemetryHealth: TelemetryHealth;
|
readonly telemetryHealth: TelemetryHealth;
|
||||||
/** Qualified passing promotion cohort (PIEXTENSIO-10 floor). */
|
|
||||||
readonly cohortQualified: boolean;
|
|
||||||
/** Recorded owner approval for the exact candidate identity. */
|
|
||||||
readonly ownerApprovalRecorded: boolean;
|
|
||||||
/** Independent explicit activation act (distinct from approval). */
|
|
||||||
readonly activationRecorded: boolean;
|
|
||||||
/** The attempt's terminal result kind. */
|
/** The attempt's terminal result kind. */
|
||||||
readonly resultKind: "judgment" | "preflight_defer" | "infrastructure_failure";
|
readonly resultKind: "judgment" | "preflight_defer" | "infrastructure_failure";
|
||||||
/** The semantic verdict when resultKind is judgment, else null. */
|
/** The semantic verdict when resultKind is judgment, else null. */
|
||||||
@@ -52,15 +47,6 @@ export function evaluateEnforceAuthority(
|
|||||||
if (state.telemetryHealth !== "healthy") {
|
if (state.telemetryHealth !== "healthy") {
|
||||||
return { kind: "defer", blockedBy: `telemetry_${state.telemetryHealth}` };
|
return { kind: "defer", blockedBy: `telemetry_${state.telemetryHealth}` };
|
||||||
}
|
}
|
||||||
if (!state.cohortQualified) {
|
|
||||||
return { kind: "defer", blockedBy: "cohort_not_qualified" };
|
|
||||||
}
|
|
||||||
if (!state.ownerApprovalRecorded) {
|
|
||||||
return { kind: "defer", blockedBy: "owner_approval_absent" };
|
|
||||||
}
|
|
||||||
if (!state.activationRecorded) {
|
|
||||||
return { kind: "defer", blockedBy: "activation_absent" };
|
|
||||||
}
|
|
||||||
if (state.resultKind !== "judgment") {
|
if (state.resultKind !== "judgment") {
|
||||||
return { kind: "defer", blockedBy: `result_${state.resultKind}` };
|
return { kind: "defer", blockedBy: `result_${state.resultKind}` };
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -28,6 +28,7 @@ const MAX_OUTPUT_TOKENS = 4_096;
|
|||||||
export type InfrastructureCode =
|
export type InfrastructureCode =
|
||||||
| "no_model"
|
| "no_model"
|
||||||
| "unsupported_api"
|
| "unsupported_api"
|
||||||
|
| "judge_model_unavailable"
|
||||||
| "timeout"
|
| "timeout"
|
||||||
| "aborted"
|
| "aborted"
|
||||||
| "model_error"
|
| "model_error"
|
||||||
|
|||||||
@@ -1,322 +0,0 @@
|
|||||||
import {
|
|
||||||
closeSync,
|
|
||||||
existsSync,
|
|
||||||
mkdirSync,
|
|
||||||
openSync,
|
|
||||||
readFileSync,
|
|
||||||
writeSync,
|
|
||||||
fsyncSync,
|
|
||||||
} from "node:fs";
|
|
||||||
import { join } from "node:path";
|
|
||||||
import { stripForbiddenKeys } from "./review";
|
|
||||||
|
|
||||||
/**
|
|
||||||
* Judge-owned promotion-gate records (ADR 0006 self-sufficiency; PIEXTENSIO-10
|
|
||||||
* promotion governance; PIEXTENSIO-21 seam).
|
|
||||||
*
|
|
||||||
* The Enforce truth table consumes three gate inputs — cohort qualification,
|
|
||||||
* owner approval, activation — that v0.1 hardcoded to false. This module is
|
|
||||||
* their storage and resolution:
|
|
||||||
*
|
|
||||||
* - **Records file**: append-only JSONL under the agent dir, fsync per
|
|
||||||
* record, same discipline as the audit log. Three record kinds:
|
|
||||||
* `cohort_qualified`, `owner_approval`, `activation`. Owner actions write
|
|
||||||
* records offline (the CLI helper in tools/); the online judge only reads.
|
|
||||||
* - **Exact-identity qualification**: a record qualifies the gate only if
|
|
||||||
* its candidate identity matches the live identity field-for-field. A
|
|
||||||
* record for any other identity is inert. This is the PIEXTENSIO-10 rule
|
|
||||||
* that approval binds to the exact candidate, and it makes promotion
|
|
||||||
* immune to identity drift after the records are written.
|
|
||||||
* - **Fail-closed everywhere**: missing file, unreadable file, malformed
|
|
||||||
* line, or shape-invalid record ⇒ the gate stays false. No record kind
|
|
||||||
* may flip any gate other than its own. Gates are resolved once at
|
|
||||||
* session start (immutable snapshot for the session), mirroring the
|
|
||||||
* reload-only config contract.
|
|
||||||
*
|
|
||||||
* Mode is intentionally *not* part of a candidate identity: `mode` is the
|
|
||||||
* authority knob (shadow/enforce), while identity is what the records
|
|
||||||
* certify. Requiring mode equality would let a shadow cohort qualify an
|
|
||||||
* enforce activation record — or vice versa — which the governance
|
|
||||||
* explicitly separates.
|
|
||||||
*/
|
|
||||||
|
|
||||||
/** The candidate-identity fields a promotion record must match exactly. */
|
|
||||||
export interface CandidateIdentity {
|
|
||||||
readonly judge: string;
|
|
||||||
readonly permissionSystem: string;
|
|
||||||
readonly provider: string;
|
|
||||||
readonly model: string;
|
|
||||||
readonly api: string;
|
|
||||||
readonly promptVersion: string;
|
|
||||||
readonly toolSchemaVersion: string;
|
|
||||||
readonly reviewSchemaVersion: string;
|
|
||||||
readonly timeoutCohort: "default" | number;
|
|
||||||
}
|
|
||||||
|
|
||||||
export type PromotionRecordKind =
|
|
||||||
| "cohort_qualified"
|
|
||||||
| "owner_approval"
|
|
||||||
| "activation";
|
|
||||||
|
|
||||||
export interface PromotionRecord {
|
|
||||||
readonly kind: PromotionRecordKind;
|
|
||||||
readonly candidateIdentity: CandidateIdentity;
|
|
||||||
/** ISO timestamp written by the record author. */
|
|
||||||
readonly recordedAt: string;
|
|
||||||
/** Human-readable basis (cohort id, report reference, approval note). */
|
|
||||||
readonly basis: string;
|
|
||||||
}
|
|
||||||
|
|
||||||
export interface PromotionRecordsSnapshot {
|
|
||||||
/** Shape-valid records parsed from the file (all identities). */
|
|
||||||
readonly records: readonly PromotionRecord[];
|
|
||||||
/** False when the file exists but is unreadable or has malformed lines. */
|
|
||||||
readonly healthy: boolean;
|
|
||||||
readonly diagnostic: string | null;
|
|
||||||
/** Absolute path of the records file (empty string when unset). */
|
|
||||||
readonly path: string;
|
|
||||||
}
|
|
||||||
|
|
||||||
export interface PromotionRecordsDeps {
|
|
||||||
/** User-global agent dir (`~/.pi/agent`). */
|
|
||||||
readonly agentDir: string;
|
|
||||||
/** Injectable reader for tests; defaults to readFileSync(utf-8). */
|
|
||||||
readonly readFile?: (path: string) => string;
|
|
||||||
}
|
|
||||||
|
|
||||||
const RECORDS_DIR_SEGMENTS = ["extensions", "pi-permission-ai-judge"];
|
|
||||||
const RECORDS_FILENAME = "promotion-records.jsonl";
|
|
||||||
|
|
||||||
const IDENTITY_KEYS = [
|
|
||||||
"judge",
|
|
||||||
"permissionSystem",
|
|
||||||
"provider",
|
|
||||||
"model",
|
|
||||||
"api",
|
|
||||||
"promptVersion",
|
|
||||||
"toolSchemaVersion",
|
|
||||||
"reviewSchemaVersion",
|
|
||||||
"timeoutCohort",
|
|
||||||
] as const;
|
|
||||||
|
|
||||||
const STRING_IDENTITY_KEYS = IDENTITY_KEYS.filter(
|
|
||||||
(key) => key !== "timeoutCohort",
|
|
||||||
);
|
|
||||||
|
|
||||||
export function promotionRecordsPath(agentDir: string): string {
|
|
||||||
return join(agentDir, ...RECORDS_DIR_SEGMENTS, RECORDS_FILENAME);
|
|
||||||
}
|
|
||||||
|
|
||||||
function identityMatches(
|
|
||||||
record: CandidateIdentity,
|
|
||||||
live: CandidateIdentity,
|
|
||||||
): boolean {
|
|
||||||
return IDENTITY_KEYS.every((key) => record[key] === live[key]);
|
|
||||||
}
|
|
||||||
|
|
||||||
function isRecordShape(value: unknown): value is PromotionRecord {
|
|
||||||
if (typeof value !== "object" || value === null || Array.isArray(value)) {
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
const record = value as Record<string, unknown>;
|
|
||||||
if (
|
|
||||||
record.kind !== "cohort_qualified" &&
|
|
||||||
record.kind !== "owner_approval" &&
|
|
||||||
record.kind !== "activation"
|
|
||||||
) {
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
if (
|
|
||||||
typeof record.recordedAt !== "string" ||
|
|
||||||
record.recordedAt.length === 0
|
|
||||||
) {
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
if (typeof record.basis !== "string") {
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
const identity = record.candidateIdentity;
|
|
||||||
if (
|
|
||||||
typeof identity !== "object" ||
|
|
||||||
identity === null ||
|
|
||||||
Array.isArray(identity)
|
|
||||||
) {
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
const identityRecord = identity as Record<string, unknown>;
|
|
||||||
// Every field except timeoutCohort is a string; timeoutCohort is
|
|
||||||
// "default" or a positive integer (config-cohort semantics).
|
|
||||||
const timeoutCohort = identityRecord.timeoutCohort;
|
|
||||||
const validTimeoutCohort =
|
|
||||||
timeoutCohort === "default" ||
|
|
||||||
(typeof timeoutCohort === "number" &&
|
|
||||||
Number.isInteger(timeoutCohort) &&
|
|
||||||
timeoutCohort > 0);
|
|
||||||
if (
|
|
||||||
!STRING_IDENTITY_KEYS.every((key) =>
|
|
||||||
typeof identityRecord[key] === "string",
|
|
||||||
) ||
|
|
||||||
!validTimeoutCohort
|
|
||||||
) {
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
return true;
|
|
||||||
}
|
|
||||||
|
|
||||||
/**
|
|
||||||
* Load and parse the records file once per session (immutable snapshot).
|
|
||||||
*
|
|
||||||
* Read-only and side-effect free. Malformed or partial state never raises —
|
|
||||||
* the snapshot reports `healthy: false`, which `resolvePromotionGates`
|
|
||||||
* turns into all-gates-closed. A missing file is the normal pre-promotion
|
|
||||||
* state: healthy with zero records.
|
|
||||||
*/
|
|
||||||
export function loadPromotionRecords(
|
|
||||||
deps: PromotionRecordsDeps,
|
|
||||||
): PromotionRecordsSnapshot {
|
|
||||||
const file = promotionRecordsPath(deps.agentDir);
|
|
||||||
const read = deps.readFile ?? ((p: string) => readFileSync(p, "utf-8"));
|
|
||||||
|
|
||||||
let raw: string;
|
|
||||||
try {
|
|
||||||
raw = read(file);
|
|
||||||
} catch {
|
|
||||||
return {
|
|
||||||
records: [],
|
|
||||||
healthy: true,
|
|
||||||
diagnostic: existsSync(file) ? `records unreadable at ${file}` : null,
|
|
||||||
path: file,
|
|
||||||
};
|
|
||||||
}
|
|
||||||
|
|
||||||
const records: PromotionRecord[] = [];
|
|
||||||
let malformed = false;
|
|
||||||
for (const line of raw.split("\n")) {
|
|
||||||
const trimmed = line.trim();
|
|
||||||
if (trimmed.length === 0) {
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
let parsed: unknown;
|
|
||||||
try {
|
|
||||||
parsed = JSON.parse(trimmed);
|
|
||||||
} catch {
|
|
||||||
malformed = true;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
if (!isRecordShape(parsed)) {
|
|
||||||
malformed = true;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
records.push(parsed);
|
|
||||||
}
|
|
||||||
|
|
||||||
if (malformed) {
|
|
||||||
return {
|
|
||||||
records: [],
|
|
||||||
healthy: false,
|
|
||||||
diagnostic: `malformed records at ${file}`,
|
|
||||||
path: file,
|
|
||||||
};
|
|
||||||
}
|
|
||||||
return { records, healthy: true, diagnostic: null, path: file };
|
|
||||||
}
|
|
||||||
|
|
||||||
/** The three Enforce promotion gates for one live candidate identity. */
|
|
||||||
export interface PromotionGateResolution {
|
|
||||||
readonly cohortQualified: boolean;
|
|
||||||
readonly ownerApprovalRecorded: boolean;
|
|
||||||
readonly activationRecorded: boolean;
|
|
||||||
}
|
|
||||||
|
|
||||||
const ALL_GATES_CLOSED: PromotionGateResolution = {
|
|
||||||
cohortQualified: false,
|
|
||||||
ownerApprovalRecorded: false,
|
|
||||||
activationRecorded: false,
|
|
||||||
};
|
|
||||||
|
|
||||||
/**
|
|
||||||
* Resolve the three Enforce promotion gates for the live candidate
|
|
||||||
* identity against a session-start records snapshot.
|
|
||||||
*
|
|
||||||
* Called per ask: the live identity carries the current model segment, so
|
|
||||||
* a mid-session model switch cannot inherit another segment's records.
|
|
||||||
* Unhealthy snapshot ⇒ all gates closed (fail-closed). A record qualifies
|
|
||||||
* only when its candidate identity matches field-for-field; records for
|
|
||||||
* any other identity are inert.
|
|
||||||
*/
|
|
||||||
export function resolvePromotionGates(
|
|
||||||
snapshot: PromotionRecordsSnapshot,
|
|
||||||
liveIdentity: CandidateIdentity,
|
|
||||||
): PromotionGateResolution {
|
|
||||||
if (!snapshot.healthy) {
|
|
||||||
return ALL_GATES_CLOSED;
|
|
||||||
}
|
|
||||||
let cohortQualified = false;
|
|
||||||
let ownerApprovalRecorded = false;
|
|
||||||
let activationRecorded = false;
|
|
||||||
for (const record of snapshot.records) {
|
|
||||||
if (!identityMatches(record.candidateIdentity, liveIdentity)) {
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
if (record.kind === "cohort_qualified") cohortQualified = true;
|
|
||||||
if (record.kind === "owner_approval") ownerApprovalRecorded = true;
|
|
||||||
if (record.kind === "activation") activationRecorded = true;
|
|
||||||
}
|
|
||||||
return { cohortQualified, ownerApprovalRecorded, activationRecorded };
|
|
||||||
}
|
|
||||||
|
|
||||||
export interface AppendPromotionRecordDeps {
|
|
||||||
/** User-global agent dir (`~/.pi/agent`). */
|
|
||||||
readonly agentDir: string;
|
|
||||||
readonly record: PromotionRecord;
|
|
||||||
/** Injectable clock for tests; defaults to ISO-now. */
|
|
||||||
readonly now?: () => string;
|
|
||||||
}
|
|
||||||
|
|
||||||
/**
|
|
||||||
* Offline owner action: append one promotion record (append + fsync, same
|
|
||||||
* write discipline as the audit log). Never imported by online modules —
|
|
||||||
* the CLI helper in tools/ is the only in-package consumer.
|
|
||||||
*
|
|
||||||
* Returns null on success or a human-readable failure reason.
|
|
||||||
*/
|
|
||||||
export function appendPromotionRecord(
|
|
||||||
deps: AppendPromotionRecordDeps,
|
|
||||||
): string | null {
|
|
||||||
const now = deps.now ?? (() => new Date().toISOString());
|
|
||||||
const dir = join(deps.agentDir, ...RECORDS_DIR_SEGMENTS);
|
|
||||||
const file = join(dir, RECORDS_FILENAME);
|
|
||||||
if (!isRecordShape(deps.record)) {
|
|
||||||
return "record is not shape-valid";
|
|
||||||
}
|
|
||||||
try {
|
|
||||||
mkdirSync(dir, { recursive: true });
|
|
||||||
} catch (error) {
|
|
||||||
return `cannot create ${dir}: ${error instanceof Error ? error.message : String(error)}`;
|
|
||||||
}
|
|
||||||
const record = JSON.stringify({
|
|
||||||
recordedAt: now(),
|
|
||||||
...stripForbiddenKeys({
|
|
||||||
kind: deps.record.kind,
|
|
||||||
candidateIdentity: deps.record.candidateIdentity,
|
|
||||||
basis: deps.record.basis,
|
|
||||||
} as Record<string, unknown>),
|
|
||||||
});
|
|
||||||
let fd: number | undefined;
|
|
||||||
try {
|
|
||||||
fd = openSync(file, "a");
|
|
||||||
writeSync(fd, `${record}\n`);
|
|
||||||
fsyncSync(fd);
|
|
||||||
return null;
|
|
||||||
} catch (error) {
|
|
||||||
return `cannot append to ${file}: ${error instanceof Error ? error.message : String(error)}`;
|
|
||||||
} finally {
|
|
||||||
if (fd !== undefined) {
|
|
||||||
try {
|
|
||||||
closeSync(fd);
|
|
||||||
} catch {
|
|
||||||
// Close failure does not un-fail the append.
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
@@ -20,9 +20,11 @@ describe("loadJudgeConfig — missing and malformed", () => {
|
|||||||
it("resolves a missing file to all defaults with one diagnostic", () => {
|
it("resolves a missing file to all defaults with one diagnostic", () => {
|
||||||
const config = loadJudgeConfig(deps());
|
const config = loadJudgeConfig(deps());
|
||||||
expect(config).toEqual({
|
expect(config).toEqual({
|
||||||
|
configVersion: 1,
|
||||||
mode: "shadow",
|
mode: "shadow",
|
||||||
timeoutMs: 15_000,
|
timeoutMs: 15_000,
|
||||||
timeoutCohort: "default",
|
timeoutCohort: "default",
|
||||||
|
judgeModel: undefined,
|
||||||
diagnostics: [
|
diagnostics: [
|
||||||
expect.objectContaining({ key: "file", fallback: "all defaults" }),
|
expect.objectContaining({ key: "file", fallback: "all defaults" }),
|
||||||
],
|
],
|
||||||
@@ -45,19 +47,76 @@ describe("loadJudgeConfig — missing and malformed", () => {
|
|||||||
});
|
});
|
||||||
});
|
});
|
||||||
|
|
||||||
|
describe("loadJudgeConfig — version selection", () => {
|
||||||
|
it("treats a missing version as v1", () => {
|
||||||
|
const config = loadJudgeConfig(deps({ [CONFIG_PATH]: '{"mode":"shadow"}' }));
|
||||||
|
expect(config.configVersion).toBe(1);
|
||||||
|
});
|
||||||
|
|
||||||
|
it("accepts version 1 explicitly", () => {
|
||||||
|
const config = loadJudgeConfig(
|
||||||
|
deps({ [CONFIG_PATH]: '{"version":1,"mode":"shadow"}' }),
|
||||||
|
);
|
||||||
|
expect(config.configVersion).toBe(1);
|
||||||
|
expect(config.diagnostics).toEqual([]);
|
||||||
|
});
|
||||||
|
|
||||||
|
it("accepts version 2", () => {
|
||||||
|
const config = loadJudgeConfig(
|
||||||
|
deps({ [CONFIG_PATH]: '{"version":2,"mode":"shadow"}' }),
|
||||||
|
);
|
||||||
|
expect(config.configVersion).toBe(2);
|
||||||
|
expect(config.diagnostics).toEqual([]);
|
||||||
|
});
|
||||||
|
|
||||||
|
it("rejects an unknown version to all defaults with a diagnostic", () => {
|
||||||
|
const config = loadJudgeConfig(
|
||||||
|
deps({ [CONFIG_PATH]: '{"version":3,"mode":"enforce"}' }),
|
||||||
|
);
|
||||||
|
expect(config).toMatchObject({ configVersion: 1, mode: "shadow" });
|
||||||
|
expect(config.diagnostics[0]?.key).toBe("version");
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
describe("loadJudgeConfig — v1 enforce migration (fail-closed)", () => {
|
||||||
|
it("downgrades v1 enforce to shadow with a migration diagnostic", () => {
|
||||||
|
for (const raw of ['{"mode":"enforce"}', '{"version":1,"mode":"enforce"}']) {
|
||||||
|
const config = loadJudgeConfig(deps({ [CONFIG_PATH]: raw }));
|
||||||
|
expect(config.mode).toBe("shadow");
|
||||||
|
expect(config.diagnostics).toEqual([
|
||||||
|
{
|
||||||
|
key: "mode",
|
||||||
|
problem: expect.stringMatching(/"version": 2/),
|
||||||
|
fallback: "shadow (v1 enforce requires explicit migration)",
|
||||||
|
},
|
||||||
|
]);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
it("keeps v1 shadow without diagnostics", () => {
|
||||||
|
const config = loadJudgeConfig(deps({ [CONFIG_PATH]: '{"mode":"shadow"}' }));
|
||||||
|
expect(config.mode).toBe("shadow");
|
||||||
|
expect(config.diagnostics).toEqual([]);
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
describe("loadJudgeConfig — mode", () => {
|
describe("loadJudgeConfig — mode", () => {
|
||||||
it("accepts shadow and enforce", () => {
|
it("accepts shadow and enforce in v2", () => {
|
||||||
expect(
|
expect(
|
||||||
loadJudgeConfig(deps({ [CONFIG_PATH]: '{"mode":"shadow"}' })).mode,
|
loadJudgeConfig(
|
||||||
|
deps({ [CONFIG_PATH]: '{"version":2,"mode":"shadow"}' }),
|
||||||
|
).mode,
|
||||||
).toBe("shadow");
|
).toBe("shadow");
|
||||||
expect(
|
expect(
|
||||||
loadJudgeConfig(deps({ [CONFIG_PATH]: '{"mode":"enforce"}' })).mode,
|
loadJudgeConfig(
|
||||||
|
deps({ [CONFIG_PATH]: '{"version":2,"mode":"enforce"}' }),
|
||||||
|
).mode,
|
||||||
).toBe("enforce");
|
).toBe("enforce");
|
||||||
});
|
});
|
||||||
|
|
||||||
it("resolves an unknown mode to shadow with a diagnostic", () => {
|
it("resolves an unknown mode to shadow with a diagnostic", () => {
|
||||||
const config = loadJudgeConfig(
|
const config = loadJudgeConfig(
|
||||||
deps({ [CONFIG_PATH]: '{"mode":"yolo"}' }),
|
deps({ [CONFIG_PATH]: '{"version":2,"mode":"yolo"}' }),
|
||||||
);
|
);
|
||||||
expect(config.mode).toBe("shadow");
|
expect(config.mode).toBe("shadow");
|
||||||
expect(config.diagnostics).toEqual([
|
expect(config.diagnostics).toEqual([
|
||||||
@@ -70,18 +129,72 @@ describe("loadJudgeConfig — mode", () => {
|
|||||||
});
|
});
|
||||||
|
|
||||||
it("resolves a missing mode to shadow without diagnostics", () => {
|
it("resolves a missing mode to shadow without diagnostics", () => {
|
||||||
const config = loadJudgeConfig(deps({ [CONFIG_PATH]: "{}" }));
|
const config = loadJudgeConfig(deps({ [CONFIG_PATH]: '{"version":2}' }));
|
||||||
expect(config.mode).toBe("shadow");
|
expect(config.mode).toBe("shadow");
|
||||||
expect(config.diagnostics).toEqual([]);
|
expect(config.diagnostics).toEqual([]);
|
||||||
});
|
});
|
||||||
});
|
});
|
||||||
|
|
||||||
|
describe("loadJudgeConfig — v2 judge model", () => {
|
||||||
|
it("parses an explicit fixed judge model", () => {
|
||||||
|
const config = loadJudgeConfig(
|
||||||
|
deps({
|
||||||
|
[CONFIG_PATH]:
|
||||||
|
'{"version":2,"mode":"enforce","model":{"provider":"openai-codex","id":"gpt-5.6-sol"}}',
|
||||||
|
}),
|
||||||
|
);
|
||||||
|
expect(config.mode).toBe("enforce");
|
||||||
|
expect(config.judgeModel).toEqual({
|
||||||
|
provider: "openai-codex",
|
||||||
|
id: "gpt-5.6-sol",
|
||||||
|
});
|
||||||
|
expect(config.diagnostics).toEqual([]);
|
||||||
|
});
|
||||||
|
|
||||||
|
it("allows a judge model in shadow mode too", () => {
|
||||||
|
const config = loadJudgeConfig(
|
||||||
|
deps({
|
||||||
|
[CONFIG_PATH]:
|
||||||
|
'{"version":2,"mode":"shadow","model":{"provider":"p","id":"m"}}',
|
||||||
|
}),
|
||||||
|
);
|
||||||
|
expect(config.mode).toBe("shadow");
|
||||||
|
expect(config.judgeModel).toEqual({ provider: "p", id: "m" });
|
||||||
|
});
|
||||||
|
|
||||||
|
it.each([
|
||||||
|
'{"version":2,"model":"openai-codex"}',
|
||||||
|
'{"version":2,"model":{"provider":"p"}}',
|
||||||
|
'{"version":2,"model":{"id":"m"}}',
|
||||||
|
'{"version":2,"model":{"provider":"","id":"m"}}',
|
||||||
|
'{"version":2,"model":{"provider":"p","id":""}}',
|
||||||
|
'{"version":2,"model":{"provider":1,"id":"m"}}',
|
||||||
|
'{"version":2,"model":[]}',
|
||||||
|
])("fails a malformed model %j closed to shadow with a diagnostic", (raw) => {
|
||||||
|
const config = loadJudgeConfig(deps({ [CONFIG_PATH]: raw }));
|
||||||
|
expect(config.mode).toBe("shadow");
|
||||||
|
expect(config.judgeModel).toBeUndefined();
|
||||||
|
expect(config.diagnostics[0]?.key).toBe("model");
|
||||||
|
});
|
||||||
|
|
||||||
|
it("ignores a model field in v1 with a diagnostic", () => {
|
||||||
|
const config = loadJudgeConfig(
|
||||||
|
deps({ [CONFIG_PATH]: '{"model":{"provider":"p","id":"m"}}' }),
|
||||||
|
);
|
||||||
|
expect(config.configVersion).toBe(1);
|
||||||
|
expect(config.judgeModel).toBeUndefined();
|
||||||
|
expect(config.diagnostics[0]?.key).toBe("model");
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
describe("loadJudgeConfig — timeout boundaries", () => {
|
describe("loadJudgeConfig — timeout boundaries", () => {
|
||||||
it.each([4_999, 30_001, 0, -5_000, 15.5, NaN, Infinity, "20000"])(
|
it.each([4_999, 30_001, 0, -5_000, 15.5, NaN, Infinity, "20000"])(
|
||||||
"rejects invalid timeoutMs %p with fallback to 15,000",
|
"rejects invalid timeoutMs %p with fallback to 15,000",
|
||||||
(value) => {
|
(value) => {
|
||||||
const config = loadJudgeConfig(
|
const config = loadJudgeConfig(
|
||||||
deps({ [CONFIG_PATH]: JSON.stringify({ timeoutMs: value }) }),
|
deps({
|
||||||
|
[CONFIG_PATH]: JSON.stringify({ version: 2, timeoutMs: value }),
|
||||||
|
}),
|
||||||
);
|
);
|
||||||
expect(config.timeoutMs).toBe(15_000);
|
expect(config.timeoutMs).toBe(15_000);
|
||||||
expect(config.timeoutCohort).toBe("default");
|
expect(config.timeoutCohort).toBe("default");
|
||||||
@@ -91,16 +204,20 @@ describe("loadJudgeConfig — timeout boundaries", () => {
|
|||||||
|
|
||||||
it("accepts the inclusive boundaries 5,000 and 30,000", () => {
|
it("accepts the inclusive boundaries 5,000 and 30,000", () => {
|
||||||
expect(
|
expect(
|
||||||
loadJudgeConfig(deps({ [CONFIG_PATH]: '{"timeoutMs":5000}' })),
|
loadJudgeConfig(
|
||||||
|
deps({ [CONFIG_PATH]: '{"version":2,"timeoutMs":5000}' }),
|
||||||
|
),
|
||||||
).toMatchObject({ timeoutMs: 5_000, timeoutCohort: 5_000 });
|
).toMatchObject({ timeoutMs: 5_000, timeoutCohort: 5_000 });
|
||||||
expect(
|
expect(
|
||||||
loadJudgeConfig(deps({ [CONFIG_PATH]: '{"timeoutMs":30000}' })),
|
loadJudgeConfig(
|
||||||
|
deps({ [CONFIG_PATH]: '{"version":2,"timeoutMs":30000}' }),
|
||||||
|
),
|
||||||
).toMatchObject({ timeoutMs: 30_000, timeoutCohort: 30_000 });
|
).toMatchObject({ timeoutMs: 30_000, timeoutCohort: 30_000 });
|
||||||
});
|
});
|
||||||
|
|
||||||
it("marks an explicit default timeout as the default cohort", () => {
|
it("marks an explicit default timeout as the default cohort", () => {
|
||||||
const config = loadJudgeConfig(
|
const config = loadJudgeConfig(
|
||||||
deps({ [CONFIG_PATH]: '{"timeoutMs":15000}' }),
|
deps({ [CONFIG_PATH]: '{"version":2,"timeoutMs":15000}' }),
|
||||||
);
|
);
|
||||||
expect(config.timeoutCohort).toBe("default");
|
expect(config.timeoutCohort).toBe("default");
|
||||||
expect(config.diagnostics).toEqual([]);
|
expect(config.diagnostics).toEqual([]);
|
||||||
@@ -108,7 +225,7 @@ describe("loadJudgeConfig — timeout boundaries", () => {
|
|||||||
|
|
||||||
it("marks a non-default timeout as a distinct cohort", () => {
|
it("marks a non-default timeout as a distinct cohort", () => {
|
||||||
const config = loadJudgeConfig(
|
const config = loadJudgeConfig(
|
||||||
deps({ [CONFIG_PATH]: '{"timeoutMs":30000}' }),
|
deps({ [CONFIG_PATH]: '{"version":2,"timeoutMs":30000}' }),
|
||||||
);
|
);
|
||||||
expect(config.timeoutCohort).toBe(30_000);
|
expect(config.timeoutCohort).toBe(30_000);
|
||||||
});
|
});
|
||||||
@@ -117,13 +234,18 @@ describe("loadJudgeConfig — timeout boundaries", () => {
|
|||||||
describe("loadJudgeConfig — snapshot immutability", () => {
|
describe("loadJudgeConfig — snapshot immutability", () => {
|
||||||
it("returns an immutable effective-config snapshot", () => {
|
it("returns an immutable effective-config snapshot", () => {
|
||||||
const config = loadJudgeConfig(
|
const config = loadJudgeConfig(
|
||||||
deps({ [CONFIG_PATH]: '{"mode":"enforce","timeoutMs":20000}' }),
|
deps({
|
||||||
|
[CONFIG_PATH]:
|
||||||
|
'{"version":2,"mode":"enforce","timeoutMs":20000,"model":{"provider":"p","id":"m"}}',
|
||||||
|
}),
|
||||||
);
|
);
|
||||||
expect(Object.isFrozen(config)).toBe(true);
|
expect(Object.isFrozen(config)).toBe(true);
|
||||||
expect(config).toEqual({
|
expect(config).toEqual({
|
||||||
|
configVersion: 2,
|
||||||
mode: "enforce",
|
mode: "enforce",
|
||||||
timeoutMs: 20_000,
|
timeoutMs: 20_000,
|
||||||
timeoutCohort: 20_000,
|
timeoutCohort: 20_000,
|
||||||
|
judgeModel: { provider: "p", id: "m" },
|
||||||
diagnostics: [],
|
diagnostics: [],
|
||||||
});
|
});
|
||||||
});
|
});
|
||||||
|
|||||||
@@ -0,0 +1,208 @@
|
|||||||
|
import { describe, expect, it } from "vitest";
|
||||||
|
import { classifyHighRisk } from "../src/highrisk";
|
||||||
|
|
||||||
|
describe("classifyHighRisk — data_loss", () => {
|
||||||
|
it.each([
|
||||||
|
"git clean -xfd",
|
||||||
|
"git clean -fdx",
|
||||||
|
"git clean -fxd",
|
||||||
|
"git clean -f -x -d",
|
||||||
|
"git clean --force -xd",
|
||||||
|
"git reset --hard",
|
||||||
|
"git reset --hard HEAD~3",
|
||||||
|
"git reset --hard origin/main",
|
||||||
|
"git checkout -- .",
|
||||||
|
"git checkout .",
|
||||||
|
"git restore .",
|
||||||
|
"git restore -- .",
|
||||||
|
"rm -rf ~",
|
||||||
|
"rm -rf /",
|
||||||
|
"rm -rf ~/*",
|
||||||
|
"rm -rf $HOME",
|
||||||
|
"rm -rf $HOME/projects",
|
||||||
|
"rm -fr ~",
|
||||||
|
])("flags %j", (command) => {
|
||||||
|
expect(classifyHighRisk(command)?.category).toBe("data_loss");
|
||||||
|
});
|
||||||
|
|
||||||
|
it.each([
|
||||||
|
"git clean -nxd",
|
||||||
|
"git clean -nd",
|
||||||
|
"git reset --soft HEAD~1",
|
||||||
|
"git checkout main",
|
||||||
|
"git checkout -- file.txt",
|
||||||
|
"git restore file.txt",
|
||||||
|
"rm -rf build/",
|
||||||
|
"rm -rf ./dist",
|
||||||
|
"rm -r build",
|
||||||
|
"rm file.txt",
|
||||||
|
"echo git clean -xfd",
|
||||||
|
"git status",
|
||||||
|
])("does not flag %j", (command) => {
|
||||||
|
expect(classifyHighRisk(command)).toBeUndefined();
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
describe("classifyHighRisk — history_rewrite", () => {
|
||||||
|
it.each([
|
||||||
|
"git push --force",
|
||||||
|
"git push -f",
|
||||||
|
"git push --force origin main",
|
||||||
|
"git push --force-with-lease origin main",
|
||||||
|
"git push origin +main",
|
||||||
|
])("flags %j", (command) => {
|
||||||
|
expect(classifyHighRisk(command)?.category).toBe("history_rewrite");
|
||||||
|
});
|
||||||
|
|
||||||
|
it.each(["git push origin main", "git push", "git push --tags"])(
|
||||||
|
"does not flag %j",
|
||||||
|
(command) => {
|
||||||
|
expect(classifyHighRisk(command)).toBeUndefined();
|
||||||
|
},
|
||||||
|
);
|
||||||
|
});
|
||||||
|
|
||||||
|
describe("classifyHighRisk — publish_deploy", () => {
|
||||||
|
it.each([
|
||||||
|
"npm publish",
|
||||||
|
"npm publish --access public",
|
||||||
|
"pnpm publish",
|
||||||
|
"yarn publish",
|
||||||
|
"cargo publish",
|
||||||
|
"terraform destroy",
|
||||||
|
"terraform destroy -auto-approve",
|
||||||
|
])("flags %j", (command) => {
|
||||||
|
expect(classifyHighRisk(command)?.category).toBe("publish_deploy");
|
||||||
|
});
|
||||||
|
|
||||||
|
it.each(["npm run publish", "npm install", "pnpm test", "terraform plan", "terraform apply"])(
|
||||||
|
"does not flag %j",
|
||||||
|
(command) => {
|
||||||
|
expect(classifyHighRisk(command)).toBeUndefined();
|
||||||
|
},
|
||||||
|
);
|
||||||
|
});
|
||||||
|
|
||||||
|
describe("classifyHighRisk — system_modify", () => {
|
||||||
|
it.each([
|
||||||
|
"sudo rm /etc/hosts",
|
||||||
|
"sudo -i",
|
||||||
|
"sudo apt-get install ripgrep",
|
||||||
|
"mkfs.ext4 /dev/sda1",
|
||||||
|
"mkfs /dev/sdb",
|
||||||
|
"dd if=/dev/zero of=/dev/sda bs=1M",
|
||||||
|
"shutdown now",
|
||||||
|
"shutdown -h now",
|
||||||
|
"reboot",
|
||||||
|
"halt",
|
||||||
|
"poweroff",
|
||||||
|
])("flags %j", (command) => {
|
||||||
|
expect(classifyHighRisk(command)?.category).toBe("system_modify");
|
||||||
|
});
|
||||||
|
|
||||||
|
it.each([
|
||||||
|
"dd if=/dev/zero of=/tmp/img bs=1M count=10",
|
||||||
|
"echo sudo",
|
||||||
|
"cat /etc/hosts",
|
||||||
|
])("does not flag %j", (command) => {
|
||||||
|
expect(classifyHighRisk(command)).toBeUndefined();
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
describe("classifyHighRisk — credential_access", () => {
|
||||||
|
it.each([
|
||||||
|
"cat ~/.ssh/id_rsa",
|
||||||
|
"cat ~/.ssh/id_ed25519",
|
||||||
|
"less ~/.ssh/id_ed25519",
|
||||||
|
"cat ~/.aws/credentials",
|
||||||
|
"head ~/.netrc",
|
||||||
|
"cat ~/.config/gcloud/application_default_credentials.json",
|
||||||
|
"rm ~/.ssh/id_ed25519",
|
||||||
|
"mv ~/.ssh/authorized_keys /tmp",
|
||||||
|
"cp ~/.gnupg/secring.gpg .",
|
||||||
|
"cat $HOME/.ssh/id_rsa",
|
||||||
|
"cat $HOME/.aws/credentials",
|
||||||
|
])("flags %j", (command) => {
|
||||||
|
expect(classifyHighRisk(command)?.category).toBe("credential_access");
|
||||||
|
});
|
||||||
|
|
||||||
|
it.each([
|
||||||
|
"echo x > ~/.aws/credentials",
|
||||||
|
"echo x >> ~/.ssh/authorized_keys",
|
||||||
|
"ssh-keygen foo 2> ~/.aws/credentials",
|
||||||
|
"printf '%s' key > ~/.ssh/id_ed25519",
|
||||||
|
"curl -s url | tee ~/.netrc",
|
||||||
|
])("flags credential replacement %j", (command) => {
|
||||||
|
expect(classifyHighRisk(command)).toMatchObject({
|
||||||
|
category: "credential_access",
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
it.each([
|
||||||
|
"cat ~/.ssh/config",
|
||||||
|
"cat ~/.ssh/known_hosts",
|
||||||
|
"cat package.json",
|
||||||
|
"ls ~/.ssh",
|
||||||
|
"cat ~/.bashrc",
|
||||||
|
"cat id_rsa",
|
||||||
|
"echo x > ~/.bashrc",
|
||||||
|
"curl -s url | tee notes.txt",
|
||||||
|
])("does not flag %j", (command) => {
|
||||||
|
expect(classifyHighRisk(command)).toBeUndefined();
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
describe("classifyHighRisk — compound inputs", () => {
|
||||||
|
it("flags when any unit matches", () => {
|
||||||
|
expect(classifyHighRisk("pnpm test && git clean -xfd")?.category).toBe(
|
||||||
|
"data_loss",
|
||||||
|
);
|
||||||
|
expect(classifyHighRisk("git push --force; echo done")?.category).toBe(
|
||||||
|
"history_rewrite",
|
||||||
|
);
|
||||||
|
expect(classifyHighRisk("npm publish | tee log")?.category).toBe(
|
||||||
|
"publish_deploy",
|
||||||
|
);
|
||||||
|
});
|
||||||
|
|
||||||
|
it("flags single-`&` background compound units", () => {
|
||||||
|
expect(classifyHighRisk("echo ready & npm publish")?.category).toBe(
|
||||||
|
"publish_deploy",
|
||||||
|
);
|
||||||
|
expect(classifyHighRisk("sleep 5 & rm -rf ~")?.category).toBe(
|
||||||
|
"data_loss",
|
||||||
|
);
|
||||||
|
});
|
||||||
|
|
||||||
|
it("does not flag separators inside quotes or escapes", () => {
|
||||||
|
expect(classifyHighRisk("printf 'x; npm publish;'")).toBeUndefined();
|
||||||
|
expect(classifyHighRisk("echo \"git push --force\"")).toBeUndefined();
|
||||||
|
expect(classifyHighRisk("echo git\\;npm\\;publish")).toBeUndefined();
|
||||||
|
expect(classifyHighRisk("echo 'rm -rf ~'")).toBeUndefined();
|
||||||
|
});
|
||||||
|
|
||||||
|
it("treats a quoted command argument as one token, not a command", () => {
|
||||||
|
// git commit -m "junk; sudo rm /" quotes an argument, not a unit.
|
||||||
|
expect(classifyHighRisk("git commit -m 'do not npm publish'")).toBeUndefined();
|
||||||
|
});
|
||||||
|
|
||||||
|
it("flags high-risk units even when another unit quotes text", () => {
|
||||||
|
expect(classifyHighRisk("echo 'a;b' && git push --force")).toMatchObject({
|
||||||
|
category: "history_rewrite",
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
it("returns undefined for all-benign compounds", () => {
|
||||||
|
expect(classifyHighRisk("pnpm test && pnpm check")).toBeUndefined();
|
||||||
|
expect(classifyHighRisk("echo a; echo b | wc -l")).toBeUndefined();
|
||||||
|
expect(classifyHighRisk("sleep 5 & echo done")).toBeUndefined();
|
||||||
|
});
|
||||||
|
|
||||||
|
it("reports the matching rule for audit observability", () => {
|
||||||
|
const match = classifyHighRisk("git clean -xfd");
|
||||||
|
expect(match).toMatchObject({
|
||||||
|
category: "data_loss",
|
||||||
|
rule: expect.stringMatching(/^git_clean/),
|
||||||
|
});
|
||||||
|
});
|
||||||
|
});
|
||||||
@@ -3,18 +3,10 @@ import {
|
|||||||
evaluateEnforceAuthority,
|
evaluateEnforceAuthority,
|
||||||
type EnforceGateState,
|
type EnforceGateState,
|
||||||
} from "../src/judge";
|
} from "../src/judge";
|
||||||
import {
|
|
||||||
loadPromotionRecords,
|
|
||||||
resolvePromotionGates,
|
|
||||||
type CandidateIdentity,
|
|
||||||
} from "../src/promotion";
|
|
||||||
|
|
||||||
const ALL_OPEN: EnforceGateState = {
|
const ALL_OPEN: EnforceGateState = {
|
||||||
auditHealthy: true,
|
auditHealthy: true,
|
||||||
telemetryHealth: "healthy",
|
telemetryHealth: "healthy",
|
||||||
cohortQualified: true,
|
|
||||||
ownerApprovalRecorded: true,
|
|
||||||
activationRecorded: true,
|
|
||||||
resultKind: "judgment",
|
resultKind: "judgment",
|
||||||
verdict: "allow",
|
verdict: "allow",
|
||||||
reviewAcknowledged: true,
|
reviewAcknowledged: true,
|
||||||
@@ -23,7 +15,10 @@ const ALL_OPEN: EnforceGateState = {
|
|||||||
};
|
};
|
||||||
|
|
||||||
describe("evaluateEnforceAuthority — every gate independently forces defer", () => {
|
describe("evaluateEnforceAuthority — every gate independently forces defer", () => {
|
||||||
it("allows only when every gate holds", () => {
|
it("allows when every gate holds — no promotion records required", () => {
|
||||||
|
// ADR 0008: the promotion gates (cohort qualification, owner
|
||||||
|
// approval, activation) are no longer authority inputs; a session
|
||||||
|
// with zero promotion records can hold Enforce authority.
|
||||||
expect(evaluateEnforceAuthority(ALL_OPEN)).toEqual({ kind: "allow" });
|
expect(evaluateEnforceAuthority(ALL_OPEN)).toEqual({ kind: "allow" });
|
||||||
});
|
});
|
||||||
|
|
||||||
@@ -37,9 +32,6 @@ describe("evaluateEnforceAuthority — every gate independently forces defer", (
|
|||||||
{ name: "telemetry disabled", patch: { telemetryHealth: "disabled" }, expectedReason: "telemetry_disabled" },
|
{ name: "telemetry disabled", patch: { telemetryHealth: "disabled" }, expectedReason: "telemetry_disabled" },
|
||||||
{ name: "telemetry write failed", patch: { telemetryHealth: "write_failed" }, expectedReason: "telemetry_write_failed" },
|
{ name: "telemetry write failed", patch: { telemetryHealth: "write_failed" }, expectedReason: "telemetry_write_failed" },
|
||||||
{ name: "telemetry integrity anomaly", patch: { telemetryHealth: "integrity_anomaly" }, expectedReason: "telemetry_integrity_anomaly" },
|
{ name: "telemetry integrity anomaly", patch: { telemetryHealth: "integrity_anomaly" }, expectedReason: "telemetry_integrity_anomaly" },
|
||||||
{ name: "cohort not qualified", patch: { cohortQualified: false }, expectedReason: "cohort_not_qualified" },
|
|
||||||
{ name: "owner approval absent", patch: { ownerApprovalRecorded: false }, expectedReason: "owner_approval_absent" },
|
|
||||||
{ name: "activation absent", patch: { activationRecorded: false }, expectedReason: "activation_absent" },
|
|
||||||
{ name: "preflight result", patch: { resultKind: "preflight_defer" }, expectedReason: "result_preflight_defer" },
|
{ name: "preflight result", patch: { resultKind: "preflight_defer" }, expectedReason: "result_preflight_defer" },
|
||||||
{ name: "infrastructure result", patch: { resultKind: "infrastructure_failure" }, expectedReason: "result_infrastructure_failure" },
|
{ name: "infrastructure result", patch: { resultKind: "infrastructure_failure" }, expectedReason: "result_infrastructure_failure" },
|
||||||
{ name: "semantic deny verdict", patch: { verdict: "deny" }, expectedReason: "verdict_deny" },
|
{ name: "semantic deny verdict", patch: { verdict: "deny" }, expectedReason: "verdict_deny" },
|
||||||
@@ -58,86 +50,3 @@ describe("evaluateEnforceAuthority — every gate independently forces defer", (
|
|||||||
});
|
});
|
||||||
}
|
}
|
||||||
});
|
});
|
||||||
|
|
||||||
describe("evaluateEnforceAuthority — production state with no promotion records", () => {
|
|
||||||
// The real post-PIEXTENSIO-21 seam: an empty records file (the normal
|
|
||||||
// pre-promotion state) closes all promotion gates, so every mode and
|
|
||||||
// telemetry state defers — mechanically identical to v0.1's hardcoded
|
|
||||||
// closure, now derived from actual storage.
|
|
||||||
const emptySnapshot = loadPromotionRecords({
|
|
||||||
agentDir: "/nonexistent-agent-dir",
|
|
||||||
});
|
|
||||||
const identity: CandidateIdentity = {
|
|
||||||
judge: "@sikongjueluo/pi-permission-ai-judge@0.0.1",
|
|
||||||
permissionSystem: "25.4.0",
|
|
||||||
provider: "openai-codex",
|
|
||||||
model: "gpt-5.6-sol",
|
|
||||||
api: "openai-codex-responses",
|
|
||||||
promptVersion: "bash-shadow-v4",
|
|
||||||
toolSchemaVersion: "report-verdict-v1",
|
|
||||||
reviewSchemaVersion: "1",
|
|
||||||
timeoutCohort: 30000,
|
|
||||||
};
|
|
||||||
|
|
||||||
it("never grants authority for any mode or telemetry state without records", () => {
|
|
||||||
const modes = ["shadow", "enforce"] as const;
|
|
||||||
const healths = ["healthy", "disabled", "write_failed", "integrity_anomaly"] as const;
|
|
||||||
for (const mode of modes) {
|
|
||||||
for (const health of healths) {
|
|
||||||
const gates = resolvePromotionGates(emptySnapshot, identity);
|
|
||||||
const outcome = evaluateEnforceAuthority({
|
|
||||||
...ALL_OPEN,
|
|
||||||
mode,
|
|
||||||
telemetryHealth: health,
|
|
||||||
...gates,
|
|
||||||
});
|
|
||||||
expect(outcome.kind).toBe("defer");
|
|
||||||
}
|
|
||||||
}
|
|
||||||
});
|
|
||||||
|
|
||||||
it("blocks enforce on the cohort gate first even with healthy audit", () => {
|
|
||||||
const gates = resolvePromotionGates(emptySnapshot, identity);
|
|
||||||
const outcome = evaluateEnforceAuthority({
|
|
||||||
...ALL_OPEN,
|
|
||||||
...gates,
|
|
||||||
});
|
|
||||||
expect(outcome).toEqual({
|
|
||||||
kind: "defer",
|
|
||||||
blockedBy: "cohort_not_qualified",
|
|
||||||
});
|
|
||||||
});
|
|
||||||
|
|
||||||
it("grants authority only when every record kind exists for the exact identity", () => {
|
|
||||||
const records = [
|
|
||||||
{
|
|
||||||
kind: "cohort_qualified",
|
|
||||||
candidateIdentity: identity,
|
|
||||||
recordedAt: "2026-08-20T12:00:00Z",
|
|
||||||
basis: "cohort test",
|
|
||||||
},
|
|
||||||
{
|
|
||||||
kind: "owner_approval",
|
|
||||||
candidateIdentity: identity,
|
|
||||||
recordedAt: "2026-08-20T12:01:00Z",
|
|
||||||
basis: "approved",
|
|
||||||
},
|
|
||||||
{
|
|
||||||
kind: "activation",
|
|
||||||
candidateIdentity: identity,
|
|
||||||
recordedAt: "2026-08-20T12:02:00Z",
|
|
||||||
basis: "activated",
|
|
||||||
},
|
|
||||||
] as const;
|
|
||||||
const snapshot = {
|
|
||||||
records,
|
|
||||||
healthy: true,
|
|
||||||
diagnostic: null,
|
|
||||||
path: "unused",
|
|
||||||
};
|
|
||||||
const gates = resolvePromotionGates(snapshot, identity);
|
|
||||||
expect(evaluateEnforceAuthority({ ...ALL_OPEN, ...gates })).toEqual({
|
|
||||||
kind: "allow",
|
|
||||||
});
|
|
||||||
});
|
|
||||||
});
|
|
||||||
|
|||||||
@@ -50,7 +50,6 @@ import {
|
|||||||
unpublishPermissionsService,
|
unpublishPermissionsService,
|
||||||
} from "@gotgenes/pi-permission-system";
|
} from "@gotgenes/pi-permission-system";
|
||||||
import extension from "../src/index";
|
import extension from "../src/index";
|
||||||
import { appendPromotionRecord, type CandidateIdentity } from "../src/promotion";
|
|
||||||
import { PROMPT_VERSION, TOOL_SCHEMA_VERSION } from "../src/prompt";
|
import { PROMPT_VERSION, TOOL_SCHEMA_VERSION } from "../src/prompt";
|
||||||
|
|
||||||
function createFakePi(): {
|
function createFakePi(): {
|
||||||
@@ -542,33 +541,39 @@ describe("AI judge lifecycle", () => {
|
|||||||
});
|
});
|
||||||
});
|
});
|
||||||
|
|
||||||
describe("AI judge Enforce authority seam (PIEXTENSIO-21)", () => {
|
describe("AI judge Enforce authority seam (PIEXTENSIO-23, ADR 0008)", () => {
|
||||||
beforeEach(() => {
|
beforeEach(() => {
|
||||||
createMockAgentDir();
|
createMockAgentDir();
|
||||||
writeFileSync(
|
|
||||||
join(mockAgentDir.dir, "pi-permission-ai-judge.config.json"),
|
|
||||||
JSON.stringify({ mode: "enforce" }),
|
|
||||||
);
|
|
||||||
});
|
});
|
||||||
|
|
||||||
/** Records identity matching the lifecycle fake model + static fields. */
|
function writeConfig(config: Record<string, unknown>): void {
|
||||||
function lifecycleIdentity(): CandidateIdentity {
|
writeFileSync(
|
||||||
return {
|
join(mockAgentDir.dir, "pi-permission-ai-judge.config.json"),
|
||||||
judge: "@sikongjueluo/pi-permission-ai-judge@0.0.1",
|
JSON.stringify(config),
|
||||||
permissionSystem: "25.4.0",
|
);
|
||||||
provider: "test-provider",
|
|
||||||
model: "test-model",
|
|
||||||
api: "openai-codex-responses",
|
|
||||||
promptVersion: PROMPT_VERSION,
|
|
||||||
toolSchemaVersion: TOOL_SCHEMA_VERSION,
|
|
||||||
reviewSchemaVersion: "1",
|
|
||||||
timeoutCohort: "default",
|
|
||||||
};
|
|
||||||
}
|
}
|
||||||
|
|
||||||
async function runAsk(): Promise<{
|
interface RunAskOptions {
|
||||||
|
/** Config file contents (written before session start). */
|
||||||
|
config: Record<string, unknown>;
|
||||||
|
/** Command unit for the ask; full command becomes `pnpm test && <unit>`. */
|
||||||
|
command?: string;
|
||||||
|
/** modelRegistry.find result for a configured judge model (when set). */
|
||||||
|
findResult?: Model<any> | undefined;
|
||||||
|
/** Whether the found model has configured auth. */
|
||||||
|
authConfigured?: boolean;
|
||||||
|
/** Overridable model verdict. */
|
||||||
|
response?: AssistantMessage;
|
||||||
|
}
|
||||||
|
|
||||||
|
async function runAsk(
|
||||||
|
options: RunAskOptions,
|
||||||
|
): Promise<{
|
||||||
verdict: { kind: string };
|
verdict: { kind: string };
|
||||||
reviews: Array<{ event: string; details?: Record<string, unknown> }>;
|
reviews: Array<{ event: string; details?: Record<string, unknown> }>;
|
||||||
|
notify: ReturnType<typeof vi.fn>;
|
||||||
|
complete: ReturnType<typeof vi.fn>;
|
||||||
|
find: ReturnType<typeof vi.fn>;
|
||||||
}> {
|
}> {
|
||||||
let authorize: Authorizer["authorize"] | undefined;
|
let authorize: Authorizer["authorize"] | undefined;
|
||||||
const service = {
|
const service = {
|
||||||
@@ -582,28 +587,45 @@ describe("AI judge Enforce authority seam (PIEXTENSIO-21)", () => {
|
|||||||
publishPermissionsService(service);
|
publishPermissionsService(service);
|
||||||
publishedService = service;
|
publishedService = service;
|
||||||
|
|
||||||
const complete = vi.fn(async () => modelResponse());
|
const complete = vi.fn(async () => options.response ?? modelResponse());
|
||||||
|
const find = vi.fn(() => options.findResult);
|
||||||
|
const hasConfiguredAuth = vi.fn(() => options.authConfigured !== false);
|
||||||
|
const notify = vi.fn();
|
||||||
|
const sessionManager = fakeSessionManager();
|
||||||
const ctx = {
|
const ctx = {
|
||||||
hasUI: true,
|
hasUI: true,
|
||||||
sessionManager: fakeSessionManager(),
|
sessionManager,
|
||||||
model: {
|
model: {
|
||||||
id: "test-model",
|
id: "session-model",
|
||||||
provider: "test-provider",
|
provider: "session-provider",
|
||||||
api: "openai-codex-responses",
|
api: "openai-codex-responses",
|
||||||
} as Model<any>,
|
} as Model<any>,
|
||||||
modelRegistry: { complete },
|
modelRegistry: { complete, find, hasConfiguredAuth },
|
||||||
ui: { notify: vi.fn() },
|
ui: { notify },
|
||||||
} as unknown as ExtensionContext;
|
} as unknown as ExtensionContext;
|
||||||
|
|
||||||
const harness = createFakePi();
|
const harness = createFakePi();
|
||||||
extension(harness.pi);
|
extension(harness.pi);
|
||||||
|
writeConfig(options.config);
|
||||||
harness.start(ctx);
|
harness.start(ctx);
|
||||||
harness.ready();
|
harness.ready();
|
||||||
expect(authorize).toBeDefined();
|
expect(authorize).toBeDefined();
|
||||||
|
|
||||||
|
const unit = options.command ?? "git push origin main";
|
||||||
|
const askDetails = ask();
|
||||||
|
askDetails.command = unit;
|
||||||
|
const request = askDetails.payload.request as { value: string };
|
||||||
|
request.value = unit;
|
||||||
|
const evidenceEntry = (
|
||||||
|
askDetails.payload as { evidence: ReadonlyArray<{ label: string; text: string }> }
|
||||||
|
).evidence[0];
|
||||||
|
if (evidenceEntry !== undefined) {
|
||||||
|
evidenceEntry.text = `pnpm test && ${unit}`;
|
||||||
|
}
|
||||||
|
|
||||||
const reviews: Array<{ event: string; details?: Record<string, unknown> }> = [];
|
const reviews: Array<{ event: string; details?: Record<string, unknown> }> = [];
|
||||||
const verdict = await authorize!(
|
const verdict = await authorize!(
|
||||||
ask(),
|
askDetails,
|
||||||
{
|
{
|
||||||
checkPermission: vi.fn(),
|
checkPermission: vi.fn(),
|
||||||
getToolPermission: vi.fn(),
|
getToolPermission: vi.fn(),
|
||||||
@@ -614,45 +636,13 @@ describe("AI judge Enforce authority seam (PIEXTENSIO-21)", () => {
|
|||||||
},
|
},
|
||||||
);
|
);
|
||||||
harness.shutdown();
|
harness.shutdown();
|
||||||
return { verdict, reviews };
|
return { verdict, reviews, notify, complete, find };
|
||||||
}
|
}
|
||||||
|
|
||||||
it("defers in enforce mode when no promotion records exist", async () => {
|
it("grants authority in v2 enforce mode with no promotion records", async () => {
|
||||||
const { verdict, reviews } = await runAsk();
|
const { verdict, reviews } = await runAsk({
|
||||||
expect(verdict).toEqual({ kind: "defer" });
|
config: { version: 2, mode: "enforce" },
|
||||||
expect(reviews).toMatchObject([
|
|
||||||
{
|
|
||||||
event: "ai_bash_judge.result",
|
|
||||||
details: expect.objectContaining({
|
|
||||||
mode: "enforce",
|
|
||||||
verdict: "allow",
|
|
||||||
effectiveVerdict: "defer",
|
|
||||||
authorityBlockedBy: "cohort_not_qualified",
|
|
||||||
}),
|
|
||||||
},
|
|
||||||
]);
|
|
||||||
});
|
});
|
||||||
|
|
||||||
it("grants authority in enforce mode only with all three exact-identity records", async () => {
|
|
||||||
const identity = lifecycleIdentity();
|
|
||||||
for (const [kind, basis] of [
|
|
||||||
["cohort_qualified", "cohort piextensio-test"],
|
|
||||||
["owner_approval", "approved for test"],
|
|
||||||
["activation", "activated for test"],
|
|
||||||
] as const) {
|
|
||||||
expect(
|
|
||||||
appendPromotionRecord({
|
|
||||||
agentDir: mockAgentDir.dir,
|
|
||||||
record: {
|
|
||||||
kind,
|
|
||||||
candidateIdentity: identity,
|
|
||||||
recordedAt: "2026-08-21T10:00:00Z",
|
|
||||||
basis,
|
|
||||||
},
|
|
||||||
}),
|
|
||||||
).toBeNull();
|
|
||||||
}
|
|
||||||
const { verdict, reviews } = await runAsk();
|
|
||||||
expect(verdict).toEqual({ kind: "allow" });
|
expect(verdict).toEqual({ kind: "allow" });
|
||||||
expect(reviews).toMatchObject([
|
expect(reviews).toMatchObject([
|
||||||
{
|
{
|
||||||
@@ -662,73 +652,227 @@ describe("AI judge Enforce authority seam (PIEXTENSIO-21)", () => {
|
|||||||
verdict: "allow",
|
verdict: "allow",
|
||||||
effectiveVerdict: "allow",
|
effectiveVerdict: "allow",
|
||||||
authorityBlockedBy: null,
|
authorityBlockedBy: null,
|
||||||
|
modelSource: "session",
|
||||||
}),
|
}),
|
||||||
},
|
},
|
||||||
]);
|
]);
|
||||||
});
|
});
|
||||||
|
|
||||||
it("defers in enforce mode when records exist for another identity", async () => {
|
it("downgrades v1 enforce to shadow with a migration diagnostic notification", async () => {
|
||||||
const identity = { ...lifecycleIdentity(), model: "other-model" };
|
const { verdict, reviews, notify, complete } = await runAsk({
|
||||||
for (const kind of [
|
config: { mode: "enforce" },
|
||||||
"cohort_qualified",
|
|
||||||
"owner_approval",
|
|
||||||
"activation",
|
|
||||||
] as const) {
|
|
||||||
appendPromotionRecord({
|
|
||||||
agentDir: mockAgentDir.dir,
|
|
||||||
record: {
|
|
||||||
kind,
|
|
||||||
candidateIdentity: identity,
|
|
||||||
recordedAt: "2026-08-21T10:00:00Z",
|
|
||||||
basis: "other identity",
|
|
||||||
},
|
|
||||||
});
|
});
|
||||||
}
|
|
||||||
const { verdict, reviews } = await runAsk();
|
|
||||||
expect(verdict).toEqual({ kind: "defer" });
|
|
||||||
expect(reviews).toMatchObject([
|
|
||||||
{
|
|
||||||
event: "ai_bash_judge.result",
|
|
||||||
details: expect.objectContaining({
|
|
||||||
effectiveVerdict: "defer",
|
|
||||||
authorityBlockedBy: "cohort_not_qualified",
|
|
||||||
}),
|
|
||||||
},
|
|
||||||
]);
|
|
||||||
});
|
|
||||||
|
|
||||||
it("never grants authority in shadow mode regardless of records", async () => {
|
|
||||||
writeFileSync(
|
|
||||||
join(mockAgentDir.dir, "pi-permission-ai-judge.config.json"),
|
|
||||||
JSON.stringify({ mode: "shadow" }),
|
|
||||||
);
|
|
||||||
const identity = lifecycleIdentity();
|
|
||||||
for (const kind of [
|
|
||||||
"cohort_qualified",
|
|
||||||
"owner_approval",
|
|
||||||
"activation",
|
|
||||||
] as const) {
|
|
||||||
appendPromotionRecord({
|
|
||||||
agentDir: mockAgentDir.dir,
|
|
||||||
record: {
|
|
||||||
kind,
|
|
||||||
candidateIdentity: identity,
|
|
||||||
recordedAt: "2026-08-21T10:00:00Z",
|
|
||||||
basis: "shadow still defers",
|
|
||||||
},
|
|
||||||
});
|
|
||||||
}
|
|
||||||
const { verdict, reviews } = await runAsk();
|
|
||||||
expect(verdict).toEqual({ kind: "defer" });
|
expect(verdict).toEqual({ kind: "defer" });
|
||||||
expect(reviews).toMatchObject([
|
expect(reviews).toMatchObject([
|
||||||
{
|
{
|
||||||
event: "ai_bash_judge.result",
|
event: "ai_bash_judge.result",
|
||||||
details: expect.objectContaining({
|
details: expect.objectContaining({
|
||||||
mode: "shadow",
|
mode: "shadow",
|
||||||
|
verdict: "allow",
|
||||||
effectiveVerdict: "defer",
|
effectiveVerdict: "defer",
|
||||||
authorityBlockedBy: "mode_shadow",
|
authorityBlockedBy: "mode_shadow",
|
||||||
}),
|
}),
|
||||||
},
|
},
|
||||||
]);
|
]);
|
||||||
|
const notified = notify.mock.calls.map((call) => String(call[0]));
|
||||||
|
expect(
|
||||||
|
notified.some((message) =>
|
||||||
|
/v1 enforce requires explicit migration|version 2/.test(message),
|
||||||
|
),
|
||||||
|
).toBe(true);
|
||||||
|
expect(complete).toHaveBeenCalledTimes(1);
|
||||||
|
});
|
||||||
|
|
||||||
|
it("defers immediately on a high-risk command in enforce mode without calling the model", async () => {
|
||||||
|
const { verdict, reviews, complete } = await runAsk({
|
||||||
|
config: { version: 2, mode: "enforce" },
|
||||||
|
command: "git clean -xfd",
|
||||||
|
});
|
||||||
|
expect(verdict).toEqual({ kind: "defer" });
|
||||||
|
expect(complete).not.toHaveBeenCalled();
|
||||||
|
expect(reviews).toMatchObject([
|
||||||
|
{
|
||||||
|
event: "ai_bash_judge.result",
|
||||||
|
details: expect.objectContaining({
|
||||||
|
resultKind: "preflight_defer",
|
||||||
|
verdict: null,
|
||||||
|
effectiveVerdict: "defer",
|
||||||
|
modelCalled: false,
|
||||||
|
code: "high_risk_override",
|
||||||
|
riskCategory: "data_loss",
|
||||||
|
riskRule: expect.any(String),
|
||||||
|
}),
|
||||||
|
},
|
||||||
|
]);
|
||||||
|
});
|
||||||
|
|
||||||
|
it("still calls the model on a high-risk command in shadow mode, records the override, and defers", async () => {
|
||||||
|
const { verdict, reviews, complete } = await runAsk({
|
||||||
|
config: { version: 2, mode: "shadow" },
|
||||||
|
command: "git clean -xfd",
|
||||||
|
});
|
||||||
|
expect(complete).toHaveBeenCalledTimes(1);
|
||||||
|
expect(verdict).toEqual({ kind: "defer" });
|
||||||
|
expect(reviews).toMatchObject([
|
||||||
|
{
|
||||||
|
event: "ai_bash_judge.result",
|
||||||
|
details: expect.objectContaining({
|
||||||
|
resultKind: "judgment",
|
||||||
|
verdict: "allow",
|
||||||
|
effectiveVerdict: "defer",
|
||||||
|
riskOverride: { category: "data_loss", rule: expect.any(String) },
|
||||||
|
}),
|
||||||
|
},
|
||||||
|
]);
|
||||||
|
});
|
||||||
|
|
||||||
|
it("resolves a configured v2 judge model instead of the session model", async () => {
|
||||||
|
const { verdict, reviews, complete, find } = await runAsk({
|
||||||
|
config: {
|
||||||
|
version: 2,
|
||||||
|
mode: "enforce",
|
||||||
|
model: { provider: "fixed-provider", id: "fixed-model" },
|
||||||
|
},
|
||||||
|
findResult: {
|
||||||
|
id: "fixed-model",
|
||||||
|
provider: "fixed-provider",
|
||||||
|
api: "openai-codex-responses",
|
||||||
|
} as Model<any>,
|
||||||
|
});
|
||||||
|
expect(find).toHaveBeenCalledWith("fixed-provider", "fixed-model");
|
||||||
|
expect(complete).toHaveBeenCalledTimes(1);
|
||||||
|
expect((complete.mock.calls[0] as unknown[])[0]).toMatchObject({
|
||||||
|
id: "fixed-model",
|
||||||
|
provider: "fixed-provider",
|
||||||
|
});
|
||||||
|
expect(verdict).toEqual({ kind: "allow" });
|
||||||
|
expect(reviews[0]?.details).toMatchObject({
|
||||||
|
modelSource: "configured",
|
||||||
|
provider: "fixed-provider",
|
||||||
|
model: "fixed-model",
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
it("defers with judge_model_unavailable when a configured model cannot be resolved, without fallback", async () => {
|
||||||
|
const { verdict, reviews, complete } = await runAsk({
|
||||||
|
config: {
|
||||||
|
version: 2,
|
||||||
|
mode: "enforce",
|
||||||
|
model: { provider: "ghost-provider", id: "ghost-model" },
|
||||||
|
},
|
||||||
|
findResult: undefined,
|
||||||
|
});
|
||||||
|
expect(verdict).toEqual({ kind: "defer" });
|
||||||
|
expect(complete).not.toHaveBeenCalled();
|
||||||
|
expect(reviews).toMatchObject([
|
||||||
|
{
|
||||||
|
event: "ai_bash_judge.result",
|
||||||
|
details: expect.objectContaining({
|
||||||
|
resultKind: "infrastructure_failure",
|
||||||
|
code: "judge_model_unavailable",
|
||||||
|
modelCalled: false,
|
||||||
|
provider: "ghost-provider",
|
||||||
|
model: "ghost-model",
|
||||||
|
}),
|
||||||
|
},
|
||||||
|
]);
|
||||||
|
});
|
||||||
|
|
||||||
|
it("keeps riskOverride on the early judge_model_unavailable row for a high-risk shadow ask", async () => {
|
||||||
|
const { reviews } = await runAsk({
|
||||||
|
config: {
|
||||||
|
version: 2,
|
||||||
|
mode: "shadow",
|
||||||
|
model: { provider: "ghost-provider", id: "ghost-model" },
|
||||||
|
},
|
||||||
|
command: "git clean -xfd",
|
||||||
|
findResult: undefined,
|
||||||
|
});
|
||||||
|
expect(reviews[0]?.details).toMatchObject({
|
||||||
|
code: "judge_model_unavailable",
|
||||||
|
riskOverride: { category: "data_loss", rule: expect.any(String) },
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
it("defers with judge_model_unavailable when a configured model lacks auth", async () => {
|
||||||
|
const { verdict, reviews, complete } = await runAsk({
|
||||||
|
config: {
|
||||||
|
version: 2,
|
||||||
|
mode: "enforce",
|
||||||
|
model: { provider: "p", id: "m" },
|
||||||
|
},
|
||||||
|
findResult: { id: "m", provider: "p", api: "openai-codex-responses" } as Model<any>,
|
||||||
|
authConfigured: false,
|
||||||
|
});
|
||||||
|
expect(verdict).toEqual({ kind: "defer" });
|
||||||
|
expect(complete).not.toHaveBeenCalled();
|
||||||
|
expect(reviews[0]?.details).toMatchObject({
|
||||||
|
resultKind: "infrastructure_failure",
|
||||||
|
code: "judge_model_unavailable",
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
it("notifies once per session in v2 enforce mode with the judge model and risk contract", async () => {
|
||||||
|
let authorize: Authorizer["authorize"] | undefined;
|
||||||
|
const service = {
|
||||||
|
registerAuthorizer: vi.fn((_name, callback) => {
|
||||||
|
authorize = callback;
|
||||||
|
return vi.fn();
|
||||||
|
}),
|
||||||
|
checkPermission: vi.fn(),
|
||||||
|
getToolPermission: vi.fn(),
|
||||||
|
} as unknown as PermissionsService;
|
||||||
|
publishPermissionsService(service);
|
||||||
|
publishedService = service;
|
||||||
|
|
||||||
|
const complete = vi.fn(async () => modelResponse());
|
||||||
|
const notify = vi.fn();
|
||||||
|
const ctx = {
|
||||||
|
hasUI: true,
|
||||||
|
sessionManager: fakeSessionManager(),
|
||||||
|
model: {
|
||||||
|
id: "session-model",
|
||||||
|
provider: "session-provider",
|
||||||
|
api: "openai-codex-responses",
|
||||||
|
} as Model<any>,
|
||||||
|
modelRegistry: { complete, find: vi.fn(), hasConfiguredAuth: vi.fn(() => true) },
|
||||||
|
ui: { notify },
|
||||||
|
} as unknown as ExtensionContext;
|
||||||
|
|
||||||
|
const harness = createFakePi();
|
||||||
|
extension(harness.pi);
|
||||||
|
writeConfig({ version: 2, mode: "enforce" });
|
||||||
|
harness.start(ctx);
|
||||||
|
harness.ready();
|
||||||
|
|
||||||
|
const enforceNotice = notify.mock.calls.filter((call) =>
|
||||||
|
/Enforce/i.test(String(call[0])),
|
||||||
|
);
|
||||||
|
expect(enforceNotice).toHaveLength(1);
|
||||||
|
expect(String(enforceNotice[0]?.[0])).toContain(
|
||||||
|
"session-provider/session-model",
|
||||||
|
);
|
||||||
|
expect(String(enforceNotice[0]?.[0])).toMatch(/risk/i);
|
||||||
|
|
||||||
|
// Two more asks must not repeat the notification.
|
||||||
|
for (let i = 0; i < 2; i += 1) {
|
||||||
|
await authorize!(ask(), { checkPermission: vi.fn(), getToolPermission: vi.fn() }, {
|
||||||
|
review: vi.fn(),
|
||||||
|
debug: vi.fn(),
|
||||||
|
});
|
||||||
|
}
|
||||||
|
expect(
|
||||||
|
notify.mock.calls.filter((call) => /Enforce/i.test(String(call[0]))),
|
||||||
|
).toHaveLength(1);
|
||||||
|
harness.shutdown();
|
||||||
|
});
|
||||||
|
|
||||||
|
it("does not show the enforce notification in shadow mode", async () => {
|
||||||
|
const { notify } = await runAsk({
|
||||||
|
config: { version: 2, mode: "shadow" },
|
||||||
|
});
|
||||||
|
expect(
|
||||||
|
notify.mock.calls.filter((call) => /Enforce/i.test(String(call[0]))),
|
||||||
|
).toHaveLength(0);
|
||||||
});
|
});
|
||||||
});
|
});
|
||||||
|
|||||||
@@ -1,260 +0,0 @@
|
|||||||
import { mkdtempSync, mkdirSync, readFileSync, rmSync, writeFileSync } from "node:fs";
|
|
||||||
import { tmpdir } from "node:os";
|
|
||||||
import { dirname, join } from "node:path";
|
|
||||||
import { afterEach, beforeEach, describe, expect, it } from "vitest";
|
|
||||||
import {
|
|
||||||
appendPromotionRecord,
|
|
||||||
loadPromotionRecords,
|
|
||||||
promotionRecordsPath,
|
|
||||||
resolvePromotionGates,
|
|
||||||
type CandidateIdentity,
|
|
||||||
type PromotionRecord,
|
|
||||||
} from "../src/promotion";
|
|
||||||
|
|
||||||
const IDENTITY: CandidateIdentity = {
|
|
||||||
judge: "@sikongjueluo/pi-permission-ai-judge@0.0.1",
|
|
||||||
permissionSystem: "25.4.0",
|
|
||||||
provider: "openai-codex",
|
|
||||||
model: "gpt-5.6-sol",
|
|
||||||
api: "openai-codex-responses",
|
|
||||||
promptVersion: "bash-shadow-v4",
|
|
||||||
toolSchemaVersion: "report-verdict-v1",
|
|
||||||
reviewSchemaVersion: "1",
|
|
||||||
timeoutCohort: 30000,
|
|
||||||
};
|
|
||||||
|
|
||||||
function record(
|
|
||||||
kind: PromotionRecord["kind"],
|
|
||||||
identity: CandidateIdentity = IDENTITY,
|
|
||||||
basis = "test basis",
|
|
||||||
): PromotionRecord {
|
|
||||||
return {
|
|
||||||
kind,
|
|
||||||
candidateIdentity: identity,
|
|
||||||
recordedAt: "2026-08-20T12:00:00Z",
|
|
||||||
basis,
|
|
||||||
};
|
|
||||||
}
|
|
||||||
|
|
||||||
function line(value: unknown): string {
|
|
||||||
return `${JSON.stringify(value)}\n`;
|
|
||||||
}
|
|
||||||
|
|
||||||
/** writeFileSync, but creating the records directory first. */
|
|
||||||
function writeRecords(dir: string, content: string): void {
|
|
||||||
const path = promotionRecordsPath(dir);
|
|
||||||
mkdirSync(dirname(path), { recursive: true });
|
|
||||||
writeFileSync(path, content);
|
|
||||||
}
|
|
||||||
|
|
||||||
describe("promotion records — loadPromotionRecords", () => {
|
|
||||||
let dir: string;
|
|
||||||
beforeEach(() => {
|
|
||||||
dir = mkdtempSync(join(tmpdir(), "ai-judge-promotion-"));
|
|
||||||
});
|
|
||||||
afterEach(() => {
|
|
||||||
rmSync(dir, { recursive: true, force: true });
|
|
||||||
});
|
|
||||||
|
|
||||||
it("treats a missing file as the healthy pre-promotion state", () => {
|
|
||||||
const snapshot = loadPromotionRecords({ agentDir: dir });
|
|
||||||
expect(snapshot).toEqual({
|
|
||||||
records: [],
|
|
||||||
healthy: true,
|
|
||||||
diagnostic: null,
|
|
||||||
path: promotionRecordsPath(dir),
|
|
||||||
});
|
|
||||||
});
|
|
||||||
|
|
||||||
it("parses well-formed records of every kind", () => {
|
|
||||||
writeRecords(
|
|
||||||
dir,
|
|
||||||
[record("cohort_qualified"), record("owner_approval"), record("activation")]
|
|
||||||
.map(line)
|
|
||||||
.join(""),
|
|
||||||
);
|
|
||||||
const snapshot = loadPromotionRecords({ agentDir: dir });
|
|
||||||
expect(snapshot.healthy).toBe(true);
|
|
||||||
expect(snapshot.records).toHaveLength(3);
|
|
||||||
expect(snapshot.records.map((r) => r.kind)).toEqual([
|
|
||||||
"cohort_qualified",
|
|
||||||
"owner_approval",
|
|
||||||
"activation",
|
|
||||||
]);
|
|
||||||
});
|
|
||||||
|
|
||||||
it("fails closed on malformed JSON lines", () => {
|
|
||||||
writeRecords(
|
|
||||||
dir,
|
|
||||||
line(record("cohort_qualified")) + "{not json\n",
|
|
||||||
);
|
|
||||||
const snapshot = loadPromotionRecords({ agentDir: dir });
|
|
||||||
expect(snapshot.healthy).toBe(false);
|
|
||||||
expect(snapshot.records).toEqual([]);
|
|
||||||
expect(snapshot.diagnostic).toContain("malformed");
|
|
||||||
});
|
|
||||||
|
|
||||||
it("fails closed on shape-invalid records", () => {
|
|
||||||
writeRecords(
|
|
||||||
dir,
|
|
||||||
line({ kind: "activation" }), // missing identity/basis/recordedAt
|
|
||||||
);
|
|
||||||
const snapshot = loadPromotionRecords({ agentDir: dir });
|
|
||||||
expect(snapshot.healthy).toBe(false);
|
|
||||||
expect(snapshot.diagnostic).toContain("malformed");
|
|
||||||
});
|
|
||||||
|
|
||||||
it("skips blank lines without failing", () => {
|
|
||||||
writeRecords(
|
|
||||||
dir,
|
|
||||||
"\n" + line(record("activation")) + "\n\n",
|
|
||||||
);
|
|
||||||
const snapshot = loadPromotionRecords({ agentDir: dir });
|
|
||||||
expect(snapshot.healthy).toBe(true);
|
|
||||||
expect(snapshot.records).toHaveLength(1);
|
|
||||||
});
|
|
||||||
});
|
|
||||||
|
|
||||||
describe("promotion records — resolvePromotionGates", () => {
|
|
||||||
let dir: string;
|
|
||||||
beforeEach(() => {
|
|
||||||
dir = mkdtempSync(join(tmpdir(), "ai-judge-promotion-"));
|
|
||||||
});
|
|
||||||
afterEach(() => {
|
|
||||||
rmSync(dir, { recursive: true, force: true });
|
|
||||||
});
|
|
||||||
|
|
||||||
it("flips only its own gate per record kind", () => {
|
|
||||||
writeRecords(dir, line(record("owner_approval")));
|
|
||||||
const snapshot = loadPromotionRecords({ agentDir: dir });
|
|
||||||
expect(resolvePromotionGates(snapshot, IDENTITY)).toEqual({
|
|
||||||
cohortQualified: false,
|
|
||||||
ownerApprovalRecorded: true,
|
|
||||||
activationRecorded: false,
|
|
||||||
});
|
|
||||||
});
|
|
||||||
|
|
||||||
it("all three gates open with all three records for the exact identity", () => {
|
|
||||||
writeRecords(
|
|
||||||
dir,
|
|
||||||
[
|
|
||||||
record("cohort_qualified", IDENTITY, "cohort id x"),
|
|
||||||
record("owner_approval", IDENTITY, "approved"),
|
|
||||||
record("activation", IDENTITY, "activated"),
|
|
||||||
]
|
|
||||||
.map(line)
|
|
||||||
.join(""),
|
|
||||||
);
|
|
||||||
const snapshot = loadPromotionRecords({ agentDir: dir });
|
|
||||||
expect(resolvePromotionGates(snapshot, IDENTITY)).toEqual({
|
|
||||||
cohortQualified: true,
|
|
||||||
ownerApprovalRecorded: true,
|
|
||||||
activationRecorded: true,
|
|
||||||
});
|
|
||||||
});
|
|
||||||
|
|
||||||
const identityDrifts: ReadonlyArray<[string, Partial<CandidateIdentity>]> = [
|
|
||||||
["provider", { provider: "other-provider" }],
|
|
||||||
["model", { model: "gpt-5.7" }],
|
|
||||||
["api", { api: "openai-responses" }],
|
|
||||||
["promptVersion", { promptVersion: "bash-shadow-v5" }],
|
|
||||||
["toolSchemaVersion", { toolSchemaVersion: "report-verdict-v2" }],
|
|
||||||
["reviewSchemaVersion", { reviewSchemaVersion: "2" }],
|
|
||||||
["timeoutCohort", { timeoutCohort: "default" }],
|
|
||||||
["permissionSystem", { permissionSystem: "25.5.0" }],
|
|
||||||
["judge package", { judge: "@sikongjueluo/pi-permission-ai-judge@0.0.2" }],
|
|
||||||
];
|
|
||||||
for (const [name, patch] of identityDrifts) {
|
|
||||||
it(`keeps every gate closed when the live identity drifts on ${name}`, () => {
|
|
||||||
writeRecords(
|
|
||||||
dir,
|
|
||||||
[
|
|
||||||
record("cohort_qualified"),
|
|
||||||
record("owner_approval"),
|
|
||||||
record("activation"),
|
|
||||||
]
|
|
||||||
.map(line)
|
|
||||||
.join(""),
|
|
||||||
);
|
|
||||||
const snapshot = loadPromotionRecords({ agentDir: dir });
|
|
||||||
expect(
|
|
||||||
resolvePromotionGates(snapshot, { ...IDENTITY, ...patch }),
|
|
||||||
).toEqual({
|
|
||||||
cohortQualified: false,
|
|
||||||
ownerApprovalRecorded: false,
|
|
||||||
activationRecorded: false,
|
|
||||||
});
|
|
||||||
});
|
|
||||||
}
|
|
||||||
|
|
||||||
it("closes every gate when the snapshot is unhealthy", () => {
|
|
||||||
writeRecords(dir, "garbage\n");
|
|
||||||
const snapshot = loadPromotionRecords({ agentDir: dir });
|
|
||||||
expect(snapshot.healthy).toBe(false);
|
|
||||||
expect(resolvePromotionGates(snapshot, IDENTITY)).toEqual({
|
|
||||||
cohortQualified: false,
|
|
||||||
ownerApprovalRecorded: false,
|
|
||||||
activationRecorded: false,
|
|
||||||
});
|
|
||||||
});
|
|
||||||
|
|
||||||
it("leaves other-identity records inert, not malformed", () => {
|
|
||||||
const other: CandidateIdentity = {
|
|
||||||
...IDENTITY,
|
|
||||||
model: "glm-5.2",
|
|
||||||
};
|
|
||||||
writeRecords(
|
|
||||||
dir,
|
|
||||||
[
|
|
||||||
record("cohort_qualified", other),
|
|
||||||
record("activation", other),
|
|
||||||
]
|
|
||||||
.map(line)
|
|
||||||
.join(""),
|
|
||||||
);
|
|
||||||
const snapshot = loadPromotionRecords({ agentDir: dir });
|
|
||||||
expect(snapshot.healthy).toBe(true);
|
|
||||||
expect(resolvePromotionGates(snapshot, IDENTITY)).toEqual({
|
|
||||||
cohortQualified: false,
|
|
||||||
ownerApprovalRecorded: false,
|
|
||||||
activationRecorded: false,
|
|
||||||
});
|
|
||||||
});
|
|
||||||
});
|
|
||||||
|
|
||||||
describe("promotion records — appendPromotionRecord", () => {
|
|
||||||
let dir: string;
|
|
||||||
beforeEach(() => {
|
|
||||||
dir = mkdtempSync(join(tmpdir(), "ai-judge-promotion-"));
|
|
||||||
});
|
|
||||||
afterEach(() => {
|
|
||||||
rmSync(dir, { recursive: true, force: true });
|
|
||||||
});
|
|
||||||
|
|
||||||
it("appends a shape-valid record that round-trips through the loader", () => {
|
|
||||||
const error = appendPromotionRecord({
|
|
||||||
agentDir: dir,
|
|
||||||
record: record("owner_approval", IDENTITY, "approved v4"),
|
|
||||||
now: () => "2026-08-21T09:00:00Z",
|
|
||||||
});
|
|
||||||
expect(error).toBeNull();
|
|
||||||
const raw = readFileSync(promotionRecordsPath(dir), "utf-8");
|
|
||||||
expect(raw).toContain('"recordedAt":"2026-08-21T09:00:00Z"');
|
|
||||||
expect(raw).toContain('"basis":"approved v4"');
|
|
||||||
const snapshot = loadPromotionRecords({ agentDir: dir });
|
|
||||||
expect(snapshot.healthy).toBe(true);
|
|
||||||
expect(resolvePromotionGates(snapshot, IDENTITY).ownerApprovalRecorded).toBe(true);
|
|
||||||
});
|
|
||||||
|
|
||||||
it("rejects a shape-invalid record without touching the file", () => {
|
|
||||||
const bad = {
|
|
||||||
kind: "activation",
|
|
||||||
candidateIdentity: { judge: "x" },
|
|
||||||
recordedAt: "",
|
|
||||||
basis: "",
|
|
||||||
} as unknown as PromotionRecord;
|
|
||||||
const error = appendPromotionRecord({ agentDir: dir, record: bad });
|
|
||||||
expect(error).toBe("record is not shape-valid");
|
|
||||||
expect(loadPromotionRecords({ agentDir: dir }).records).toEqual([]);
|
|
||||||
});
|
|
||||||
});
|
|
||||||
@@ -1,167 +0,0 @@
|
|||||||
/**
|
|
||||||
* PIEXTENSIO-21 owner action CLI (offline, package-local).
|
|
||||||
*
|
|
||||||
* Appends one promotion record to the Judge-owned records file
|
|
||||||
* (~/.pi/agent/extensions/pi-permission-ai-judge/promotion-records.jsonl).
|
|
||||||
* Never imported by online modules; no npm bin. The three record kinds map
|
|
||||||
* one-to-one to the Enforce truth-table promotion gates:
|
|
||||||
*
|
|
||||||
* cohort_qualified — after a declared replacement cohort meets the frozen
|
|
||||||
* floor, citing the cohort id + report reference
|
|
||||||
* owner_approval — the owner's explicit approval of that exact candidate
|
|
||||||
* activation — the distinct, explicit activation act
|
|
||||||
*
|
|
||||||
* Fail-closed contract: a record qualifies only its exact candidate
|
|
||||||
* identity; the judge re-derives identity from the live runtime, so a
|
|
||||||
* mismatched or malformed record is inert and malformed files close all
|
|
||||||
* gates. There is no CLI to *remove* authority records by design —
|
|
||||||
* rollback to Shadow is the config `mode` switch, and the append-only
|
|
||||||
* trail keeps the promotion history auditable.
|
|
||||||
*
|
|
||||||
* Usage:
|
|
||||||
* npx tsx tools/promotion-record.ts --kind cohort_qualified \
|
|
||||||
* --provider openai-codex --model gpt-5.6-sol --api openai-codex-responses \
|
|
||||||
* [--timeout-cohort 30000] --basis "cohort <id>; report docs/testing/..."
|
|
||||||
*
|
|
||||||
* judge/permission-system/prompt/tool-schema/review-schema versions are
|
|
||||||
* taken from the live package (src/prompt.ts + package constants) so the
|
|
||||||
* recorded identity cannot drift from the code that will check it.
|
|
||||||
*/
|
|
||||||
|
|
||||||
import { getAgentDir } from "@earendil-works/pi-coding-agent";
|
|
||||||
import {
|
|
||||||
appendPromotionRecord,
|
|
||||||
promotionRecordsPath,
|
|
||||||
type CandidateIdentity,
|
|
||||||
type PromotionRecordKind,
|
|
||||||
} from "../src/promotion";
|
|
||||||
import { PROMPT_VERSION, TOOL_SCHEMA_VERSION } from "../src/prompt";
|
|
||||||
|
|
||||||
const JUDGE_IDENTITY = "@sikongjueluo/pi-permission-ai-judge@0.0.1";
|
|
||||||
const PERMISSION_SYSTEM_VERSION = "25.4.0";
|
|
||||||
const REVIEW_SCHEMA_VERSION = "1";
|
|
||||||
|
|
||||||
interface CliOptions {
|
|
||||||
kind: PromotionRecordKind;
|
|
||||||
provider: string;
|
|
||||||
model: string;
|
|
||||||
api: string;
|
|
||||||
timeoutCohort: "default" | number;
|
|
||||||
basis: string;
|
|
||||||
}
|
|
||||||
|
|
||||||
function usage(): string {
|
|
||||||
return [
|
|
||||||
"usage: promotion-record --kind <cohort_qualified|owner_approval|activation>",
|
|
||||||
" --provider <p> --model <m> --api <api>",
|
|
||||||
" [--timeout-cohort default|<ms>] --basis <text>",
|
|
||||||
].join("\n");
|
|
||||||
}
|
|
||||||
|
|
||||||
function parseArgs(argv: readonly string[]): CliOptions | { error: string } {
|
|
||||||
const args = argv.slice(2);
|
|
||||||
const opts: Partial<CliOptions> = {};
|
|
||||||
for (let i = 0; i < args.length; i += 1) {
|
|
||||||
const arg = args[i] as string;
|
|
||||||
const value = args[i + 1];
|
|
||||||
if (
|
|
||||||
arg === "--kind" ||
|
|
||||||
arg === "--provider" ||
|
|
||||||
arg === "--model" ||
|
|
||||||
arg === "--api" ||
|
|
||||||
arg === "--basis"
|
|
||||||
) {
|
|
||||||
if (value === undefined) return { error: `${arg} requires a value` };
|
|
||||||
if (arg === "--kind") {
|
|
||||||
if (
|
|
||||||
value !== "cohort_qualified" &&
|
|
||||||
value !== "owner_approval" &&
|
|
||||||
value !== "activation"
|
|
||||||
) {
|
|
||||||
return { error: `unknown kind: ${value}` };
|
|
||||||
}
|
|
||||||
opts.kind = value;
|
|
||||||
}
|
|
||||||
if (arg === "--provider") opts.provider = value;
|
|
||||||
if (arg === "--model") opts.model = value;
|
|
||||||
if (arg === "--api") opts.api = value;
|
|
||||||
if (arg === "--basis") opts.basis = value;
|
|
||||||
i += 1;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
if (arg === "--timeout-cohort") {
|
|
||||||
if (value === undefined) return { error: `${arg} requires a value` };
|
|
||||||
if (value === "default") {
|
|
||||||
opts.timeoutCohort = "default";
|
|
||||||
} else {
|
|
||||||
const n = Number(value);
|
|
||||||
if (!Number.isInteger(n)) {
|
|
||||||
return { error: `--timeout-cohort must be default or an integer` };
|
|
||||||
}
|
|
||||||
opts.timeoutCohort = n;
|
|
||||||
}
|
|
||||||
i += 1;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
return { error: `unknown option: ${arg}\n${usage()}` };
|
|
||||||
}
|
|
||||||
if (
|
|
||||||
opts.kind === undefined ||
|
|
||||||
opts.provider === undefined ||
|
|
||||||
opts.model === undefined ||
|
|
||||||
opts.api === undefined ||
|
|
||||||
opts.basis === undefined ||
|
|
||||||
opts.basis.length === 0
|
|
||||||
) {
|
|
||||||
return { error: usage() };
|
|
||||||
}
|
|
||||||
return {
|
|
||||||
kind: opts.kind,
|
|
||||||
provider: opts.provider,
|
|
||||||
model: opts.model,
|
|
||||||
api: opts.api,
|
|
||||||
timeoutCohort: opts.timeoutCohort ?? "default",
|
|
||||||
basis: opts.basis,
|
|
||||||
};
|
|
||||||
}
|
|
||||||
|
|
||||||
function main(): number {
|
|
||||||
const parsed = parseArgs(process.argv);
|
|
||||||
if ("error" in parsed) {
|
|
||||||
process.stderr.write(`${parsed.error}\n`);
|
|
||||||
return 1;
|
|
||||||
}
|
|
||||||
const identity: CandidateIdentity = {
|
|
||||||
judge: JUDGE_IDENTITY,
|
|
||||||
permissionSystem: PERMISSION_SYSTEM_VERSION,
|
|
||||||
provider: parsed.provider,
|
|
||||||
model: parsed.model,
|
|
||||||
api: parsed.api,
|
|
||||||
promptVersion: PROMPT_VERSION,
|
|
||||||
toolSchemaVersion: TOOL_SCHEMA_VERSION,
|
|
||||||
reviewSchemaVersion: REVIEW_SCHEMA_VERSION,
|
|
||||||
timeoutCohort: parsed.timeoutCohort,
|
|
||||||
};
|
|
||||||
const agentDir = getAgentDir();
|
|
||||||
const error = appendPromotionRecord({
|
|
||||||
agentDir,
|
|
||||||
record: {
|
|
||||||
kind: parsed.kind,
|
|
||||||
candidateIdentity: identity,
|
|
||||||
recordedAt: new Date().toISOString(),
|
|
||||||
basis: parsed.basis,
|
|
||||||
},
|
|
||||||
});
|
|
||||||
if (error !== null) {
|
|
||||||
process.stderr.write(`${error}\n`);
|
|
||||||
return 1;
|
|
||||||
}
|
|
||||||
process.stdout.write(
|
|
||||||
`appended ${parsed.kind} record to ${promotionRecordsPath(agentDir)}\n` +
|
|
||||||
`identity: ${JSON.stringify(identity)}\n` +
|
|
||||||
`note: records load at session start; restart sessions to pick this up\n`,
|
|
||||||
);
|
|
||||||
return 0;
|
|
||||||
}
|
|
||||||
|
|
||||||
process.exit(main());
|
|
||||||
Reference in New Issue
Block a user