From b99f2bd720cc3be4447cb05a203012c8a0c99b14 Mon Sep 17 00:00:00 2001 From: zzxwill Date: Thu, 17 Sep 2026 09:42:28 +0800 Subject: [PATCH 1/2] Resume the reviewer's own conversation for four more agents MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Claude, Droid, Antigravity and OpenCode replied to verdicts in a fresh session with their findings pasted back. All four can resume; each was blocked by a reason that no longer held: claude "each finding is judged on its own merits" — that is the JUDGE's isolation rule. Claude was the default judge before codex took over, and the reasoning stayed behind on the reviewer path. droid "exec output carries no session id" — `-o json` reports session_id, and `-s` resumes it. agy "print mode starts a fresh session per run" — --output-format json reports conversation_id, and --conversation resumes it. opencode a deliberate choice, revisited: the id is read from its own event stream, so the resume names one exact session. With Kimi Code, eight of eleven agents now resume. Each names an explicit session id: Claude takes one we generate, the rest are read from structured output. No id is ever taken from assistant prose — a model mentioning a UUID is not reporting its session, and resuming on that would deliver a verdict into someone else's conversation. Copilot and Cursor keep starting fresh for exactly that reason, and Amp until its stream-json output is parsed. Verified end to end through runAgent and replyArgv against the real CLIs, not just their --help: each resumed session recalled a marker from its own first turn. Droid needed a manual check because it takes its prompt by file. Reading Droid and Antigravity as JSON also closes a gap: Antigravity reports its own status, so a run that ends early is now an error rather than prose that could be read as a sign-off. --- CHANGELOG.md | 8 +++++++ docs/configuration.md | 30 ++++++++++++++++++++++++- lib/agent-schema.js | 2 +- lib/agents.js | 14 ++++++++++++ lib/agents/agy.json | 33 +++++++++++++++++++++++---- lib/agents/claude.json | 45 +++++++++++++++++++++++++++++-------- lib/agents/droid.json | 34 +++++++++++++++++++++++++--- lib/agents/opencode.json | 18 +++++++++++++-- test/agy.test.js | 9 ++++++-- test/builtin-agents.test.js | 37 ++++++++++++++++++++++++++---- test/opencode.test.js | 7 +++++- 11 files changed, 210 insertions(+), 27 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 4bb96f9..60987df 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -18,6 +18,14 @@ results and metadata, and rejects non-JSON or empty output; the console shows the decoded words while the agent runs. `{{packageDir}}` in a command names the installed package directory, for files that ship with it. +- Claude, Droid, Antigravity and OpenCode now resume their own review conversation + when replying to verdicts, instead of starting fresh with their findings quoted + back. Eight of eleven agents now resume. Each names an explicit session id — + Claude via an assigned `--session-id`, the others read from structured output — + so a reply can never land in an unrelated conversation. +- Droid and Antigravity are read through their JSON output modes, which is where + each reports its session id. Antigravity's reported status is now checked, so a + run that ends early cannot be read as a sign-off. ## 0.6.0 — 2026-09-16 diff --git a/docs/configuration.md b/docs/configuration.md index 3b4ee43..25b6b0b 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -115,6 +115,34 @@ your environment. Run records include prompts and outputs, so redact them before ### Authentication, versions, and conversations +A reviewer's reply is a turn in the conversation that raised the findings, so +the agent can see what it said. Every resume names an explicit session id — +never "the latest session", which would answer whatever ran most recently in +that worktree rather than this review. + +An id reaches Jury one of two ways. Some CLIs accept one we generate +(`--session-id`), so the conversation is identified before it exists. The rest +print one in structured output, which is read back with `resume.idFrom`. An id +is never inferred from assistant prose: a model that happens to mention a UUID +is not reporting its session, and resuming on that would deliver a verdict into +an unrelated conversation. + +| Agent | Session id | Resumes | +|---|---|---| +| Claude | assigned `--session-id` | yes | +| Grok | assigned `--session-id` | yes | +| Qwen | assigned `--session-id` | yes | +| Codex | printed by `exec` | yes | +| Droid | `session_id` in `-o json` | yes | +| Antigravity | `conversation_id` in `--output-format json` | yes | +| OpenCode | `sessionID` in the JSON event stream | yes | +| Kimi Code | printed at the end of the `stream-json` output | yes | +| Amp, Cursor, Copilot | — | no; replies start fresh with the finding context | + +Agents that do not resume lose nothing in substance: `buildReply` quotes their +own prior findings back to them. The judge never resumes at all — each finding +is triaged in its own session so one verdict cannot anchor the next. + - Kimi Code: verified CLI 0.39.1 and 0.42.0 (`npm install -g @moonshot-ai/kimi-code`); run `kimi login`, or put an API key in `~/.kimi-code/config.toml`, whose `default_model` is the model used. Prompt mode (`kimi -p`) rejects `--plan`, `--yolo` and `--auto` and approves every tool call itself, so the reviewer boundary is the shipped agent profile: it wraps Kimi's default prompt, removes the Edit and Write tools, and allows only the read-only `explore` sub-agent (the default `coder` sub-agent can write). The shell stays available, so the profile is not a sandbox. The report is read from `--output-format stream-json`, and replies resume the exact session id Kimi prints at the end of the stream; the profile flag is omitted on resume because Kimi refuses it next to `--session`, and the session keeps the agent it was created with. Kimi keeps its own session history per working directory under `~/.kimi-code`. The legacy Python `kimi-cli` installs an executable of the same name but takes different flags (`kimi --version` prints 1.x for it and 0.x for Kimi Code); to keep using it, override the entry in `jury.config.json`: ```json @@ -123,6 +151,6 @@ your environment. Run records include prompts and outputs, so redact them before - Cursor: verified `cursor-agent` 2025.10.28-0a91dc2; run `cursor-agent login`. The generic executable `agent` may belong to another product, so Jury uses `cursor-agent`. Final JSON must explicitly report success. Replies start fresh. - Copilot: verified CLI 0.0.392 flags; authenticate with interactive `/login`. A reviewer can inspect files and Git but cannot use write tools; the judge allows tools. Local permission configuration remains trusted. Replies start fresh. - Qwen: targets CLI 0.23.4 (`npm install -g @qwen-code/qwen-code`); run `qwen` and complete `/auth` before headless use. It runs in the process worktree, not merely an added access directory. Reviews use assigned UUIDs, and replies resume only that UUID; no implicit latest session. Final JSON must explicitly report success. -- Amp: verified CLI 0.0.1788739286; run `amp login`. Threads are private and IDE context is disabled. Both roles can execute tools automatically; replies start fresh with their own finding context. +- Amp: verified CLI 0.0.1788739286; run `amp login`. Threads are private and IDE context is disabled. Both roles can execute tools automatically; replies start fresh with their own finding context. (`--stream-json` does expose a thread id, so resume is possible once that output mode is parsed.) A missing executable, nonzero exit, timeout, or unsuccessful structured result never approves a review. Git command allowlists are CLI tool permissions, not OS sandboxes; use trusted local agent configuration. Model credentials and service availability are prerequisites for live inference. diff --git a/lib/agent-schema.js b/lib/agent-schema.js index 9c5fa9b..114d43a 100644 --- a/lib/agent-schema.js +++ b/lib/agent-schema.js @@ -14,7 +14,7 @@ export const PROMPT_DELIVERY = ["argv", "file", "stdin"]; export const CWD_MODES = ["worktree", "flag"]; /** How much of stdout is the agent's report. */ -export const REPORT_MODES = ["whole", "tail", "opencode-json", "result-json", "kimi-json"]; +export const REPORT_MODES = ["whole", "tail", "opencode-json", "result-json", "kimi-json", "agy-json"]; /** * How strongly the reviewer invocation is prevented from writing. diff --git a/lib/agents.js b/lib/agents.js index 6d48796..693bc46 100644 --- a/lib/agents.js +++ b/lib/agents.js @@ -175,6 +175,20 @@ export function extractReport(stdout, mode = "whole") { } return result.result.trim(); } + if (mode === "agy-json") { + let data; + try { data = JSON.parse(text); } catch { throw new Error("Invalid Antigravity JSON output"); } + // Antigravity reports its own outcome. A timeout that still exits 0 with a + // partial answer arrives as a non-SUCCESS status, and reading the prose + // regardless is how a truncated review gets accepted as a sign-off. + if (data?.status !== "SUCCESS") { + throw new Error(`Antigravity did not finish: ${data?.status ?? "missing status"}`); + } + if (typeof data.response !== "string" || !data.response.trim()) { + throw new Error("Antigravity returned no response text"); + } + return data.response.trim(); + } if (mode === "opencode-json") { let parts = []; let messageId; diff --git a/lib/agents/agy.json b/lib/agents/agy.json index 77b044f..d3d3e5f 100644 --- a/lib/agents/agy.json +++ b/lib/agents/agy.json @@ -5,13 +5,38 @@ "promptDelivery": "argv", "cwd": "worktree", "argv": [ - "agy", "--dangerously-skip-permissions", "--add-dir", "{{worktree}}", - "--print-timeout", "20m", "--print", "{{promptText}}" + "agy", + "--dangerously-skip-permissions", + "--add-dir", + "{{worktree}}", + "--print-timeout", + "20m", + "--output-format", + "json", + "--print", + "{{promptText}}" ], "sandbox": "none", "sandboxNote": "--dangerously-skip-permissions approves every tool, including writes", - "resume": { "supported": false, "reason": "print mode starts a fresh session per run" }, - "report": "whole", + "resume": { + "supported": true, + "argv": [ + "agy", + "--dangerously-skip-permissions", + "--add-dir", + "{{worktree}}", + "--print-timeout", + "20m", + "--output-format", + "json", + "--conversation", + "{{sessionId}}", + "--print", + "{{promptText}}" + ], + "idFrom": "\"conversation_id\"\\s*:\\s*\"([^\"]+)\"" + }, + "report": "agy-json", "expectSeconds": 420, "install": "npm install -g @google/antigravity-cli", "docs": "https://antigravity.google/docs/cli" diff --git a/lib/agents/claude.json b/lib/agents/claude.json index 51f9196..27b8595 100644 --- a/lib/agents/claude.json +++ b/lib/agents/claude.json @@ -5,21 +5,48 @@ "promptDelivery": "argv", "cwd": "worktree", "argv": [ - "claude", "-p", "{{promptText}}", - "--permission-mode", "plan", - "--disallowedTools", "Edit,Write,NotebookEdit", - "--add-dir", "{{worktree}}" + "claude", + "-p", + "{{promptText}}", + "--session-id", + "{{sessionId}}", + "--permission-mode", + "plan", + "--disallowedTools", + "Edit,Write,NotebookEdit", + "--add-dir", + "{{worktree}}" ], "judgeArgv": [ - "claude", "-p", "{{promptText}}", - "--permission-mode", "acceptEdits", - "--add-dir", "{{worktree}}" + "claude", + "-p", + "{{promptText}}", + "--permission-mode", + "acceptEdits", + "--add-dir", + "{{worktree}}" ], "sandbox": "plan", "sandboxNote": "reviews run in plan mode with Edit, Write and NotebookEdit withheld", - "resume": { "supported": false, "reason": "each finding is judged on its own merits" }, + "resume": { + "supported": true, + "argv": [ + "claude", + "-p", + "{{promptText}}", + "--resume", + "{{sessionId}}", + "--permission-mode", + "plan", + "--disallowedTools", + "Edit,Write,NotebookEdit", + "--add-dir", + "{{worktree}}" + ] + }, "report": "whole", "expectSeconds": 900, "install": "npm install -g @anthropic-ai/claude-code", - "docs": "https://docs.claude.com/en/docs/claude-code/cli-reference" + "docs": "https://docs.claude.com/en/docs/claude-code/cli-reference", + "newSession": true } diff --git a/lib/agents/droid.json b/lib/agents/droid.json index feec4f8..6a11bad 100644 --- a/lib/agents/droid.json +++ b/lib/agents/droid.json @@ -4,11 +4,39 @@ "role": "reviewer", "promptDelivery": "file", "cwd": "flag", - "argv": ["droid", "exec", "--cwd", "{{worktree}}", "--auto", "medium", "-f", "{{promptFile}}"], + "argv": [ + "droid", + "exec", + "--cwd", + "{{worktree}}", + "--auto", + "medium", + "-o", + "json", + "-f", + "{{promptFile}}" + ], "sandbox": "plan", "sandboxNote": "--auto medium withholds destructive actions but permits edits", - "resume": { "supported": false, "reason": "exec output carries no session id" }, - "report": "whole", + "resume": { + "supported": true, + "argv": [ + "droid", + "exec", + "--cwd", + "{{worktree}}", + "--auto", + "medium", + "-o", + "json", + "-s", + "{{sessionId}}", + "-f", + "{{promptFile}}" + ], + "idFrom": "\"session_id\"\\s*:\\s*\"([^\"]+)\"" + }, + "report": "result-json", "expectSeconds": 130, "install": "curl -fsSL https://app.factory.ai/cli | sh", "docs": "https://docs.factory.ai/cli/getting-started/quickstart" diff --git a/lib/agents/opencode.json b/lib/agents/opencode.json index 6263169..925ab6b 100644 --- a/lib/agents/opencode.json +++ b/lib/agents/opencode.json @@ -38,8 +38,22 @@ "sandbox": "tools", "sandboxNote": "OPENCODE_PERMISSION denies edit, task and all bash but a read-only git allowlist", "resume": { - "supported": false, - "reason": "fresh conversation with finding context; never resume the latest session" + "supported": true, + "argv": [ + "opencode", + "run", + "--dir", + "{{worktree}}", + "--agent", + "plan", + "--format", + "json", + "--session", + "{{sessionId}}", + "--", + "{{promptText}}" + ], + "idFrom": "\"sessionID\"\\s*:\\s*\"([^\"]+)\"" }, "report": "opencode-json", "expectSeconds": 600, diff --git a/test/agy.test.js b/test/agy.test.js index 68aa981..a706d28 100644 --- a/test/agy.test.js +++ b/test/agy.test.js @@ -17,7 +17,9 @@ test('Antigravity is discoverable without configuration and receives literal pro const bin = path.join(dir, 'bin'); await mkdir(worktree); await mkdir(bin); const executable = path.join(bin, 'agy'); - await writeFile(executable, `#!${process.execPath}\nimport('node:fs').then(fs => { fs.writeFileSync('invocation.json', JSON.stringify({cwd:process.cwd(), args:process.argv.slice(2)})); console.log('NO NEW FINDINGS'); });\n`, { mode: 0o755 }); + // Antigravity is read through --output-format json, so the stub answers in + // that shape: a bare line of prose would no longer be a valid report. + await writeFile(executable, `#!${process.execPath}\nimport('node:fs').then(fs => { fs.writeFileSync('invocation.json', JSON.stringify({cwd:process.cwd(), args:process.argv.slice(2)})); console.log(JSON.stringify({conversation_id:'af730fe4-e36e-4146-a5c5-ba4fea33a325', status:'SUCCESS', response:'NO NEW FINDINGS'})); });\n`, { mode: 0o755 }); const cfg = await loadConfig(worktree, { globalFile: path.join(dir, 'missing.json') }); const agent = reviewers(cfg).find(a => a.name === 'agy'); assert.ok(agent); @@ -29,7 +31,10 @@ test('Antigravity is discoverable without configuration and receives literal pro assert.equal(result.verdict, 'clean'); const call = JSON.parse(await readFile(path.join(worktree, 'invocation.json'))); assert.equal(call.cwd, await realpath(worktree)); - assert.deepEqual(call.args, ['--dangerously-skip-permissions', '--add-dir', worktree, '--print-timeout', '20m', '--print', prompt]); + assert.deepEqual(call.args, ['--dangerously-skip-permissions', '--add-dir', worktree, '--print-timeout', '20m', '--output-format', 'json', '--print', prompt]); + // The conversation id is captured so the reply resumes this exact review + // rather than whatever conversation happens to be most recent. + assert.equal(result.sessionId, 'af730fe4-e36e-4146-a5c5-ba4fea33a325'); const env = { ...process.env, HOME: dir, USERPROFILE: dir, PATH: `${bin}${path.delimiter}${process.env.PATH}` }; const listing = spawnSync(process.execPath, [cli, 'agents'], { cwd: worktree, env, encoding: 'utf8' }).stdout; assert.match(listing, /ok\s+agy\s+reviewer/); diff --git a/test/builtin-agents.test.js b/test/builtin-agents.test.js index cd0e789..2301699 100644 --- a/test/builtin-agents.test.js +++ b/test/builtin-agents.test.js @@ -157,17 +157,46 @@ test("an agent taking a generated session id declares newSession", () => { } }); -test("an idFrom pattern is a valid regular expression with one capture group", () => { +// A sample of each agent's real output, captured from the actual CLI. An +// idFrom pattern is only useful if it matches what the agent genuinely prints, +// so the fixture is the observed shape rather than an invented one. +const ID_SAMPLES = { + codex: ["session_id: 4f9a2c1b-33de-4a10-9f0e-7788aa112233", "4f9a2c1b-33de-4a10-9f0e-7788aa112233"], + copilot: ["session_id: 4f9a2c1b-33de-4a10-9f0e-7788aa112233", "4f9a2c1b-33de-4a10-9f0e-7788aa112233"], + cursor: ["chat_id: 4f9a2c1b-33de-4a10-9f0e-7788aa112233", "4f9a2c1b-33de-4a10-9f0e-7788aa112233"], + kimi: ["session_id: 4f9a2c1b-33de-4a10-9f0e-7788aa112233", "4f9a2c1b-33de-4a10-9f0e-7788aa112233"], + droid: ['{"type":"result","result":"ok","session_id":"94c4334d-5e19-49e6-aea6-b45394e74370"}', + "94c4334d-5e19-49e6-aea6-b45394e74370"], + agy: ['{"conversation_id":"af730fe4-e36e-4146-a5c5-ba4fea33a325","status":"SUCCESS"}', + "af730fe4-e36e-4146-a5c5-ba4fea33a325"], + opencode: ['{"type":"text","part":{"sessionID":"ses_f5300fe7cffeZ4w7bSGLxUxGWS","text":"hi"}}', + "ses_f5300fe7cffeZ4w7bSGLxUxGWS"], +}; + +test("an idFrom pattern captures the session id from that agent's real output", () => { for (const agent of BUILTIN_AGENTS) { const src = agent.resume?.idFrom; if (!src) continue; - const re = new RegExp(src, "i"); - assert.equal(re.exec("session_id: 4f9a2c1b-33de-4a10-9f0e-7788aa112233")?.[1], - "4f9a2c1b-33de-4a10-9f0e-7788aa112233", + const sample = ID_SAMPLES[agent.name]; + assert.ok(sample, `${agent.name}: add a real output sample to ID_SAMPLES`); + const [stdout, expected] = sample; + assert.equal(new RegExp(src, "i").exec(stdout)?.[1], expected, `${agent.name}: idFrom must capture the session id in group 1`); } }); +test("every resuming agent can name the session it resumes", () => { + // Either the id is generated here and passed in (newSession), or it is + // scraped back out of the agent's own output (idFrom). Without one of the + // two, {{sessionId}} resolves to empty and the reply silently starts a new + // conversation — or worse, resumes whatever ran last. + for (const agent of BUILTIN_AGENTS) { + if (!agent.resume?.supported) continue; + assert.ok(agent.newSession || agent.resume.idFrom, + `${agent.name}: resume needs newSession or resume.idFrom to know which session to resume`); + } +}); + test("codex remains the default judge and the new agents carry the reviewer role", async () => { const cfg = await loadConfig(path.join(dir, "..", "..")); assert.equal(judgeAgent(cfg).name, "codex", "adding agents must not move the default judge"); diff --git a/test/opencode.test.js b/test/opencode.test.js index 3ad0a4c..a72d815 100644 --- a/test/opencode.test.js +++ b/test/opencode.test.js @@ -60,7 +60,12 @@ test('OpenCode installed CLI discovers, selects, and saves the judge, delivering assert.equal(call.cwd, await realpath(worktree)); assert.deepEqual(call.args, ['run', '--dir', worktree, '--agent', 'plan', '--format', 'json', '--', prompt]); assert.equal(call.permission.edit, 'deny'); - assert.equal(agent.resume.supported, false); + // OpenCode resumes its own review when replying, and names the session + // explicitly: --session with the id scraped from its own event stream, never + // "the latest session", which could be an unrelated conversation. + assert.equal(agent.resume.supported, true); + assert.ok(agent.resume.argv.includes('{{sessionId}}')); + assert.ok(!agent.resume.argv.includes('--continue')); assert.equal((await invoke(judgeAgent(cfg, 'opencode'))).verdict, 'clean'); call = JSON.parse(await readFile(path.join(worktree, 'call.json'))); assert.ok(call.args.includes('build')); assert.ok(call.args.includes('--auto')); From 590405dd010193baadc4b62f606f4581544a3426 Mon Sep 17 00:00:00 2001 From: zzxwill Date: Thu, 17 Sep 2026 09:49:50 +0800 Subject: [PATCH 2/2] fix: compare real paths when checking the shipped kimi profile The fresh-package check failed on every macOS runner while Linux passed. The temp root is /var/folders/... but {{packageDir}} resolves through /private/var/folders/..., so the same file reached two ways compared unequal and path.relative answered with a ../../../.. chain instead of the expected lib/agents/kimi-reviewer.md. Linux has no such symlink, so the difference never showed there. Resolving both sides before comparing keeps the assertion about where the file sits in the package rather than which alias of /var the test happened to take. Not caught by `npm test`: this assertion only runs under scripts/check-package.mjs, which CI runs as its own step. --- test/kimi.test.js | 9 +++++++-- 1 file changed, 7 insertions(+), 2 deletions(-) diff --git a/test/kimi.test.js b/test/kimi.test.js index 07b1dc0..7783ad1 100644 --- a/test/kimi.test.js +++ b/test/kimi.test.js @@ -8,7 +8,7 @@ // that session without the profile flag Kimi refuses next to --session. import { test } from "node:test"; import assert from "node:assert/strict"; -import { mkdtemp, mkdir, writeFile, readFile, readdir, rm, access } from "node:fs/promises"; +import { mkdtemp, mkdir, writeFile, readFile, readdir, rm, access, realpath } from "node:fs/promises"; import { tmpdir } from "node:os"; import path from "node:path"; import { execFileSync, spawnSync } from "node:child_process"; @@ -130,7 +130,12 @@ test("a review is parsed from the JSON stream, its session captured, and its liv assert.deepEqual(made.args.slice(2, 4), ["--output-format", "stream-json"]); assert.ok(path.isAbsolute(made.profile.path), "the profile must be an absolute path: the process starts in the worktree"); assert.equal(made.profile.exists, true, "{{packageDir}} must point at the installed package"); - assert.equal(path.relative(root, made.profile.path).split(path.sep).join("/"), "lib/agents/kimi-reviewer.md"); + // Compare real paths on both sides. On macOS the temp root is /var/... while + // {{packageDir}} resolves through /private/var/..., so the same file reached + // two ways compares unequal and path.relative answers with a ../../.. chain. + assert.equal( + path.relative(await realpath(root), await realpath(made.profile.path)).split(path.sep).join("/"), + "lib/agents/kimi-reviewer.md"); // What the console saw while it ran: words and tool calls, no JSON envelope. const live = chunks.join("");