Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,13 @@

## Unreleased

- Configure the model each agent runs with: `jury agents model <agent> <model>` saves a default in
`~/.jury/config.json`, a `model` field per agent in `jury.config.json` sets one per repository,
and `--model <agent>=<model>` sets one for a run, each overriding the one before. The model is
passed as a flag or environment variable only when Jury runs the agent, on first runs, resumed
rounds, replies and judging alike; the agent's own configuration is never changed. Supported for
codex, claude, grok, opencode, qwen, copilot and kimi; a configured model for any other agent
stops the run with an error. Runs record the models used, and `jury agents` shows them (#99).
- Add `jury agents jury` to save default reviewers globally in `~/.jury/config.json`, beside the
global judge: `jury agents jury claude,droid`, `--reset`, and a checkbox picker when run with no
arguments in a terminal. Runs without `--jury` use them; repository `reviewer` roles and
Expand Down
9 changes: 9 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -180,6 +180,15 @@ off automatic role assignment, as `--jury` does. The judge is left out of its ow
reviewer is later uninstalled or disabled in a repository, the run stops with an error naming the
saved setting rather than reviewing with fewer agents.

Pick the model an agent runs with, without touching its own configuration:

```bash
jury agents model claude opus # saved default
jury review <pr-url> --model codex=gpt-5.5 # this run only
```

See [models](docs/configuration.md#models) for each agent's flag and the precedence rules.

### Built-in agents

| name | product | install | default |
Expand Down
65 changes: 62 additions & 3 deletions bin/jury.js
Original file line number Diff line number Diff line change
Expand Up @@ -13,8 +13,8 @@ import { readFileSync } from "node:fs";
import { promisify } from "node:util";
import path from "node:path";
import { automaticRoles, automaticReviewers } from "../lib/roles.js";
import { loadConfig, reviewers, judgeAgent, knownAgents, readGlobalConfig, saveGlobalJudge, saveGlobalReviewers, defaultReviewers, savedReviewersNote, globalConfigPath } from "../lib/config.js";
import { runAgent, probe } from "../lib/agents.js";
import { loadConfig, reviewers, judgeAgent, knownAgents, readGlobalConfig, saveGlobalJudge, saveGlobalReviewers, defaultReviewers, savedReviewersNote, globalConfigPath, applyModelFlags, assertModelsSupported, modelsUsed, saveGlobalModel, modelProblem } from "../lib/config.js";
import { runAgent, probe, supportsModel } from "../lib/agents.js";
import { threadFor, buildReply, replyArgv } from "../lib/reply.js";
import { buildPrompt } from "../lib/prompt.js";
import { serve } from "../lib/server.js";
Expand Down Expand Up @@ -64,6 +64,7 @@ Common flags
--reviewer <name> only these reviewers, repeatable or comma-separated (default: configured or saved reviewers)
--jury <name> same as --reviewer
--judge codex one agent that triages and fixes (default: configured, auto for 1–2 installed CLIs, then codex)
--model <agent>=<model> this run's model for an agent, repeatable (default: jury.config.json, then jury agents model)
--push <true|false> commit and push fixes (default: true)
--web <true|false> open the browser console (default: true)

Expand All @@ -87,6 +88,7 @@ const USAGE_FULL = `jury — review a pull request with multiple AI reviewers un
jury agents check which configured agents are installed
jury agents judge <agent> set the global default judge
jury agents jury <agents> set the default reviewers (picker in a terminal)
jury agents model <agent> set the model jury starts an agent with
jury version

Related PRs: jury review <pr-url-1> <pr-url-2>
Expand All @@ -106,6 +108,7 @@ review (triages, fixes, commits, and pushes automatica
--reviewer <name> only these reviewers, repeatable or comma-separated (default: configured or saved reviewers)
--jury <name> same as --reviewer
--judge <agent> one agent that triages and fixes (default: configured, auto for 1–2 installed CLIs, then codex)
--model <agent>=<model> this run's model for an agent, repeatable (default: jury.config.json, then jury agents model)
--resume <slug> continue an existing run instead of starting a new one
--web <true|false> open the console; stays up after review (default: true)
--web-only view saved reviews without running agents
Expand Down Expand Up @@ -470,6 +473,7 @@ async function cmdAgent(argv) {
rounds: { type: "string", default: "10" },
...reviewerOptions,
judge: { type: "string" },
model: { type: "string", multiple: true },
push: { type: "boolean", default: true },
resume: { type: "string" },
"dry-run": { type: "boolean", default: false },
Expand Down Expand Up @@ -519,6 +523,7 @@ async function cmdAgent(argv) {
// A local ignored config belongs to the requested checkout. An automatic
// clone intentionally starts from the repository's committed/default config.
const cfg = await loadConfig(group ? group.targets[0].worktree : resolved ? worktree : requestedWorktree);
applyModelFlags(cfg, values.model);
const git = await describe(group ? group.targets[0].worktree : worktree);

// Both asked for rather than assumed: the trunk from the remote's own HEAD,
Expand Down Expand Up @@ -605,6 +610,8 @@ async function cmdAgent(argv) {
const pool = roles ? automaticReviewers(cfg, roles)
: selectReviewers(cfg, requestedReviewers(values), judge.name);
if (!pool.length) throw new Error("no reviewers configured after excluding the judge");
assertModelsSupported([judge, ...pool]);
target.models = modelsUsed([judge, ...pool]);
if (!values["dry-run"]) {
const checks = await Promise.all(pool.map(probe));
const missing = checks.filter(a => !a.ok);
Expand Down Expand Up @@ -637,6 +644,9 @@ async function cmdAgent(argv) {
console.log(st.field("judge", st.agent(judge.name)));
console.log(st.field("juries", pool.map((a) => st.agent(a.name)).join(", ")
+ (values["dry-run"] ? st.warn(" (dry run)") : "")));
if (target.models) {
console.log(st.field("models", Object.entries(target.models).map(([n, m]) => `${st.agent(n)} ${m}`).join(", ")));
}
console.log(st.field("rounds", st.muted(`${first}..${first + maxRounds - 1}, ${MAX_TURNS} turns per finding`)));
console.log(st.field("fixes", st.muted(values.push
? `committed and pushed to ${group ? group.targets.map(t => t.branch).join(", ") : pushTarget.branch}` : "committed to the worktree only")));
Expand Down Expand Up @@ -1072,12 +1082,14 @@ async function cmdReply(argv) {
run: { type: "string" },
dir: { type: "string" },
...reviewerOptions,
model: { type: "string", multiple: true },
"dry-run": { type: "boolean", default: false },
},
});

const worktree = await commandDirectory(values.dir);
const cfg = await loadConfig(worktree);
applyModelFlags(cfg, values.model);
const dir = await resolveRun(values.run, worktree);
const events = await readEvents(dir);
const judge = currentJudge(events);
Expand All @@ -1094,6 +1106,7 @@ async function cmdReply(argv) {
return t.length && t.some((f) => f.status !== "open");
});
if (!pool.length) throw new Error("nothing to reply about — resolve some findings first");
assertModelsSupported(pool);

console.log(`judge ${judge}`);
console.log(`replying ${pool.map((a) => a.name).join(", ")} (separate conversations)`);
Expand Down Expand Up @@ -1198,6 +1211,7 @@ async function cmdRuns(argv) {

async function cmdAgents(args = []) {
if (["jury", "reviewer", "reviewers"].includes(args[0])) return cmdDefaultReviewers(args.slice(1));
if (["model", "models"].includes(args[0])) return cmdAgentModels(args.slice(1));
if (args.length) {
if (args[0] !== "judge" || args.length > 2) throw new Error("Usage: jury agents judge [<agent>|--reset] | jury agents jury [<agent>,...|--reset]");
const name = args[1];
Expand Down Expand Up @@ -1239,7 +1253,8 @@ async function cmdAgents(args = []) {
console.log(
`${p.ok ? "ok " : "MISSING"} ${p.name.padEnd(9)} ${role.padEnd(8)} ${p.path ?? p.bin}`
+ (byName.get(p.name)?.enabled === false ? " (opt-in)" : "")
+ (saved.has(p.name) ? " (default reviewer)" : ""),
+ (saved.has(p.name) ? " (default reviewer)" : "")
+ (byName.get(p.name)?.model ? ` (model ${byName.get(p.name).model})` : ""),
);
}
if (cfg.savedReviewers) {
Expand Down Expand Up @@ -1281,6 +1296,50 @@ async function cmdAgents(args = []) {
}
}

/**
* `jury agents model`: the model each agent is started with when jury runs it.
*
* Saved in ~/.jury/config.json and passed to the agent as a flag or variable
* on each run; the agent's own configuration is never touched, so running the
* CLI outside jury keeps its normal default.
*/
async function cmdAgentModels(args) {
const usage = "Usage: jury agents model [<agent> <model>|<agent> --reset]";
if (args.includes("--help") || args.includes("-h")) {
process.stdout.write(commandHelp("agents", USAGE_FULL));
return;
}
const cfg = await loadConfig();
const known = knownAgents(cfg);
if (!args.length) {
for (const a of known) {
const setting = a.model ? `${a.model} (${a.modelSource})` : supportsModel(a) ? "CLI default" : "CLI default (no per-run model)";
console.log(`${a.name.padEnd(9)} ${setting}`);
}
return;
}
if (args.length !== 2) throw new Error(usage);
const [name, model] = args;
const agent = known.find(a => a.name === name);
if (!agent) throw new Error(`Unknown or disabled agent "${name}". Choose: ${known.map(a => a.name).join(", ")}`);
if (model === "--reset") {
await saveGlobalModel(name, null);
console.log(`${name}: saved model removed; the repository setting or the CLI's own default applies.`);
return;
}
const problem = modelProblem(model);
if (problem) throw new Error(problem);
if (!supportsModel(agent)) {
throw new Error(`${name} cannot select a model per run: ${agent.modelNote ?? "its command has no {{modelArgs}} slot"}`);
}
await saveGlobalModel(name, model);
console.log(`${name}: model ${model} (${globalConfigPath()})`);
if (agent.modelSource && agent.modelSource !== "~/.jury/config.json") {
console.log(`${agent.modelSource} sets ${agent.model} for ${name} here and takes precedence.`);
}
console.log("Used only when jury runs this agent; its own configuration is unchanged. --model <agent>=<model> overrides it per run.");
}

/**
* `jury agents jury`: the saved default reviewers, used when a run names none.
*
Expand Down
41 changes: 41 additions & 0 deletions docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -57,6 +57,47 @@ A judge may specify separate `judgeArgv`. A custom `argv` overrides the built-in
unless you also specify `judgeArgv`. Optional session configuration is illustrated by the built-in
definitions in `lib/agents/`.

## Models

Jury can start an agent with a chosen model. The model is passed only on the command line or in
the environment of the process Jury starts; Jury never edits an agent's own configuration, so
running `claude`, `codex` and the others outside Jury keeps their normal default.

```bash
jury agents model # the model each agent starts with
jury agents model claude opus # save a default in ~/.jury/config.json
jury agents model claude --reset # back to the CLI's own default
jury review <pr-url> --model claude=sonnet --model codex=gpt-5.5 # this run only
```

A repository can set one per agent in `jury.config.json`:

```json
{ "agents": [{ "name": "claude", "model": "opus" }] }
```

Precedence: `--model <agent>=<model>`, then `model` in `jury.config.json`, then the saved model,
then the CLI's own default. The same model is used on the first run, on resumed rounds and
replies, and when the agent judges. The run records the models it used (`models` in `run.json`),
and `jury agents` shows each agent's model. An agent with a configured model but no way to take
one stops the run before any agent starts rather than silently using its default.

| Agent | How the model is passed | Verified against |
| --- | --- | --- |
| `codex` | `--model` after `exec` (and `exec resume`) | codex-cli 0.156.1 |
| `claude` | `--model` | Claude Code 2.1.281 |
| `grok` | `GROK_MODEL` environment variable | grok-cli README |
| `opencode` | `--model provider/model` after `run` | opencode 1.18.32 |
| `qwen` | `--model` | qwen 0.24.4 |
| `copilot` | `--model` | Copilot CLI 1.0.88 |
| `kimi` | `--model`, an alias defined in `~/.kimi-code/config.toml` | Kimi Code 2.1.1 |
| `amp` | not supported: `--mode` selects model, prompt and tools together | — |
| `droid`, `agy`, `cursor`, `trae` | not yet verified; set the model in the CLI's own configuration | — |

A built-in definition supports models by marking the flag's position with a `"{{modelArgs}}"`
element in every command it runs and giving `modelArgs` (for example `["--model", "{{model}}"]`),
or by giving `modelEnv`. A custom agent can use `{{model}}` directly in its `argv`.

## Adding a built-in agent

Built-in agents are data, not code: each is one JSON file in `lib/agents/`, listed in
Expand Down
22 changes: 22 additions & 0 deletions lib/agent-schema.js
Original file line number Diff line number Diff line change
Expand Up @@ -69,6 +69,12 @@ const FIELDS = {
why: "an object with a boolean `supported`" },
expectSeconds: { required: false, check: (v) => Number.isFinite(v) && v > 0,
why: "a positive number of seconds" },
modelArgs: { required: false, check: (v) => isArgv(v) && v.some((a) => a.includes("{{model}}")),
why: 'the arguments that select a model, naming it as {{model}} (e.g. ["--model", "{{model}}"])' },
modelEnv: { required: false, check: (v) => isEnv(v) && Object.values(v).some((a) => a.includes("{{model}}")),
why: 'environment variables that select a model, naming it as {{model}} (e.g. {"GROK_MODEL": "{{model}}"})' },
modelNote: { required: false, check: (v) => typeof v === "string" && v.trim().length > 0,
why: "why this agent cannot select a model per run, shown when a model is configured for it" },
enabled: { required: false, check: (v) => typeof v === "boolean", why: "a boolean" },
install: { required: false, check: (v) => typeof v === "string" && v.trim().length > 0,
why: "the command that installs this agent" },
Expand Down Expand Up @@ -135,6 +141,22 @@ export function validateAgent(agent, { source = "agent" } = {}) {
problems.push(`${source}: argv[0] must be the executable name, not a template`);
}

// A model slot in one command but not another would switch models between a
// review and its resumed reply. Every variant carries it, or none does.
const variants = [["argv", agent.argv], ["judgeArgv", agent.judgeArgv], ["replyArgv", agent.replyArgv],
["resume.argv", agent.resume?.supported ? agent.resume.argv : null]].filter(([, v]) => isArgv(v));
const slotted = variants.filter(([, v]) => v.includes("{{modelArgs}}"));
if (Object.hasOwn(agent, "modelArgs")) {
for (const [key] of variants.filter((v) => !slotted.includes(v))) {
problems.push(`${source}: modelArgs is set, so ${key} must contain a "{{modelArgs}}" element marking where it goes`);
}
} else if (slotted.length) {
problems.push(`${source}: "{{modelArgs}}" appears in ${slotted.map(([k]) => k).join(", ")} but modelArgs is not defined`);
}
if ((agent.modelArgs || agent.modelEnv) && agent.modelNote) {
problems.push(`${source}: modelNote explains a missing model setting, but this agent has one`);
}

return problems;
}

Expand Down
39 changes: 37 additions & 2 deletions lib/agents.js
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,41 @@ function subst(argv, vars) {
return argv.map((a) => a.replace(/\{\{(\w+)\}\}/g, (_, k) => vars[k] ?? ""));
}

/**
* Put the configured model on a command line.
*
* A definition marks where its model flag belongs with a standalone
* "{{modelArgs}}" element and says what goes there in `modelArgs` (codex:
* ["-m", "{{model}}"]). The slot disappears when no model is set, so a run
* without one is exactly the command it always was, and the agent's own
* default applies. Nothing is ever written to the agent's own config.
*/
export function withModel(argv, agent) {
const model = agent.model;
return argv.flatMap((a) => {
if (a !== "{{modelArgs}}") return [a.replaceAll("{{model}}", model ?? "")];
return model && agent.modelArgs ? agent.modelArgs.map((m) => m.replaceAll("{{model}}", model)) : [];
});
}

/** Environment that carries the model, for agents that take it that way. */
export function modelEnv(agent) {
if (!agent.model || !agent.modelEnv) return {};
return Object.fromEntries(Object.entries(agent.modelEnv).map(([k, v]) => [k, v.replaceAll("{{model}}", agent.model)]));
}

/**
* Whether this invocation can carry a model. Every command variant must: a
* model that reached the first round but was dropped from a resumed one would
* switch models halfway through a conversation.
*/
export function supportsModel(agent) {
if (agent.modelEnv) return true;
const variants = [agent.argv, agent.resume?.supported ? agent.resume.argv : null].filter(Boolean);
const carries = (argv) => argv.some((a) => a === "{{modelArgs}}" && agent.modelArgs || a.includes("{{model}}"));
return variants.length > 0 && variants.every(carries);
}

/**
* Run one reviewer against a worktree. Never throws for a failing agent — a
* dead reviewer is a result, not a crash, and the round should still report the
Expand Down Expand Up @@ -60,7 +95,7 @@ export async function runAgent(agent, { worktree, prompt, stopToken, dryRun, onL
worktree, promptFile: promptFile ?? "", promptText: prompt,
sessionId: sessionId ?? assigned ?? "", packageDir: PACKAGE_DIR,
};
const [cmd, ...args] = subst(commandArgv, vars);
const [cmd, ...args] = subst(withModel(commandArgv, agent), vars);
const cwd = agent.cwd === "worktree" ? worktree : process.cwd();

// The caller already prints the agent's name at the head of this line, so
Expand All @@ -77,7 +112,7 @@ export async function runAgent(agent, { worktree, prompt, stopToken, dryRun, onL
// stdin must be closed, not an open pipe. A pipe that never delivers and
// never ends leaves an agent waiting on input forever: codex sat at 0% CPU
// for over an hour before this was fixed.
child = spawn(cmd, args, { cwd, env: { ...process.env, PWD: cwd, ...agent.env }, stdio: ["ignore", "pipe", "pipe"] });
child = spawn(cmd, args, { cwd, env: { ...process.env, PWD: cwd, ...agent.env, ...modelEnv(agent) }, stdio: ["ignore", "pipe", "pipe"] });
} catch (err) {
resolve({ code: -1, stdout: "", stderr: String(err) });
return;
Expand Down
1 change: 1 addition & 0 deletions lib/agents/agy.json
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,7 @@
],
"idFrom": "\"conversation_id\"\\s*:\\s*\"([^\"]+)\""
},
"modelNote": "no per-run model flag verified for this CLI yet; set its model in the CLI's own configuration",
"report": "agy-json",
"expectSeconds": 420,
"install": "npm install -g @google/antigravity-cli",
Expand Down
1 change: 1 addition & 0 deletions lib/agents/amp.json
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@
"supported": false,
"reason": "execute mode prints no thread id to resume by"
},
"modelNote": "Amp has no per-run model flag: --mode (low, medium, high, ultra) selects the model, system prompt and tools together",
"report": "whole",
"expectSeconds": 400,
"install": "npm install -g @sourcegraph/amp",
Expand Down
7 changes: 7 additions & 0 deletions lib/agents/claude.json
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@
"cwd": "worktree",
"argv": [
"claude",
"{{modelArgs}}",
"-p",
"{{promptText}}",
"--session-id",
Expand All @@ -19,6 +20,7 @@
],
"judgeArgv": [
"claude",
"{{modelArgs}}",
"-p",
"{{promptText}}",
"--permission-mode",
Expand All @@ -32,6 +34,7 @@
"supported": true,
"argv": [
"claude",
"{{modelArgs}}",
"-p",
"{{promptText}}",
"--resume",
Expand All @@ -44,6 +47,10 @@
"{{worktree}}"
]
},
"modelArgs": [
"--model",
"{{model}}"
],
"report": "whole",
"expectSeconds": 900,
"install": "npm install -g @anthropic-ai/claude-code",
Expand Down
Loading
Loading