Skip to content
nlink-jpPublic

About

CLI agent runtime on a local LLM (LM Studio, mlx-serve, Ollama) for work that should stay off a cloud API — the same minimal, auditable loop as gem-agent, defended by sandbox-exec containment plus operator approval. Drop-in with existing projects: it reads their AGENTS.md, CLAUDE.md and .mcp.json unchanged

Resources

Contributing

Stars

0 stars

Watchers

1 watching

Forks

Repository files navigation

lagent

Sandboxed CLI agent runtime on a local LLM, served over an OpenAI-compatible API (LM Studio, mlx-serve, Ollama). File read/write, sandboxed shell commands and MCP servers, with mutating calls gated by the operator.

lagent is for work that should not go to a cloud API: offline environments, confidential projects, and work where per-token spend is unwanted. It is a separate product line from gem-agent, built on the same design: the same auditable minimal loop (read / edit / shell / MCP / approval), the same drop-in reading of a project's AGENTS.md / CLAUDE.md / .mcp.json and Claude Code-format skills, the same session records.

Released. Install with brew install nlink-jp/tap/lagent or from the releases page (Developer ID signed, Apple-notarized, macOS arm64) — the releases page is the authority on the current version. It is a cli-series tool: interface stability is a promise, and breaking changes go through the org's breaking-change process (ADR-0023). See the RFP for the specification.

Japanese: README.ja.md

Requirements

  • macOS on Apple silicon (isolation is built on sandbox-exec)
  • A local LLM server with an OpenAI-compatible API: LM Studio (the tested backend, model google/gemma-4-26b-a4b-qat), mlx-serve (provider = "mlxserve"; measured with Qwen 3.6 35B-A3B — bind it to 127.0.0.1 or set its API key, since it listens on every interface by default, and raise its prefix cache memory cap (8 GB measured) so long sessions are not re-read every turn: ADR-0031) or Ollama
  • No credentials

Configuration

~/.config/lagent/config.toml — see config.example.toml. Precedence: flags > LAGENT_* environment > file > defaults; unknown keys are errors.

[llm]
provider = "lmstudio"        # lmstudio | mlxserve | ollama | openai
base_url = "http://localhost:1234/v1"
model    = "google/gemma-4-26b-a4b-qat"

[model]
context_window = 0           # 0 = detect from the provider

Quickstart

Start LM Studio, load the model, and run lagent in the project directory:

cd ~/work/my-project
lagent

The first start asks whether to trust the project's own AGENTS.md / CLAUDE.md / .mcp.json. Mutating tools ask before running; --auto lets rule-tier Safe calls run unasked, and with [approval].model_tier = "shell" a judge model decides write-lane shell commands too (about half of them ran unasked on a month of real prompts, and no exfiltration or injection case passed — ADR-0032); -p "…" runs one prompt and exits. /help lists the slash commands.

What it does

  • Tools: list_files, list_tree, search_files, read_file, file_info, view_image, show_image, write_file, edit_file, shell_exec, ask_user, and every tool of the MCP servers in .mcp.json — shown to the model as a catalog, and advertised per server once it calls mcp_load (a local model cannot afford 243 schemas on every turn; [mcp].preload and [mcp].advertise = "all" are the operator's levers).
  • Partial results are reachable: read_file reads by line window or by byte offset/length (a negative offset counts from the end), so the tail of a long single-line file — a tool result saved to the work directory — is in reach; a read cut at 200 KB says where to read on (ADR-0029). An MCP result too large to hold inline is saved whole to the session work directory and previewed by its first 600 and last 200 characters — metadata a server appends, such as "truncated": true, comes last — with the byte spans shown and the route to the rest. shell_exec output past 20,000 bytes keeps its first 15,000 and last 5,000 bytes — a script's totals come last — and is saved whole to the work directory (up to 32 MiB, private; not in the operator lane or without the sandbox, where commands may read credentials). search_files counts every file it did not search — over 2 MB (named), binary, image, unreadable — so "no match" says what was not looked at.
  • Confinement: file tools stay inside the project (and the session work directory), and a credential file — .env, a private key, a token store — is read only when you approve it, every time, never unattended: those reads run in a sandboxed child that cannot open credential material at all, so the list raises the question and the kernel is what refuses; shell_exec runs under sandbox-exec in the lane it declares — read runs unasked (inspection, and Go builds and tests, whose cache lives in the session scratch), write and operator ask.
  • Terminal safety: text the model writes, and the other text from outside the runtime the TUI shows, reaches the terminal with its control characters removed, so a prompt-injected reply cannot retitle the window, write the clipboard, move the cursor or disguise a command in an approval dialog; -p output stays byte-for-byte to a pipe or a file (ADR-0024).
  • Hooks: [[hooks.pre_tool_use]] runs your guard script before a model tool call, on Claude Code's contract, so the same script that guards Claude Code guards this runtime; a deny is final, whatever the approval mode, and the reason goes back to the model (ADR-0012). session_start, user_prompt_submit and session_end hooks share the mechanism: their output reaches the model as quoted data, and a prompt hook can refuse a prompt (ADR-0014).
  • Memory: /remember <name> <fact> keeps a short fact for every later session in this project (global for every project); the model can propose one with save_memory, which asks you. Memories are recalled in the runtime facts, the place this model acts on, so a memory that names a file or a command works as a pointer (ADR-0013).
  • Sessions: a JSONL transcript per session; --continue and --resume; usage records in the shape gem-usage-lens reads for both runtimes.
  • Skills: Claude Code's SKILL.md format, read as-is from ~/.config/lagent/skills/<name>/ and a trusted project's .claude/skills/<name>/. One line per skill tells the model what each is for; it loads one with load_skill, and you invoke one by hand with /skill <name> (ADR-0011). A skill arrives as instructions; everything a tool returns — a knowledge-vault server's notes included — arrives as data the model is told never to obey, so a procedure you want followed belongs in a skill, not in a note read through MCP (gem-agent's integration reference, "Where an operator's procedures go").
  • Inline images: the model shows you a picture with show_image, and you ask for one with /show <path> (PNG or JPEG, up to 2 MiB; the path may hold spaces, be quoted, or be a file dragged into the window) — drawn in the terminal when it can draw (iTerm2 or kitty; [tui].images, auto by default, off inside tmux and screen). view_image is the other direction: that is the model looking at an image, not you (ADR-0020/0021/0022).
  • Diagrams as pictures: a mermaid fence in a reply (flowchart / graph, sequenceDiagram, erDiagram, pie, stateDiagram, gantt, mindmap, CJK labels included) is drawn as a picture where the terminal draws images, one em of its text per terminal line; anything that cannot be drawn right stays source with a one-line note. Elsewhere flowchart, sequence and ER fences are drawn as box art by the same engine, and the other types stay source (ADR-0027). The plain REPL and -p always show the source. [tui.diagram] picks the font. Narrowing the window keeps the pictures on screen (ADR-0026).
  • Not here: web search and fetch, media uploads, audit-log export, history compaction — see the RFP and ADR-0002.

Attachments

@<path> attaches a project file or directory; @<image> attaches an image from anywhere (absolute or ~ paths), and an image path dropped on the terminal attaches without the @ — escaped spaces included. A reference that cannot be read is reported to you and told to the model.

Install

Apple Silicon Mac, via the nlink-jp Homebrew tap (the signed and notarized release archive, installed as-is):

brew tap nlink-jp/tap
brew install nlink-jp/tap/lagent

Build

make build      # → dist/lagent
make test
make check      # vet + lint + test + docs mirror + release gate + build

Documentation

License

MIT

About

CLI agent runtime on a local LLM (LM Studio, mlx-serve, Ollama) for work that should stay off a cloud API — the same minimal, auditable loop as gem-agent, defended by sandbox-exec containment plus operator approval. Drop-in with existing projects: it reads their AGENTS.md, CLAUDE.md and .mcp.json unchanged

Resources

Contributing

Stars

0 stars

Watchers

1 watching

Forks

Releases

Contributors

Languages