A production-grade agentic AI configuration for OpenCode — 31 specialised subagents, a hybrid retrieval plane, declarative event hooks, a 14-hub command system, and a 1,086-test suite that asserts on behaviour rather than on source text.
This repository is the agent's entire operating system. It is not a prompt collection: it is a runtime with policy enforcement, an incremental knowledge graph, a token budget, and an evaluation harness that fails when any of them stop working.
Most agent configurations are a markdown file and good intentions. The failure modes are predictable and none of them crash — they are silent:
- A caching layer that writes but is never read, so every turn re-pays for retrieval.
- A knowledge graph that reports
18,000 edgesand contains 88 edges pointing at nodes that do not exist, soimpactanalysis answers with confident nonsense. - A rule that says "verify your own work", which depends on the agent remembering it.
- A skill that four hub commands delegate to, which has no frontmatter and is therefore invisible to the loader.
- A test suite that passes against code that never runs.
Every subsystem here is built to make one of those failures loud. rules/efficiency-first.md
is the standing constraint that shapes the whole design.
┌──────────────────────────────────────────────┐
user input ──────▶│ plugins/hooks/hooks.ts (event interceptor) │
│ 8 hook points · modes · focus · telemetry │
└───────┬──────────────────────────────┬─────────┘
│ │
┌───────────────▼───────────┐ ┌────────────▼─────────────┐
│ retrieval (always local) │ │ plugins/command-hooks/ │
│ graph (SQL) → vectors │ │ declarative shell hooks │
│ → rerank → token budget │ │ injects into next turn │
└───────────────┬───────────┘ └────────────┬─────────────┘
│ │
┌───────▼──────────────────────────────▼─────────┐
│ 31 subagents · 125 skills · 46 tools │
│ 14 hubs / 183 subcommands · 20 rules │
└──────────────────────────────────────────────┘
| Layer | Implementation | Scale (measured) |
|---|---|---|
| Event interception | plugins/hooks/ |
3,987 LOC, 8 hook points |
| Declarative hooks | plugins/command-hooks/ |
1,726 LOC, 7 modules |
| TUI | plugins/hubs-tui/ |
899 LOC |
| Graph retrieval | skills/graph-context/scripts/graphlib.ts |
2,708 LOC |
| Vector retrieval | skills/vectorize-context/scripts/veclib.ts |
1,484 LOC |
| Tooling | tools/ |
46 tools, 11,785 LOC |
| Evals | tests/ |
1,086 tests, 13 suites |
| Knowledge store | SQLite | 3,744 nodes / 22,551 edges |
A local knowledge plane with no inference call on the hot path.
- Graph layer — SQLite entity/relation store, incrementally derived from the symbol sidecar. 3,744 nodes, 22,551 edges, zero dangling edges, zero edgeless nodes. Enforced as invariants, not maintained by hand.
- Vector layer — local ONNX embeddings (384-dim BGE-small), 2,396 code chunks, 1,315 context chunks, 2,660 symbols, 70,868 identifier postings for reverse "which files mention X" lookups.
- Reranking — local BGE cross-encoder, degrading to distance ordering on failure.
- No daemon — embedder, reranker, and gate classifier all run in-process on ONNX Runtime. An unreachable model is a diagnostic, not an empty result set.
- Negative caching — a query that matches nothing is remembered, so a miss is not re-run.
Graph-only recall is pure SQL: < 5 ms, no child process, no model load.
→ memory-system.md ·
knowledge-plane.md
Attach shell commands to tool and session events; inject their output into the agent's context.
Two design decisions carry this:
- Injection is queued, not prompted. Results enter the next turn's system transform via the
session queue. Upstream implementations use
client.session.promptAsync, which costs a full inference request per result. A test asserts the fallback is never taken. injectOn: "failure"means a green check costs zero tokens. This is why the hook gets used rather than switched off.
14 topical hubs, 183 subcommands, two-tier routing. Bare hub → slim identity slice; explicit
/hub subcommand → full spec in one response. One LLM round-trip instead of two.
Every subcommand description is ≤ 80 characters and self-contained, because the TUI dialog shows a flat list with no hub name. Descriptions that name their hub are rejected by a test.
No test in this repository asserts that source text contains a substring. 1,086 tests drive the actual runtime:
- Behavioural — hooks are invoked and their side effects observed.
- Database-level — invariants are SQL queries over a live graph, not regexes.
- Adversarial — a test asserts the injection path makes zero
promptAsynccalls. - Regression-pinned — the flaky wall-clock assertions that measured machine load were replaced with assertions on mechanism.
| Technique | Effect |
|---|---|
| Session-scoped positive + negative cache | A repeat turn costs 0 retrieval calls |
| Long-lived query server | 1 child process serves N queries; no per-call model load |
| Read-only query path | Query never writes — the sync child is the sole writer |
| Node-first runtime resolution | Bun crashes on better-sqlite3; children run under Node by default |
| Per-turn token budget | A reminder under 600 chars instead of a wall of guidance |
| Readonly probe before write lock | No exclusive DB lock when there is nothing to write |
| Graph-only recall path | Pure SQL fallback when vectors are cold |
# Verify the whole system
bun run test:run # 1086 tests
npx tsc --noEmit -p tsconfig.json # config + plugins
npx tsc --noEmit -p skills/tsconfig.json # skills
# Rebuild the knowledge plane
npx tsx skills/graph-context/scripts/graph.ts build
npx tsx skills/graph-context/scripts/regen-index.tsRequires Bun and Node 20+. No daemon and no network after a one-time model fetch:
npx tsx skills/vectorize-context/scripts/prefetch-models.ts # ~340 MB, onceEvery model is then local: embeddings, reranking, and the prompt-queue classifier.
| Page | Covers |
|---|---|
| Architecture | Layer-by-layer design, data flow, extension points |
| Event Interception | plugins/hooks/ — modes, focus, telemetry, caching, vectorize |
| Command Hooks | plugins/command-hooks/ — declarative shell hooks on events |
| TUI Plugin | plugins/hubs-tui/ — native dialogs, generated menu bundle |
| Prompt Queue | plugins/prompt-queue/ — FIFO follow-ups, released only when a turn ends without a question |
| ONNX Runtime | Local embedder, reranker, and gate classifier — no daemon, pooling, measured thresholds |
| Memory System | Durable context, state vs context, the wiki layer |
| Knowledge Plane | Graph + vector retrieval, invariants, incremental rebuilds |
| Orchestration | 31 subagents, delegation, the task tool |
| Hub Command System | 14 hubs, 183 subcommands, two-tier routing |
| Rules & Skills | Rule loading, 125 skills, 46 tools |
| Efficiency | Request and token budgets |
| Testing Strategy | 1,086 behavioural tests and what they protect |
├── agents/ 31 subagent definitions
├── skills/ 125 skills (+ vectorize-context, graph-context)
├── rules/ 20 rules — 11 preloaded, 9 on demand
├── tools/ 46 TypeScript tools
├── plugins/ hooks · command-hooks · prompt-queue · hubs-tui
├── tests/ 1,086 tests across 13 suites
├── command-hooks.jsonc global declarative hooks
├── .documentation/ 14 pages — architecture, plugins, memory, evals
├── opencode.jsonc plugin, MCP, instructions, references
├── .documentation/ this documentation set
└── .opencode/ state (ephemeral) + context (durable)
- Nothing loads silently. A skill, rule, agent, or subcommand that is referenced but unreachable fails a test. Four were found and fixed this way.
- A hook never breaks a tool call. Hook failures are contained and reported.
- Derived state is verifiable. Graph edges are checked against live nodes; a rebuild that changes structure without a reason is impossible by construction.
- Every regeneration is idempotent. A settled rebuild reports zero work; a second build skips. A derivation-rule change is versioned so it cannot skip.
- Tests observe effects. If a subsystem could be broken and the suite stay green, the test is the bug.
This configuration was built and hardened against its own failure modes. Numbers quoted above
are measured by the test suite or by graph.ts stats, not estimated.
Keywords: OpenCode plugin, agentic AI, agent harness, agent orchestration, multi-agent systems, hybrid retrieval, RAG, knowledge graph, vector search, MCP (Model Context Protocol), context engineering, prompt caching, token optimization, LLM evals, TypeScript, Bun, SQLite, deterministic tooling.
{ "id": "typecheck-after-task", "when": { "phase": "after", "tool": "task" }, "run": "npx tsc --noEmit", "injectOn": "failure", "inject": "Type errors after the subagent finished:\n{stdout}{stderr}" }