Skip to content

feat: local models with their real windows; useLocalModel; refreshed demo models (Kimi, Groq, Mistral, OpenRouter, Gemini Nano, transformers.js); fast answers; release 0.0.13 - #19

Merged
siarheidudko merged 3 commits into
mainfrom
claude/agents-review-improvements-2m2406
Oct 5, 2026
Merged

siarheidudko merged 3 commits into
mainfrom
claude/agents-review-improvements-2m2406

Conversation

@siarheidudko

Copy link
Copy Markdown
Member

1. Context window (bug)

Report: a 13B model on the MCP tab failed with "Prompt tokens exceed context window size: 4124; 4096" while the Window setting said 8k.

Cause: WebLLM loads every prebuilt model with a 4096-token window.

Fix:

  • useWebLLMModel({ contextWindowTokens }) loads a model with a larger window. A different window is a different load.

  • contextWindow reports the window the model was loaded with.

  • Demo: each local model loads with its own window through engineConfig.appConfig (withWebLLMContextWindow from agent-web 0.0.22):

    Window Models
    32k Qwen3.5 0.8B / 2B
    16k Qwen3.5 4B / 9B, Llama 3.2 1B
    8k the rest

    Each model's note gives the VRAM it takes.

  • The agent (agent-web 0.0.22) fits compaction, tool lists and tool results into that window.

  • Window defaults to "auto" (the model's own).

2. useLocalModel: any on-device runtime

  • useLocalModel(id, engine) loads any on-device runtime: WebLLM, the browser's built-in model, transformers.js, or your own. It covers progress, switching and unloading.
  • An engine is: create, plus optional warmUp, unload, supported and contextWindowOf.
  • When switching, the previous model is freed by the engine that built it.
  • useWebLLMModel is this hook with createWebLLMEngine. Its API is unchanged.

3. Demo models (October 2026), as chosen in the session

Cloud. Each was checked to answer a page directly (CORS).

Provider Models
Gemini 3.8 Flash, 3.5 Flash-Lite, 3.1 Pro (3.1-flash-lite-preview is shut down)
Claude Haiku 4.5, Sonnet 5.5, Opus 5.5, Fable 5.1
GPT-6 Luna, Sol, Astra
Kimi K2.6, K3
Groq / Cerebras (fast) GPT-OSS 20B/120B, Qwen3.8 27B with images
Mistral Small 4, Medium 3.5
OpenRouter Kimi K3, or any model id you type

Local.

Runtime Models
WebLLM Gemma 3 1B, Llama 3.2 1B/3B, Qwen3.5 0.8B–9B, Ministral 3 3B, Phi-4 mini, Llama 3.1 8B, Phi-3.5 Vision (images)
Chrome built-in Gemini Nano: no download, takes images
transformers.js SmolVLM 256M, Gemma 4 E2B (images; Gemma 4 is marked experimental)

Llama 2 13B (old, 4k window, 12 GB of VRAM) is removed.

Bundle.

  • Every cloud provider is built in demo/src/cloud.ts with a literal dynamic import, so each one is its own chunk fetched on first use. The page and the chess-analyst worker share this builder.
  • Local runtimes load the same way (local-engines.ts).
  • The main bundle shrinks from 7.4 MB to 6.8 MB.
  • transformers.js and its 27 MB wasm load only for its models.

The published library gains no dependencies. The new packages belong to the demo only. The demo's .npmrc skips install scripts: transformers.js pulls in onnxruntime-node, a native Node runtime the page never loads.

4. Speed

"Fast answers" in the agent settings turns off the replanner and the separate final answer: 2 model calls per turn. The step's own reply is the answer (agent-web 0.0.22 fixed the stock "Done — the changes have been applied.").

Checks

  • typecheck, format:check, build, npm test 43/43.
  • Demo typecheck and build. npm run test:e2e 17/17, with 2 new tests:
    • Kimi and a typed OpenRouter model reach their APIs with the user's key and model id;
    • Fast answers makes no synthesizer call.
  • useLocalModel driven in Chromium with fake engines, 18 checks:
    • switch and switch back;
    • free before load;
    • unload;
    • superseded and double loads;
    • a window change is a new load;
    • an engine switch frees the model with its own engine.
  • Chess analysts in a Web Worker build their model through the shared lazy builder: the worker's own calls reach (mocked) Gemini and the agent plays its move.
  • Real WebLLM in Chromium (SwiftShader WebGPU, SmolLM2-360M q4f32):
    • with defaults, the engine reports 4096 and the agent, configured for 128k, uses 4096;
    • loaded with withWebLLMContextWindow(…, 8192), the engine reports 8192.
  • Not verifiable in this sandbox (no WebGPU f16, no built-in model):
    • the q4f16 WebLLM models;
    • Gemini Nano;
    • Gemma 4 E2B.

🤖 Generated with Claude Code

https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5


Generated by Claude Code

claude added 3 commits October 5, 2026 21:24
… release 0.0.13

The demo assumed a window the local models don't have: WebLLM loads every
prebuilt model with 4096 tokens (its default, to save VRAM), so a 13B model
on the MCP tab overflowed ("Prompt tokens exceed context window size: 4124;
4096") while the Window setting said 8k.

- useWebLLMModel({ contextWindowTokens }) passes the window to the factory;
  a different window is a different load (the old engine is freed first);
  `contextWindow` reports the window the model was loaded with.
- Demo: every local model is loaded with its own window through
  engineConfig.appConfig (withWebLLMContextWindow): Qwen3.5 0.8B/2B 32k,
  4B/9B 16k, Llama 3.x 8k, Llama 2 13B its native 4k, with the VRAM each
  takes in its note. The Window setting is "auto" (the model's) by default;
  larger choices say they are capped at the model's. The agent (agent-web
  0.0.22) fits compaction, tool lists and tool results into the window.
- A local multimodal model: Phi-3.5 Vision (WebLLM), 8k window.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
…(Kimi, Groq, Mistral, OpenRouter, Gemini Nano, transformers.js); fast answers

- useLocalModel(id, engine): load any on-device runtime — WebLLM, the
  browser's built-in model, transformers.js or your own — with progress,
  switching (the previous model is freed by the engine that built it) and
  unloading. useWebLLMModel is it with createWebLLMEngine.
- Demo models (October 2026): Gemini 3.8 Flash / 3.5 Flash-Lite / 3.1 Pro
  (3.1-flash-lite-preview is shut down), Claude Haiku 4.5 / Sonnet 5.5 /
  Opus 5.5 / Fable 5.1, GPT-6 Luna / Sol / Astra; new: Kimi K2.6 and K3,
  Groq and Cerebras (fast GPT-OSS, Qwen3.8 with images), Mistral Small 4 and
  Medium 3.5, OpenRouter with any typed model id. Local: Gemma 3 1B, Llama
  3.2 1B, Ministral 3 3B, Phi-4 mini (WebLLM, each with its window), Gemini
  Nano (Chrome's built-in model, no download) and SmolVLM 256M / Gemma 4 E2B
  (transformers.js, images); Llama 2 13B removed.
- Every cloud provider is built in cloud.ts with a literal dynamic import, so
  each is its own chunk fetched on first use (the page and the analyst worker
  share it); local runtimes load the same way (local-engines.ts). The main
  bundle shrinks (7.4 MB → 6.8 MB); transformers.js and its 27 MB wasm load
  only for its models. The demo skips install scripts (.npmrc): transformers.js
  pulls onnxruntime-node, a native Node runtime the page never loads.
- "Fast answers" in the agent settings: no replanner and no separate final
  answer — 2 model calls per turn; the step's own reply is the answer.
- e2e: Kimi and a typed OpenRouter model reach their APIs with the user's key;
  fast answers make no synthesizer call.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
@siarheidudko
siarheidudko marked this pull request as ready for review October 5, 2026 21:56
@siarheidudko
siarheidudko merged commit 6aa12ac into main Oct 5, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants