Repository navigation
feat: local models with their real windows; useLocalModel; refreshed demo models (Kimi, Groq, Mistral, OpenRouter, Gemini Nano, transformers.js); fast answers; release 0.0.13 - #19
Merged
Conversation
… release 0.0.13
The demo assumed a window the local models don't have: WebLLM loads every
prebuilt model with 4096 tokens (its default, to save VRAM), so a 13B model
on the MCP tab overflowed ("Prompt tokens exceed context window size: 4124;
4096") while the Window setting said 8k.
- useWebLLMModel({ contextWindowTokens }) passes the window to the factory;
a different window is a different load (the old engine is freed first);
`contextWindow` reports the window the model was loaded with.
- Demo: every local model is loaded with its own window through
engineConfig.appConfig (withWebLLMContextWindow): Qwen3.5 0.8B/2B 32k,
4B/9B 16k, Llama 3.x 8k, Llama 2 13B its native 4k, with the VRAM each
takes in its note. The Window setting is "auto" (the model's) by default;
larger choices say they are capped at the model's. The agent (agent-web
0.0.22) fits compaction, tool lists and tool results into the window.
- A local multimodal model: Phi-3.5 Vision (WebLLM), 8k window.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
…(Kimi, Groq, Mistral, OpenRouter, Gemini Nano, transformers.js); fast answers - useLocalModel(id, engine): load any on-device runtime — WebLLM, the browser's built-in model, transformers.js or your own — with progress, switching (the previous model is freed by the engine that built it) and unloading. useWebLLMModel is it with createWebLLMEngine. - Demo models (October 2026): Gemini 3.8 Flash / 3.5 Flash-Lite / 3.1 Pro (3.1-flash-lite-preview is shut down), Claude Haiku 4.5 / Sonnet 5.5 / Opus 5.5 / Fable 5.1, GPT-6 Luna / Sol / Astra; new: Kimi K2.6 and K3, Groq and Cerebras (fast GPT-OSS, Qwen3.8 with images), Mistral Small 4 and Medium 3.5, OpenRouter with any typed model id. Local: Gemma 3 1B, Llama 3.2 1B, Ministral 3 3B, Phi-4 mini (WebLLM, each with its window), Gemini Nano (Chrome's built-in model, no download) and SmolVLM 256M / Gemma 4 E2B (transformers.js, images); Llama 2 13B removed. - Every cloud provider is built in cloud.ts with a literal dynamic import, so each is its own chunk fetched on first use (the page and the analyst worker share it); local runtimes load the same way (local-engines.ts). The main bundle shrinks (7.4 MB → 6.8 MB); transformers.js and its 27 MB wasm load only for its models. The demo skips install scripts (.npmrc): transformers.js pulls onnxruntime-node, a native Node runtime the page never loads. - "Fast answers" in the agent settings: no replanner and no separate final answer — 2 model calls per turn; the step's own reply is the answer. - e2e: Kimi and a typed OpenRouter model reach their APIs with the user's key; fast answers make no synthesizer call. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
siarheidudko
marked this pull request as ready for review
October 5, 2026 21:56
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
1. Context window (bug)
Report: a 13B model on the MCP tab failed with "Prompt tokens exceed context window size: 4124; 4096" while the Window setting said 8k.
Cause: WebLLM loads every prebuilt model with a 4096-token window.
Fix:
useWebLLMModel({ contextWindowTokens })loads a model with a larger window. A different window is a different load.contextWindowreports the window the model was loaded with.Demo: each local model loads with its own window through
engineConfig.appConfig(withWebLLMContextWindowfrom agent-web 0.0.22):Each model's note gives the VRAM it takes.
The agent (agent-web 0.0.22) fits compaction, tool lists and tool results into that window.
Window defaults to "auto" (the model's own).
2.
useLocalModel: any on-device runtimeuseLocalModel(id, engine)loads any on-device runtime: WebLLM, the browser's built-in model, transformers.js, or your own. It covers progress, switching and unloading.create, plus optionalwarmUp,unload,supportedandcontextWindowOf.useWebLLMModelis this hook withcreateWebLLMEngine. Its API is unchanged.3. Demo models (October 2026), as chosen in the session
Cloud. Each was checked to answer a page directly (CORS).
3.1-flash-lite-previewis shut down)Local.
Llama 2 13B (old, 4k window, 12 GB of VRAM) is removed.
Bundle.
demo/src/cloud.tswith a literal dynamic import, so each one is its own chunk fetched on first use. The page and the chess-analyst worker share this builder.local-engines.ts).The published library gains no dependencies. The new packages belong to the demo only. The demo's
.npmrcskips install scripts: transformers.js pulls in onnxruntime-node, a native Node runtime the page never loads.4. Speed
"Fast answers" in the agent settings turns off the replanner and the separate final answer: 2 model calls per turn. The step's own reply is the answer (agent-web 0.0.22 fixed the stock "Done — the changes have been applied.").
Checks
npm test43/43.npm run test:e2e17/17, with 2 new tests:useLocalModeldriven in Chromium with fake engines, 18 checks:withWebLLMContextWindow(…, 8192), the engine reports 8192.🤖 Generated with Claude Code
https://claude.ai/code/session_01A7jHa5Pim4ufrdgxg561G5
Generated by Claude Code