Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,3 +1,13 @@
# 0.0.13 / 2026-10-05

### :tada: Enhancements
- `useLocalModel(id, engine)`: load any on-device runtime (WebLLM, the browser's built-in model, transformers.js, your own) with progress, switching and unloading; `useWebLLMModel` is it with `createWebLLMEngine`
- `useWebLLMModel({ contextWindowTokens })` loads a local model with a larger window than WebLLM's 4096 (a different window is a different load); `contextWindow` reports the loaded window
- Demo — windows: every local model is loaded with its own window and the agent fits its runs into it — no more "Prompt tokens exceed context window size"; the Window setting is "auto" (the model's) by default
- Demo — models: current Gemini (3.8 Flash, 3.5 Flash-Lite, 3.1 Pro), Claude (Haiku 4.5, Sonnet 5.5, Opus 5.5, Fable 5.1) and GPT-6 (Luna, Sol, Astra); new providers Kimi (K2.6, K3), Groq and Cerebras (fast), Mistral, OpenRouter (any model id); local Gemma 3 1B, Llama 3.2 1B, Ministral 3 3B, Phi-4 mini, Phi-3.5 Vision (WebLLM), Gemini Nano (Chrome's built-in model) and SmolVLM / Gemma 4 E2B (transformers.js); Llama 2 13B removed. Every provider and runtime loads on first use.
- Demo — speed: "Fast answers" (no replanner and no synthesizer: 2 model calls per turn)
- Updated dependencies: @dudko.dev/agent-web 0.0.22 (runs fitted to the model's window, local vision models, provider capabilities)

# 0.0.12 / 2026-10-03

### :tada: Enhancements
Expand Down
3 changes: 2 additions & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,8 @@ The repo also contains a **Vite demo** (`demo/`) that is auto-deployed to
- `src/state.ts` — the pure `agentStateReducer` + `createInitialAgentState`.
- `src/types.ts` — `AgentUiState`, `ChatMessage`, `StepView`, `ToolCallView`, …
- `src/hooks/` — `use-agent` (the hook), `use-credentials` (vault),
`use-webllm-model`, `use-mcp` (remote MCP + OAuth round-trip),
`use-local-model` (any on-device runtime as an engine) + `use-webllm-model`
(its WebLLM engine), `use-mcp` (remote MCP + OAuth round-trip),
`use-mcp-servers` (several servers), `use-chat-history`, `use-virtual-files`,
`use-speech-to-text`.
- `src/chat-history.ts` — `ChatHistoryStore` (raw IndexedDB, memory fallback).
Expand Down
46 changes: 45 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,9 @@ app with a single hook and (optionally) a set of pre-styled components:
`useVirtualFiles`, `useSpeechToText`, `useMcpServers`.
- 🔑 **`useCredentials`** — store BYOK API keys **encrypted at rest**
(WebCrypto + IndexedDB).
- 🖥️ **`useWebLLMModel`** — load a local WebGPU model with download progress.
- 🖥️ **`useWebLLMModel`** / **`useLocalModel`** — load an on-device model (WebLLM,
the browser's built-in model, transformers.js, or your own runtime) with
download progress, switching and unloading.
- 🎛️ **Headless-first** — the event→UI logic is a pure, exported reducer
(`agentStateReducer`); the components are optional sugar on top.

Expand Down Expand Up @@ -229,6 +231,48 @@ function LocalAgent() {
> yet. The previous model stays in memory (`loadedModelId`; switching back is
> instant) until the new one is loaded — that frees it first, so two models
> never share the GPU — or until `unload()`.
>
> **Context window.** WebLLM loads its models with a 4096-token window, far
> below what Qwen3 or Llama 3.x were trained on. Pass `contextWindowTokens` to
> load with more (it costs KV-cache VRAM, not a new download); your `create`
> factory receives it and applies it with the core's `withWebLLMContextWindow`
> (see the demo's `providers.ts`). `local.contextWindow` is the window the model
> was loaded with, and the agent fits its runs into it on its own — compaction,
> tool lists and tool results are sized from it.
>
> **Vision.** `Phi-3.5-vision-instruct-q4f16_1-MLC` takes images (paste, drop
> or attach them) — a multimodal model that never sends them anywhere.

### Any on-device runtime — `useLocalModel`

`useWebLLMModel` is `useLocalModel` with the WebLLM engine. Describe another
runtime as an engine — `create`, and optionally `warmUp` (download now, with
progress), `unload`, `supported`, `contextWindowOf` — and the hook gives the
same `load` / `unload` / progress / switching:

```tsx
import { useLocalModel, createWebLLMEngine, type LocalModelEngine } from '@dudko.dev/agent-web-react'

const builtIn: LocalModelEngine = {
create: async () => (await import('@browser-ai/core')).browserAI('text', {
expectedInputs: [{ type: 'text' }, { type: 'image' }],
}),
warmUp: async (m, { onProgress }) => {
await m.createSessionWithProgress((p) => onProgress({ progress: p, text: 'Downloading' }))
},
supported: () => 'LanguageModel' in globalThis, // Chrome's Prompt API
contextWindowOf: (m) => m.getContextWindow(),
}
const webllm = createWebLLMEngine({ create: createLocalModel })

// One hook, any runtime: switching frees the previous model with its own engine.
const local = useLocalModel(option.model, option.builtIn ? builtIn : webllm, {
contextWindowTokens: option.contextWindow,
})
```

The demo runs WebLLM, Chrome's Gemini Nano and transformers.js this way — see
[`demo/src/local-engines.ts`](demo/src/local-engines.ts).

## Models in a bundler (Vite, Next, CRA)

Expand Down
4 changes: 4 additions & 0 deletions demo/.npmrc
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
# @huggingface/transformers depends on onnxruntime-node (a native Node runtime the
# browser never loads) whose install script downloads binaries; nothing the demo
# uses needs an install script, so skip them all.
ignore-scripts=true
23 changes: 21 additions & 2 deletions demo/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,17 @@

A Vite + React app showcasing [`@dudko.dev/agent-web-react`](../): an in-browser
LLM agent driving tools. Pick a cloud model (bring your own key, stored
encrypted) or load a local WebGPU model — everything runs in the browser.
encrypted) or an on-device one — everything runs in the browser.

**Models** (October 2026): Gemini (3.8 Flash, 3.5 Flash-Lite, 3.1 Pro), Claude
(Haiku 4.5, Sonnet 5.5, Opus 5.5, Fable 5.1), GPT-6 (Luna, Sol, Astra), Kimi
(K2.6, K3), Groq and Cerebras for speed (GPT-OSS, Qwen3.8 with images), Mistral
(Small 4, Medium 3.5) and OpenRouter (type any model id). Every one of them
answers a page directly with your key (CORS). On the device: WebLLM (Gemma 3 1B,
Llama 3.2 1B/3B, Qwen3.5 0.8B–9B, Ministral 3 3B, Phi-4 mini, Llama 3.1 8B and
Phi-3.5 Vision for images), Chrome's built-in Gemini Nano (no download), and
transformers.js (SmolVLM 256M and Gemma 4 E2B, images). Each provider and
runtime is fetched the first time you pick one of its models.

Three tabs, sharing one "Agent settings" panel:

Expand Down Expand Up @@ -44,10 +54,19 @@ and in total, thoughts, subagents and consent prompts; the composer has
attachments, "/" commands, a run timer, the running-agents count, the model +
thinking chip, the consent chip and a speech-to-text mic.

For quick answers: a fast provider (Groq, Cerebras, Flash-Lite), a small local
model with Thinking "none", and **Fast answers** in the agent settings (no
replanner, no separate final answer: 2 model calls per turn instead of 3+).

Local models: pick one and press "Download & load". Switching to another one
shows it as not loaded (the loaded one stays in memory, so switching back is
instant); loading it frees the previous model's GPU memory first, and
"Unload" frees it on demand.
"Unload" frees it on demand. Each is loaded with its own context window —
WebLLM's default is 4096 tokens; Qwen3.5 0.8B/2B get 32k, 4B/9B 16k, Llama 3.2
1B 16k, the rest 8k (each note says the VRAM it takes) — and the agent
fits its runs into it: compaction, tool lists and tool results are sized from
the window ("Window: auto"). **Phi-3.5 Vision** is a local multimodal model:
paste or drop an image and ask about it.

**Live:** https://dudko-dev.github.io/agent-web-react/

Expand Down
Loading
Loading