Skip to content

perf: faster browser tasks: settling, hedged and warmed Jev calls, quorum decisions, early pages, fewer rescues (openhuman#7000) - #77

Merged
senamakel merged 44 commits into
tinyhumansai:mainfrom
YellowSnnowmann:perf/7000-faster-browser-tasks
Oct 8, 2026
Merged

senamakel merged 44 commits into
tinyhumansai:mainfrom
YellowSnnowmann:perf/7000-faster-browser-tasks

Conversation

@YellowSnnowmann

@YellowSnnowmann YellowSnnowmann commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This makes browser tasks faster without losing reliability, for tinyhumansai/openhuman#7000. Each change was measured on live runs before it went in. The branch then had a final review, whose fixes are 13 commits, and a live run of ten utility sites, whose one regression is fixed in the commit after them.

Since then, Stage 4 adds four more speed changes, each first measured by replaying the journals of 245 live runs, and two jev_journal fixes. A second review, by an Opus model, found no high-severity defect in them; its fixes are the last five commits. Stage 4 has not been run live yet: an A/B of the module before and after it is the next step.

Measuring (Stage 0)

  • The journal now records planning, rescues and waits for a person, so a task's whole time is accounted for.
  • jev_journal --split breaks a task's time down into planning, Jev, settling, actions and rescues.
  • jev_journal --compare sets two batches side by side.

Less waiting (Stage 1, now default)

  • browser.settle = "prompt": after an action, the page is read once the requests that change it are done and it stops changing, about 0.6 s on an idle page instead of 1.6 s.
  • browser.prelaunch = true: a browser-only task's browser opens while its plan is drafted, which takes the first launch from 2.8 s to 0.6 s.
  • "steady" and false restore the old behaviour.
  • A switch to plan without reasoning was tried and removed: every extra rescue in its A/B traced back to those plans.

Shorter Jev waits (Stage 2)

  • Hedged framings. A framing still running after 4 s (5 s for a request of 32 KB or more) gets a copy, and the first answer counts. On 7 October, 21 of 36,459 calls stalled 12–32 s until the gateway or the 30 s timeout gave up, while their siblings answered in ~0.6 s; roughly one run in four hit one.
  • 10 s attempt timeout. A Jev attempt now gives up after 10 s unless jev.timeout_ms says otherwise.
  • Network quiet counts only page changes. Settling counts only the requests that can change the page, for at most 1 s (agent-browser networkquiet). Amazon waited ~2.4 s after every page-opening action: 100+ requests over 2 s, some of which never report finishing to the page's session. Its content shows after 0.7–1.2 s, and settling now takes 0.67–1.47 s a page load.
  • Planner retries. Planner, rescuer and shaper calls are tried again on a passing failure (1, 2, 4 s apart). One HTTP 502 had failed a whole task at its plan.

Fixes from the live checks

  • Late suggestion waits. A place box's wait for late suggestions now ends as soon as the page changes (Surface::await_change). A still page ends the waiting: BlazeDemo's address and city boxes went from four waits of 1.2–2.1 s to two of 1.0 s.
  • Shadow-root banners. Sight now reads one shadow root's controls from the tree under its host, labelled with the layer they draw. Lenskart's consent banner sat in a display: contents host that sight never saw.
  • Covered presses. A press the page refuses as covered is answered by pressing the front layer's least committal control ("Allow Selection"), else Escape, as before.
  • Price of one item. The flow guide tells planners to read the price of one item before raising its count. A cart line then shows the total for both.

Fewer rescues (Stage 3)

  • Already-done check. A do step whose last three actions changed nothing first asks whether the screen already shows its result, and ends AlreadyDone when it clearly does. 25 of 189 rescues on 7 October were such steps, ~18 s each.
  • Element memory in task_live. task_live remembers learned elements between runs (TASK_MEMORY); it is off unless set.

Fixes from the final review

  • networkquiet waits for what the action sent (agent-browser). A click's own navigation or fetch starts before a separate wait can subscribe, so the wait never saw it, and a page slower than 500 ms could be read as it was before the click. agent-browser now keeps the page's unfinished page-changing requests, and the wait starts with those sent in the last 3 s.
  • Deadlines on the settle and change watches. Their scripts had no deadline of their own, sent right after actions that replace pages; such a call waited out the browser's 30 s live. Both watches now also see changes inside open shadow roots.
  • Shadow hosts read once. The tree's reading of a shadow host also holds the host and its slotted content, which sight had read too; sight now leaves them to the tree.
  • The covered target's own layer is never closed, nor a layer the step names.
  • Hedging: no copy for Sage (its calls take 5–6 s as a rule), at most two copies in flight, and the journal names the answer that counted.
  • Early browsers let go while a task waits for values, and never opened for a task cancelled first.
  • The stall check asks only with the completion loop on.
  • The journal builds nothing when it is off, journals a resume only once it is accepted, and times a plan on the planner alone.
  • task_live makes the memory file's folder, and refuses switch values it cannot read.
  • jev_journal's split.rs and task_live/main.rs are split under 400 lines, and stale docs are brought up to date.

Fix from the live run of the utility sites

  • A place box waits past its own rows. Live on Uber, the pickup box first listed rows of its own ("Allow location access", "Search in a different city"). The prompt settle reads the page as soon as it goes still, so it looked before the matches were fetched, and a box with any new row was not looked at again: the pickup was never set. A place box is now looked at again, still up to LATE_LOOKS waits that each end at the page's first change, until a row names the text typed.

Waiting less for the first decision and the first page (Stage 4)

  • Settle briefly after a launch or Escape. These fetch nothing: live, 3% of launches and 12% of Escapes settled with a page request still running, against 39% of fills and 70% of clicks. They now wait only while the page changes, 120–400 ms instead of about 0.65 s (Surface::settle_briefly; the desktop and steady settling settle in full as before). Typing, Return and every other action keep the full settle.
  • Warm Jev's connections while the plan is drafted. A task's first decision took 330 ms more a call than its later ones (median over 84 runs), opening a connection per framing. JevRuntime::warm now opens one for each call of a step's first turn (its judging and grounding's opening, each in every framing: 2 × votes, at most 18) with a one-question evaluation, journaled as warm-up. The task never waits for it, Sage is not warmed, nor is a task whose budget caps its Jev calls.
  • End a decision once all but two framings agree plainly. Asked 7 ways or more, a decision is merged without its last 2 framings once those in settle every question: the same option first in every Choice and Score at 0.9 or more, every yes/no at 0.97 or more (or 0.08 or less) in each, and the page kind alike. Replayed over 13,669 decisions from 245 live runs, this ended 12% of them, a median 0.14 s sooner, about 2 s a run, and every one read the same on every threshold and evidence gate as all seven did, but for two whose mean moved under 0.002 across the band's edge. A quorum of four changed what was read in 0.8% of the decisions it ended, so five it is. The framings left run to their end and count as calls.
  • Load the page a task goes to while its plan is drafted. The first page took 1.9 s to load (p90 6.4 s) and 1.1 s to settle, after a 10–30 s plan; in 56 of 57 live plans the first step browsed exactly the page the task named. When a browser-only task's words send its browser to one address ("go to", "open", "visit", "start at", "on", …) with no query or fragment, it now loads in the early browser in the same wait, for at most 10 s, journaled as open_page; the first step to the same place, in that session, finds it loaded. It rides on browser.prelaunch.
  • jev_journal: --split and --compare price Jev's input tokens at $0.042 per million (jev_cost_micro_usd), as the evals do; a quorum's late framings and the warm-up's calls count in no round.

Together, by those replays, about 6.5–7 s a run.

Fixes from the second review

  • A decision a quorum ended now widens past every framing it asked. Widening counted the framings asked by the ballot, which a quorum leaves two short: before an irreversible press, which always widens, it asked the sixth and seventh framings again and never the eighth and ninth. A regression test fails without the fix.
  • The warm-up covers a whole first turn, as above, and is skipped under a call cap.
  • --split matches a quorum's late framings to their decision by question, rather than taking the next exchanges whichever decision they were, and leaves the warm-up out of the call statistics.
  • The early page loads only where the task goes, not an address it only mentions or one whose query could carry a one-time token; within 10 s; its mark names its session; the reuse check reads within READ_TIMEOUT.
  • Docs that still said every action settles in full, and three tests that passed whatever the code did, are fixed.

Checked and skipped, with data

Idea Why skipped
Return early on a decided vote The last framing costs 0.05 s a round. Stage 4's quorum ends only decisions the first five framings settle plainly, where waiting is longest
Batch grounding every turn ~60% of the speculative calls would be wasted
Smaller requests Repeated text is 7% of the bytes
Adaptive framings Matches 7 framings at threshold level only 99.34% of the time, and saves calls, not time
Send only the relevant part of a big page when finding an element Saved 0.02 s a call, and answers changed in 1–2% of cases, beyond Jev's own variation
One request a turn (judge and target search together, every turn) Only 28% of later turns go on to a target search; batching cost 0.7 s more a round and about 4× the tokens
Decide while the page settles Only 24% of clicks and 41% of fills settled on an idle page; at best 1.2 s a run, for about 47 more Jev calls (+15%) that slow the real ones

Related issue

Refs tinyhumansai/openhuman#7000.

Depends on tinyhumansai/agent-browser#2. The vendor/agent-browser gitlink pins its head, ebf10fb, which CI fetches through that pull request. It must merge first, by a merge commit; if it is squash-merged, the gitlink must move to the merged commit before this one merges.

API or behavior changes

  • Module configuration: new browser.settle (prompt default, or steady) and browser.prelaunch (true default). jev.timeout_ms now defaults to 10 s, from the client's 30 s.
  • Behaviour:
    • prompt settling, prelaunch, and hedged framings;
    • late suggestion waits end on a still page;
    • covered presses close a front layer's closer before Escape;
    • stalled steps can end AlreadyDone;
    • hosted model calls retry passing failures.
  • Surface trait: new await_change (pause and assume a change) and settle_briefly (settle in full) methods with defaults, so existing implementations compile unchanged.
  • FlowRunner: new warm and open_page, which do nothing by default; JevRuntime::warm(votes); BrowserSurface::open_at(url).
  • Stage 4 behaviour: a brief settle after a launch or Escape; Jev warmed while a task is planned; decisions ended on a quorum of five plain framings; the page a browser-only task goes to loaded while it is planned.
  • FlowRunner::journal takes its fields as a closure, built only when written; JevRuntime::journal_event likewise, and JevRuntime::journaling says whether the journal is on.
  • Journal: new plan, rescue, resume, hedge and open_page events; a wait action notes "nothing changed"; a decision says how many framings it did not wait for (left); the warm-up's exchanges are journaled under the step warm-up.
  • jev_journal: --split and --compare, with Jev's cost.
  • Contract docs only: TaskBudget.max_model_calls notes that a capped task is not warmed, and JevMetrics says which counts take in the framings a quorum did not wait for.
  • task_live: TASK_PLAN=in-task, TINYCOMPUTER_BROWSER_SETTLE, TINYCOMPUTER_BROWSER_PRELAUNCH, TASK_MEMORY.

None of this breaks the wire contract.

Validation

Commands run on the final commit (Rust 1.98, macOS), with their outcomes; the first review's round ran on 1.99:

  • cargo fmt --all -- --check: clean.
  • cargo clippy --workspace --exclude tinycomputer-accessibility --all-targets --all-features -- -D warnings: clean. tinycomputer-accessibility already fails clippy 1.99 on macOS on main (its macOS-only files), and this branch does not touch it.
  • cargo build --workspace --exclude tinycomputer-accessibility --all-targets --all-features: ok (built by the test and coverage runs).
  • cargo test --workspace --exclude tinycomputer-accessibility --all-features: 1,164 passed, 0 failed.
  • RUSTDOCFLAGS="-D warnings" cargo doc --no-deps --all-features --workspace --exclude tinycomputer-accessibility: clean.
  • cargo llvm-cov on core, bus, browser, engine and the module crate, with the CI gate's per-file filter: all 183 files at or above 90%; every source file changed since main at 100% but escalate/mod.rs (98.4%, a defensive fallback).
  • Live fixture tests (TINYCOMPUTER_LIVE_BROWSER=1, real Chrome): all 11 sight live tests pass, including the shadow-root banner and its slotted content.

Live A/B batches (task_live, planned in-task as OpenHuman does):

Batch Result
Stage 1, 16 runs Wait per action 1.69 → 1.28 s; launch 2.77 → 0.64 s. No failure traced to either switch: they were the Myntra irreversible gate, a store's own error page, and planner variance
Stage 2, 12 runs Amazon 2/2 in 61–66 s (was 69–89 s); BlazeDemo 2/2; Zepto 2/2; median wait per action 0.65 s. Myntra failed on the irreversible gate (kept by decision); Blinkit on its own error page and a log-in wall. Lenskart failed on the shadow-root read that the agent-browser pierce fix addresses
Stage 3, 10 runs 7/10, median 101 s a run, 0 rescues, launch 0.64 s. Amazon 2/2 at checkout (63–77 s); Zepto 2/2 at its sign-in wall; Lenskart 1/2 (90 s, through the consent banner); BlazeDemo 1/2; Blinkit 1/2
Ten utility sites, final code (headed, a person at the terminal) Worked: Amazon (checkout, 87 s, 0 rescues), Blinkit (cart, 140 s), Zepto (cart, then its sign-in wall, 0 rescues), BigBasket (checkout after a sign-in, 0 rescues). Stopped at a safety gate: Myntra ("Add to Bag" is named with "Buy Now"), Ola ("Continue", declined). Failed: Flipkart (its own "Something went wrong" page on the cart), Lenskart and BookMyShow (a step ended as finished without pressing; follow-ups), and Uber, the regression the last commit fixes. The gateway was slow all evening: per-call p50 1.1–1.3 s, against 0.66 s at 13:00

Stage 4 and the second review's fixes have not been run live: they are checked by the replays above and by tests. The final review's fixes ran in the live run of the utility sites. The last commit's place-box fix is covered by a simulator test that fails without it; Uber has not been run again with it. The open Stage 3 failures are on this branch's list of follow-ups, not caused by it: BlazeDemo's stop_before not finding "Purchase Flight", Lenskart's consent banner asking for a second choice, and Blinkit adding extra items.

Tests

  • New simulator tests:
    • hedging: stalled copy, in time, large request, failures, Sage, the copy cap, and the journal's won;
    • late suggestions: rows one and two waits late, a still page (and that it is not settled), and a box's own rows waited past;
    • a covered click closes a consent banner with its own button;
    • a stalled step whose result shows ends AlreadyDone;
    • prelaunch on browser-only planning, with release on plan failure and while values are asked for.
  • New unit tests:
    • settle defaults and the module's settle/prelaunch keys;
    • prompt settle and the change watch's scripts, and their deadlines with a stalled browser;
    • a closed surface opens no browser early;
    • which layer a covered press may close;
    • workspace await_change delegation;
    • the planner retry classification and schedule;
    • the 10 s attempt timeout;
    • TASK_MEMORY merging and its folder, and task_live's switch values;
    • journal split and compare, and a resume journaled only once accepted.
  • Stage 4 and its review:
    • a brief settle waits only while the page changes, the desktop's settles in full, and the simulator's trail;
    • warm-ups: one small question for each call of a first turn, at most 18, Sage not warmed, given up on after 10 s, journaled with the task (also through the module runner), and none under a call cap;
    • quorum: what settles a decision (labels, open answers, short quorums), ending on the fifth plain answer while the rest still run, waiting for all when one dissents, and widening past every framing asked;
    • the early page: which address a task goes to (and which it only mentions), loaded once in the session it was opened in, within 10 s, opened after the browser, journaled as open_page;
    • --split: cost, late framings told apart by question, warm-ups out of the rounds.
  • New live fixture tests: a shadow-root banner in a display: contents host; a host whose slotted button, hidden dialog and aria-labelledby are read once and right.

Documentation

  • configuration.md, MODULE.md and .env.example cover the new keys and variables.
  • interacting.md, sight.md, surface.md and specs/browser-sight.md describe settling, the change watch and shadow roots.
  • decision-thresholds.md lists HEDGE_AFTER, HEDGE_COPIES, LATE_LOOK_MS and LATE_LOOKS, and the stall's STALL_TURNS and DONE.
  • jev-journal.md documents the new events and how a task_live run journals to TASK_OUT/journal.
  • the-do-loop.md, catching-mistakes.md, decision-loops.md and jev-questions.md describe the stall check; decision-loops.md the covered press.
  • planner.md covers the retries, tasks.md prelaunch, and live-tasks.md the task_live variables.
  • Stage 4: surface.md, interacting.md and surfaces-and-screens.md (the brief settle and the early page), jev-runtime.md and tasks.md (warming, the early page), voting-and-briefing.md and decision-thresholds.md (the quorum), jev-journal.md (left, warm-up, open_page, cost).
  • decision-loops.md was already 509 lines on main; this branch leaves it at 508.

Checklist

  • No new #[allow(...)], #[ignore], or relaxed lints
  • No secrets, tokens, or .env contents in the diff or the description
  • The change is focused on one logical change: it is one effort (openhuman#7000's speed work), split into focused commits.

A task's journal covered only its flows: the planner's call before the
first flow, each rescue between flows, and a person's answer to a pause
left gaps that read as unexplained time (38% of a live batch). The task
controller now times them and hands plan, rescue, and resume events to a
new FlowRunner::journal hook; the module's runner writes them into the
task's file, or for PlanTask, which plans before a task exists, into a
run of its own. Planner::plan_measured and Rescuer::guide_measured count
the model calls each took, repairs of refused answers included.

jev_journal gains --split, which reads whole tasks by their timestamps
(each run restarts elapsed_ms) and splits wall time into planning,
rescues, waits, and flows (Jev, settling, acting, reading), with per-call
and per-decision latency and what the slowest call adds to each round;
and --compare, the medians of two sets of tasks side by side.

No change to what a task does: the hook does nothing by default, and
nothing when the journal is off.
Picks up 5efb633 on the local agent-browser branch tinycomputer/network-quiet
(not yet pushed to tinyhumansai/agent-browser): waitforloadstate gains
networkquiet, which counts its 500 ms of quiet from the start.
Each is a setting so it can be compared against the current behaviour on
the same tasks before it becomes the default (openhuman#7000, Stage 1).

- browser.settle = "prompt" (TINYCOMPUTER_BROWSER_SETTLE): after an
  action, wait for the network to go quiet counted from the start
  (networkquiet), then only while the page still changes: no DOM change
  for 120 ms and no finite CSS animation running, over two frames, at
  most 400 ms. An idle page is read again after about 0.6 s instead of
  1.6 s; a busy one still waits for its requests.
- planner.plan_reasoning = "off" (TINYCOMPUTER_PLAN_REASONING): the
  planner asks its model not to reason first, with OpenRouter's
  reasoning.enabled=false. Live, a plan took about 3.5 s instead of 16 to
  20 s; reasoning_effort and thinking were ignored on that route.
- browser.prelaunch (TINYCOMPUTER_BROWSER_PRELAUNCH): StartTask with a
  plain-language task opens a browser-only task's browser while it is
  planned (FlowRunner::prepare, BrowserSurface::open), and lets it go if
  planning fails. task_live's TASK_PLAN=in-task plans inside StartTask,
  as OpenHuman does, so this can be measured.
planner.plan_reasoning = "off" cut a plan from 16-20 s to about 3.5 s,
but in the Stage 1 A/B every extra rescue traced back to a plan written
without reasoning: card fields filled before the stop, "press Enter"
where only a button searches, a location step that stalled, a verify
step that failed. The time saved was lost again to rescues, so plans
are drafted with reasoning again and the switch is removed.
browser.settle now defaults to prompt and browser.prelaunch to true. In
44 live runs over the two Stage 1 A/B batches, no failure traced back to
either: Myntra's size showed selected straight after its click, and the
failures were the irreversible-press gate, a store's own error overlay,
and plans drafted without reasoning. Settling promptly cut the wait after
an action by 0.2-0.7 s on three of four stores (Amazon's network never
goes quiet, so it still waits about 2.3 s), and opening the browser while
the plan is drafted cut the first step's launch from 2.8 s to 0.6 s.

Both stay switchable: "settle": "steady" and "prelaunch": false restore
the old behaviour, and task_live takes TINYCOMPUTER_BROWSER_SETTLE=steady
and TINYCOMPUTER_BROWSER_PRELAUNCH=0.
Sight hands a page to the accessibility tree when it sees a control inside
a shadow root, which a CSS selector from the page cannot address. It only
counted a shadow root whose host itself showed. On Lenskart the consent
banner's host is laid out as `display: contents`, so it has no box, sight
skipped it, and the banner's "Allow Selection" and "Allow all" buttons
never reached Jev while the banner lay over "Add To Cart": every press
came back covered, and all four runs failed after five rescues.

A shadow root now counts once its host or any of its controls shows. The
live fixture test reproduces the banner (sight read only "Add To Cart"
before the fix), checks a block host whose banner is fixed and draws no
box either, keeps a hidden banner out, and checks that the tree read
instead offers the banner's buttons.
After text goes into a place box (an address, a city, a pickup), the flow
looked twice more for suggestions the page lists late, each look a fixed
500 ms pause plus a full settle. A plain form box lists none, so every
address and city field on BlazeDemo's passenger form paid 2.3 s (4.2 s
with steady settling) for a list that never comes, four waits a run.

Surface::await_change waits for the surface to change by itself and says
whether it did. The browser watches the page with a MutationObserver
that ends at the first change of its own (never sight's data-tc- marks),
or after LATE_LOOK_MS = 1000 ms with none; a surface that cannot watch,
the desktop's, pauses as a Wait did and says it may have changed. The
flow looks again after each wait, as before, but a wait that saw the
page stay still is not settled, notes "nothing changed", and ends the
looking: a still page lists nothing more. Rows that arrive late are
looked at as soon as they are drawn.

The simulated ride form now draws its rows a number of waits late; the
new flow tests pick rows that come one and two waits late, and wait
once, not twice, beside a box that lists nothing.
Asked for the price of one packet after raising its count to 2, Zepto
runs read the cart line's ₹40, the total for both: the cart shows no
price for one, and the product page's ₹20 lay behind the cart drawer.
The flow guide, which the planner and the rescuer both read, now says to
read the price of one before the count is raised.
One HTTP 502 from Tiny Humans' gateway failed a whole BlazeDemo task at
its plan, 3 s in: a hosted model call that failed was never tried again.
It now is, up to 4 times, 1, 2, then 4 seconds apart, when the failure
can pass: a server error, a rate limit that is not a spending cap, or a
dropped connection (tinyinference-llm's own classification). A refused
key, a rejected request, and the test guard's refusal are not retried.
In the 7 October runs, 21 of 36,459 Jev calls needed a retry: one framing
of a burst stalled while its siblings answered in about 0.6 s, until the
gateway gave up after ~10 s (12 s with the retry) or nothing came back
before the client's 30 s timeout (31.6 s). Every decision waits for all of
its framings, so about one run in four lost 12-32 s to one of them.

A framing that has not answered after HEDGE_AFTER (2.5 s; 3.5 s for a
request of 32 KB or more, whose p99 was 3.1 s) is now sent once more, and
whichever copy answers first counts; a copy that fails gives way to the
other. Fewer than 0.3% of calls ran that long otherwise, so the copies
cost little. The answer that counts carries the extra attempt, and the
journal records a `hedge` event. Every framing goes through it, the
evidence gate's included.

A Jev attempt now gives up after 10 s unless jev.timeout_ms says
otherwise (the slowest answer seen took 8.6 s), so a request both copies
of which stall is retried after 10 s rather than 30 s.
The networkquiet wait now counts only requests that can change what the
page shows (the page's own document, scripts, stylesheets, XHR and fetch
data), so analytics pings and other frames' documents that never report
finishing no longer hold every settle to its cap.
On Amazon every action that opened a page sent 100+ requests for over
2 s, so prompt settling always ran its network wait to the 2 s cap, then
its 400 ms stillness cap, about 2.4 s an action. What the task needed was
on screen by then for a long time: the results after 0.87-1.04 s, a
product's title and Add to Cart after 0.96-1.24 s, the cart's subtotal
after 0.66-0.81 s (six measured page loads).

Settle::Prompt now waits at most QUIET_MS = 1 s for the requests that
change the page, then for the page to stop changing as before, so such a
page is read after about 1.4 s. Steady settling keeps its 2 s.
In the live check after the rule first went in, 3 of 4 plans read the
price of one before raising the count, but one Zepto plan still read it
in the cart and got ₹40, the line for two. The rule now also sits in the
`read` row of the step table, where a planner looks when it writes a
read.
Handing the whole page to the tree whenever a shadow root showed
controls made Lenskart readable but worse to work with. While its consent
banner was up, the tree read the search box unnamed, so Enter did not
search; it cut the page at its element budget; and it read an image
carousel's dots as sizes. The banner reached Jev in 1,330 calls of one
run and was still never closed.

Sight now keeps reading the page. For a shadow root that shows controls
(its host's box or any control's), it marks the host and names the layer
the shadow content draws, from its fixed or dialog element's label,
heading, or first words: `popover "We value your privacy"`. The surface
reads the tree under that host alone and adds its controls and text after
everything sight read, under that label, so the digest sees a layer in
front and the attention pass a privacy card with "Allow Selection" to
close it with. A tree snapshot's refs last until the next snapshot, so
with two such shadow roots, or when the subtree cannot be read, the tree
reads the whole page as before.
A press the page refused as covered was answered with Escape and one
more try. Escape leaves a consent banner where it is: on Lenskart every
press of "Add To Cart" and "Frame Size" under the banner was refused
until the run failed.

When a layer in front holds a control that closes it, the least committal
one ("Allow Selection" before "Allow all", as the attention pass ranks
them) is now pressed instead, as `click (uncover)`, and the same target
tried again; otherwise Escape, as before. The simulator gains a consent
banner Escape does not close.
On a slow evening the Tiny Humans route answered calls in up to 3.4 s
(median 0.91 s, against 0.72 s earlier in the day), and a copy sent at
2.5 s lost the race to its original 15 times in 16: 5% more calls for
nothing. A framing now gets its copy after 4 s, or 5 s for a request of
32 KB or more (p99.9 3.9 s), past what a slow but live answer takes; a
stuck one still answers in about 5 s rather than 12-32 s.
A `do` step whose last three actions changed nothing failed outright,
even when the page had already done its work: "press Enter to search"
pressed Enter three times over results a live search had listed as the
query was typed. On 7 October, 25 of 189 rescues found such a step's
work already done, each ~18 s after the failure.

Such a step now asks one question first: does the screen already show
the result the step is meant to bring about? When it clearly does (DONE),
the step ends AlreadyDone; otherwise it fails with the same note as
before, so a rescue still skips rather than retries it.
A task's report lists the elements it grounded (`learned`), meant to be
passed back as StartTask's `memory` so a later run confirms a remembered
element with one yes or no rather than searching for it. Neither
task_live nor OpenHuman passed any.

TASK_MEMORY names a JSON file of grounding hints: the task starts with
them, and what it learns is saved back, a newer hint replacing an older
one for the same element. Unset, every run starts fresh, as benchmark
batches should.
…w root

Reading one shadow root's controls beside sight snapshots the tree under
its host. On Lenskart that snapshot failed (agent-browser described the
host's subtree without piercing its shadow root, so no accessibility
node was found), and the surface fell back to reading the whole page
through the tree, as before the change. agent-browser d8e93d8 describes
the subtree through shadow roots. The live fixture now puts a stylesheet
beside the banner in its shadow root, as the live banner had.
agent-browser now keeps each page session's page-changing requests that
have not finished, and a networkquiet wait starts with those sent in the
last 3 s: a click's own navigation or fetch, which starts before the wait
can subscribe, is waited for, up to QUIET_MS, instead of the page being
read as it was before the action. The bump also brings the networkquiet
docs, help and MCP entries, and the subtree doc comment fix.

Settle's docs say so, and give their waits as values: rustdoc refuses
links from a public enum to private constants, which failed the docs job.
The prompt settle's stillness script and await_change's watch were sent
with no deadline of their own, right after the actions that replace
pages, and an evaluate sent while a page is replaced waited out the
browser's 30 s live. Each call now has its own cap plus 500 ms
(watch::deadline): a quiet wait given up on still lets the page be
watched, and a watch given up on says the page may have changed.

Both watches now observe the page's open shadow roots as well as the
document, so a suggestion list a web component draws counts as a change.
The stillness script stops asking for frames once it has resolved, and
await_change with no page open opens none.

The scripts move from operations.rs to their own watch.rs, and the test
engine can stall a command to prove the deadlines hold.
With one shadow root showing controls, the tree reads its host's whole
subtree beside sight: the host itself and the light-DOM children the page
puts in its slots, which sight had already read. Each was then offered
twice, as seen:N and eN, and a host that slots the whole page (an app
shell) would have listed the page twice. Sight now leaves the first such
host and everything under it to the tree.

Only the first host is labelled (with two, the tree reads the page
anyway); a layer that is not shown, such as a hidden dialog template,
names none; and a banner's aria-labelledby resolves inside its shadow
root. A reading the tree takes over no longer reports sight's denoising
counts. A live fixture test covers the slotted button, the hidden dialog,
and the label.
A browser-only task whose plan asks for values published needs_input with
the browser it opened while planning still open, holding one of the
browser's sessions for as long as the person took, or for good if the
host never answered. It is now released first; the run that follows opens
one again.

A task cancelled while its early open had not begun yet could leak a
session: release found none to close, and the open then launched one
nobody owned. A closed BrowserSurface now refuses an early open, checked
under the session lock, so either the open sees the close or the close
waits for the open and closes what it made.
Sage's calls take 5-6 s as a rule, past any hedge delay: almost every
framing would have been copied, adding half again to its calls and cost
for answers that rarely came first. Sage now gets no copy.

Many framings outliving their delay together means a slow or failing
gateway, where the client may be waiting out its own retry delay; a copy
of each only loaded it more. A runtime now has at most HEDGE_COPIES (2)
copies in flight; past that, a framing waits for its own answer.

The hedge event's won now names the answer that counted (first, copy, or
neither), not the copy that finished first. Hedging moves from decide.rs
into its own module, and the harness doc's settle row says how the
browser settles now.
When a press was refused as covered, front_closer took the least committal
closer of the first layer in front, asked with no intent and never told
the target. A toast over a size popover's rows could get the popover's
own close pressed, and a chat bubble over a consent banner's Accept all
the banner's Reject all: either closes the layer the step works in, and
the retried press then fails.

It now skips the layer the target sits in, the target itself, and any
layer the step's intent names, as the attention pass does. The covered
press moves into act/uncover.rs, and its docs and the attention module's
say what it presses; decision-loops.md is tightened around it so the file
does not grow.
Whether a step is done is the completion loop's question, and every other
completion question in the do loop is gated on it; the stalled-step check
asked holds() even when a run turned that loop off. Such a run now fails
a stalled step outright, as before the check.

decision-loops.md, decision-thresholds.md (STALL_TURNS and DONE) and
jev-questions.md still described a stall as an outright failure, and now
say what happens instead.
The simulator now keeps a trail of its settle and await_change calls. The
late-suggestion tests check that a wait that saw a change is settled like
any action, and one that saw a still page is not.
The plan, rescue and resume events were built for every task, journal on
or off, and journal_event drew a run id even when off: the journal's own
rule is to build nothing then. FlowRunner::journal and journal_event now
take a closure, called only when an event is written, and the module's
runner and JevRuntime::journal_event return before anything is built when
the journal is off (JevRuntime::journaling).

A resume was journaled before ContinueTask's answer was checked, so a
refused answer, or one still missing values, recorded a wait the task
had not left, and --split counted it twice. It is journaled once the task
has left the wait. A plan's wall_ms is timed on the planner alone: a
browser slower to open than the plan was counted as planning. ModelUse's
calls are documented as excluding a hosted call's own retries.
TASK_MEMORY's own example, target/task-live/memory/amazon.json, names a
folder nothing made: the whole task ran, then saving what it learned
failed and the run ended with an error before its pass or fail line. The
folder is now made when missing.

TASK_PLAN other than in-task, and TINYCOMPUTER_BROWSER_PRELAUNCH other
than 0 or 1, were silently ignored; they are now refused. A FLOW_FILE run
with TASK_PLAN=in-task no longer saves the given flow as a plan drafted
inside the task. What a run keeps beside its report moves into its own
module, so main.rs stays under 400 lines.
split.rs had grown to 537 lines. Building a split stays in split/mod.rs;
rendering it for a terminal moves to split/render.rs, and reading an
event's time and merging spans of it to split/time.rs. Nothing changes in
what it prints.
The docs said task_live writes its journal under TASK_OUT/journal, and
showed try-* folders. Neither task_live nor tasks/run sets the journal's
folder: a run journals there only when started with
TINYCOMPUTER_JEV_JOURNAL=$TASK_OUT/journal, and try-* was one private
script's naming. The docs, jev_journal's help and runs.rs now say so.
Docs still said the browser settles on network idle (decision-loops.md,
the-do-loop.md), that a place box waits a fixed beat between its late
looks (filling-forms.md), and that a Jev attempt's timeout defaults to
the client's (JevConfig::timeout_ms, jev-runtime.md); it is the module's
10 s. Three paragraphs of decision-loops.md are reflowed so the file does
not grow past its length on main.
@coderabbitai

coderabbitai Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: d7f121d8-2a15-4525-b4a9-5318c80603d4
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@tinysweeper

tinysweeper Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

⚠️ Review failed for 7bc42399764a. the review of #77 did not finish within 900s

A place box was looked at again only while it listed no new row at all.
Live on Uber, the pickup box first listed rows of its own ("Allow
location access", "Search in a different city"), and the prompt settle,
which reads the page as soon as it goes still, looked before the matches
were fetched: no row named the place, Jev rightly chose none, and the
pickup was never set, so no ride options showed. The slower settle had
hidden this by looking later.

A place box is now looked at again, as before up to LATE_LOOKS waits that
each end at the page's first change, until a row names the text typed.
The simulator's ride form can list such rows of its own while its matches
are pending, and the new test fails on the old condition.
The flow settles the surface after every successful action. On the
browser, prompt settling waits for the requests that change the page
(at least its 500 ms of quiet, at most 1 s) and then for the page to go
still. A launch leaves an open page as it is and Escape closes a layer:
live, 3% of launches and 12% of Escapes settled with a request of the
page still running, against 39% of fills (fetching the suggestions the
next look reads) and 70% of clicks.

Those two now call Surface::settle_briefly. The browser waits only while
the page changes, 120-400 ms instead of about 0.65 s on an idle page;
steady settling, and the desktop, settle in full as before. Typing,
Return and every other action keep the full settle.
A decision asks its framings all at once, each on an HTTP/1.1 connection
of its own, and opening one to the gateway costs a handshake: over 84
live runs, a task's first decision took 330 ms more a call (median) than
its later decisions of the same size.

JevRuntime::warm(votes) sends one one-question evaluation per framing
(at most MAX_VOTES), all at once, through evaluate, so each is journaled
under the step "warm-up"; answers are dropped and calls still out after
10 s are given up on. Sage is not warmed. A planned task starts it with
the planner through FlowRunner::warm, which does nothing by default, and
never waits for it; the module runner warms its Jev runtime, journaled
with the task.
A decision waits for its slowest framing. Asked 7 ways or more, it is now
merged without its last 2 framings once those in settle every question:
every Choice and Score ranks the same option first at 0.9 or more, every
yes/no is at 0.97 or more in each framing or at 0.08 or less in each (the
evidence band beyond every yes/no threshold), and the page kind reads alike.

Replayed over 13,669 decisions asked seven ways in 245 live runs, such a
quorum ended 12% of them, a median 0.14 s sooner, about 2 s a run, and every
one read the same on every threshold, choice floor and evidence gate as all
seven framings did, but for two whose mean moved by under 0.002 across the
band's edge. A quorum of four agreed only 99.2% of the time: its stragglers
dissented.

Framings are taken in the order they finish, and an answer already in when
the quorum is reached is used. Those left are not cancelled: they run to
their end, so their connections go back to the pool, count as calls, and
journal their exchanges after the decision, whose event now says how many
it did not wait for (left).
The framings a quorum did not wait for journal their exchanges after
their decision, so jev_journal --split read them as the next decision's
round, and its "slowest call adds" figure took a straggler for that
round's slowest. A decision's `left` exchanges that follow it now count
in no round; they still count as calls.
openhuman#7000 asks the timing harness to report cost beside wall time,
decisions and input tokens. A split now prices Jev's answers the way the
repository's evals do: their input tokens at $0.042 per million, output
free. It is kept as jev_cost_micro_usd (millionths of a dollar, so the
medians stay whole numbers), printed on the tokens line and in the table,
and compared. Another model's calls (Sage bills by units) are not priced,
nor planning and rescues, whose tokens are not journaled.
A browser-only task's browser already opens while its plan is drafted
(browser.prelaunch). Its first step then browses to the site the task
names: over 40 live runs that page took 1.9 s to load (p90 6.4 s) and
1.1 s to settle, while the plan before it took 10-30 s. In 56 of 57 live
plans the first step browsed exactly the page the task's text named.

When the task's text names one web address, written out with https:// or
http://, the runner now loads it in the early browser in the same wait
(FlowRunner::open_page, BrowserSurface::open_at). Until the page is first
read or another address loads, a navigation to the same place (scheme,
www. and a trailing slash aside) finds it loaded and loads nothing. A task
naming several addresses or none loads none; a page not drawn yet, a page
that would not load, or a surface let go meanwhile is loaded as asked, and
the session's allowed origins apply as to any navigation. It rides on
browser.prelaunch: off, nothing opens early.
Widening asked framings from the ballot's length on, but a decision ended
on a quorum holds only the answers it waited for. After one asked seven
ways, a widening sent the sixth and seventh framings again, which had been
asked and left, put their answers in the ballot twice, and never heard the
eighth and ninth; with nine votes it asked again where it would have asked
nothing. FlowRun now keeps how many framings each question was asked in,
set by each decision and raised by each widening, and widens from there. A
press nothing undoes is vouched for with a widening every time.

SURE_YES and SURE_NO keep each yes/no beyond every threshold one answer is
read against, not beyond the band of a belief that pairs two answers (a
yes/no with its negation, or a coverage): such a belief can still be
widened, as it would be after all the framings. The docs now say so and
give the quorum of four's figure from the same replay as the rest (0.8%),
and JevMetrics says which of its counts take in the framings a quorum did
not wait for.
A step's first turn asks its judging and grounding's opening together,
each in every framing, so warming one connection per framing left half of
that turn's calls opening their own. The warm-up now opens one for each
(FIRST_TURN, 2, times the votes; at most 18).

A task whose budget caps its Jev calls (max_model_calls) is not warmed:
the warm-up's calls would spend from it unseen. The module runner's warm
test now checks that the warm-up's calls reach the task's journal, with a
runtime whose calls give up before any request can reach Jev.
jev_journal --split counted the warm-up's calls with the decisions' (calls,
their percentiles, failed calls) and in the first round's slowest call;
they are still priced. A quorum's left framings were skipped by position,
the next exchanges after their decision whichever decision they belonged
to, so in a batch a grounding framing could be skipped and the judging's
straggler counted in grounding's round. A late exchange is now matched to
its decision by its questions, which are among the decision's; a journal
without question ids reads as before.
The early load took any one address a task's text wrote out, though a task
may only mention one: to check it, read it out, or pass it on, and a link's
query can carry a one-time token a load would spend. It now loads an
address only when the words right before it send the browser there ("go
to", "open", "visit", "start at", "on", ...) and it carries no query or
fragment.

The early navigation waits for its page's load at most 10 s
(EARLY_LOAD_MS, beyond nine in ten live first pages), and is journaled as
open_page (wall_ms, loaded), since the plan's outcome waits for it. The
check that a navigation finds the early page still shown reads it within
READ_TIMEOUT, and the mark of the early page names its session and is set
only while the surface is not let go, so it can match no other session.
…erwise

The desktop surface relies on Surface::settle_briefly's default running
its own settle after a launch or Escape; a test now counts that it does.
The settle rustdoc, the do loop's page and the surfaces page no longer say
that every action settles in full.
@YellowSnnowmann YellowSnnowmann changed the title perf: faster browser tasks, settling, hedged Jev calls and fewer rescues (openhuman#7000) perf: faster browser tasks: settling, hedged and warmed Jev calls, quorum decisions, early pages, fewer rescues (openhuman#7000) Oct 8, 2026
@senamakel
senamakel merged commit da56e9d into tinyhumansai:main Oct 8, 2026
5 of 6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants