Repository navigation
perf: faster browser tasks: settling, hedged and warmed Jev calls, quorum decisions, early pages, fewer rescues (openhuman#7000) - #77
Merged
senamakel merged 44 commits intoOct 8, 2026
Conversation
A task's journal covered only its flows: the planner's call before the first flow, each rescue between flows, and a person's answer to a pause left gaps that read as unexplained time (38% of a live batch). The task controller now times them and hands plan, rescue, and resume events to a new FlowRunner::journal hook; the module's runner writes them into the task's file, or for PlanTask, which plans before a task exists, into a run of its own. Planner::plan_measured and Rescuer::guide_measured count the model calls each took, repairs of refused answers included. jev_journal gains --split, which reads whole tasks by their timestamps (each run restarts elapsed_ms) and splits wall time into planning, rescues, waits, and flows (Jev, settling, acting, reading), with per-call and per-decision latency and what the slowest call adds to each round; and --compare, the medians of two sets of tasks side by side. No change to what a task does: the hook does nothing by default, and nothing when the journal is off.
Picks up 5efb633 on the local agent-browser branch tinycomputer/network-quiet (not yet pushed to tinyhumansai/agent-browser): waitforloadstate gains networkquiet, which counts its 500 ms of quiet from the start.
Each is a setting so it can be compared against the current behaviour on the same tasks before it becomes the default (openhuman#7000, Stage 1). - browser.settle = "prompt" (TINYCOMPUTER_BROWSER_SETTLE): after an action, wait for the network to go quiet counted from the start (networkquiet), then only while the page still changes: no DOM change for 120 ms and no finite CSS animation running, over two frames, at most 400 ms. An idle page is read again after about 0.6 s instead of 1.6 s; a busy one still waits for its requests. - planner.plan_reasoning = "off" (TINYCOMPUTER_PLAN_REASONING): the planner asks its model not to reason first, with OpenRouter's reasoning.enabled=false. Live, a plan took about 3.5 s instead of 16 to 20 s; reasoning_effort and thinking were ignored on that route. - browser.prelaunch (TINYCOMPUTER_BROWSER_PRELAUNCH): StartTask with a plain-language task opens a browser-only task's browser while it is planned (FlowRunner::prepare, BrowserSurface::open), and lets it go if planning fails. task_live's TASK_PLAN=in-task plans inside StartTask, as OpenHuman does, so this can be measured.
planner.plan_reasoning = "off" cut a plan from 16-20 s to about 3.5 s, but in the Stage 1 A/B every extra rescue traced back to a plan written without reasoning: card fields filled before the stop, "press Enter" where only a button searches, a location step that stalled, a verify step that failed. The time saved was lost again to rescues, so plans are drafted with reasoning again and the switch is removed.
browser.settle now defaults to prompt and browser.prelaunch to true. In 44 live runs over the two Stage 1 A/B batches, no failure traced back to either: Myntra's size showed selected straight after its click, and the failures were the irreversible-press gate, a store's own error overlay, and plans drafted without reasoning. Settling promptly cut the wait after an action by 0.2-0.7 s on three of four stores (Amazon's network never goes quiet, so it still waits about 2.3 s), and opening the browser while the plan is drafted cut the first step's launch from 2.8 s to 0.6 s. Both stay switchable: "settle": "steady" and "prelaunch": false restore the old behaviour, and task_live takes TINYCOMPUTER_BROWSER_SETTLE=steady and TINYCOMPUTER_BROWSER_PRELAUNCH=0.
Sight hands a page to the accessibility tree when it sees a control inside a shadow root, which a CSS selector from the page cannot address. It only counted a shadow root whose host itself showed. On Lenskart the consent banner's host is laid out as `display: contents`, so it has no box, sight skipped it, and the banner's "Allow Selection" and "Allow all" buttons never reached Jev while the banner lay over "Add To Cart": every press came back covered, and all four runs failed after five rescues. A shadow root now counts once its host or any of its controls shows. The live fixture test reproduces the banner (sight read only "Add To Cart" before the fix), checks a block host whose banner is fixed and draws no box either, keeps a hidden banner out, and checks that the tree read instead offers the banner's buttons.
After text goes into a place box (an address, a city, a pickup), the flow looked twice more for suggestions the page lists late, each look a fixed 500 ms pause plus a full settle. A plain form box lists none, so every address and city field on BlazeDemo's passenger form paid 2.3 s (4.2 s with steady settling) for a list that never comes, four waits a run. Surface::await_change waits for the surface to change by itself and says whether it did. The browser watches the page with a MutationObserver that ends at the first change of its own (never sight's data-tc- marks), or after LATE_LOOK_MS = 1000 ms with none; a surface that cannot watch, the desktop's, pauses as a Wait did and says it may have changed. The flow looks again after each wait, as before, but a wait that saw the page stay still is not settled, notes "nothing changed", and ends the looking: a still page lists nothing more. Rows that arrive late are looked at as soon as they are drawn. The simulated ride form now draws its rows a number of waits late; the new flow tests pick rows that come one and two waits late, and wait once, not twice, beside a box that lists nothing.
Asked for the price of one packet after raising its count to 2, Zepto runs read the cart line's ₹40, the total for both: the cart shows no price for one, and the product page's ₹20 lay behind the cart drawer. The flow guide, which the planner and the rescuer both read, now says to read the price of one before the count is raised.
One HTTP 502 from Tiny Humans' gateway failed a whole BlazeDemo task at its plan, 3 s in: a hosted model call that failed was never tried again. It now is, up to 4 times, 1, 2, then 4 seconds apart, when the failure can pass: a server error, a rate limit that is not a spending cap, or a dropped connection (tinyinference-llm's own classification). A refused key, a rejected request, and the test guard's refusal are not retried.
In the 7 October runs, 21 of 36,459 Jev calls needed a retry: one framing of a burst stalled while its siblings answered in about 0.6 s, until the gateway gave up after ~10 s (12 s with the retry) or nothing came back before the client's 30 s timeout (31.6 s). Every decision waits for all of its framings, so about one run in four lost 12-32 s to one of them. A framing that has not answered after HEDGE_AFTER (2.5 s; 3.5 s for a request of 32 KB or more, whose p99 was 3.1 s) is now sent once more, and whichever copy answers first counts; a copy that fails gives way to the other. Fewer than 0.3% of calls ran that long otherwise, so the copies cost little. The answer that counts carries the extra attempt, and the journal records a `hedge` event. Every framing goes through it, the evidence gate's included. A Jev attempt now gives up after 10 s unless jev.timeout_ms says otherwise (the slowest answer seen took 8.6 s), so a request both copies of which stall is retried after 10 s rather than 30 s.
The networkquiet wait now counts only requests that can change what the page shows (the page's own document, scripts, stylesheets, XHR and fetch data), so analytics pings and other frames' documents that never report finishing no longer hold every settle to its cap.
On Amazon every action that opened a page sent 100+ requests for over 2 s, so prompt settling always ran its network wait to the 2 s cap, then its 400 ms stillness cap, about 2.4 s an action. What the task needed was on screen by then for a long time: the results after 0.87-1.04 s, a product's title and Add to Cart after 0.96-1.24 s, the cart's subtotal after 0.66-0.81 s (six measured page loads). Settle::Prompt now waits at most QUIET_MS = 1 s for the requests that change the page, then for the page to stop changing as before, so such a page is read after about 1.4 s. Steady settling keeps its 2 s.
In the live check after the rule first went in, 3 of 4 plans read the price of one before raising the count, but one Zepto plan still read it in the cart and got ₹40, the line for two. The rule now also sits in the `read` row of the step table, where a planner looks when it writes a read.
Handing the whole page to the tree whenever a shadow root showed controls made Lenskart readable but worse to work with. While its consent banner was up, the tree read the search box unnamed, so Enter did not search; it cut the page at its element budget; and it read an image carousel's dots as sizes. The banner reached Jev in 1,330 calls of one run and was still never closed. Sight now keeps reading the page. For a shadow root that shows controls (its host's box or any control's), it marks the host and names the layer the shadow content draws, from its fixed or dialog element's label, heading, or first words: `popover "We value your privacy"`. The surface reads the tree under that host alone and adds its controls and text after everything sight read, under that label, so the digest sees a layer in front and the attention pass a privacy card with "Allow Selection" to close it with. A tree snapshot's refs last until the next snapshot, so with two such shadow roots, or when the subtree cannot be read, the tree reads the whole page as before.
A press the page refused as covered was answered with Escape and one
more try. Escape leaves a consent banner where it is: on Lenskart every
press of "Add To Cart" and "Frame Size" under the banner was refused
until the run failed.
When a layer in front holds a control that closes it, the least committal
one ("Allow Selection" before "Allow all", as the attention pass ranks
them) is now pressed instead, as `click (uncover)`, and the same target
tried again; otherwise Escape, as before. The simulator gains a consent
banner Escape does not close.
On a slow evening the Tiny Humans route answered calls in up to 3.4 s (median 0.91 s, against 0.72 s earlier in the day), and a copy sent at 2.5 s lost the race to its original 15 times in 16: 5% more calls for nothing. A framing now gets its copy after 4 s, or 5 s for a request of 32 KB or more (p99.9 3.9 s), past what a slow but live answer takes; a stuck one still answers in about 5 s rather than 12-32 s.
A `do` step whose last three actions changed nothing failed outright, even when the page had already done its work: "press Enter to search" pressed Enter three times over results a live search had listed as the query was typed. On 7 October, 25 of 189 rescues found such a step's work already done, each ~18 s after the failure. Such a step now asks one question first: does the screen already show the result the step is meant to bring about? When it clearly does (DONE), the step ends AlreadyDone; otherwise it fails with the same note as before, so a rescue still skips rather than retries it.
A task's report lists the elements it grounded (`learned`), meant to be passed back as StartTask's `memory` so a later run confirms a remembered element with one yes or no rather than searching for it. Neither task_live nor OpenHuman passed any. TASK_MEMORY names a JSON file of grounding hints: the task starts with them, and what it learns is saved back, a newer hint replacing an older one for the same element. Unset, every run starts fresh, as benchmark batches should.
…w root Reading one shadow root's controls beside sight snapshots the tree under its host. On Lenskart that snapshot failed (agent-browser described the host's subtree without piercing its shadow root, so no accessibility node was found), and the surface fell back to reading the whole page through the tree, as before the change. agent-browser d8e93d8 describes the subtree through shadow roots. The live fixture now puts a stylesheet beside the banner in its shadow root, as the live banner had.
agent-browser now keeps each page session's page-changing requests that have not finished, and a networkquiet wait starts with those sent in the last 3 s: a click's own navigation or fetch, which starts before the wait can subscribe, is waited for, up to QUIET_MS, instead of the page being read as it was before the action. The bump also brings the networkquiet docs, help and MCP entries, and the subtree doc comment fix. Settle's docs say so, and give their waits as values: rustdoc refuses links from a public enum to private constants, which failed the docs job.
The prompt settle's stillness script and await_change's watch were sent with no deadline of their own, right after the actions that replace pages, and an evaluate sent while a page is replaced waited out the browser's 30 s live. Each call now has its own cap plus 500 ms (watch::deadline): a quiet wait given up on still lets the page be watched, and a watch given up on says the page may have changed. Both watches now observe the page's open shadow roots as well as the document, so a suggestion list a web component draws counts as a change. The stillness script stops asking for frames once it has resolved, and await_change with no page open opens none. The scripts move from operations.rs to their own watch.rs, and the test engine can stall a command to prove the deadlines hold.
With one shadow root showing controls, the tree reads its host's whole subtree beside sight: the host itself and the light-DOM children the page puts in its slots, which sight had already read. Each was then offered twice, as seen:N and eN, and a host that slots the whole page (an app shell) would have listed the page twice. Sight now leaves the first such host and everything under it to the tree. Only the first host is labelled (with two, the tree reads the page anyway); a layer that is not shown, such as a hidden dialog template, names none; and a banner's aria-labelledby resolves inside its shadow root. A reading the tree takes over no longer reports sight's denoising counts. A live fixture test covers the slotted button, the hidden dialog, and the label.
A browser-only task whose plan asks for values published needs_input with the browser it opened while planning still open, holding one of the browser's sessions for as long as the person took, or for good if the host never answered. It is now released first; the run that follows opens one again. A task cancelled while its early open had not begun yet could leak a session: release found none to close, and the open then launched one nobody owned. A closed BrowserSurface now refuses an early open, checked under the session lock, so either the open sees the close or the close waits for the open and closes what it made.
Sage's calls take 5-6 s as a rule, past any hedge delay: almost every framing would have been copied, adding half again to its calls and cost for answers that rarely came first. Sage now gets no copy. Many framings outliving their delay together means a slow or failing gateway, where the client may be waiting out its own retry delay; a copy of each only loaded it more. A runtime now has at most HEDGE_COPIES (2) copies in flight; past that, a framing waits for its own answer. The hedge event's won now names the answer that counted (first, copy, or neither), not the copy that finished first. Hedging moves from decide.rs into its own module, and the harness doc's settle row says how the browser settles now.
When a press was refused as covered, front_closer took the least committal closer of the first layer in front, asked with no intent and never told the target. A toast over a size popover's rows could get the popover's own close pressed, and a chat bubble over a consent banner's Accept all the banner's Reject all: either closes the layer the step works in, and the retried press then fails. It now skips the layer the target sits in, the target itself, and any layer the step's intent names, as the attention pass does. The covered press moves into act/uncover.rs, and its docs and the attention module's say what it presses; decision-loops.md is tightened around it so the file does not grow.
Whether a step is done is the completion loop's question, and every other completion question in the do loop is gated on it; the stalled-step check asked holds() even when a run turned that loop off. Such a run now fails a stalled step outright, as before the check. decision-loops.md, decision-thresholds.md (STALL_TURNS and DONE) and jev-questions.md still described a stall as an outright failure, and now say what happens instead.
The simulator now keeps a trail of its settle and await_change calls. The late-suggestion tests check that a wait that saw a change is settled like any action, and one that saw a still page is not.
The plan, rescue and resume events were built for every task, journal on or off, and journal_event drew a run id even when off: the journal's own rule is to build nothing then. FlowRunner::journal and journal_event now take a closure, called only when an event is written, and the module's runner and JevRuntime::journal_event return before anything is built when the journal is off (JevRuntime::journaling). A resume was journaled before ContinueTask's answer was checked, so a refused answer, or one still missing values, recorded a wait the task had not left, and --split counted it twice. It is journaled once the task has left the wait. A plan's wall_ms is timed on the planner alone: a browser slower to open than the plan was counted as planning. ModelUse's calls are documented as excluding a hosted call's own retries.
TASK_MEMORY's own example, target/task-live/memory/amazon.json, names a folder nothing made: the whole task ran, then saving what it learned failed and the run ended with an error before its pass or fail line. The folder is now made when missing. TASK_PLAN other than in-task, and TINYCOMPUTER_BROWSER_PRELAUNCH other than 0 or 1, were silently ignored; they are now refused. A FLOW_FILE run with TASK_PLAN=in-task no longer saves the given flow as a plan drafted inside the task. What a run keeps beside its report moves into its own module, so main.rs stays under 400 lines.
split.rs had grown to 537 lines. Building a split stays in split/mod.rs; rendering it for a terminal moves to split/render.rs, and reading an event's time and merging spans of it to split/time.rs. Nothing changes in what it prints.
The docs said task_live writes its journal under TASK_OUT/journal, and showed try-* folders. Neither task_live nor tasks/run sets the journal's folder: a run journals there only when started with TINYCOMPUTER_JEV_JOURNAL=$TASK_OUT/journal, and try-* was one private script's naming. The docs, jev_journal's help and runs.rs now say so.
Docs still said the browser settles on network idle (decision-loops.md, the-do-loop.md), that a place box waits a fixed beat between its late looks (filling-forms.md), and that a Jev attempt's timeout defaults to the client's (JevConfig::timeout_ms, jev-runtime.md); it is the module's 10 s. Three paragraphs of decision-loops.md are reflowed so the file does not grow past its length on main.
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configuration
Comment |
YellowSnnowmann
marked this pull request as ready for review
October 7, 2026 14:43
Tiny Sweeper review
|
A place box was looked at again only while it listed no new row at all.
Live on Uber, the pickup box first listed rows of its own ("Allow
location access", "Search in a different city"), and the prompt settle,
which reads the page as soon as it goes still, looked before the matches
were fetched: no row named the place, Jev rightly chose none, and the
pickup was never set, so no ride options showed. The slower settle had
hidden this by looking later.
A place box is now looked at again, as before up to LATE_LOOKS waits that
each end at the page's first change, until a row names the text typed.
The simulator's ride form can list such rows of its own while its matches
are pending, and the new test fails on the old condition.
14 tasks
The flow settles the surface after every successful action. On the browser, prompt settling waits for the requests that change the page (at least its 500 ms of quiet, at most 1 s) and then for the page to go still. A launch leaves an open page as it is and Escape closes a layer: live, 3% of launches and 12% of Escapes settled with a request of the page still running, against 39% of fills (fetching the suggestions the next look reads) and 70% of clicks. Those two now call Surface::settle_briefly. The browser waits only while the page changes, 120-400 ms instead of about 0.65 s on an idle page; steady settling, and the desktop, settle in full as before. Typing, Return and every other action keep the full settle.
A decision asks its framings all at once, each on an HTTP/1.1 connection of its own, and opening one to the gateway costs a handshake: over 84 live runs, a task's first decision took 330 ms more a call (median) than its later decisions of the same size. JevRuntime::warm(votes) sends one one-question evaluation per framing (at most MAX_VOTES), all at once, through evaluate, so each is journaled under the step "warm-up"; answers are dropped and calls still out after 10 s are given up on. Sage is not warmed. A planned task starts it with the planner through FlowRunner::warm, which does nothing by default, and never waits for it; the module runner warms its Jev runtime, journaled with the task.
A decision waits for its slowest framing. Asked 7 ways or more, it is now merged without its last 2 framings once those in settle every question: every Choice and Score ranks the same option first at 0.9 or more, every yes/no is at 0.97 or more in each framing or at 0.08 or less in each (the evidence band beyond every yes/no threshold), and the page kind reads alike. Replayed over 13,669 decisions asked seven ways in 245 live runs, such a quorum ended 12% of them, a median 0.14 s sooner, about 2 s a run, and every one read the same on every threshold, choice floor and evidence gate as all seven framings did, but for two whose mean moved by under 0.002 across the band's edge. A quorum of four agreed only 99.2% of the time: its stragglers dissented. Framings are taken in the order they finish, and an answer already in when the quorum is reached is used. Those left are not cancelled: they run to their end, so their connections go back to the pool, count as calls, and journal their exchanges after the decision, whose event now says how many it did not wait for (left).
The framings a quorum did not wait for journal their exchanges after their decision, so jev_journal --split read them as the next decision's round, and its "slowest call adds" figure took a straggler for that round's slowest. A decision's `left` exchanges that follow it now count in no round; they still count as calls.
openhuman#7000 asks the timing harness to report cost beside wall time, decisions and input tokens. A split now prices Jev's answers the way the repository's evals do: their input tokens at $0.042 per million, output free. It is kept as jev_cost_micro_usd (millionths of a dollar, so the medians stay whole numbers), printed on the tokens line and in the table, and compared. Another model's calls (Sage bills by units) are not priced, nor planning and rescues, whose tokens are not journaled.
A browser-only task's browser already opens while its plan is drafted (browser.prelaunch). Its first step then browses to the site the task names: over 40 live runs that page took 1.9 s to load (p90 6.4 s) and 1.1 s to settle, while the plan before it took 10-30 s. In 56 of 57 live plans the first step browsed exactly the page the task's text named. When the task's text names one web address, written out with https:// or http://, the runner now loads it in the early browser in the same wait (FlowRunner::open_page, BrowserSurface::open_at). Until the page is first read or another address loads, a navigation to the same place (scheme, www. and a trailing slash aside) finds it loaded and loads nothing. A task naming several addresses or none loads none; a page not drawn yet, a page that would not load, or a surface let go meanwhile is loaded as asked, and the session's allowed origins apply as to any navigation. It rides on browser.prelaunch: off, nothing opens early.
Widening asked framings from the ballot's length on, but a decision ended on a quorum holds only the answers it waited for. After one asked seven ways, a widening sent the sixth and seventh framings again, which had been asked and left, put their answers in the ballot twice, and never heard the eighth and ninth; with nine votes it asked again where it would have asked nothing. FlowRun now keeps how many framings each question was asked in, set by each decision and raised by each widening, and widens from there. A press nothing undoes is vouched for with a widening every time. SURE_YES and SURE_NO keep each yes/no beyond every threshold one answer is read against, not beyond the band of a belief that pairs two answers (a yes/no with its negation, or a coverage): such a belief can still be widened, as it would be after all the framings. The docs now say so and give the quorum of four's figure from the same replay as the rest (0.8%), and JevMetrics says which of its counts take in the framings a quorum did not wait for.
A step's first turn asks its judging and grounding's opening together, each in every framing, so warming one connection per framing left half of that turn's calls opening their own. The warm-up now opens one for each (FIRST_TURN, 2, times the votes; at most 18). A task whose budget caps its Jev calls (max_model_calls) is not warmed: the warm-up's calls would spend from it unseen. The module runner's warm test now checks that the warm-up's calls reach the task's journal, with a runtime whose calls give up before any request can reach Jev.
jev_journal --split counted the warm-up's calls with the decisions' (calls, their percentiles, failed calls) and in the first round's slowest call; they are still priced. A quorum's left framings were skipped by position, the next exchanges after their decision whichever decision they belonged to, so in a batch a grounding framing could be skipped and the judging's straggler counted in grounding's round. A late exchange is now matched to its decision by its questions, which are among the decision's; a journal without question ids reads as before.
The early load took any one address a task's text wrote out, though a task
may only mention one: to check it, read it out, or pass it on, and a link's
query can carry a one-time token a load would spend. It now loads an
address only when the words right before it send the browser there ("go
to", "open", "visit", "start at", "on", ...) and it carries no query or
fragment.
The early navigation waits for its page's load at most 10 s
(EARLY_LOAD_MS, beyond nine in ten live first pages), and is journaled as
open_page (wall_ms, loaded), since the plan's outcome waits for it. The
check that a navigation finds the early page still shown reads it within
READ_TIMEOUT, and the mark of the early page names its session and is set
only while the surface is not let go, so it can match no other session.
…erwise The desktop surface relies on Surface::settle_briefly's default running its own settle after a launch or Escape; a test now counts that it does. The settle rustdoc, the do loop's page and the surfaces page no longer say that every action settles in full.
8 of 10 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This makes browser tasks faster without losing reliability, for tinyhumansai/openhuman#7000. Each change was measured on live runs before it went in. The branch then had a final review, whose fixes are 13 commits, and a live run of ten utility sites, whose one regression is fixed in the commit after them.
Since then, Stage 4 adds four more speed changes, each first measured by replaying the journals of 245 live runs, and two
jev_journalfixes. A second review, by an Opus model, found no high-severity defect in them; its fixes are the last five commits. Stage 4 has not been run live yet: an A/B of the module before and after it is the next step.Measuring (Stage 0)
jev_journal --splitbreaks a task's time down into planning, Jev, settling, actions and rescues.jev_journal --comparesets two batches side by side.Less waiting (Stage 1, now default)
browser.settle = "prompt": after an action, the page is read once the requests that change it are done and it stops changing, about 0.6 s on an idle page instead of 1.6 s.browser.prelaunch = true: a browser-only task's browser opens while its plan is drafted, which takes the first launch from 2.8 s to 0.6 s."steady"andfalserestore the old behaviour.Shorter Jev waits (Stage 2)
jev.timeout_mssays otherwise.networkquiet). Amazon waited ~2.4 s after every page-opening action: 100+ requests over 2 s, some of which never report finishing to the page's session. Its content shows after 0.7–1.2 s, and settling now takes 0.67–1.47 s a page load.Fixes from the live checks
Surface::await_change). A still page ends the waiting: BlazeDemo's address and city boxes went from four waits of 1.2–2.1 s to two of 1.0 s.display: contentshost that sight never saw.Fewer rescues (Stage 3)
dostep whose last three actions changed nothing first asks whether the screen already shows its result, and endsAlreadyDonewhen it clearly does. 25 of 189 rescues on 7 October were such steps, ~18 s each.TASK_MEMORY); it is off unless set.Fixes from the final review
jev_journal'ssplit.rsandtask_live/main.rsare split under 400 lines, and stale docs are brought up to date.Fix from the live run of the utility sites
LATE_LOOKSwaits that each end at the page's first change, until a row names the text typed.Waiting less for the first decision and the first page (Stage 4)
Surface::settle_briefly; the desktop and steady settling settle in full as before). Typing, Return and every other action keep the full settle.JevRuntime::warmnow opens one for each call of a step's first turn (its judging and grounding's opening, each in every framing: 2 × votes, at most 18) with a one-question evaluation, journaled aswarm-up. The task never waits for it, Sage is not warmed, nor is a task whose budget caps its Jev calls.open_page; the first step to the same place, in that session, finds it loaded. It rides onbrowser.prelaunch.jev_journal:--splitand--compareprice Jev's input tokens at $0.042 per million (jev_cost_micro_usd), as the evals do; a quorum's late framings and the warm-up's calls count in no round.Together, by those replays, about 6.5–7 s a run.
Fixes from the second review
--splitmatches a quorum's late framings to their decision by question, rather than taking the next exchanges whichever decision they were, and leaves the warm-up out of the call statistics.READ_TIMEOUT.Checked and skipped, with data
Related issue
Refs tinyhumansai/openhuman#7000.
Depends on tinyhumansai/agent-browser#2. The
vendor/agent-browsergitlink pins its head,ebf10fb, which CI fetches through that pull request. It must merge first, by a merge commit; if it is squash-merged, the gitlink must move to the merged commit before this one merges.API or behavior changes
browser.settle(promptdefault, orsteady) andbrowser.prelaunch(truedefault).jev.timeout_msnow defaults to 10 s, from the client's 30 s.AlreadyDone;Surfacetrait: newawait_change(pause and assume a change) andsettle_briefly(settle in full) methods with defaults, so existing implementations compile unchanged.FlowRunner: newwarmandopen_page, which do nothing by default;JevRuntime::warm(votes);BrowserSurface::open_at(url).FlowRunner::journaltakes its fields as a closure, built only when written;JevRuntime::journal_eventlikewise, andJevRuntime::journalingsays whether the journal is on.plan,rescue,resume,hedgeandopen_pageevents; awaitaction notes "nothing changed"; adecisionsays how many framings it did not wait for (left); the warm-up's exchanges are journaled under the stepwarm-up.jev_journal:--splitand--compare, with Jev's cost.TaskBudget.max_model_callsnotes that a capped task is not warmed, andJevMetricssays which counts take in the framings a quorum did not wait for.TASK_PLAN=in-task,TINYCOMPUTER_BROWSER_SETTLE,TINYCOMPUTER_BROWSER_PRELAUNCH,TASK_MEMORY.None of this breaks the wire contract.
Validation
Commands run on the final commit (Rust 1.98, macOS), with their outcomes; the first review's round ran on 1.99:
cargo fmt --all -- --check: clean.cargo clippy --workspace --exclude tinycomputer-accessibility --all-targets --all-features -- -D warnings: clean.tinycomputer-accessibilityalready fails clippy 1.99 on macOS onmain(its macOS-only files), and this branch does not touch it.cargo build --workspace --exclude tinycomputer-accessibility --all-targets --all-features: ok (built by the test and coverage runs).cargo test --workspace --exclude tinycomputer-accessibility --all-features: 1,164 passed, 0 failed.RUSTDOCFLAGS="-D warnings" cargo doc --no-deps --all-features --workspace --exclude tinycomputer-accessibility: clean.cargo llvm-covon core, bus, browser, engine and the module crate, with the CI gate's per-file filter: all 183 files at or above 90%; every source file changed sincemainat 100% butescalate/mod.rs(98.4%, a defensive fallback).TINYCOMPUTER_LIVE_BROWSER=1, real Chrome): all 11 sight live tests pass, including the shadow-root banner and its slotted content.Live A/B batches (
task_live, planned in-task as OpenHuman does):piercefix addressesStage 4 and the second review's fixes have not been run live: they are checked by the replays above and by tests. The final review's fixes ran in the live run of the utility sites. The last commit's place-box fix is covered by a simulator test that fails without it; Uber has not been run again with it. The open Stage 3 failures are on this branch's list of follow-ups, not caused by it: BlazeDemo's
stop_beforenot finding "Purchase Flight", Lenskart's consent banner asking for a second choice, and Blinkit adding extra items.Tests
won;AlreadyDone;settle/prelaunchkeys;await_changedelegation;TASK_MEMORYmerging and its folder, and task_live's switch values;open_page;--split: cost, late framings told apart by question, warm-ups out of the rounds.display: contentshost; a host whose slotted button, hidden dialog andaria-labelledbyare read once and right.Documentation
HEDGE_AFTER,HEDGE_COPIES,LATE_LOOK_MSandLATE_LOOKS, and the stall'sSTALL_TURNSandDONE.TASK_OUT/journal.left,warm-up,open_page, cost).decision-loops.mdwas already 509 lines onmain; this branch leaves it at 508.Checklist
#[allow(...)],#[ignore], or relaxed lints.envcontents in the diff or the description