Skip to content

fix(examples): run task_live standalone on Tiny Humans' routes - #75

Merged
senamakel merged 2 commits into
tinyhumansai:mainfrom
YellowSnnowmann:fix/task-live-standalone
Oct 5, 2026
Merged

senamakel merged 2 commits into
tinyhumansai:mainfrom
YellowSnnowmann:fix/task-live-standalone

Conversation

@YellowSnnowmann

Copy link
Copy Markdown
Contributor

Summary

Two things stopped task_live from running on its own on a Mac.

  • Startup race in the example host. The module claims its bus name while its setup is still running, and the loader refuses a reconfiguration until setup returns. Host::load reconfigured as soon as the name appeared, so every task_live run failed straight away with module must be ready before reinitialization. It now waits until the loader reports the module ready. If the module ends in a terminal state instead, the error names that state.
  • Tiny Humans routes for task_live. TINYHUMANS_TOKEN can replace OPENROUTER_API_KEY. It takes a Tiny Humans session token, or an API key with the inference scope. With it set, Jev goes through tiny_humans_open_router and the planner through the tiny_humans gateway.
    • The gateway refuses the engine's OpenRouter model ids (HTTP 400 "not available"), so agentic-v1 plans, rescues and shapes unless a model variable names another model.
    • Sage still takes the decisions on either route when asked.
    • TASK_HEADED=1 shows the launched browser, for a run on the host that someone wants to watch.

Related issue

None in this repo. Found while running tinycomputer standalone for tinyhumansai/openhuman#7000.

API or behavior changes

None for the module; the changes are to the examples only.

  • task_live gains two optional env vars, TINYHUMANS_TOKEN and TASK_HEADED.
  • With neither set, behavior is unchanged, except that the missing-key error now names both OPENROUTER_API_KEY and TINYHUMANS_TOKEN.

Validation

Commands actually run, on macOS (arm64):

  • cargo fmt --all -- --check: pass
  • cargo clippy --all-targets --all-features -- -D warnings: pass, run as -p tinycomputer-examples -p tinycomputer-engine
  • cargo build --all-targets --all-features: not run as its own step. The clippy and test builds cover the changed crate, and the release build of task_live succeeded.
  • cargo test --all-features -p tinycomputer-examples: 36 passed, 0 failed

Live runs:

  • Before the host fix, every run failed at load with the reinitialization error.
  • After it, a read-only task through TINYHUMANS_TOKEN (a prod API key) planned, ran and printed PASS: "list the top 3 stories on news.ycombinator.com", 19 s headed and 28 s headless.

Tests

task_live/main_tests.rs is new. It covers:

  • the Tiny Humans route and its agentic-v1 defaults;
  • a named model overriding a default, with blank values ignored;
  • the OpenRouter route, unchanged when no bearer is set;
  • the missing-key error naming both variables;
  • Sage on either route, and its own missing-key error.

The routes are read through a variable lookup, so the tests never touch the process environment.

The host's ready-wait is only exercised by the live runs above, because it needs a loaded module.

Documentation

  • .env.example
  • docs/crates/tinycomputer-examples/live-tasks.md
  • docs/crates/tinycomputer-examples/running-the-examples.md

Checklist

  • The change is focused on one logical change: running task_live standalone. The two commits can be reviewed separately.
  • No new #[allow(...)], #[ignore], or relaxed lints
  • No secrets, tokens, or .env contents in the diff or the description

The module claims its bus name while its setup is still running, and the
loader refuses a reconfiguration until that setup returns. Host::load
reconfigured as soon as the name appeared, so task_live could fail at
once with "module must be ready before reinitialization" (seen on macOS).
It now waits for the loader to report the module ready, and names the
state the module stopped in when it never gets there.
TINYHUMANS_TOKEN, a Tiny Humans bearer (a session token, or an API key with
the inference scope), sends Jev and the planner through Tiny Humans' routes
in place of OPENROUTER_API_KEY. The gateway refuses the engine's OpenRouter
model ids, so agentic-v1 plans, rescues and shapes unless a model variable
names another. Sage still takes the decisions on either route when asked.

TASK_HEADED=1 shows the browser the task launches, for a host run someone
wants to watch.

The routes are read through a variable lookup, so they are tested without
touching the process environment.
@coderabbitai

coderabbitai Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 0ecdd917-ebe3-4e21-b82a-3f408c436644
📥 Commits

Reviewing files that changed from the base of the PR and between 816ddf2 and 1f67db2.

📒 Files selected for processing (6)
  • .env.example
  • crates/tinycomputer-examples/src/bin/task_live/main.rs
  • crates/tinycomputer-examples/src/bin/task_live/main_tests.rs
  • crates/tinycomputer-examples/src/host/mod.rs
  • docs/crates/tinycomputer-examples/live-tasks.md
  • docs/crates/tinycomputer-examples/running-the-examples.md
 _______________________________________________________________
< You're one `console.log` away from enlightenment. Keep going. >
 ---------------------------------------------------------------
  \
   \   \
        \ /\
        ( )
      .( o ).
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@senamakel
senamakel marked this pull request as ready for review October 5, 2026 21:01
@senamakel
senamakel merged commit 3837706 into tinyhumansai:main Oct 5, 2026
5 checks passed
@tinysweeper

tinysweeper Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

Tiny Sweeper reviewed this change across 6 lane(s) and found 1 active actionable finding(s). Detailed lane evidence and any incomplete work are listed below.

State: Ready for maintainer review
Priority: medium
Reviewed head: 1f67db20083d
Updated: 1791234210 (Unix time)

Review snapshot

Change surface Files Review signal Count
Production 3 Active findings 1
Tests 1 Noted findings 0
Documentation 2 Resolved findings 0
Configuration 0 Pending checks/questions 0

Completeness: Complete
Test assessment: No supported feature-to-test mapping was available; this does not mean tests are absent or passed.

What changed

The review could not produce a supported behavioral summary; inspect the cited changed surface and lane details below.

Features

None identified with supported citations.

Tests

No supported feature-to-test mapping was produced. Test execution is not inferred.

Findings

  • medium · tests · Test wait_until_ready's failure paths, not just its comment — This new function is the behavioural heart of the host-side change: the comment claims the loader refuses reconfiguration until setup returns, and the code asserts that claim only (crates/tinycomputer\-examples/src/host/mod\.rs:509)

Before merge

None.

How this fits together

flowchart LR
  n0["main<br/>changed"]:::changed
  n1["module_config<br/>changed"]:::changed
  n2["plan<br/>changed"]:::changed
  n3["Host<br/>changed<br/>1 finding"]:::flagged
  n4["wait_for_module<br/>changed<br/>1 finding"]:::flagged
  n5["LabError"]:::impacted
  n6["env"]:::impacted
  n7["Value"]:::impacted
  n8["Err"]:::impacted
  n9["data"]:::impacted
  n10["other"]:::impacted
  n0 -->|calls| n1
  n0 -->|calls| n2
  n0 -->|uses| n5
  n0 -->|calls| n6
  n0 -->|uses| n6
  n0 -->|calls| n8
  n1 -->|uses| n5
  n1 -->|uses| n6
  n1 -->|uses| n7
  n2 -->|uses| n3
  n2 -->|uses| n5
  n4 -->|uses| n5
  n4 -->|calls| n10
  n9 -->|uses| n5
  n9 -->|calls| n8
  n9 -->|calls| n10
  classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
  classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
  classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
  classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading
Agent review details

critique

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: Reviewed 6 files; 1 finding. _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

security

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: Reviewed 4 files; 0 findings. 2 files were not security-reviewed: docs/crates/tinycomputer-examples/live-tasks.md (prose or tabular data), docs/crates/tinycomputer-examples/running-the-examples.md (prose or tabular data). _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

tests

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The routes() refactor is well tested — five behavioural tests pin the bearer/OpenRouter split, defaults, and error paths, and they would fail on regression. The gap is wait_until_ready in host/mod.rs: a new polling function with Ready, terminal-state, and timeout paths, none of which any test in the diff or repository exercises; its correctness rests entirely on the comment's claim. Everything else in the diff is documentation or wiring that the existing tests cover adequately. _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._
  • Evidence: crates/tinycomputer\-examples/src/host/mod\.rs — Test wait_until_ready's failure paths, not just its comment

commits

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: Nothing sensitive found in what this pull request commits.

description

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The description accurately matches the diff: the ready-wait fix in `Host::load`, the `TINYHUMANS_TOKEN`/`TASK_HEADED` routes with `agentic-v1` defaults, the updated docs, and the new tests are all present and consistent. The change looks sound and safe to merge. _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

e2e

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: No end-to-end harness in this repository: no e2e test files and no e2e workflow.
Evidence and run details
  • Models: gpt-5.6-luna, glm-5.3-flash
  • Spend: $0.002673
  • Tokens: 157566 input · 9213 output · 12152 cached · 0 embedding
Head State Pass summary
1f67db20083d ready for maintainer review 1 active finding(s), 0 resolved finding(s) (at 1791234210)

tinysweeper 0.1.0

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

             $0.0027 · 157,566 in / 9,213 out · 12,152 cached (8%) · gpt-5.6-luna, glm-5.3-flash
critique:    $0.0009 · 74,577 in  / 3,996 out · 8,314 cached (11%) · gpt-5.6-luna, glm-5.3-flash
security:    $0.0005 · 42,571 in  / 1,419 out · 3,774 cached (9%)  · gpt-5.6-luna
tests:       $0.0002 · 20,583 in  / 1,523 out · 0 cached (0%)      · glm-5.3-flash
description: $0.0001 · 10,451 in  / 187 out   · 64 cached (1%)     · glm-5.3-flash

/// Waits until the loader reports the module ready for calls. The module
/// claims its bus name while its setup is still running, and the loader
/// refuses a reconfiguration until that setup has returned.
async fn wait_until_ready(module_host: &ModuleHost) -> Result<(), LabError> {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority medium tests likely

Test wait_until_ready's failure paths, not just its comment

This new function is the behavioural heart of the host-side change: the comment claims the loader refuses reconfiguration until setup returns, and the code asserts that claim only in prose. Its three outcomes — returning Ok on ModuleState::Ready | ModuleState::Serving, returning a descriptive error on Rejected | Unresolved | Faulted | Failed | Stopped | Disabled, and the 30-second timeout — are exercised by no test in the diff, and the repository requires failure paths to be covered and at least 90% line coverage per source file. A regression that flips the match arms (e.g. treating Faulted as retryable, or timing out instantly) would ship silently. If constructing a ModuleHost in a test is impractical, at minimum factor the state-classification into a pure function over Option<ModuleState> and test that.

[RULE] untested-error-path ·

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants