Skip to content

Keep reconnecting the live screen after it has recovered - #651

Open
kevin9327 wants to merge 2 commits into
CopilotKit:mainfrom
kevin9327:fix/live-screen-retry-budget
Open

kevin9327 wants to merge 2 commits into
CopilotKit:mainfrom
kevin9327:fix/live-screen-retry-budget

Conversation

@kevin9327

Copy link
Copy Markdown
Contributor

What this changes

A live screen that loses its socket reconnects on its own: five tries, half a second apart at first and eight seconds apart at the end, then it stops and asks for Retry. The count of tries was never reset after a reconnect worked. So a screen left open for an afternoon, through five short drops that each healed within a second (a laptop's Wi-Fi hopping access points, a proxy recycling idle sockets), gave up for good on the sixth drop and showed "The live screen is disconnected. Retry to reconnect." while the computer was fine.

The count now starts over when a reconnected socket delivers a frame, the one message that proves the stream is back. A socket that opens and closes again at once never shows a frame, so it still uses up the tries and a computer that is really gone still ends in Retry after five failures in a row. A reconnect still checks who holds the computer before input is ready, since that check keys off the tries taken so far, which a drop increments before it reconnects.

Where it runs

OpenBot is deployed as several server processes behind a load balancer, serving a whole company.
Consecutive requests from the same person reach different processes, and the process that answered a
WebSocket upgrade is rarely the one that answers the next call on that conversation.

State that outlives a single request therefore has to be shared, or the change works on one machine
and stops working the moment there are two, without saying so. That failure is worse than not
shipping the feature: it passes review, passes CI, passes a local demo, and only surfaces as a Bot
that forgets, a question nobody can answer, or a boundary that never fires.

Answer these even when the answer is "none":

  • New state that outlives a request? None. The retry count is the browser tab's, as before.
  • What happens on the second replica? The same. A reconnect may reach any process; this only changes when the tab stops trying.
  • Anything serialised? None.
  • Anything fanned out to a browser? None.
  • New listener, port, or schedule? None.

Boundary and audit

  • Every acting call still goes through the gateway: resolve, decide, audit, then act. Untouched; this is the viewer's socket lifecycle.
  • New refusals and new failures each write a row. None added.
  • Nothing new is trusted from the client that the server can resolve itself. Ownership is still re-read on every reconnect.

Changelog

  • A line in CHANGELOG.md under Unreleased.

Proof

New test in app/tests/screen-connection.test.ts: six drops, each after the screen was showing frames again. On unmodified main:

- Expected  - 5
+ Received  + 4
(fail) a reconnect that shows the screen again restores the whole retry budget

The waits were 500, 1000, 2000, 4000, 8000 and the sixth drop asked for Retry. After the fix: bun test tests/screen-connection.test.ts in app → 5 pass, 0 fail, including the existing test that five failures in a row still end in Retry. tsc --noEmit in app exits 0. Biome check clean on the changed files.

🤖 Generated with Claude Code

The retry count behind the live screen's backoff was never reset, so
drops a viewer had already recovered from used up the five retries for
good. A frame on a reconnected socket now starts the count over.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

davidmckayv
davidmckayv previously approved these changes Sep 28, 2026
Move the unchanged changelog entry to its own existing anchor.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants