Skip to content

feat: a contributor PR digest, and every scheduled job on the host from one package - #65

Open
Bilb wants to merge 87 commits into
mainfrom
refactor/ops-platform
Open

Bilb wants to merge 87 commits into
mainfrom
refactor/ops-platform

Conversation

@Bilb

@Bilb Bilb commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

Two changes in one PR, because the second replaces most of the first's scaffolding before either has reached the host. They were #59 and #65, #59 is closed into this one.

  1. A daily digest of contributor pull requests, posted to Discord.
  2. Every scheduled job on the self-hosted box: Zendesk, the PR digest, the Crowdin sync and duplicate report, the snode list and release stats. They run on systemd timers, built from one package, with one HTTP transport and Discord alerts on failure. GitHub Actions keeps only tests.yml once the host has taken over.

It reviews in commit order, one phase at a time.

The contributor PR digest

Each weekday morning, one message lists the open PRs across session-foundation's own repositories whose author is not a maintainer and which have moved in the last three days.

**Contributor pull requests** · last 3 days
🟢 **2** new · ✏️ **1** updated
**36** open from contributors across the org.

**session-desktop**
🟢 [#2001](…) @example-contributor-1 · 12h · 💬1 · Fix a crash when pasting an image
✏️ [#1990](…) @example-contributor-2 · 3h · 💬2 · feat: add a proxy setting

🟢 is a PR the digest has not reported in the past year. ✏️ is one it has, which has moved since. A PR that hasn't moved is left out of the message but still counts in the backlog line.

  • Window and state. The timer runs Mon..Fri 09:30 Australia/Melbourne, so Monday's run has to cover the weekend: hence a 72h window, which overlaps itself by two days. --state absorbs the overlap. It records which PRs reached Discord and each one's updated_at. Only what Discord accepted is recorded, so a run failing on its second message re-reports that message's PRs. Any failure to read the state file treats the window as new.
  • Dedup is keyed on updated_at alone. That is a deliberate trade: a label sweep resurfaces every PR it touched. The accurate alternative, head SHA plus comment counts, costs a request per PR that moved.
  • One search fetches every open PR and the window is applied to the result, which buys the backlog line for one query. When the search comes back short of its own total, the header says the counts are a floor. That happens past GitHub's 1000-result ceiling, when GitHub flags incomplete_results, or when a PR updated between page fetches shifts the pages.
  • Maintainers are a hand-written list. Neither signal GitHub offers works. Org membership covers six accounts, two of them outside the review loop. Push access is held by a dozen outside collaborators, several of them contractors whose PRs are exactly what this is for.
  • Bots are dropped on GitHub's account type. Forks, archived and private repositories are excluded by checking results against the org's repository list, so a repository created today is covered today. Private repositories are never reported, whatever the token can see.

Operations and flags are in docs/jobs/github-prs-digest.md.

Phases

phase commits
digest The digest, then its move onto shared helpers b0702ef … 8e30ddd
0 Test and lint CI (the base had none) da30f81, c7aff0b
1 Success stamps and a silence checker f0ca8c2
2 One session_ops package: pyproject.toml + uv.lock, src/ layout, entry points, one venv 2f995e8
3 One transport (shared/http.py: urllib3 Retry + token bucket); Crowdin on crowdin-api-client 331b16c
4 Crowdin duplicate report from webhooks, reconciled daily 7bc559f
5 One page per job in docs/jobs/ (its proofreader token went with crowdin-approve-strings in c12ac33) 31fe984
6 jobs.toml registry, session-ops run, one templated unit, three-layer alerts, install.sh 092505f
6 Crowdin sync, snode list and release stats on the host; publishing via git + REST, GitHub App tokens c690b4b
6 Fixes from the live Crowdin runs and the container rehearsal; CroQL on 6518314, 56938a1, 0aa7867
— triage.py split into zendesk/api.py, claude_cli.py, transcript.py: a verbatim move, then the comment-fetcher merge and the --model alias fix on their own 8c2671f, 3e27afa, ff60382, bfc2e25
— Review fixes: GitHub's rate-limited 403 and its reset, search paging, a year of digest state, a posting test 70699dc … 5082b19
— Second review: see below cb11d17 … b78efa4
— Third review: see below bd6378e … 40e0eda
— Fourth review: see below 60d9d6c … 3d9ecf7
6 Delete the replaced workflows after two host runs of each

How it was checked

  • Goldens. Only the systemd drop-ins generated from jobs.toml are pinned, so a schedule or sandbox change shows as a diff. Job output is checked against main instead (below).

  • The digest against the live org and a test webhook:

    106 open PRs in session-foundation, 36 from contributors across 34 repos
    run 1  no state          31 new,  0 changed,  0 unchanged   → posted, 31 recorded
    run 2  state present      0 new,  0 changed, 31 unchanged   → header only
    run 3  one entry rewound  0 new,  1 changed, 30 unchanged   → one ✏️ line
    

    A busy-day render split into two messages, 3830 and 743 characters, on a repository boundary.

  • Release stats. In Python, fed the same live releases as the TypeScript it replaces, release-stats writes byte-identical CSVs.

  • Snode list. A live session-ops run snode-list --dry-run fetched the list, sparse-cloned session-ios and found the change to publish.

  • systemd. The unit template and every generated drop-in pass systemd-analyze verify, and each schedule passes systemd-analyze calendar.

  • Publishing. It is tested with real git against bare repositories on disk.

  • Crowdin, live and read-only. The CroQL narrowing caught a planted duplicate in fr that a full scan also found, and found nothing anywhere else. That matches the old report's last full cycle. crowdin-sync --dry-run downloaded and generated all three platforms. The localization module has no changes against main. Android and iOS differ from the workflow's own bot branches only by that planted translation.

  • Install. deploy/install.sh ran in a systemd container on Ubuntu 24.04: a fresh install, a second run once the env files were filled, then an idempotent third. Every job started under its own account. Failures alerted once from the run while the backstop stayed quiet, and the backstop reported a forced timeout. The Zendesk relay answered correctly under its sandbox.

Against main

Every job that moved was run on both sides against the same live data, with nothing pushed or posted. main's side replays each workflow step by step with its own scripts, including the artefact hand-offs between jobs; this branch's side runs each job with --dry-run. Both started from the same commit of each target repository.

job result
Crowdin sync: session-android and session-ios PRs, session-localization push patches and resulting files identical
Crowdin download and parse identical
Static snode list: session-ios PR identical
Release stats both CSVs identical
Duplicate translations: old report vs. the ported one, all 80 locales findings and Discord payload identical: a planted fr duplicate and 3 in kmr
Duplicate translations: reconciliation from no state the same 4 slots
Zendesk resolver dry run, digest --dump-batch identical

What the patches do not show:

  • Commit author. The workflows commit as github-actions[bot]; the jobs commit as PUBLISH_GIT_AUTHOR.
  • A locale Crowdin drops. The workflow extracts the generated files over a fresh checkout, so the dropped locale's strings.xml stays; crowdin-sync deletes it.
  • Android's Gradle check. The workflow runs generatePlayDebugResources before opening the PR; crowdin-sync leaves validation to session-android's own CI.

Second review

  • Alerts: a dead Claude login is named in the run's own alert (4baa241); an optional target's failure whose alert can't post fails the run (b61568b); webhook tokens stay out of the journal (b7531af); no role mention anywhere (0028edb).
  • install.sh: it refuses any clone but a root-owned /opt/session-ops (f908c86), stops old units before copying their state and unschedules stale jobs (9b8321e), and warns while alerts.env is empty (4402649).
  • Crowdin: duplicates on deleted strings or removed locales resolve (cb11d17); reconciliation refuses to run unseeded (c88d8a8); crowdin-sync and snode-list move to 13:00 Melbourne, clear of the workflows they replace (bf4d4a1); the real generators are under test (9e2c2a9); crowdin-approve-strings is removed (c12ac33).
  • Shared: a long rate-limit reset comes back at once instead of burning the timeout (5101d53); a corrupt state file is a cache miss (1f8ab9b); a rejected push quotes the server's reason (9395906).
  • Zendesk: replies keep their line breaks (971241d), transcripts use the resolved model id (affa809), a failed user lookup no longer fails the digest (34166d8).
  • Docs: deploy/README.md cut from 543 to 331 lines (4030647); stale paths and the apt repository fixed (6faccfc, e61def7, 7d34cd7).

Third review

  • Discord: no digest message can ping anyone (c587878); a PR title's links show as their URL, never as masked text (a8e8ed8).
  • Zendesk: a command whose author cannot be looked up fails the run and stays claude-queued, rather than reading as refused and being dropped (bd6378e).
  • Deploy: a crash-looping relay reaches its start limit and alerts (e2deeeb); install.sh refuses a tree any account but root can write to (c179739); each job's state is private to its account, directories 0711 so the backstop can still read the alerted marker (fe7af48).
  • Crowdin: a missing duplicates state is refused, not recreated empty (dc89c88); reconciliation's run time is the measured 9 minutes, 85 without CroQL (aede503); the relay is gone, leaving the daily reconciliation as the only writer, holding the state's lock for the whole run (40e0eda).
  • Publishing: snode-list refuses an empty or malformed list (015222a); a relative work directory clones where it names (744a873).
  • Shared: Discord's embed limits and packing live in shared.discord (44ee932).
  • Removed: the recorded-response goldens and their tests (a5e0e80); ZENDESK_NOTE_AUTHORS and ZENDESK_NOTE_MODEL (ea894b7); example handles that are real accounts (e4a8f7b).
  • Rehearsal: a host with /etc/session-ops/rehearsal publishes [Rehearsal] pull requests from rehearsal/ branches, opens a pull request instead of pushing to session-localization's main, and solves no Zendesk tickets (2ed2741).

Fourth review

  • Discord: a PR title's own backslashes are escaped, so \[x\](url) cannot re-open a masked link, and an uppercase scheme is wrapped too (60d9d6c).
  • Rehearsal: a marker that cannot be read fails the run rather than publishing for real, and a rehearsing run says so (bfba06e); it writes no English transcripts to live tickets (e78406c); a rehearsal with nothing to publish retires only its own branch, now tested (49ed2dc); ending one stops its timers, and the marker is removed only to promote the host (a6bab44).
  • Zendesk: claude-queued stays only on a command a retry can serve: an author Zendesk no longer has is a refusal, and a handled command needs no lookup (680f5ef).
  • Deploy: install.sh starts a relay that hit its start limit (9fd804d); the README's one-off runs keep the units' state permissions (474dd44).
  • Docs: the snode guard, a CroQL check that actually runs, and the sogs-perms link (0be3d07).
  • Deploy README: down to 110 lines; what each env file takes is in deploy/env/<name>.env.example, which install.sh copies beside it (3d9ecf7).

Not checked live

  • Anything on the real host, and a GitHub push with a real token.
  • A real GitHub rate limit. The 403 and reset handling is tested against a local server only.

Decisions worth a look

  • No Crowdin relay. Duplicates are posted by the daily reconciliation alone, within a day of appearing and before the weekly export. That leaves no public endpoint whose URL is a secret, and no second writer to merge with.
  • Opening PRs. PRs are opened over the REST API rather than the gh CLI; pushes are plain git.
  • The triage.py split. 3e27afa moves code and nothing else. Each of the 112 top-level definitions keeps its source text, checked by comparing every definition before and after. Against the live Zendesk account, the digest's --dump-batch and the resolver's dry run are byte-identical before and after the whole series. Names are unchanged, so callers differ only in which module they name.
  • Publishing needs the GitHub App. GITHUB_PUBLISH_TOKEN is gone; without GITHUB_APP_ID and its key, crowdin-sync and snode-list exit naming what is missing (38a9069).
  • The PR digest's window reaches back to the last fully posted search, capped at the state's 365-day retention, so a DST weekend, an outage or a partial post loses nothing (e37e388). The Zendesk digest keeps its fixed 72 hours.
  • 403 retries. The transport retries a 403 only when it carries Retry-After or x-ratelimit-remaining: 0, GitHub's two rate-limit signals. It waits for x-ratelimit-reset only in the second case, because GitHub sends that header on every response.

Bilb added 30 commits September 24, 2026 11:43
One message each weekday morning listing the open PRs across session-foundation's
own repositories whose author is not a maintainer and which have moved in the
last three days.

A single search fetches every open PR in the org and the window is applied to
the result, so the header can carry the total contributor backlog for the cost
of one query. Forks and archived repos are excluded by checking results against
the org's repository list rather than by name, and bot accounts on GitHub's own
account type.

Weekdays means Monday has to cover the weekend, so the window is 72h and
overlaps itself by two days on every run. A state file absorbs the overlap: 🟢
is a PR never reported, ✏️ one that has moved since it was, and a PR that has
not moved is left out. Only what Discord accepted is recorded, so a run that
fails partway re-reports that message's PRs rather than losing them, and every
way of failing to read the file treats the window as new — noisy once, never
wrong.

Dedup is keyed on updated_at, which moves on any change at all, so an edit
touching several PRs at once resurfaces all of them. The accurate alternative
needs a request per PR that moved; this is the cheaper half of that trade.

Maintainers are a hand-written list. Org membership covers six accounts, two of
which are not in the review loop, and push access is held by a dozen more as
outside collaborators — several of them contractors whose PRs are the point of
the digest. Both signals would get it wrong in both directions.

Runs on the box that already hosts the Zendesk digest, under its own user, env
file and venv — the venv because the two jobs pin requests differently.
A block over the text budget was posted whole and rejected, and since nothing in
a rejected message is recorded, the same block was rebuilt on every run until the
PRs aged out of the window. Blocks are now split under a repeated heading, each
sized to fit beside the header.
…s in the docs

Exclusion is by repository property, so nothing private is named anywhere in the
repo. The token guidance drops the option of a `repo` scope with it: the digest
posts to Discord, and nothing about a private repository belongs there.
…eads

Forks, archived and private repositories are always left out; there was no run
that wanted them back in.
The envelope's `result` carries the reason whichever way the CLI exits, and
the exit-1 path already reads it. The exit-0 path reported only the subtype
and the status code, which does not say what went wrong.
…ers into shared/

The pull request digest was about to carry its own copy of each; crowdin's
report already does. One copy, imported by every script out of the same clone.

triage.py keeps get_env, request_with_retry, clip and post_to_discord in its
namespace: note_reply, resolve_reviews and the alert reach them there, and the
tests patch them there.
The digest carried its own copy of the retry loop, the dedup state file, the
Components V2 constants, the chunker and the webhook posting. It now imports
them from shared/, and keeps only what is its own: the record it stores per
PR, its per-message block cap, and the rendering.

The tests that covered the copies move with the code; what stays here covers
the wrappers and the rendering.
…imeout

RFC 9110 allows Retry-After in either form. Crowdin's report already parsed
the date form, so the shared loop has to before it can replace that copy; a
date already gone reads as no header, since time.sleep() rejects a negative.

The timeout was fixed at 30s; Crowdin's scan runs on 60.
…k code

The third copy of the retry loop and the Discord posting. What stays is
Crowdin's own contract: ten attempts at a 60s timeout, a 4xx that raises so
callers can read the body unchecked, and the plain-text warning posted before
the run exits on a rejected embed.
…ponses

Each golden replays a recording of API responses through the unchanged script and
compares its output byte for byte, so the packaging and transport changes that
follow cannot alter a payload unnoticed.

- digest: a live --dry-run over a 720 h window (2026-09-25), trimmed to the fields
  the digest reads; its replay is byte-identical to the live output. A second case
  adds a state file, covering the new/changed/unchanged split.
- report: --locales de over a synthetic project with duplicates in plain and
  plural slots, both the Discord payload and the --json findings.
- download: the export payloads and the files written, invoked as the sync
  workflow invokes it.

RecordedSession matches requests by method, URL, query and body rather than
order, because both Crowdin scripts fan requests out across threads.
One job per suite, since two requirements files pin requests differently.
sogs_moderation runs on the system interpreter with a venv that can see
python3-session-util, which ships as a deb rather than a wheel.

ruff is limited to pyflakes and syntax errors, the class of mistake a large
refactor introduces. The one finding, an unused import in relay.py, is removed.
…ails

OnFailure= reports a run that failed and nothing about one that never happened: a
timer left disabled, a unit renamed, a host down through a whole schedule.

Each digest unit now touches /var/lib/session-ops/stamps/<name> from a `+`
ExecStartPost=, which a oneshot runs only after every ExecStart= succeeded and
which runs outside the sandbox, so neither job gains a write path. silence.py,
on an hourly timer under an account of its own, compares each stamp's age with
deploy/jobs.toml and posts one message for every job past its max_age_hours,
repeated daily while it stays quiet. A missing stamp is timed from the first check
that found it missing, so installing the checker alerts on nothing.

The stamps are root's and world-readable, so the checker needs no access to the
Zendesk state directory and the ticket data in it. test_silence.py fails when a
shipped timer has no registry entry or no stamp line.
pyproject.toml and uv.lock replace the five requirements files, which pinned
requests two different ways and so needed a venv each. The code moves to
src/session_ops/{shared,github_prs,zendesk,crowdin,sogs,monitor} with absolute
imports, so no module inserts anything on sys.path; tests move to tests/ in the
same shape and run as one discover.

Every script is a console entry point, and the units call those out of one venv
at /opt/zendesk/.venv, built with the system Python: a uv-managed interpreter
would sit under /home/zendesk, which ProtectHome=yes hides from the units that do
not run as zendesk. test_units.py fails when a unit names an entry point the
package does not declare. The relay runs note_reply with -m rather than by path.

download_translations_from_crowdin.py parsed its arguments at import; that moves
into main(argv) so it can be an entry point. babel stays pinned exactly, since its
CLDR data is what the generated language names come from.

CI runs one locked environment instead of a matrix. tests/sogs skips itself
without session_util; its own job checks the import first, so the skip cannot hide
a broken suite. The Crowdin workflows install the package and run modules until
they move to the host.

All 586 tests and the four goldens pass unchanged, and a live digest dry run
through the new entry point matches its golden apart from PR ages.
shared/http.py replaces shared/retry.py. Session is a requests.Session with a
urllib3 Retry mounted, so retries happen below the caller, which sees only the
final answer: 429 and every 5xx retried for all methods, 1 s doubling to 30 s,
Retry-After in either form capped at a minute, GitHub's x-ratelimit-reset when
Retry-After is absent, and an unusable header falling back to backoff. It adds a
default timeout, an optional token bucket, and a per-call attempts= for optional
data, carried to the adapter in a ContextVar because requests has no way to pass
one down. Its tests run against a local server, so the retries tested are
urllib3's own.

The Crowdin scripts move to crowdin-api-client through crowdin/sdk.py. The SDK
never retries a 429 and retries 5xx on a fixed 100 ms, so its session is swapped
for per-thread http.Sessions drawing on one 30/s bucket, and its own loop is off.
Its decoder also turns every ISO timestamp into a datetime, which json.dump cannot
write: the download's project and glossary files would have failed on the first
real payload. Responses stay plain JSON.

The recordings now carry timestamps as Crowdin sends them, and approve_strings.py
gains a golden. Every golden was regenerated by the previous code from the same
recordings, and the new code matches it byte for byte; only request keys the SDK
spells differently (offset=0, a query moved into params) were edited. A live
digest dry run matches its golden apart from PR ages.

Caller tests that queued a retry's worth of failures now queue one, and assert the
budget they ask for instead.
… daily

The sharded daily scan saw each locale once every eight days and could say nothing
about slots that had been resolved. This keeps the open slots, keyed (string,
locale, plural category), and posts only what changed.

- crowdin-relay.service takes Crowdin's suggestion events, acknowledges at once and
  re-checks the one (string, locale) named. Crowdin signs nothing, so the secret is
  the last path segment, and a wrong or unset one is a 404. It is a process and
  account of its own rather than a route on the Zendesk relay, which can write
  public comments and has no business holding the Crowdin token.
- crowdin-reconcile-duplicates, daily, judges every string of every locale and posts
  new and resolved slots, or nothing. Crowdin never retries a webhook, so this is
  what keeps the state correct. --croql narrows each locale to the strings CroQL
  counts two translations for; it stays off until that query is checked live.
- Both write the state under a file lock. The relay records when it checked each
  (string, locale), and a scan's older view of that scope is ignored, so a scan
  that began before a suggestion landed cannot resolve what the event opened.
- The state is written only once every message landed; a run that fails to post
  repeats its whole diff next time. --seed records the current slots without
  posting them.

report_multiple_translations.py and its workflow stay until a reconciliation cycle
has run clean on the host; the two now share the slot logic, and its golden is
unchanged.
Each job's section of the README moves to docs/jobs/<unit name>.md, opening with
where it runs, which secrets it reads, how to dry-run it, how to re-run it and where
it logs. The text moves as it was; only headings and relative links change. The
README becomes the index, and tests/test_docs.py fails when a job in the registry
has no page. crowdin-duplicates and session-ops-silence are new pages; the SOGS
tools, which nothing schedules, move to docs/tools/.

crowdin-approve-strings is the only script that writes to Crowdin, and it read the
same token as every read-only job. It now reads CROWDIN_PROOFREADER_TOKEN, or the
keyring's proofreader-api-token, and never falls back to CROWDIN_API_TOKEN, which
can therefore be read-only everywhere else.
…iled

src/session_ops/jobs.toml is the registry: each job's entry point, arguments,
account, env files, required variables, schedule and allowed silence. From it,
`session-ops units` writes each job's drop-ins over deploy/session-ops@.service
and .timer, which hold the hardening once; `session-ops run <job>` runs a job the
way its timer does; and the silence checker watches every scheduled one.

A run checks the variables its job needs, gives it a scratch directory in the
unit's private /tmp, and reports its own failure: the job, the host, the step it
was on, one sentence of error with secrets and webhook URLs scrubbed, each target's
result for a job with several, and the commands to read the journal and re-run it.
The traceback stays in the journal. Having alerted, it records its invocation id,
and the OnFailure= backstop stays quiet for that invocation, speaking only for what
a run cannot report: killed, timed out, Discord unreachable, and the relays.
ALERT_DISCORD_ROLE_ID, which the GitHub notifier mentioned, is mentioned by both.

The Zendesk resolver and digest become one job with two targets, the resolver
optional: its failure is reported and the digest still runs and succeeds, which
is what the `||` in the old unit did. The tests that read that unit now assert the
same order and failure handling on the job itself.

deploy/install.sh replaces the numbered install steps and is idempotent: accounts,
venv, env files created empty at 0600 (systemd reads them as root, so no account
can read another's), tmpfiles, units and drop-ins, then it enables each job whose
env file has content. It moves the previous layout's env and state files across
where the new ones do not exist yet, and retires the per-job units. Code now lives
in /opt/session-ops, owned by root; kept run copies under
/var/lib/session-ops/<job>/runs are pruned after 14 days by tmpfiles.
crowdin-sync is one process where the workflow was eight jobs passing artefacts:
download, parse and validate, then Android, iOS and the localization module as
independent targets, so one failing to publish does not stop the others and the
alert says which. Each platform is a shallow, sparse checkout of only the paths its
generator writes. It keeps today's publishing: a pull request from
feature/update-crowdin-translations rebuilt from dev and force-pushed, and a commit
straight onto session-localization's main that overwrites the generated files
without removing others, which is what the workflow's misdirected `rm` amounted to.
The Gradle validation step is gone; Android's CI checks the pull request. The run's
downloads, parsed JSON and validation report are kept under runs/ for 14 days.

snode-list copies the dynamic-assets list into session-ios, byte for byte and only
once it parses as JSON. release-stats is the TypeScript script in Python, on demand;
fed the same live releases, both wrote byte-identical CSVs, which are now its golden.
Node leaves the repo with it.

Publishing replaces peter-evans/create-pull-request and git-auto-commit-action with
git and the REST API: an unchanged tree is not pushed again, and nothing left to
merge closes the pull request and deletes the branch. It authenticates as a GitHub
App when one is set up, minting a one-hour token scoped to the repositories the run
pushes to, and otherwise with GITHUB_PUBLISH_TOKEN, the PAT the workflows used. The
App's key reaches only the unit, through LoadCredential=. The token reaches git in
GIT_CONFIG_* variables, never on a command line.

The workflows these replace keep their schedules until the host has run each twice.
A dry run printed an empty stat: it diffed the index against HEAD after
committing. It now shows the commit.

A failed git command quoted the first line of stderr, which is whatever a wrapper
or credential helper printed first. It now quotes git's fatal: or error: line.
Checked live on 2026-09-25. Across all 80 locales CroQL returned only plural strings,
which it cannot tell from duplicates, until a second suggestion was planted on a
singular string in fr: that string was then the one singular candidate, and the
narrowed check and a full scan of fr found the same slot. A locale now costs about
50 requests instead of 1,371.
Found by installing and running every unit in a systemd container.

- The OnFailure= backstop quoted the unit's last journal line whatever the run: a
  run killed by its timeout before logging anything was reported with the previous,
  successful run's "Written to ...". It now reads only the failed invocation, and says
  how the unit failed (timed out, killed, out of memory).
- An error was cut at its first newline, so GitHub's pretty-printed 401 body read
  as `{`. Errors are now whitespace-collapsed and clipped.
- A Crowdin error read as the SDK exception's repr; it is now "Crowdin 401:
  Unauthorized".
- A job that names no step no longer fails "during starting".
…y the digest

note_reply reaches it too, and has no --batch-size to lower.
…t of triage.py

note_reply, resolve_reviews and the failure alert imported the 1,900-line
classifier to reach a ticket fetch, a marker or the CLI's name. Those now live in
modules of their own:

  zendesk/api.py         the session, search and its 1000-result ceiling, a ticket,
                         its comments and users, requester activity, the markers,
                         the customer's side of a conversation, store reviews
  zendesk/claude_cli.py  the CLI call, the model aliases, a failure worth quoting
  zendesk/transcript.py  the English transcript: detect, render, write, attach
  shared/text.py         undash_english and squash, which name no service

triage.py keeps the queries, the taxonomy, the dedup, the review filter, the
analysis and the rendering: 1,160 lines. resolve_reviews and the alert no longer
import it at all.

Nothing is renamed or rewritten. Each of the 112 top-level definitions in the old
triage.py is in the new tree with its source text unchanged, checked by comparing
every definition's source segment before and after. Callers and tests change only
which module they name, and a patch now lands on the module where the name is
looked up. Against the live Zendesk account, the previous commit and this one
produce a byte-identical --dump-batch and the same resolver dry run.
…ir own

A command's read, the transcript's and the description hydration's differed only
in page size, order, retry budget and what a failure does. Those are now
fetch_comments' parameters; optional= is the enrichment's budget and failure mode,
a note and None rather than an exit.

Each caller sends the request it sent before: test_api.py pins all three, and
passes against the previous commit too. The one visible change is that hydration
now prints a note when Zendesk answers it with an error, where it skipped the
ticket silently.
note_reply passed --model to the CLI as given, so `opus` meant whatever the CLI
calls opus that week, while the digest maps the same word to a pinned id. Both go
through resolve_api_model now; a full id still passes through.
…n the quota is spent

GitHub answers a rate limit with a 403 as often as a 429, so the transport never
retried it. It also sends x-ratelimit-reset on every response, so a 5xx without
Retry-After waited for the rate-limit window, capped at a minute per attempt,
instead of backing off. The reset is read now only when x-ratelimit-remaining is 0,
and a 403 is retried only when it carries Retry-After or an exhausted quota.
…mplete search as a floor

The search is sorted by update time, so a PR updated between the two page fetches
moves onto page one after it was read and pushes that page's last item onto page
two: that PR was listed twice and the updated one missed. Results are now deduped
by id, and a search GitHub flags incomplete_results reads as a floor like the
1000-result ceiling does. The page loop is bounded by the ceiling rather than by
how many items have arrived, which a repeating page would never grow.
…eads as updated

State entries were pruned 30 days after a PR was last reported, so a contributor
PR that went quiet for a month and then got a push came back as 🟢, which the docs
define as never reported. The default retention is now 365 days, a few KB of
state, and the docs say what 🟢 means under it.
The property the state file exists for lived only in main(), which no test ran
past a dry run. These run main() from the fetched PRs to the state file with
Discord faked to accept all, some or none of the messages.
sparse_clone runs git clone from the parent directory, which resolved a relative destination against that parent again: work/session-ios landed in work/work/session-ios and the sparse checkout then failed. The runner always passes an absolute directory; a hand run with a relative SESSION_OPS_WORK_DIR did not.
The usage text called the bare invocation what the timer runs, contradicting the next example; the timer passes --window-hours 72 and --state.
A full --croql run against the live project took 509 s across 80 locales, not the 20 minutes stated, and a locale scanned without CroQL took 62 s, about 85 minutes for all of them rather than an hour.
container_message sent no allowed_mentions, so Discord parsed every mention in a Components V2 message's text. The PR digest quotes titles and logins written by outside contributors, so a PR titled "@everyone please review", or opened by an account named everyone or here, pinged the whole channel every weekday it moved. The Zendesk digest had the same gap with ticket text. The runner and backstop alerts already sent {"parse": []}.

Every container message now sends it too. Nobody needs pinging from a digest; the channel is read when people can.
A contributor could title a PR "[Download the build](https://example.com)" and the digest rendered it as a masked link carrying the digest's authority. Square brackets in titles are now escaped, so a masked link shows literally, and every URL is wrapped in <> so Discord shows it plainly without a preview.
The first real runs of crowdin-sync and snode-list would otherwise push to the bot branches the GitHub workflows still own and to session-localization's main, and zendesk-digest would solve live tickets, so the only way to see what the host does was to let it do it.

While /etc/session-ops/rehearsal exists, the runner marks every job as a rehearsal. Pull requests then come from rehearsal/ branches with a [Rehearsal] title and a do-not-merge note, a direct push opens a pull request against its branch instead, and the Zendesk resolver only reports. deploy/README.md covers the rest of a rehearsal host's setup and the clean-up.
Comment thread deploy/README.md
Comment thread tests/goldens/digest/dry-run-with-state.txt Outdated
Comment thread tests/goldens/digest/dry-run-with-state.txt Outdated
Comment thread tests/goldens/digest/responses.json Outdated
Comment thread docs/jobs/github-prs-digest.md Outdated
Comment thread tests/goldens/reconcile/dry-run.json Outdated
Comment thread tests/goldens/release-stats/session-android.csv Outdated
Comment thread tests/goldens/release-stats/session-desktop.csv Outdated
Comment thread tests/goldens/report/dry-run.json Outdated
Comment thread tests/goldens/report/findings.json Outdated
… needs

The digest's example and tests named @octocat and @monalisa, both real GitHub accounts; they are now example-contributor-1 and -2, which do not exist. The digest page loses its case against org membership and push access as the maintainer list, and the SOGS ban page its walkthrough for letting a test account post.
…dule

The duplicates reconciliation and the old report each defined Discord's per-message embed count and text budget, plus identical embed_len and pack_embeds, beside a 3800-character description cap that was a margin rather than Discord's 4096. shared.discord now holds the limits Discord enforces and the packing, and both callers use them; a description is split at Discord's own limit.
ZENDESK_NOTE_AUTHORS narrowed who may command the relay below any agent or admin, and ZENDESK_NOTE_MODEL changed the note model's default; neither is used. The command check is the role check alone, and --model still overrides the model per run.
The reconciliation and relay tests read their Crowdin exchanges from the report golden's responses.json, and one compared its payload with the reconcile golden. The synthetic exchanges now live beside the tests, and that test checks the payload's counts, sections and strings instead.
…op-ins

The digest, Crowdin report, download and reconcile goldens replayed recorded API responses and compared the output byte for byte; the digest's recorded real contributors' logins and titles. They and their tests are gone, along with the docs test. What they protected is covered by each job's unit tests, and the output was compared against main's workflows on live data before this. The drop-ins golden stays: it makes a schedule or sandbox change visible in review.
…icates state alone

The relay only made a new duplicate show up within seconds instead of by the next daily run, and a duplicate only matters at the weekly export. It cost a public endpoint whose URL is the secret, a Crowdin webhook, a unit and an nginx route, and a two-writer merge (per-scope check times, remember/forget) that is where this PR's reviews kept finding races.

Reconciliation is now the only writer. apply() judges the scopes a scan saw, the state holds just the open slots (an older file's check times are dropped on load), and a run holds the state's lock throughout, so one started while another is scanning exits instead of overlapping. install.sh stops and removes an installed crowdin-relay.service.
@Bilb
Bilb force-pushed the refactor/ops-platform branch from ee3657f to 40e0eda Compare September 28, 2026 07:49
Bilb added 11 commits September 28, 2026 17:58
…ed link

show_links escaped brackets but left backslashes alone, so a title written as \[text\](url) came out as \\[text\\](url): Discord reads the doubled backslash as a literal one and the bracket after it as live, and the masked link was back. Backslashes are now escaped first. The URL match also ignores case, so HTTPS:// is wrapped like https://.
A rehearsal gated the resolver but not the digest's transcripts: with production's zendesk.env copied over, ZENDESK_ENGLISH_FIELD_ID made the digest write its English rendering onto live tickets. A rehearsal now passes no field, so attach_english writes nothing. rehearsing() moves to shared.env, since the Zendesk jobs read it too and platforms.publish is GitHub publishing.
…when it does

The runner checked the rehearsal marker with os.path.exists, which is False when the directory cannot be read as well as when the file is missing. So /etc/session-ops locked down to 0700 made a rehearsal host publish for real, force-pushing production's bot branches and solving live tickets, with nothing in the journal to say so.

A marker that cannot be read now fails the run before the job, through the usual alert. A rehearsing run prints one line saying so. The test records what the job itself sees while it runs, not the environment after it.
The no-change path closes a pull request and deletes a branch, and no test pinned which one during a rehearsal: had the prefixing moved below the changed() check, a rehearsal would have retired production's open translations PR with every test passing.
…ess promoting

The clean-up closed the rehearsal pull requests but left every timer running, so the next snode-list or crowdin-sync run opened them again, and its closing line read like a teardown step when removing the marker turns the host into a second publisher racing production on the same branches and main.
A failed author lookup exited to keep the tag, but fetch_user answered {} for a 404 as for a timeout, so a deleted author left the ticket queued for good. And the done-marker check ran after the lookup, so a lookup failing on the run Claude's own outcome note triggers left a served command looking stuck.

api.lookup_user tells a user Zendesk no longer has ({}, read as a refusal) from a lookup that could not be made (None, which exits and keeps the tag). A command already carrying its done marker returns before any lookup; its author was allowed when it ran.
With the five-minute start-limit window, re-running install.sh soon after the relay crash-looped had systemctl start refuse it, and set -e stopped the script before the alerts.env warning. The limit is cleared before the relay is started.
The seed and dry-run commands set StateDirectory= without StateDirectoryMode=, which resets crowdin-duplicates' directory to the default 0755 while the job keeps 0711, and they and the Zendesk catch-up run wrote their files under umask 022, readable by every account. They now pass the same StateDirectoryMode=0711 and UMask=0077 as the units.
…t runs

The snode-list docstring and page still said the list was published once it parsed as JSON; it now has to hold service nodes. The documented CroQL check ran through session-ops run, which always passes --croql and posts for real; it is now two dry runs of one locale, with and without CroQL. The README pointed sogs-perms at a page that no longer covers it; it now links the tool's own docstring.
…ile's keys beside it

The README had grown back to 366 lines, most of them restating install.sh or listing every env variable. What each env file takes now lives in deploy/env/<name>.env.example, which install.sh copies next to the file in /etc/session-ops; the README keeps the install, the one-off steps install.sh cannot do (the Claude login, certbot, the Crowdin seed, now with --croql), a short check list, rehearsal, updating and recovery.
@Bilb Bilb self-assigned this Sep 30, 2026
@Bilb
Bilb marked this pull request as ready for review September 30, 2026 07:10
@Bilb Bilb assigned mpretty-cyro and unassigned Bilb Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants