Skip to content

Latest commit

 

History

History
334 lines (274 loc) · 21.6 KB

File metadata and controls

334 lines (274 loc) · 21.6 KB

Pithead Integration

The contract between a RigForge worker and the Pithead dashboard. RigForge works against any RandomX pool, but it's built as the companion miner for Pithead, a self-hosted Monero + P2Pool + Tari mining stack.

There are two connections between a worker and the stack:

  1. Mining: the worker → the stack's stratum proxy on :3333.
  2. Stats: the dashboard → the worker's XMRig HTTP API on :8080.

1. Mining connection (:3333)

Point a pool at the stack: { "pools": [{ "url": "your-stack:3333" }] }, the stack's xmrig-proxy endpoint (its proxy listens on 3333). The stack handles pool selection, payouts, and the P2Pool/XvB split centrally, so the worker config stays minimal and you never put a wallet address in it.

  • The XMRig pool user field is just a label for the rig. It defaults to the hostname (set pools[].user to name it) so you can tell workers apart on the dashboard.
  • Point as many workers as you like at the same stack endpoint; the stack aggregates them.
  • LAN workers talk to the pool over plain Stratum and need no Tor. An .onion pool needs a Tor SOCKS proxy that the DIY operator installs and runs; RigForge only configures its address.
  • The endpoint must be reachable from the worker; if the stack host has a firewall, allow the Stratum port (3333) on the LAN.

Stratum authentication (optional)

By default the stack's :3333 is open: any rig that can reach it may mine, and the pool pass is ignored (RigForge defaults it to "x"). If the operator turns authentication on by setting p2pool.stratum_password on the stack, the proxy then rejects any rig whose pass doesn't match: XMRig logs Permission denied and the rig won't mine. Put that same secret in the rig's pool pass:

// config.json — set "pass" to the stack's p2pool.stratum_password
{
    "pools": [
        { "url": "your-stack:3333", "pass": "the-stratum-password" }
    ]
}

Then ./rigforge.sh apply (or setup) regenerates the worker config with the new password.

On a fresh rig you don't need to edit anything: the interactive first run (setup with no config.json yet) prompts for the stratum password and writes it for you — press Enter to skip it if your stack doesn't use one.

  • It's the same secret on every rig. The operator finds it on the stack side: it's printed after pithead apply/setup, stored in the stack's .env as PROXY_STRATUM_PASSWORD, and shown by pithead status.
  • The password travels cleartext over your LAN's plain Stratum, so this is access control (who may mine), not encryption. Keep :3333 on a trusted LAN (the stack's p2pool.stratum_bind / a firewall do the rest) — or add stratum over TLS below for confidentiality.
  • This is unrelated to the DONATION knob (that's this rig's dev-fee donation) and to the optional API ACCESS_TOKEN below (that gates the read-only stats API on :8080, which is open by default).

Rotating the password

  1. On the stack: change or regenerate p2pool.stratum_password, run pithead apply, and read the new secret from pithead status.
  2. On each rig: set pools[].pass in config.json to the new secret and run sudo ./rigforge.sh apply.
  3. Until step 2 lands on a rig, it logs Permission denied and drops off the dashboard — that's the expected signal that it still has the old secret, not a fault (see Troubleshooting).

Sister API (optional, :8081)

Set "api": "enabled" (+ sudo ./rigforge.sh apply) and the worker serves a second read-only HTTP endpoint the stack can consume for data XMRig doesn't know (the enriched feed for pithead#235):

  • GET /1/summary and GET /2/summary — XMRig's own body passed through verbatim (a strict superset: everything the :8080 probe returns is here unchanged), a top-level UTC generated_at timestamp for freshness checks, plus one namespaced rigforge object: version/xmrig_version/xmrig_commit (provenance), tune (applied overrides, last run's target/best/candidates, the autotune schedule), power (RAPL watts and hashrate-per-watt over a 1s window; null when unmeasurable), health (the doctor probes as JSON: HugePages, MSR state, governor, RAM channels/speeds, XMP and SMT state, throttling), watchdog (armed state, thermal-hold, max_temp_c), config (the effective writable config — exactly the control-path allowlist, pool secrets masked to {"__secret__": true}; see the prefill note in §3), and config_meta ({revision, changed_at, source, last_change_id} — revision is a content hash of the writable config that changes iff that config changes, so a poller can detect a change made directly on the rig; source is control/local/restore; see §3), control (the last control-path outcome, {change_id, status, reason} mirrored from the control status file, so a poller that missed a slow rollback catches the terminal outcome here without dialing the control port; null when the rig has never taken a control change or the status file is unreadable), and control_history (the same {change_id, status, reason} shape, one entry per outcome in the control path's own changes/ ring — up to the last ~20, unordered — so a caller with only this feed and no control token can still resolve an OLDER change_id that a later change on the same rig has since overwritten in control; [] when the rig has never taken a control change).
  • GET /health and GET /tune — the rigforge.health / rigforge.tune objects bare; /health carries the same generated_at timestamp as the summaries from that refresh pass.
  • When XMRig's own API is unreachable the response is still 200 with "rigforge": {..., "xmrig_api": "unreachable"} — a down miner is exactly when the health data matters.

When a token is set, :8081 accepts either the exact bearer for existing clients or the lowercase hex HMAC-SHA256 derived with that token as the key and rigforge:api-read:v1 as the message. The derived bearer grants reads only: :8082 rejects it and continues to require an exact control token. Derivation requires a random token of at least 32 ASCII characters; generate one with openssl rand -hex 16. This rule applies per token: ACCESS_TOKEN and every named ACCESS_TOKENS entry each get their own derived read bearer, and :8081/:8082 accept any of them (see one token per stack below). When no token is set at all, :8081 remains open. The release gate's network phase enforces the boundary on the wire: the miner's only TCP peers are the configured pool, :8081 exists exactly while enabled, and no response byte ever contains ACCESS_TOKEN or a pool pass. :8080 stays the canonical Pithead summary probe; :8081 is additive. Port/bind are api_port/api_bind. Architecture mirrors XMRig's own API: a tiny persistent server ships pre-computed state (a request costs microseconds — polling cannot shave hashrate), refreshed every 15s by a wall-clock timer. Refresh work uses normal CPU priority with idle disk I/O to give control-change publication a normal CPU share during mining; the cached HTTP server remains at idle CPU priority. Consumers should still age generated_at: it distinguishes a current report from a cached response when a rig or its refresh job is unhealthy.

Stratum over TLS (optional)

The stack can serve TLS on the same :3333 (p2pool.stratum_tls, default off): its proxy detects TLS vs plain per connection, so nothing re-points and a mixed fleet migrates one rig at a time — rigs still on cleartext keep mining throughout. On the first pithead apply with TLS on, the stack generates a self-signed certificate and prints its SHA-256 fingerprint (64 lowercase hex chars); pithead status repeats it. That fingerprint is what the rig pins:

// config.json — TLS on, the stack's cert pinned by its SHA-256 fingerprint
{
    "pools": [
        { "url": "your-stack:3333", "tls": true, "tls-fingerprint": "<the fingerprint pithead prints>" }
    ]
}

Then sudo ./rigforge.sh apply. The first-run prompt doesn't ask about TLS — these two fields are edited into config.json by hand. The same fields work against any TLS stratum endpoint, e.g. a public pool's TLS port.

The trust model, plainly: XMRig does no CA validation on stratum TLS. With "tls": true and no fingerprint, the link is encrypted but not authenticated — fine against passive snooping, no defense against an active man-in-the-middle. The fingerprint pin IS the server authentication: a pinned rig refuses anything that doesn't hold the stack's exact certificate. pithead status is the canonical source for the pin; to read it off the wire instead (XMRig compares case-insensitively, so the uppercase openssl output works as-is):

echo | openssl s_client -connect your-stack:3333 2>/dev/null \
    | openssl x509 -noout -fingerprint -sha256 | cut -d= -f2 | tr -d ':'
  • TLS is confidentiality; the stratum password (above) is access control. They're orthogonal — set both on an untrusted network.
  • Rotation: the operator deletes the two files in the stack's proxy-tls data directory and re-runs pithead apply (new certificate, new fingerprint); then update tls-fingerprint on each TLS rig and run apply (same runbook shape as the password). A stale pin shows up as Failed to verify server certificate fingerprint in the XMRig log; cleartext rigs are unaffected throughout.

2. Stats connection — the Worker API (:8080)

Each worker exposes XMRig's HTTP API so Pithead's dashboard can show per-rig stats (hashrate, shares, uptime). RigForge configures the API to match Pithead's contract exactly, so there's nothing to set up stack-side:

Setting Value Why
Port 8080 Pithead reads GET http://<rig>:8080/1/summary; the port is fixed dashboard-side.
Bind 0.0.0.0 (all interfaces) The dashboard polls each worker from the stack host over the LAN.
Mode restricted: true (read-only) The API can be read but not used to control the miner remotely.
Auth token none (open) by default; set ACCESS_TOKEN to require a Bearer token Pithead's stock probe is no-auth, so an open, read-only API works without extra config. Setting ACCESS_TOKEN turns auth on; see below. This port takes only the master ACCESS_TOKEN — XMRig allows exactly one http.access-token, so named ACCESS_TOKENS entries are not accepted here.

Pithead discovers workers from the stratum proxy's connection list (the pool user label, which is the rig name), so there's nothing to register stack-side. Workers on a trusted LAN need no Tor; an .onion pool requires an operator-managed Tor SOCKS proxy.


The token rule (important)

NOTE: By default the worker API is open (read-only, no token), which matches Pithead's default probe (workers.api_auth: none). Nothing to coordinate. Leave ACCESS_TOKEN unset and it works.

If you do want a token (e.g. you don't fully trust the LAN), set ACCESS_TOKEN here. Existing raw token clients remain compatible:

  • a single shared token → Pithead workers.api_auth: token + workers.api_token: <the token>;
  • the rig name as the token (ACCESS_TOKEN = the first pool's user) → Pithead workers.api_auth: name.

Pithead 2.0 derives a separate read bearer for an adopted RigForge 1.17.2+ rig and keeps the raw token on the host for control. An adopted token-protected 1.17.0/1.17.1 rig must be upgraded before its enriched feed can be read without giving the dashboard its control capability; use the remote upgrade when enabled or upgrade locally otherwise.

One token per stack

A rig can serve more than one Pithead stack — production, which it mines to, and a bench that borrows it for hardware tests. Give each stack its own token with ACCESS_TOKENS instead of sharing ACCESS_TOKEN, so rotating one stack's credential never touches the other's and no stack has to be handed the rig's master token:

{
  "ACCESS_TOKEN": "<master, 32+ random hex>",
  "ACCESS_TOKENS": { "prod": "<prod's own>", "bench": "<bench's own>" }
}

Which token a stack must hold depends only on what it reads:

The stack reads Token it must hold
XMRig's own :8080/1/summary (the stock Pithead probe) ACCESS_TOKEN, the master. XMRig takes exactly one http.access-token, so a named entry is not accepted there.
The enriched sister feed on :8081 ACCESS_TOKEN or any named ACCESS_TOKENS entry — raw, or that same token's derived read bearer.
The writable control path on :8082 ACCESS_TOKEN or any named ACCESS_TOKENS entry, raw (derived read bearers are rejected here, as always).

Stack-side this needs no new Pithead setting: workers.list[].token is already per worker per stack, so each stack simply carries the value you gave it. A stack whose probe still hits :8080 directly stays on the master; move it to the :8081 feed first if you want it off the master token.

Two consequences worth knowing before you rely on it:

  • ACCESS_TOKEN is still privileged. Beyond :8080, it is the credential RigForge's own liveness/rollback probe uses when the control path restores pools after a failed apply. Enabling control therefore still requires ACCESS_TOKEN (and api_allow_from); a rig cannot run the control path on named tokens alone.
  • A named entry is a full control credential, not a read-only one. Any entry that authenticates :8081 also authenticates :8082, which is the point for a bench that must apply config during a test run. Hand out a derived read bearer instead when a consumer should only ever read.

Tokens are read once when the services start, so adding or revoking an entry takes effect on the next sudo rigforge.sh apply. sudo rigforge.sh doctor reports how many tokens are configured and which one XMRig carries.

Likewise, don't bind the API to localhost only and don't change the port without matching it on the stack side (workers.api_port): a non-8080 port, or a worker reachable at a different host than the one it connects from, also need matching configuration on both sides (Pithead #171 / #172).


3. Writable control path (:8082, producer for Worker Inspect)

Off by default. When you set "control": "enabled" (plus the required ACCESS_TOKEN and api_allow_from), the rig serves a separate authenticated write endpoint that lets the stack apply config changes through RigForge — the RigForge-side producer for pithead's Worker Inspect (pithead #185). It is deliberately independent of the read API: a POST :8082/apply of an allowlisted change (pools, DONATION, autotune, watchdog, watchdog_interval_min, max_temp_c) returns 202 Accepted; RigForge validates, snapshots the old config, applies it, and rolls back anything that doesn't come back live. Changing only watchdog_interval_min and/or max_temp_c never restarts XMRig (#381) — those two are proven not to reach XMRig's config or unit, so RigForge reconciles just the watchdog timer instead of running the full apply pipeline, landing in about a second instead of the ~60s a restart costs. Any other key, alone or mixed with those two, takes the full, XMRig-restarting path. The stack reads the new effective config back from :8081/2/summary and polls :8082/status for the outcome. The write path is pinned to the stack host by api_allow_from (mandatory) — the miner never accepts a config from anywhere else. Full mechanics and the security model: Operations › Writable control path and ADR 0001.

Polling a specific change (?change_id). The no-arg GET :8082/status returns the most recent change's outcome, which a concurrent change (another dashboard edit, a local apply, an autotune restart) can step on between your POST and your poll. To avoid that race, poll GET :8082/status?change_id=<the 16-hex id from the 202> — it returns that change's recorded outcome (for a config change: applied/rejected/rolled_back/failed + reason/backup/changed_keys/warnings), or 404 if it isn't among the last ~20 recorded. Same bearer auth as /status. The same endpoint serves remote-upgrade change ids with a richer vocabulary (non-terminal started, terminal noop/throttled alongside the above — #320); see Operations › Remote upgrade.

Between your POST and a terminal outcome, ?change_id reads status: "pending" with an accepted_at stamp — recorded the instant /apply accepted the change, not once the (possibly tens-of-seconds) apply pipeline gets around to it — so a poll during that window resolves to pending instead of the same 404 an id that was never issued gets (#344). A change superseded by a newer one before ever being picked up stays pending forever rather than being guessed into a fake outcome; a claimed run that dies mid-apply instead records a terminal failed naming the loss (#509) — an EXIT floor armed the moment the oneshot claims the change, since nothing else will re-drive it. A remote upgrade that dies mid-fetch or mid-build gets the same floor instead of stopping at started (#535). age_seconds (see next paragraph) growing without bound on a still-pending id is the tell that it isn't coming back — don't treat pending as automatically transient. Every /status response, ?change_id or no-arg alike, carries a derived age_seconds next to whichever timestamp it has (accepted_at while pending, applied_at once terminal), computed fresh on each request — this is what closed a real rig's first no-arg /status after enabling control surfacing an 11-day-old record with no way to tell it wasn't current (#344). Pair it with config_meta.revision on the read feed to confirm the effective config actually moved.

Prefill from a live read (rigforge.config). The enriched feed exposes the rig's current writable config as rigforge.config on :8081/1/summary (= /2/summary) — exactly the keys /apply accepts (pools, DONATION, autotune, watchdog, watchdog_interval_min, max_temp_c), read the same way RigForge parses them (canonical strings, e.g. perf → performance). Worker Inspect can prefill its editor from this live read instead of its own last-applied record, and it's served even when the miner is down (it comes from config.json, not XMRig). Pool secrets are masked: pools[].pass and any tls-fingerprint are served as {"__secret__": true} when set and omitted when not — the value itself never leaves the rig. Send the marker back unchanged (or leave the key out) and the rig keeps the secret it already holds, matching the pool by url + user; send a real string to replace it. A marker for a pool the rig has no matching secret for is refused with unresolvable-secret-marker, so a re-sent pools array can never silently blank a credential — which is what it used to do, quietly re-keying the rig to XMRig's throwaway x and reporting success (#415).

The control path is a tuning channel, not a safety-removal one. A POST /apply that would disable the watchdog or unset / set an out-of-band max_temp_c (a rig's thermal cutoff) is refused with 400 — change thermal protection locally on the rig with rigforge.sh if that's really intended. Any change touching watchdog/max_temp_c is surfaced in :8082/status as a warnings[] entry, so the dashboard can require an extra confirmation. The control token is write-capable and travels in cleartext HTTP; api_allow_from scopes the source but doesn't protect the token in flight, so isolate the mining LAN — see Security › what RigForge exposes.


Troubleshooting

Symptom Fix
Rig won't mine / XMRig logs Permission denied at login The stack has stratum authentication on (p2pool.stratum_password); set the pool pass to that secret. See Stratum authentication.
XMRig logs Failed to verify server certificate fingerprint The tls-fingerprint pin doesn't match the server's certificate (rotated cert or a typo). Re-pin from pithead status (or the openssl one-liner in Stratum over TLS) and apply.
Worker missing from the dashboard The dashboard discovers rigs from their stratum user label; confirm the worker is actually connected to the pool and mining.
Rig shows as connected but no stats By default the API is open and the dashboard reads it with no token. If you set an ACCESS_TOKEN here, the dashboard must match it (workers.api_auth: token + workers.api_token, or name if the token is the rig name); otherwise clear ACCESS_TOKEN and re-run setup.
Stats unreachable from the stack host Confirm the worker's :8080 is reachable from the stack host over the LAN (firewall, correct IP). RigForge binds 0.0.0.0 by default.

See also