Skip to content

perf(runtime): read schema graph and subschema caches without locking - #132

Merged
tanmaykm merged 1 commit into
mainfrom
tan/lockfree-schema-cache
Oct 2, 2026
Merged

tanmaykm merged 1 commit into
mainfrom
tan/lockfree-schema-cache

Conversation

@tanmaykm

@tanmaykm tanmaykm commented Oct 1, 2026

Copy link
Copy Markdown
Member

Closes #128.

  • _schema_at took spec.graph_lock on every lookup to read the plain Dict spec.subschemas, which serialized concurrent decoders.
  • There was also a data race: _schema_graph called haskey(spec.graphs, ...) outside the lock while another thread could be writing to it under the lock.

Both caches now live in a SchemaCache with @atomic fields that hold immutable Dict snapshots:

  • Cache hits are lock-free atomic loads.
  • A miss takes graph_lock, re-reads the current snapshot, and returns any entry a racing writer has already published. Otherwise it copies the snapshot, inserts, and publishes the copy. A published Dict is never mutated.
  • Graphs are still compiled once per direction (checked again under the lock).
  • The Spec keyword constructor that generated modules call is unchanged, so no contract bump is needed.

Rough numbers on Julia 1.12 with 4 threads:

  • An uncontended _schema_at hit went from 36 ns to 24 ns.
  • With 4 tasks contending, the amortized cost per call went from 118–146 ns to 9–11 ns.

Tests

  • Tests for cache identity across directions, and a cold-cache multi-threaded stress test over the runtime and a generated client's decoding.
  • CI change: JULIA_NUM_THREADS: 4 is now set for the test job, so the stress test runs threaded (Pkg.test forwards the thread count). This applies to every test file. AGENTS.md documents it.
  • Full suites pass on Julia 1.12 and 1.10 with 4 threads.

Possible follow-ups (minor review notes)

  • Add a test showing that a cached hit completes while graph_lock is held by another task. That would catch a future change that adds the lock back to the hit path.
  • Use acquire loads and release stores instead of the default seq_cst ones.

`_schema_at` took `spec.graph_lock` on every lookup to read the plain
`Dict` `spec.subschemas`, which serialized concurrent decoders. Separately,
`_schema_graph` read `spec.graphs` outside the lock while another thread
could write it under the lock, which is a data race on a plain `Dict`.

Both caches now live in a `SchemaCache` with `@atomic` fields that hold
immutable `Dict` snapshots. Hits are lock-free atomic loads. A miss takes
`graph_lock`, re-reads the current snapshot, returns an entry a racing
writer already published, and otherwise copies, inserts, and publishes
the copy. Graphs are still compiled once per direction (double-checked
under the lock). The `Spec` keyword constructor that generated modules
call is unchanged.

With 4 threads on Julia 1.12, an uncontended `_schema_at` hit dropped
from 36 to 24 ns, and 4 contending tasks went from 118-146 to 9-11 ns
per call amortized.

Add tests for cache identity across directions and a cold-cache
multi-threaded stress test over the runtime and a generated client's
decoding. CI now sets JULIA_NUM_THREADS=4 so the stress test runs threaded.

Closes #128
@tanmaykm
tanmaykm marked this pull request as ready for review October 1, 2026 06:02
@tanmaykm
tanmaykm requested review from krynju and nkottary October 1, 2026 06:53
@tanmaykm
tanmaykm merged commit b49a3aa into main Oct 2, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Remove graph_lock from the _schema_at hot path

2 participants