Skip to content

gh-159137: Keep surviving objects in order in the garbage collector - #159152

Open
pablogsal wants to merge 7 commits into
python:mainfrom
pablogsal:gh-159137-gc-keep-list-order
Open

pablogsal wants to merge 7 commits into
python:mainfrom
pablogsal:gh-159137-gc-keep-list-order

Conversation

@pablogsal

@pablogsal pablogsal commented Oct 10, 2026 •

Copy link
Copy Markdown
Member

move_unreachable() moves objects with zero gc_refs out of the generation and appends them back when they turn out to be reachable. That leaves the survivors in the order the collector reached them, so every later walk of the generation misses the cache at each step. This flags those objects in place and traverses the reachable ones from a small stack, so the generation keeps its allocation order.

Benchmarks

GC time from gc.get_stats(). Medians of four runs per executable (nine for the AST heap), in baseline/candidate/candidate/baseline order.

Workload GC before GC after Less GC time
Sphinx building the CPython docs 28.902 s 21.684 s 25.0%
Pylint on five stdlib packages 2.516 s 1.979 s 21.3%
Parso parsing the standard library 1.674 s 0.657 s 60.8%
sqlglot parsing and transpiling 4,000 queries 4.762 s 3.690 s 22.5%
Beautiful Soup parsing 90 HTML pages 2.443 s 1.764 s 27.8%
Black checking three stdlib packages 1.098 s 1.019 s 7.3%
Full collection of 720,000 live AST objects 292.4 ms 113.0 ms 61.3%
Geometric mean 35.5%

Collector cycles on the Pylint run. Width is the share before the change, blue frames take fewer cycles after it:

pylint

@pablogsal
pablogsal marked this pull request as ready for review October 10, 2026 22:47
Comment thread Python/gc.c
@nascheme

Copy link
Copy Markdown
Member

LGTM. Some comments, not blockers for merging:

  • are we quite confident allocating reachable_state on the stack is okay (it's 8 KB)? When GC is called from the eval breaker, it's probably okay. What about calling gc.collect() explicitly? Perhaps we should just heap allocate this. If that fails, you could fallback to a zero size stack (or maybe a small C stack allocated one). If heap allocated, you could make it a bit larger than 1024 slots, assuming benchmarking justifies that.
  • for the two-ended sweep, a comment about why the backward step can't read an already spliced node would be good. The code is correct but it's suble.
  • gc.get_objects() generally returns them "in that order" but only if the stack is large enough. Should perhaps be careful about making it should like something that can be relied on.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants