Crash report
What happened?
sys._current_frames() materializes a PyFrameObject for the running frame of every thread of every interpreter, not just the calling one. The materialization happens on the calling thread, so the frame object is allocated from the calling interpreter's obmalloc state — but the object is attached to the other interpreter's _PyInterpreterFrame (frame->frame_obj = f), and that interpreter is left holding it.
Since PEP 684 gave each interpreter its own arenas (PyInterpreterState grew struct _obmalloc_state obmalloc in 3.12), the block is eventually returned to a different interpreter's arena than the one it was taken from. The heap is corrupted and the process dies in glibc, typically with free(): invalid size, malloc(): unaligned tcache chunk detected, or double free or corruption.
The relevant loop is explicit about the scope — Python/pystate.c, _PyThread_CurrentFrames():
/* for i in all interpreters:
* for t in all of i's thread states:
* if t's frame isn't NULL, map t's id to its frame
*/
_PyEval_StopTheWorldAll(runtime);
HEAD_LOCK(runtime);
PyInterpreterState *i;
for (i = runtime->interpreters.head; i != NULL; i = i->next) {
_Py_FOR_EACH_TSTATE_UNLOCKED(i, t) {
...
PyObject *frameobj = (PyObject *)_PyFrame_GetFrameObject(frame);
_PyFrame_GetFrameObject() → _PyFrame_MakeAndSetFrameObject() → _PyFrame_New_NoTrack() → PyObject_GC_NewVar(), which allocates from the current thread state's interpreter — the caller's, not i's.
No profiler, no C extension, and no third-party code is involved; this is stdlib only.
Reproducer
import sys, threading
from concurrent import interpreters
stop = threading.Event()
def sampler():
while not stop.is_set():
sys._current_frames() # the bug is the call, not walking the result
threading.Thread(target=sampler, daemon=True).start()
interp = interpreters.create()
try:
interp.exec("def f():\n return 1\nfor _ in range(500000):\n f()\n")
finally:
interp.close()
stop.set()
print("survived")
Aborts within seconds. Note that the returned dict is never read — creating the frame objects is sufficient.
For 3.12 and 3.13, replace the concurrent.interpreters calls with the private API (_xxsubinterpreters / _interpreters, create() + run_string()); the behaviour is the same.
Versions
| Version |
Build |
Result |
| 3.12.7 |
GIL |
aborts, 10/10 — free(): invalid size, SIGABRT within ~2s |
| 3.13.7 |
GIL |
aborts, 10/10 — free(): invalid size, SIGABRT within ~2s |
| 3.14.4 |
GIL |
aborts, 10/10 — free(): invalid size / double free or corruption (out), SIGABRT within ~3s |
| 3.14.7 |
free-threaded |
clean |
| 3.11.9 |
GIL |
not testable with this reproducer (see below) |
Linux x86_64, glibc. On 3.12 the abort still occurs with the loop reduced to 5000 iterations, so it does not depend on a long run.
3.11 could not be measured. _xxsubinterpreters.create() never returns while another thread is calling sys._current_frames() in a loop — it hangs at creation, before any workload runs, at every iteration count I tried. With the sampler thread removed the same script finishes instantly. That looks like a separate lock interaction in 3.11 rather than this bug, and since 3.11 is in security-only maintenance I have not pursued it. I mention it only to be clear that the 3.11 row below is reasoning, not measurement.
The measured boundary is consistent with the mechanism:
- Free-threaded builds are clean (measured) —
PYOBJ_ALLOC is mimalloc rather than obmalloc, so there are no per-interpreter arenas and no wrong arena to free into.
- 3.11 should be unaffected (reasoning, not measured) — it has no per-interpreter obmalloc state, so every interpreter allocates from the same arenas. I could not confirm this directly for the reason above.
- The affected set is what remains: GIL builds from 3.12 on, which is exactly where per-interpreter arenas exist.
What I am less sure about
The crash, the version boundary, and the allocation domain are all observed or read directly from the source. The part I am inferring is which interpreter performs the final free — my reading is that when the owning interpreter pops the frame, take_ownership() leaves it holding the last reference after the caller has dropped the returned dict, so the deallocation happens under the owner. A core dev may well recognise a shorter path to the same outcome.
Related, but not duplicates
Operating systems tested on
Linux (x86_64, glibc).
Linked PRs
Crash report
What happened?
sys._current_frames()materializes aPyFrameObjectfor the running frame of every thread of every interpreter, not just the calling one. The materialization happens on the calling thread, so the frame object is allocated from the calling interpreter's obmalloc state — but the object is attached to the other interpreter's_PyInterpreterFrame(frame->frame_obj = f), and that interpreter is left holding it.Since PEP 684 gave each interpreter its own arenas (
PyInterpreterStategrewstruct _obmalloc_state obmallocin 3.12), the block is eventually returned to a different interpreter's arena than the one it was taken from. The heap is corrupted and the process dies in glibc, typically withfree(): invalid size,malloc(): unaligned tcache chunk detected, ordouble free or corruption.The relevant loop is explicit about the scope —
Python/pystate.c,_PyThread_CurrentFrames():_PyFrame_GetFrameObject()→_PyFrame_MakeAndSetFrameObject()→_PyFrame_New_NoTrack()→PyObject_GC_NewVar(), which allocates from the current thread state's interpreter — the caller's, noti's.No profiler, no C extension, and no third-party code is involved; this is stdlib only.
Reproducer
Aborts within seconds. Note that the returned dict is never read — creating the frame objects is sufficient.
For 3.12 and 3.13, replace the
concurrent.interpreterscalls with the private API (_xxsubinterpreters/_interpreters,create()+run_string()); the behaviour is the same.Versions
free(): invalid size, SIGABRT within ~2sfree(): invalid size, SIGABRT within ~2sfree(): invalid size/double free or corruption (out), SIGABRT within ~3sLinux x86_64, glibc. On 3.12 the abort still occurs with the loop reduced to 5000 iterations, so it does not depend on a long run.
3.11 could not be measured.
_xxsubinterpreters.create()never returns while another thread is callingsys._current_frames()in a loop — it hangs at creation, before any workload runs, at every iteration count I tried. With the sampler thread removed the same script finishes instantly. That looks like a separate lock interaction in 3.11 rather than this bug, and since 3.11 is in security-only maintenance I have not pursued it. I mention it only to be clear that the 3.11 row below is reasoning, not measurement.The measured boundary is consistent with the mechanism:
PYOBJ_ALLOCis mimalloc rather than obmalloc, so there are no per-interpreter arenas and no wrong arena to free into.What I am less sure about
The crash, the version boundary, and the allocation domain are all observed or read directly from the source. The part I am inferring is which interpreter performs the final free — my reading is that when the owning interpreter pops the frame,
take_ownership()leaves it holding the last reference after the caller has dropped the returned dict, so the deallocation happens under the owner. A core dev may well recognise a shorter path to the same outcome.Related, but not duplicates
sys._current_frames()use-after-free, but free-threaded builds only, and about concurrently reading a returned frame. This one affects default GIL builds and is about the allocation domain of the frame object itself.PyThreadState_GetFrame()thread-safety is unclear (and inconsistent withsys._current_frames()) #148589 —PyThreadState_GetFrame()thread-safety, adjacent to the same materialization path.Operating systems tested on
Linux (x86_64, glibc).
Linked PRs