Conversation
Keep the placeholder structure private to lazyimportobject.c. Move the resolution and attribute lookup code there so it can use the structure without exposing its fields to the eval loop or import code.
Let importlib handle module locks instead of holding the global import lock while resolving a placeholder. Keep active placeholders in a set on the thread and remove them on every exit, including allocation failures. Use the same cycle and recursion checks when a hook returns a placeholder. Keep the source references alive until all attribute lookups finish.
Keep an existing cause or context when resolution fails. Add the import location as a note in that case, without adding the same note twice. Errors without an existing chain still get the declaration as their cause.
Only reuse a concrete attribute from a module that has finished loading. Leave lazy attributes for resolution so module hooks still get a chance. Check the module again after reading its spec, and keep interrupts visible.
Normalize the fromlist before calling the filter and check its items before registering imports. Use one cleanup path and the existing filter accessor. Pass the same five arguments to custom lazy hooks as to __import__.
Resolve lazy __getattr__ and __dir__ hooks before calling them. Treat a hook already being resolved as unavailable so it can import a sibling. Bind a loaded child before removing its pending entry. Recheck the module dict when another thread may have completed the load. Reuse modules already in sys.modules during package cycles, and simplify child registration.
Remember the namespace where lookup found the placeholder and use one helper to resolve and replace it. Only replace a binding that still holds the same placeholder, preserving assignments and deletions during import. Check and replace ordinary dict entries atomically. Keep the mapping protocol for custom namespaces and allow reads from readonly namespaces. Reuse the global lookup helper in the eval loop and remove the duplicate.
|
@brittanyrey this cleans up a bit of the lazy import code by sharing the resolve and replace logic and keeping the placeholder details in one file. It also removes the global import lock around resolution. Could you take a look? |
Loading a child can replace the parent binding with the module before we fetch the imported attribute. Keep the module returned by the normal importer so the shared helper can replace it in its actual parent namespace. Other values and deletions still prevent replacement, and custom hooks keep control of their assignments.
|
Marking a release blocker for @hugovk's consideration, per #157714 (comment) . I'll start the review. But, I must say that I feel the process is broken somewhere. From my side it, subjectively, looks like the issues that led here were ignored (1, 2) and argued as unnecessary by the PEP authors for months, so |
|
🤖 New build scheduled with the buildbot fleet by @hugovk for commit ad34fae 🤖 Results will be shown at: https://buildbot.python.org/all/#/grid?branch=refs%2Fpull%2F158282%2Fmerge If you want to schedule another build, you need to add the 🔨 test-with-buildbots label again. |
|
Thanks for taking on the review, Petr. I’m conscious of the extra work this creates for you and Hugo, especially at this stage of the release, and I appreciate the time you’re both putting into it.
I think that characterization is a bit unfair to the people doing this work. Many of us have day jobs (some of us have started a new one) and cannot make arbitrary amounts of time available for a particular feature. I have personally authored triaged and reviewed a lot of pull requests and fixes and worked on the main implementation so you cannot just claim we have neglected the change. We also have other responsibilities within CPython, including releases, bug fixes, mentoring contributors, and reviewing other people’s changes. These all compete for the same limited time. It is reasonable to ask why something took months to address. It is not reasonable to treat that delay alone as evidence that the concerns were ignored. Please account for those constraints when describing how we got here.
I didn’t participate in the linked discussion. Dino did, but he acknowledged the use cases people raised and agreed to keep
I understand the concern about the timing. My reason for proposing this refactor for 3.15 is to avoid maintaining two substantially different implementations and making subsequent bug fixes harder to backport. I recognise that this adds to Hugo’s release workload as well as the review burden but in the other hand the process is working as intended as we are trying to do the better fix before the release. Certainly it could have been better early but this is the best we can do now. I agree that we should discuss the process after the release. Clearer priorities, ownership, and earlier escalation are all worth discussing. But any proposed improvement needs to work with the time and people we actually have available. Assuming the authors could simply have found more time does not give us a workable process. |
One point on the size: that figure makes this look larger than it is. About 400 changed lines are generated files, and a substantial part is moving existing resolution code into The locking and rebinding changes still need careful review, but the raw line count overstates how much new logic is being introduced. |
|
I apologize for my angry tone. That was inappropriate. No need to thank me; it's my job, especially since it's clear that this is on the critical path for the release. Yes, let's discuss later. For now: I didn't want to ask you for work, I'd prefer if you asked me for work. I misread the signals as “there's a small rough edge that's not worth discussing”, not “this needs help” :( (And presently, it doesn't help that a cold took a chunk of time I could give to this :( )
Yes. Also, it's definitely a simplification. Thank you! Unfortunately I haven't looked at lazy imports before, so most of the logic is mostly new to me. I have two questions so far:
Will continue tomorrow, or sooner if I find the time. I can merge #157714 with this -- @brittanyrey, let me know if you've started on that or want to do it yourself. If there's a better way to help, let me know. |
Every `lazy import` and `lazy from ... import` is a plain import once more, still at the top of its module. On CPython 3.15.0rc2, resolving a lazy import holds the global import lock for the whole import, and deadlocks with an ordinary import of a module on its path in another thread: kpip hung that way when a revalidation thread resolved urllib3 while the main thread imported json. python/cpython#158282 removes the lock, and is not merged. The compiled kpip already imported eagerly, since Nuitka does not yet implement PEP 810, and every module was checked to import that way, so this is how the binary has been running all along. Putting `lazy` back is adding the keyword to these lines. Reading the interpreter's facts at startup, the workaround for the deadlock, goes with it.
Every `lazy import` and `lazy from ... import` from #222 is a plain import once more, still at the top of its module. On CPython 3.15.0rc2, resolving a lazy import holds the global import lock for the whole import, and deadlocks with an ordinary import of a module on its path in another thread: kpip hung that way when a revalidation thread resolved urllib3 while the main thread imported json. python/cpython#158282 removes the lock, and is not merged. The compiled kpip already imported eagerly, since Nuitka does not yet implement PEP 810, and every module was checked to import that way, so this is how the binary has been running all along. Putting `lazy` back is adding the keyword to these lines.
| int matches = PyUnicode_Tailmatch(root->lz_from, name, dot + 1, end, 1); | ||
| if (matches <= 0) { | ||
| return matches; | ||
| } |
There was a problem hiding this comment.
To avoid warnings:
| int matches = PyUnicode_Tailmatch(root->lz_from, name, dot + 1, end, 1); | |
| if (matches <= 0) { | |
| return matches; | |
| } | |
| Py_ssize_t matches = PyUnicode_Tailmatch(root->lz_from, name, dot + 1, end, 1); | |
| if (matches <= 0) { | |
| return matches ? -1 : 0; | |
| } |
The current implementation has some problems. The eval loop and module attribute lookup repeat the code that resolves a lazy import and replaces its placeholder. These copies behave differently. We also hold the global import lock during resolution, which can deadlock with another thread importing a package.
This moves the placeholder structure and resolution code into
lazyimportobject.c. One helper now resolves the import and updates the namespace where we found the name. It only replaces the value if it is still the same placeholder or the child module that the normal importer just put there. Assignments and deletions made during the import are preserved. For ordinary dicts, the check and replacement are atomic.We also remove the global import lock around resolution. Importlib still locks modules while loading them. Each thread keeps a set of placeholders being resolved to detect cycles. Custom import hooks can now run concurrently.
This removes duplicate lookup and cleanup code. The eval loop and module attribute lookup use the same helper, so they no longer need separate implementations of these rules.
With this design, a placeholder records the module to import and where the import was declared. A name taken from that import holds a reference to the original placeholder and records the attribute to look up. When the name is first used, the resolver follows those references, imports the module and performs the attribute lookups in order. The shared helper then replaces the placeholder in the namespace where it was found, provided that binding still holds the placeholder or that child module.