Skip to content

gh-142349: Simplify lazy import resolution - #158282

Open
pablogsal wants to merge 8 commits into
python:mainfrom
pablogsal:simplify-lazy-imports-c
Open

pablogsal wants to merge 8 commits into
python:mainfrom
pablogsal:simplify-lazy-imports-c

Conversation

@pablogsal

@pablogsal pablogsal commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

The current implementation has some problems. The eval loop and module attribute lookup repeat the code that resolves a lazy import and replaces its placeholder. These copies behave differently. We also hold the global import lock during resolution, which can deadlock with another thread importing a package.

This moves the placeholder structure and resolution code into lazyimportobject.c. One helper now resolves the import and updates the namespace where we found the name. It only replaces the value if it is still the same placeholder or the child module that the normal importer just put there. Assignments and deletions made during the import are preserved. For ordinary dicts, the check and replacement are atomic.

We also remove the global import lock around resolution. Importlib still locks modules while loading them. Each thread keeps a set of placeholders being resolved to detect cycles. Custom import hooks can now run concurrently.

This removes duplicate lookup and cleanup code. The eval loop and module attribute lookup use the same helper, so they no longer need separate implementations of these rules.

With this design, a placeholder records the module to import and where the import was declared. A name taken from that import holds a reference to the original placeholder and records the attribute to look up. When the name is first used, the resolver follows those references, imports the module and performs the attribute lookups in order. The shared helper then replaces the placeholder in the namespace where it was found, provided that binding still holds the placeholder or that child module.

Keep the placeholder structure private to lazyimportobject.c. Move the
resolution and attribute lookup code there so it can use the structure
without exposing its fields to the eval loop or import code.
Let importlib handle module locks instead of holding the global import
lock while resolving a placeholder. Keep active placeholders in a set on
the thread and remove them on every exit, including allocation failures.

Use the same cycle and recursion checks when a hook returns a placeholder.
Keep the source references alive until all attribute lookups finish.
Keep an existing cause or context when resolution fails. Add the import
location as a note in that case, without adding the same note twice.
Errors without an existing chain still get the declaration as their cause.
Only reuse a concrete attribute from a module that has finished loading.
Leave lazy attributes for resolution so module hooks still get a chance.
Check the module again after reading its spec, and keep interrupts visible.
Normalize the fromlist before calling the filter and check its items
before registering imports. Use one cleanup path and the existing filter
accessor. Pass the same five arguments to custom lazy hooks as to __import__.
Resolve lazy __getattr__ and __dir__ hooks before calling them. Treat a
hook already being resolved as unavailable so it can import a sibling.

Bind a loaded child before removing its pending entry. Recheck the module
dict when another thread may have completed the load. Reuse modules already
in sys.modules during package cycles, and simplify child registration.
Remember the namespace where lookup found the placeholder and use one
helper to resolve and replace it. Only replace a binding that still holds
the same placeholder, preserving assignments and deletions during import.

Check and replace ordinary dict entries atomically. Keep the mapping
protocol for custom namespaces and allow reads from readonly namespaces.
Reuse the global lookup helper in the eval loop and remove the duplicate.
@pablogsal

pablogsal commented Sep 27, 2026 •

Copy link
Copy Markdown
Member Author

@brittanyrey this cleans up a bit of the lazy import code by sharing the resolve and replace logic and keeping the placeholder details in one file. It also removes the global import lock around resolution. Could you take a look?

Loading a child can replace the parent binding with the module before we fetch the imported attribute. Keep the module returned by the normal importer so the shared helper can replace it in its actual parent namespace. Other values and deletions still prevent replacement, and custom hooks keep control of their assignments.
@encukou

encukou commented Sep 30, 2026

Copy link
Copy Markdown
Member

Marking a release blocker for @hugovk's consideration, per #157714 (comment) . I'll start the review.

But, I must say that I feel the process is broken somewhere. From my side it, subjectively, looks like the issues that led here were ignored (1, 2) and argued as unnecessary by the PEP authors for months, so +884/-946 PR this late in the RC period is quite a surprise. I'm not sure what to do so we don't get this situation again. But, that's a discussion for after the release.

@hugovk hugovk added release-blocker 🔨 test-with-buildbots Test PR w/ buildbots; report in status section labels Sep 30, 2026
@bedevere-bot

Copy link
Copy Markdown

🤖 New build scheduled with the buildbot fleet by @hugovk for commit ad34fae 🤖

Results will be shown at:

https://buildbot.python.org/all/#/grid?branch=refs%2Fpull%2F158282%2Fmerge

If you want to schedule another build, you need to add the 🔨 test-with-buildbots label again.

@bedevere-bot bedevere-bot removed the 🔨 test-with-buildbots Test PR w/ buildbots; report in status section label Sep 30, 2026
@pablogsal

pablogsal commented Sep 30, 2026 •

Copy link
Copy Markdown
Member Author

Thanks for taking on the review, Petr. I’m conscious of the extra work this creates for you and Hugo, especially at this stage of the release, and I appreciate the time you’re both putting into it.

the issues that led here were ignored

I think that characterization is a bit unfair to the people doing this work. Many of us have day jobs (some of us have started a new one) and cannot make arbitrary amounts of time available for a particular feature. I have personally authored triaged and reviewed a lot of pull requests and fixes and worked on the main implementation so you cannot just claim we have neglected the change. We also have other responsibilities within CPython, including releases, bug fixes, mentoring contributors, and reviewing other people’s changes. These all compete for the same limited time.

It is reasonable to ask why something took months to address. It is not reasonable to treat that delay alone as evidence that the concerns were ignored. Please account for those constraints when describing how we got here.

argued as unnecessary by the PEP authors for months

I didn’t participate in the linked discussion. Dino did, but he acknowledged the use cases people raised and agreed to keep sys.lazy_modules. I don’t think that thread supports the broader claim being made here.

this late in the RC period

I understand the concern about the timing. My reason for proposing this refactor for 3.15 is to avoid maintaining two substantially different implementations and making subsequent bug fixes harder to backport. I recognise that this adds to Hugo’s release workload as well as the review burden but in the other hand the process is working as intended as we are trying to do the better fix before the release. Certainly it could have been better early but this is the best we can do now.

I agree that we should discuss the process after the release. Clearer priorities, ownership, and earlier escalation are all worth discussing. But any proposed improvement needs to work with the time and people we actually have available. Assuming the authors could simply have found more time does not give us a workable process.

@pablogsal

Copy link
Copy Markdown
Member Author

so +884/-946 PR this late in the RC period is quite a surprise.

One point on the size: that figure makes this look larger than it is. About 400 changed lines are generated files, and a substantial part is moving existing resolution code into lazyimportobject.c. We also replace duplicated logic in the eval loop and module attribute lookup with a shared helper.

The locking and rebinding changes still need careful review, but the raw line count overstates how much new logic is being introduced.

@encukou

encukou commented Sep 30, 2026

Copy link
Copy Markdown
Member

I apologize for my angry tone. That was inappropriate.

No need to thank me; it's my job, especially since it's clear that this is on the critical path for the release.

Yes, let's discuss later. For now: I didn't want to ask you for work, I'd prefer if you asked me for work. I misread the signals as “there's a small rough edge that's not worth discussing”, not “this needs help” :(

(And presently, it doesn't help that a cold took a chunk of time I could give to this :( )


One point on the size: that figure makes this look larger than it is.

Yes. Also, it's definitely a simplification. Thank you!

Unfortunately I haven't looked at lazy imports before, so most of the logic is mostly new to me.
All I found so far are nitpicks, but also I have a list of things to check further.

I have two questions so far:

  • It looks like the distinction between PyExc_ImportCycleError and PyExc_ImportError is not useful any more. Should PyExc_ImportCycleError stay?
  • Are the terms reify and resolve synonyms?
    From lazyimportobject.c, it would seem that resolve only does the import, and reify also binds the name. But then, moduleobject.c's module_get_resolved_dict_item also binds. Am I looking for a distinction that isn't there?

Will continue tomorrow, or sooner if I find the time. I can merge #157714 with this -- @brittanyrey, let me know if you've started on that or want to do it yourself.

If there's a better way to help, let me know.

KRRT7 added a commit to warpforgedotco/kpip that referenced this pull request Sep 30, 2026
Every `lazy import` and `lazy from ... import` is a plain import once more,
still at the top of its module. On CPython 3.15.0rc2, resolving a lazy import
holds the global import lock for the whole import, and deadlocks with an
ordinary import of a module on its path in another thread: kpip hung that
way when a revalidation thread resolved urllib3 while the main thread
imported json. python/cpython#158282 removes the lock, and is not merged.

The compiled kpip already imported eagerly, since Nuitka does not yet
implement PEP 810, and every module was checked to import that way, so this
is how the binary has been running all along. Putting `lazy` back is adding
the keyword to these lines. Reading the interpreter's facts at startup, the
workaround for the deadlock, goes with it.
KRRT7 added a commit to warpforgedotco/kpip that referenced this pull request Sep 30, 2026
Every `lazy import` and `lazy from ... import` from #222 is a plain import
once more, still at the top of its module. On CPython 3.15.0rc2, resolving a
lazy import holds the global import lock for the whole import, and deadlocks
with an ordinary import of a module on its path in another thread: kpip hung
that way when a revalidation thread resolved urllib3 while the main thread
imported json. python/cpython#158282 removes the lock, and is not merged.

The compiled kpip already imported eagerly, since Nuitka does not yet
implement PEP 810, and every module was checked to import that way, so this
is how the binary has been running all along. Putting `lazy` back is adding
the keyword to these lines.
Comment on lines +489 to +492
int matches = PyUnicode_Tailmatch(root->lz_from, name, dot + 1, end, 1);
if (matches <= 0) {
return matches;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To avoid warnings:

Suggested change
int matches = PyUnicode_Tailmatch(root->lz_from, name, dot + 1, end, 1);
if (matches <= 0) {
return matches;
}
Py_ssize_t matches = PyUnicode_Tailmatch(root->lz_from, name, dot + 1, end, 1);
if (matches <= 0) {
return matches ? -1 : 0;
}

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Development

Successfully merging this pull request may close these issues.

4 participants