Skip to content

build: consolidate toolchain helpers - #2

Draft
juenglin wants to merge 13 commits into
consolidate-toolchain-helpersfrom
consolidate-toolchain-helpers2
Draft

juenglin wants to merge 13 commits into
consolidate-toolchain-helpersfrom
consolidate-toolchain-helpers2

Conversation

@juenglin

@juenglin juenglin commented Oct 2, 2026

Copy link
Copy Markdown
Owner

Add cuda_bindings/_build_shared.py as the single source of truth for the
toolchain helpers (moved verbatim out of the duplicated block in both
build_hooks.py files) and for the flag set unified in the previous commit.
cuda_core/_build_shared.py is a symlink to it. Both packages already use
backend-path = ["."], so the module is importable while the PEP 517 backend
runs, and Python does not dereference symlinks in file, so
Path(file)-relative locations still resolve under each package.

resolve_toolchain() takes the per-package choices as arguments: a required
cxx_std (no shared default; bindings stays on c++14, core on c++17),
warnings_as_errors (core only, from CUDA_PYTHON_WERROR), and an optional
tweak hook. Each build_hooks.py keeps a thin _resolve_toolchain() declaring
just that. cuda-bindings moves -Wno-deprecated-declarations into its tweak.

No behavior change: the resulting compiler/linker flags and CC/CXX/
LDCXXSHARED environment are identical across platform x toolchain x debug x
coverage x werror, apart from -Wno-deprecated-declarations now being last
for bindings. The Cython cache helpers stay in the synced block for now.

Also include _build_shared.py in both sdists (MANIFEST.in), add it to
known-first-party for isort, load it ahead of build_hooks.py in the tests
and the tests/cython/build_tests.py scripts, and note the symlink in the
package AGENTS.md files.

Reuse the wheel Cython cache in tests/cython/build_tests.py so the
new test stage can hit CUDA_PYTHON_CYTHON_CACHE_DIR.

(cherry picked from commit 392178690f26c7d7cb8117a5556a70f2c728c321)
Close the audit from NVIDIA#1882. The two packages' Linux and MSVC flag sets had
drifted for legacy reasons; both now go through the same structure.

Changes in cuda-bindings:
- Drop -fpermissive and -fno-var-tracking-assignments (gcc-only; not
  needed by current Cython-generated C++; confirmed by a rebuild).
- Change -O3 to -O2 (consistent with cuda-core; -O3 was not the source
  of the measured launch_{256,512}_args latency difference vs c++17).
- Add /std:c++14 and /O2 on MSVC (was missing entirely).
- Split '-std=c++14 -Wno-deprecated-declarations' onto separate lines
  with an explanatory comment on each.

Changes in cuda-core:
- Add /O2 on MSVC (modern setuptools no longer forces /Ox).
- Add comments at both the MSVC and Linux flag sites explaining why
  c++17 is required (structured bindings / if constexpr in _cpp/).

Tests:
- Update the bindings TestResolveToolchain: replace test_gnu_keeps_gcc_only_flags
  with test_linux_opt_flag_set (asserts -O2, not -O3; no gnu-only flags);
  add test_msvc_opt_flag_set; update test_gnu_sets_env_and_flags assertions.
- Add test_linux_opt_flag_set and test_msvc_opt_flag_set to the core
  TestResolveToolchain; update the comment in test_gnu_sets_env_and_flags.
Add cuda_bindings/_build_shared.py as the single source of truth for the
toolchain helpers (moved verbatim out of the duplicated block in both
build_hooks.py files) and for the flag set unified in the previous commit.
cuda_core/_build_shared.py is a symlink to it. Both packages already use
backend-path = ["."], so the module is importable while the PEP 517 backend
runs, and Python does not dereference symlinks in __file__, so
Path(__file__)-relative locations still resolve under each package.

resolve_toolchain() takes the per-package choices as arguments: a required
cxx_std (no shared default; bindings stays on c++14, core on c++17),
warnings_as_errors (core only, from CUDA_PYTHON_WERROR), and an optional
tweak hook. Each build_hooks.py keeps a thin _resolve_toolchain() declaring
just that. cuda-bindings moves -Wno-deprecated-declarations into its tweak.

No behavior change: the resulting compiler/linker flags and CC/CXX/
LDCXXSHARED environment are identical across platform x toolchain x debug x
coverage x werror, apart from -Wno-deprecated-declarations now being last
for bindings. The Cython cache helpers stay in the synced block for now.

Also include _build_shared.py in both sdists (MANIFEST.in), add it to
known-first-party for isort, load it ahead of build_hooks.py in the tests
and the tests/cython/build_tests.py scripts, and note the symlink in the
package AGENTS.md files.
…d.py

The Cython cache helpers were duplicated verbatim in both build_hooks.py
files, kept identical by a pre-commit sync check. The stamp mechanics
(_BUILD_DIR, _abi_stamp_path, the force_build_ext flag, check/record of a
stamped key) were near copies.

Move them into _build_shared.py and delete toolshed/check_build_hooks_sync.py
with its hook. Each build_hooks.py keeps only what is per package: its stamp
file and key (toolchain for cuda-bindings; CUDA major, toolchain, debug and
coverage for cuda-core). setup.py is unchanged: build_hooks re-exports
force_build_ext through a module-level __getattr__.

The Cython test builders now load _build_shared.py directly instead of
build_hooks.py. The tests patch force_build_ext and sysconfig on
_build_shared.
…d ones

The toolchain, linker-command and stamp tests were duplicated verbatim in
cuda_bindings/tests/test_build_hooks.py and cuda_core/tests/test_build_hooks.py,
although they exercise code that now lives in the one _build_shared.py.

Move them into mixins in cuda_python_test_helpers/build_shared.py, which each
package mixes in against the _build_shared module it loaded. Each package keeps
only its own choices: bindings' c++14 and -Wno-deprecated-declarations, core's
c++17 and CUDA_PYTHON_WERROR wiring, and the key each one stamps.

Prune what a wheel build already proves: the llvm preflight being a no-op for
the default toolchain, and the default toolchain's name and compiler binaries.

Add tests for behavior that had none: the Werror flags per platform, the
check_build_key/record_build_key protocol, and that build_hooks re-exports
force_build_ext from _build_shared without holding a copy. Patching the flag on
build_hooks instead would leave a plain attribute there on teardown that
shadows the re-export for later tests, so the tests patch _build_shared.
cuda_core/_build_shared.py is a symlink into cuda_bindings, so a Windows
source build fails confusingly when git symlinks were off at clone time.
CONTRIBUTING.md now documents the setting; link to it from the cuda-bindings
and cuda-core source-install notes.
_import_get_cuda_path_or_home() and _get_cuda_path() were byte-identical
copies in the two build_hooks.py files, held together by "keep in sync"
comments. Move them into the shared helper module and import them from
both backends.

The tests that patch build_hooks._get_cuda_path or call its cache_clear()
keep working: the name is bound in each build_hooks module, and the
cached function object is the same.
coverage.yml and precommit-windows still cloned without core.symlinks=true,
so cuda_core/_build_shared.py would land as a text stub. Drop the unused
_fake_sysconfig leftovers from the mixin move.
@juenglin
juenglin force-pushed the consolidate-toolchain-helpers branch from 8e06f2c to a04079f Compare October 2, 2026 17:23
rwgk and others added 4 commits October 2, 2026 17:45
* test: fix Orin coverage with CUDA 13.4

* test: preserve partial NVML coverage on Orin

* test: address Orin review feedback

* fix(cuda.core): handle partial NVML support on Orin

Preserve UUIDs reported without a prefix and propagate unsupported NVLink queries before reading unpopulated field results. Keep zero-count behavior for devices without link zero, and cover UUID normalization, lookup errors, and NVLink error propagation with regressions.

* test(cuda.core): verify unsupported NVML caller behavior

Replace the UUID-based device-wide skip with narrow checks of the actual API result. Exercise supported Orin queries, validate precise unsupported errors, preserve CUDA visibility and MIG coverage, and require successful UUID matching when a CUDA counterpart exists.

* test(cuda.core): restore skips for unsupported NVML APIs

Restore the conservative device capability gate and the original device and event test bodies instead of accepting unsupported calls as successful tests. Keep positive UUID mapping checks, report unsupported PCI validation as a separate skipped subtest, restore the process-name NotFound skip, and preserve affinity cleanup and fan serialization.

* fix(cuda.core): remove NVLink state preflight

Read the NVLink count from its own field API without requiring support for the link-zero state API. Replace the preflight regressions with valid zero-count and typed per-field error assertions, regenerate the public stub, and retain only the verified UUID release note with the requested wording.

* test(cuda.core): share VMM leak test allocation size

Use one module constant for the cached warm-up and uncached measured allocations. Preserve eight measured iterations and the strict one-allocation leak threshold.
…updates (NVIDIA#2994)

* ci: add standalone pre-commit checks and quarterly updates

* ci: retain manual testing after workflow bootstrap

* ci: use Dependabot for quarterly pre-commit hook updates

* ci: run pre-commit with Python 3.14

* ci: align Dependabot schedules and bot PR metadata
build: unify compiler flags across cuda-bindings and cuda-core (NVIDIA#2998)

Close the flag audit in NVIDIA#1882. Both packages' Linux and MSVC flag sets had
drifted for legacy reasons and now follow the same structure.

cuda-bindings: drop the gcc-only -fpermissive and
-fno-var-tracking-assignments, use -O2 instead of -O3, and add /std:c++14
and /O2 on MSVC.
cuda-core: add /O2 on MSVC and document why C++17 is required.

Also let the Cython test-extension builders reuse the wheel Cython cache.
…in-helpers2

Keep the _build_shared refactor where it overlapped NVIDIA#2998's inlined flag
and test copies now on main.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants