Skip to content

fix(cuda-macros,codegen): keep all-generic library .oxart via #72 anchor - #1372

Open
dundysm wants to merge 1 commit into
NVIDIA:mainfrom
dundysm:fix/1365-all-generic-cuda-module-anchor
Open

dundysm wants to merge 1 commit into
NVIDIA:mainfrom
dundysm:fix/1365-all-generic-cuda-module-anchor

Conversation

@dundysm

@dundysm dundysm commented Oct 1, 2026

Copy link
Copy Markdown

Summary

Fixes #1365.

An all-generic #[cuda_module] in a library that monomorphizes locally still embeds .oxart with an artifact anchor, but the macro skipped the #72 rlib keep-alive under the assumption that generics never produce an artifact in the defining crate. Nothing referenced the anchor, so the linker dropped the archive member and load_all_ptx_bundles_merged returned NoModules.

Adding one unused concrete kernel "fixed" it only because concrete kernels restored the keep-alive. This PR restores the handshake for generics without reverting #222 merge loading.

Root cause

Shape Library .oxart? Macro anchor ref? Result
Concrete in lib (#72) Yes Yes Works
Generics mono only in binary (#222) Usually no No (was skipped) Works via merge
Generics mono in library, thin binary (#1365) Yes No (bug) NoModules

Approach

Preserves the intentional design layers:

  1. ModuleNotFound after I move a #[cuda_module] into a separate crate #72 archive handshake: keep referencing the package (or v2) artifact anchor from load_named.
  2. Panic DriverError(500, "named symbol not found") when calling generic kernels from a separate crate #222 merge load: generic modules still call load_all_ptx_bundles_merged (no named-only regression).
  3. No false undefined symbols: when the defining crate has merge markers but kernel_count == 0, the backend emits a strong anchor-only .oxlink stub (no fake empty PTX payload).

Macro (cuda_module_artifact_anchor_references)

  • Remove the early return that skipped when every kernel is generic.
  • Emit the same cfg-guarded black_box anchor refs for generic kernels via effective_cfg_attrs.
  • Keep owner-filter / missing CARGO_PKG_* skips.

Backend (rustc-codegen-cuda)

  • Extract cgus_require_ptx_bundle_merge so markers are scanned even without device code.
  • When owner-selected, markers present, and there is no device codegen path (kernel_count == 0 / no device fns), emit a strong primary anchor-only object (v2 + weak legacy when owner filter is active).

Tests / example

Orthogonal to #1367 / #1166 (finalizer expected-kernel inventory).

Test plan

  • cargo test -p oxide-artifacts --features object --lib (24 passed, including new strong-anchor tests)
  • scripts/sync-example-locks.sh --check
  • scripts/check-example-smoketest-contract.sh
  • cargo test -p cuda-macros --lib (blocked on this box: no CUDA 13 toolkit for cuda-bindings)
  • cargo oxide run generic_mono_in_lib and --verify-bundles (no GPU / no toolkit here)
  • cargo oxide run cuda_module_in_lib and cross_crate_embedded (same)

CI on this PR should cover the CUDA-backed lanes.

Reviewers

cc @nihalpasham @xavierforge (prior #72 / from-anchor / embedding owners; CODEOWNERS currently declares no path owners in-tree)

… anchor

All-generic #[cuda_module] skipped the rlib artifact-anchor keep-alive
because the macro assumed "no artifact here". When the library itself
monomorphizes, codegen still embeds .oxart with an anchor; nothing
referenced it, so the linker dropped the archive member and
load_all_ptx_bundles_merged returned NoModules (issue NVIDIA#1365).

Emit the same cfg-guarded black_box anchor refs for generic kernels.
When owner-selected with PTX-merge markers but kernel_count == 0, emit
a strong anchor-only .oxlink stub so the NVIDIA#222 (mono only in the binary)
shape still links. Keep merge loading; do not invent empty PTX.

Signed-off-by: Dundy Pasupuleti <dundysm@gmail.com>
@copy-pr-bot

copy-pr-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

all-generic #[cuda_module] in a library gives "load module: NoModules"

1 participant