Skip to content

feat(finalizer): fail a link that drops an expected kernel - #1367

Open
lucifer1004 wants to merge 1 commit into
NVIDIA:mainfrom
lucifer1004:feat/final-link-kernel-check
Open

lucifer1004 wants to merge 1 commit into
NVIDIA:mainfrom
lucifer1004:feat/final-link-kernel-check

Conversation

@lucifer1004

Copy link
Copy Markdown
Contributor

First PR of the three-PR stack discussed in #1071 (final-link safety → dormant device-runtime link contract → first public Graph API).

Problem

nvJitLink can report success while dropping a module's kernels. An unresolved device-runtime symbol makes it discard the module, and the link-time optimizer strips a kernel it finds unreachable. The result is a well-formed but empty image that is_valid_cubin accepts, so the failure only surfaces at launch time, if at all.

Change

Every finalizer link now takes the kernels its output must define. After the link it lists the output's entry points and fails if any expected kernel is missing:

  • cubin: FUNC symbols with STO_CUDA_ENTRY (st_other & 0x10) from the ELF symbol table;
  • PTX: .entry callables through ptx-parse, as cuda-host's entry registry already does.

New errors: FinalizerError::MissingKernels { output, missing, present } and UnreadableEntryInventory.

Kernels stay rooted through the @llvm.used the NVVM exporter already emits. nvJitLink's -kernels-used is deliberately not passed, because it also lets the link drop kernels that are not listed. The link inputs and options are unchanged, so artifact digests and the recipe stay the same.

Where the expected kernels come from

Route Source
Build-time materialization (rustc-codegen-cuda) the collector's kernel export names
Embedded bundles finalized at load (cuda-host) the bundle's Kernel entries
Sidecar files (cargo oxide interop, cuda-host file loader) a new <module>.kernels sidecar

The compiler writes <module>.kernels next to the .ll, before the completing .target. Its first line is cuda-oxide expected-kernels v1, followed by one export name per line.

  • An artifact the compiler emitted (its .target carries the compile-options marker) must have the sidecar. A missing or malformed one fails closed with a rebuild hint.
  • A manual .ll is checked only against a sidecar that is present.

The format lives in cuda-artifact-finalizer rather than oxide-artifacts, so this PR needs no oxide-artifacts release.

Cached cubins are re-checked on a hit, because a stored image may predate the check.

API changes

  • Finalizer::{materialize_nvvm_ir, materialize_nvvm_ir_with_report, link_ltoir, link_ltoir_with_report}, LtoLinker::{link_ltoir, link_ltoir_with_report, link_ptx_to_cubin}, and the public cuda_host::ltoir::{build_*, link_*} functions take expected_kernels: &[&str]. An empty list means the module has no launchable entries.
  • cargo oxide clean also removes .kernels; .gitignore and the book mention the new sidecar.
  • The finalizer now depends on ptx-parse, which adds that edge to every example lockfile. scripts/sync-example-locks.sh --check passes.

Tests

  • Unit tests:
    • synthetic empty and partial cubins;
    • PTX with only device functions;
    • unreadable images;
    • the sidecar format.
  • A live test: an unrooted kernel that nvJitLink drops now fails for both cubin and PTX output.
  • The live file-cache test in cuda-host now roots its kernel. Before, it linked an empty module.
  • cargo test --workspace, the rustc-codegen-cuda tests, and clippy are clean.
  • End to end on SM120: cargo oxide run vecadd --materialize-cubin --arch sm_120 and cargo oxide run libdevice_math pass.
  • An external interop build (364 kernels) also passes the check.

🤖 Generated with Claude Code

@copy-pr-bot

copy-pr-bot Bot commented Sep 30, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

nvJitLink can report success while dropping a module's kernels: an
unresolved device-runtime symbol makes it discard the module, and the
link-time optimizer strips a kernel it finds unreachable. The result is a
well-formed but empty image, which is_valid_cubin accepts.

Every finalizer link now takes the kernels its output must define and
checks the output's entry points after the link: cubin entries from the
ELF symbol table (FUNC symbols with STO_CUDA_ENTRY), PTX entries from
ptx-parse. A missing kernel is FinalizerError::MissingKernels, and an
image whose entries cannot be read is UnreadableEntryInventory. Rooting
stays with the @llvm.used the exporter already emits; nvJitLink's
-kernels-used is not passed, since it also lets the link drop kernels it
does not list.

The expected kernels reach every finalization route:
- build-time materialization: the collector's kernel export names;
- embedded bundles finalized at load: the bundle's Kernel entries;
- sidecar files (cargo oxide interop, cuda-host's file loader): a new
  <module>.kernels sidecar the compiler writes next to the .ll, before the
  completing .target. An artifact the compiler emitted must have it; a
  manual .ll is checked against one only when present.

Cached cubins are checked on a hit too, since a stored image may predate
the check. The cuda-host build/link functions and the finalizer entry
points take the expected kernels as a new argument.

Tests: synthetic empty and partial cubins, PTX with only device
functions, unreadable inventories, the sidecar format, and a live link
whose unrooted kernel nvJitLink drops (cubin and PTX); the live file-cache
test now roots its kernel, which it previously linked away.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant