Skip to content

Select compatible hardware for Gzip decompression - #77

Open
tpn wants to merge 6 commits into
codex/compression-gzip-20260929from
codex/compression-hardware-20260929
Open

tpn wants to merge 6 commits into
codex/compression-gzip-20260929from
codex/compression-hardware-20260929

Conversation

@tpn

@tpn tpn commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Native FITS Gzip decoding can use nvCOMP's hardware decompression engine with CUDA fallback. The default remains auto; decompression_backend="cuda" and --xdr-decompression-backend cuda select CUDA explicitly. Raw DEFLATE and non-pooled Gzip use CUDA. Non-pooled scratch allocations cannot use the engine, so selecting CUDA directly avoids a failed hardware attempt on every read.

The benchmark JSON now reports device decompression capability and maximum chunk size after timing. These fields describe device support; nvCOMP logs establish which engine ran. The docs explain buffer eligibility and how small tiles and larger decode batches provide more parallel work.

Earlier five-system qualification used two FITS files with six image planes (390.9 MB decoded), with byte-identical results. Complete reads, including planning, transfer and pixel restoration, were 1.80–1.82× faster on B200, 1.66–1.71× on B300 and 1.55–1.56× on GB300 with automatic selection. Separate resident nvCOMP probes measured larger gains and are not whole-reader speedups. H100 and H200 used CUDA; all five systems used CUDA for 8 MiB tiles. Vendor logs confirmed those routes.

Stacked on #76, with #74/#76/#77 reconciled against the fixes already on main. Current validation: 986 local CPU tests, 26 native allocation/copy/launch fault cases, and all 447 xDR tests on B200 with no skips. Matched baseline/candidate builds produced identical bytes throughout the original-pair and small/oversized-tile comparisons. Non-pooled automatic reads made zero hardware attempts, versus ten failed attempts in the baseline; pooled hardware decoding remained active. The original-pair medians stayed within 0.9% of baseline; a reversed-order small-case repeat stayed within 0.6%. No new whole-reader speedup is claimed for the non-pooled fix.

@copy-pr-bot

copy-pr-bot Bot commented Sep 29, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@coderabbitai

coderabbitai Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (1)
  • ^codex/013-.*$

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository: NVIDIA/cuPhoton/.coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: c3cf9b5a-7f52-4db6-beab-e17403065208

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: NVIDIA/cuPhoton/.coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 123a22b7-632f-466c-84c4-247ee3ec65d6

📥 Commits

Reviewing files that changed from the base of the PR and between fc5843d and 2a9b407.

📒 Files selected for processing (1)
  • docs/components/xdr.md
🚧 Files skipped from review as they are similar to previous changes (1)
  • docs/components/xdr.md

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 7 remain after this review.


📝 Walkthrough

Walkthrough

xDR adds a configurable decompression backend for FITS reads. Python APIs and CLI options pass the setting to the native decoder. The benchmark accepts the setting and records it in JSON options. Documentation describes defaults, overrides, and backend capabilities.

Changes

xDR Decompression Backend

Layer / File(s) Summary
Runtime options and input precedence
src/cuphoton/core/cli/fits.py, src/cuphoton/core/fits_io.py, src/cuphoton/core/fits_options.py, docs/README.md, docs/cli.md, docs/components/xfit.md, docs/components/xpois.md, docs/components/xrep.md, docs/components/xscan.md, README.md, CHANGELOG.md, docs/components/xdr.md, docs/troubleshooting.md, tests/core/test_cli_contract.py
Adds the decompression backend choice and CLI flag. Documentation describes xDR controls, supported values, defaults, manifest precedence, and capability diagnostics.
Python decode and batch propagation
src/cuphoton/xdr/convenience.py, src/cuphoton/xdr/prefetch.py, src/cuphoton/xdr/reader.py, tests/core/test_fits_io.py, tests/xdr/test_frontend.py
Adds the option to reader and batch APIs, validates or normalizes it, and forwards it through Python decode and batch paths.
Native backend selection and validation
src/cuphoton/xdr/nvcomp_batch.py, src/cuphoton/xdr/src/nvcomp_batch_ext.cpp, tests/xdr/test_nvcomp_batch.py, tests/xdr/test_gzip_gpu.py
The native extension accepts auto or cuda. Gzip uses the selected nvCOMP backend; raw DEFLATE continues to use CUDA. Tests cover invalid values, capability checks, and backend forwarding.
Benchmark integration and reporting
src/cuphoton/xdr/benchmark_fits.py, tests/xdr/test_benchmark_fits.py, docs/components/xdr.md
The benchmark passes the backend setting to both batch-load phases and records the requested value in JSON options. Tests cover CLI forwarding and the default report value.

Priority: ➖ Normal

Merge Risk: ⚪ Minimal · up to 2a9b4

The hardware-support clarification presents no established merge-blocking issue. The documented CUDA override is supported by the rebuilt native extension.

🚥 Pre-merge checks | ✅ 2
✅ Passed checks (2 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Comment @coderabbitai help to get the list of available commands.

@tpn
tpn force-pushed the codex/compression-gzip-20260929 branch from ad75c65 to edc5b80 Compare September 30, 2026 01:55
@tpn
tpn force-pushed the codex/compression-hardware-20260929 branch 2 times, most recently from aeb78c6 to fc5843d Compare September 30, 2026 02:51
@tpn
tpn force-pushed the codex/compression-gzip-20260929 branch from edc5b80 to 0bdabba Compare September 30, 2026 02:51
@tpn
tpn marked this pull request as ready for review September 30, 2026 03:22
@tpn
tpn requested a review from melo-gonzo as a code owner September 30, 2026 03:22
@tpn tpn added the ai-review Request a focused CodeRabbit review label Sep 30, 2026
@tpn

tpn commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@tpn

tpn commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@melo-gonzo melo-gonzo left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Recommend deferring until the hardware qualification matrix is complete. All 191 selected CPU checks passed, but this review did not reproduce the B200 result or qualify hardware decompression across the supported configurations.

tpn added 6 commits September 30, 2026 21:47
Let nvCOMP choose the Gzip decompression backend automatically, using
CUDA when the device or buffers cannot use a hardware engine.

Signed-off-by: Trent Nelson <trentn@nvidia.com>
Signed-off-by: Trent Nelson <trentn@nvidia.com>
Explain automatic defaults, capability fallback, and explicit CLI and
Python selection. Add workflow examples and distinguish requested
benchmark settings from observed decompression hardware.

Signed-off-by: Trent Nelson <trentn@nvidia.com>
Document per-call engine eligibility and nvCOMP diagnostics, including
non-pooled allocation fallback. Keep requested choices distinct from
observed hardware execution, and cover all native backend entry points.

Signed-off-by: Trent Nelson <trentn@nvidia.com>
Signed-off-by: Trent Nelson <trentn@nvidia.com>
Select CUDA directly for non-pooled Gzip scratch allocations. Keep
automatic engine selection for pooled calls, report device capability
outside benchmark timings, and explain tile-size tradeoffs.

Signed-off-by: Trent Nelson <trentn@nvidia.com>
@tpn
tpn force-pushed the codex/compression-hardware-20260929 branch from 2a9b407 to 6a5b551 Compare October 1, 2026 05:05
@tpn
tpn force-pushed the codex/compression-gzip-20260929 branch from 0bdabba to e33dc4d Compare October 1, 2026 05:05

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ai-review Request a focused CodeRabbit review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants