Conversation
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. 🗂️ Base branches to auto review (1)
Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Repository: NVIDIA/cuPhoton/.coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: NVIDIA/cuPhoton/.coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (1)
🚧 Files skipped from review as they are similar to previous changes (1)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 7 remain after this review. 📝 WalkthroughWalkthroughxDR adds a configurable decompression backend for FITS reads. Python APIs and CLI options pass the setting to the native decoder. The benchmark accepts the setting and records it in JSON options. Documentation describes defaults, overrides, and backend capabilities. ChangesxDR Decompression Backend
Priority: ➖ Normal Merge Risk: ⚪ Minimal · up to The hardware-support clarification presents no established merge-blocking issue. The documented CUDA override is supported by the rebuilt native extension. 🚥 Pre-merge checks | ✅ 2✅ Passed checks (2 passed)
Comment |
ad75c65 to
edc5b80
Compare
aeb78c6 to
fc5843d
Compare
edc5b80 to
0bdabba
Compare
|
@coderabbitai review |
✅ Action performedReview finished.
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
melo-gonzo
left a comment
There was a problem hiding this comment.
Recommend deferring until the hardware qualification matrix is complete. All 191 selected CPU checks passed, but this review did not reproduce the B200 result or qualify hardware decompression across the supported configurations.
Let nvCOMP choose the Gzip decompression backend automatically, using CUDA when the device or buffers cannot use a hardware engine. Signed-off-by: Trent Nelson <trentn@nvidia.com>
Signed-off-by: Trent Nelson <trentn@nvidia.com>
Explain automatic defaults, capability fallback, and explicit CLI and Python selection. Add workflow examples and distinguish requested benchmark settings from observed decompression hardware. Signed-off-by: Trent Nelson <trentn@nvidia.com>
Document per-call engine eligibility and nvCOMP diagnostics, including non-pooled allocation fallback. Keep requested choices distinct from observed hardware execution, and cover all native backend entry points. Signed-off-by: Trent Nelson <trentn@nvidia.com>
Signed-off-by: Trent Nelson <trentn@nvidia.com>
Select CUDA directly for non-pooled Gzip scratch allocations. Keep automatic engine selection for pooled calls, report device capability outside benchmark timings, and explain tile-size tradeoffs. Signed-off-by: Trent Nelson <trentn@nvidia.com>
2a9b407 to
6a5b551
Compare
0bdabba to
e33dc4d
Compare
Native FITS Gzip decoding can use nvCOMP's hardware decompression engine with CUDA fallback. The default remains
auto;decompression_backend="cuda"and--xdr-decompression-backend cudaselect CUDA explicitly. Raw DEFLATE and non-pooled Gzip use CUDA. Non-pooled scratch allocations cannot use the engine, so selecting CUDA directly avoids a failed hardware attempt on every read.The benchmark JSON now reports device decompression capability and maximum chunk size after timing. These fields describe device support; nvCOMP logs establish which engine ran. The docs explain buffer eligibility and how small tiles and larger decode batches provide more parallel work.
Earlier five-system qualification used two FITS files with six image planes (390.9 MB decoded), with byte-identical results. Complete reads, including planning, transfer and pixel restoration, were 1.80–1.82× faster on B200, 1.66–1.71× on B300 and 1.55–1.56× on GB300 with automatic selection. Separate resident nvCOMP probes measured larger gains and are not whole-reader speedups. H100 and H200 used CUDA; all five systems used CUDA for 8 MiB tiles. Vendor logs confirmed those routes.
Stacked on #76, with #74/#76/#77 reconciled against the fixes already on main. Current validation: 986 local CPU tests, 26 native allocation/copy/launch fault cases, and all 447 xDR tests on B200 with no skips. Matched baseline/candidate builds produced identical bytes throughout the original-pair and small/oversized-tile comparisons. Non-pooled automatic reads made zero hardware attempts, versus ten failed attempts in the baseline; pooled hardware decoding remained active. The original-pair medians stayed within 0.9% of baseline; a reversed-order small-case repeat stayed within 0.6%. No new whole-reader speedup is claimed for the non-pooled fix.