Conversation
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. 🗂️ Base branches to auto review (1)
Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Repository: NVIDIA/cuPhoton/.coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: NVIDIA/cuPhoton/.coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (15)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 9 remain after this review. 📝 WalkthroughWalkthroughThe xDR decoding pipeline now supports ChangesxDR Gzip Decoding
Priority: ➖ Normal Merge Risk: ⚪ Minimal · up to The decoder selection and forwarding changes have no established merge-blocking issue. Merge after normal build and test checks. 🚥 Pre-merge checks | ✅ 2✅ Passed checks (2 passed)
Comment |
ad75c65 to
edc5b80
Compare
edc5b80 to
0bdabba
Compare
|
@coderabbitai review |
✅ Action performedReview finished.
|
melo-gonzo
left a comment
There was a problem hiding this comment.
All 459 selected decoder/FITS tests passed on RTX 6000 Ada using this PR’s CI-built native extension, but the exception cleanup below needs fixing before approval. Recommend landing after 0.1.3; the ARM wheel check also needs rerunning after its CFITSIO download timeout.
| else cp.empty(total, dtype=cp.uint8) | ||
| ) | ||
| if aligned_owners and keepalive is None: | ||
| d_out = _retain_decode_inputs(d_out, aligned_owners) |
There was a problem hiding this comment.
issue: With keepalive=None, an allocation or owner-wrapping failure after _align_deflate_inputs queues its kernel can release the input and repack buffers while GPU work is still pending. Owners must survive until that work completes, including on exceptions. Preserve #81’s exception-time synchronization around repacking and subsequent allocation/decode, including partial metadata-table construction, and add a failure-injection regression test. Also retain #81’s exception-safe pooled capsule construction when combining the native changes.
Signed-off-by: Trent Nelson <trentn@nvidia.com>
Align raw DEFLATE payloads on the device before decoding and retain those buffers through asynchronous work. Keep automatic selection compatible with older native extensions. Signed-off-by: Trent Nelson <trentn@nvidia.com>
Signed-off-by: Trent Nelson <trentn@nvidia.com>
Bind Python DEFLATE decoding to the current device and CUDA stream. Retain that stream until its cached codec is destroyed, ordering decode after compressed-input uploads and alignment copies. Keep Gzip dispatch and alignment covered by CPU checks and a delayed producer GPU regression. Signed-off-by: Trent Nelson <trentn@nvidia.com>
a0f728e to
3c7e2c0
Compare
0bdabba to
e33dc4d
Compare
FITS tile loading now validates common fixed Gzip headers in batches and passes intact streams to nvCOMP. An isolated two-file comparison measured full-reader time of 73.71 → 63.11 ms (14.4% lower), with identical decoded bytes. Most of that gain comes from host-side header validation; resident Gzip and aligned DEFLATE decoding took similar time.
gzip_decoder="auto"selects native Gzip when available. Python APIs, FITS descriptors and--xdr-gzip-decoderalso acceptgzipto require it ordeflateto select raw DEFLATE. Raw payloads are aligned on the GPU; older extensions remain usable in automatic mode. Decode failures propagate.Stacked on #74. Validation includes byte-for-byte Astropy parity, optional-header and alignment checks, CPU CLI/report coverage, native builds and type checks. The timing above measures the combined adapter and header-validation change; it is not a codec-only speedup.