Fix xDR Deflate alignment and nvCOMP output buffers - #81
Conversation
|
Important Review skippedAuto reviews are limited based on label configuration. 🏷️ Required labels (at least one) (1)
Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Repository: NVIDIA/cuPhoton/.coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Comment |
melo-gonzo
left a comment
There was a problem hiding this comment.
LGTM. All 343 selected decoder and FITS tests passed on RTX 6000 Ada using this PR’s CI-built native extension, including alignment, pooled buffers, and non-default streams. No blocking findings in the buffer ownership and cleanup review.
Align raw payloads before decoding and retain their device buffers until stream work completes. Supply status and actual-size scratch for nvCOMP compatibility, including pooled and failure paths. Signed-off-by: Trent Nelson <trentn@nvidia.com>
b781aa2 to
9629507
Compare
Gzip header removal can leave raw DEFLATE payloads at addresses that do not meet nvCOMP’s four-byte input alignment. The existing native call also omits output buffers needed by nvCOMP 5.2; the GPU regression tests reproduced an invalid-value error without the actual-size buffer.
Repack unaligned payloads, retain their buffers through asynchronous decoding, and supply native status and actual-size scratch with cleanup on failure. This extracts the decoder correctness fixes from #76 while preserving the existing CUDA backend and public API.
Validation: 397 decoder and FITS integration tests passed with a rebuilt native extension on an RTX PRO 6000 Blackwell. The CPU xDR suite passed 286 tests (11 GPU-only skips);
make lintand all public hooks passed. Coverage includes optional Gzip headers, raw payloads, unaligned views, pooled allocations, and non-default streams.