Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
61 changes: 58 additions & 3 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,9 @@ jobs:
python-version: "3.12"

- name: Install dependencies
run: uv sync --all-extras --dev
# --no-extra cuda: the CuPy wheel is ~1 GB and useless without a GPU.
# The `cuda` job below installs it on the self-hosted runner instead.
run: uv sync --all-extras --no-extra cuda --dev

- name: Lint with Ruff
run: uv run ruff check .
Expand Down Expand Up @@ -57,7 +59,9 @@ jobs:
python-version: ${{ matrix.python-version }}

- name: Install dependencies
run: uv sync --all-extras --dev
# --no-extra cuda: the CuPy wheel is ~1 GB and useless without a GPU.
# The `cuda` job below installs it on the self-hosted runner instead.
run: uv sync --all-extras --no-extra cuda --dev

- name: Run pytest
run: uv run pytest --cov
Expand All @@ -76,11 +80,62 @@ jobs:
python-version: "3.12"

- name: Install dependencies
run: uv sync --all-extras --dev
# --no-extra cuda: the CuPy wheel is ~1 GB and useless without a GPU.
# The `cuda` job below installs it on the self-hosted runner instead.
run: uv sync --all-extras --no-extra cuda --dev

# Fails if a gated metric exceeds its ceiling in benchmarks/baselines.json.
# Memory and layout numbers are deterministic and gated tightly; timing is
# recorded as a ratio against scipy in the same process, and gated loosely,
# because a shared runner's absolute speed means nothing. See benchmarks/README.md.
- name: Run benchmark gate
run: uv run python -m benchmarks.run --set fast

cuda:
name: CUDA Tests and Benchmarks
# Any self-hosted runner, with no extra labels to register or keep in sync.
# That means the job can land on a machine without a usable GPU, which the
# "Assert a CUDA device is visible" step below turns into a loud failure
# rather than a silent skip -- the same check that has to be there anyway.
# Forked PRs never see this runner, and `pull_request` from a fork cannot
# reach a self-hosted runner regardless -- but the guard is explicit so an
# untrusted PR can never run code on our own hardware.
if: github.event_name == 'push' || github.event.pull_request.head.repo.full_name == github.repository
runs-on: self-hosted
steps:
- name: Checkout code
uses: actions/checkout@v7

- name: Install uv
uses: astral-sh/setup-uv@v10.0.1
with:
enable-cache: true
python-version: "3.12"

# Informational, and deliberately non-fatal: on a runner with no GPU at
# all this would otherwise fail first, with a bare "nvidia-smi: not
# found" instead of the actionable message the assert step below prints.
- name: Show the device the job landed on
continue-on-error: true
run: nvidia-smi --query-gpu=name,driver_version,compute_cap,memory.total --format=csv

- name: Install dependencies
run: uv sync --all-extras --dev

# Load-bearing, not belt-and-braces: `runs-on: self-hosted` carries no
# promise of a GPU, and tests/test_cuda.py skips itself when CuPy cannot
# see one -- so without this the job would pass green having tested
# nothing at all.
- name: Assert a CUDA device is visible
run: uv run python -c "import vsparse, sys; sys.exit(0 if vsparse.cuda_is_available() else 'no CUDA device visible to CuPy')"

- name: Run the CUDA tests
run: uv run pytest -m cuda

# Same gate as the CPU benchmark job, over benchmarks/baselines.json.
# `cuda_device_bytes_ratio_vs_csr` is deterministic and gated tightly;
# the timing ratio is against cuSPARSE on the same device in the same
# process, which cancels most of the difference between GPU models, and
# is gated loosely for what it does not. See benchmarks/README.md.
- name: Run the CUDA benchmark gate
run: uv run python -m benchmarks.run --set cuda
38 changes: 37 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,14 @@ pip install vsparse
uv add vsparse
```

For the CUDA support described below (needs an NVIDIA GPU; pulls in CuPy):

```sh
pip install "vsparse[cuda]"
# or with uv
uv add "vsparse[cuda]"
```

From source:

```sh
Expand Down Expand Up @@ -140,14 +148,42 @@ adata_norm = vsparse.load_and_normalize(
)
```

### CUDA (`to_gpu`)

A normalized view moves onto an NVIDIA GPU with `.to_gpu()`, keeping its value-compressed layout there rather than expanding to one float per nonzero — so a matrix that only fits in device memory compressed still fits. Everything floating-point on the device is float32.

```python
gpu = vsparse.VCSRArray.from_scipy(counts).normalized("parafac2").to_gpu()
out = gpu @ B # float32 CuPy array; B @ gpu likewise
```

cuSPARSE reads only CSR/CSC, so walking the VCS layout on the device needs custom kernels; `vsparse._cuda` supplies four, one per (format, direction). Against the alternative — expanding to a sorted CuPy CSR and calling cuSPARSE — they measured (RTX 5080, 60k x 2k, 6M nonzeros, width 16):

| direction | kernel / cuSPARSE | device bytes vs CSR |
| --- | --- | --- |
| `VCSR @ B` | 0.71x–0.86x | 0.62x |
| `VCSC @ B` | 0.83x | 0.51x |
| `B @ VCSC` | 0.32x | 0.51x |
| `B @ VCSR` | 0.23x | 0.62x |

`CudaNormalizedView.to_cupy_sparse()` builds that CSR baseline if you want it.

See `docs/` for full usage guides and API documentation.

## Development

```sh
uv sync --all-extras --dev
uv sync --all-extras --dev # add --no-extra cuda to skip the ~1 GB CuPy wheel
uv run pytest
uv run ruff check .
uv run ty check
uv run sphinx-build -b html docs docs/_build/html
```

The CUDA tests and benchmarks need an NVIDIA GPU; without one `tests/test_cuda.py`
skips itself. On a GPU machine:

```sh
uv run pytest -m cuda
uv run python -m benchmarks.run --set cuda
```
38 changes: 36 additions & 2 deletions benchmarks/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ hand.
```sh
uv run python -m benchmarks.run --set fast # run + compare (what CI does)
uv run python -m benchmarks.run --set slow # the larger cases
uv run python -m benchmarks.run --set cuda # the GPU cases (needs a device)
uv run python -m benchmarks.run --case matvec_vs_scipy
uv run python -m benchmarks.run --set fast --record # rewrite baselines.json
```
Expand Down Expand Up @@ -40,11 +41,44 @@ recorded for context.

Memory is no longer measured here; see below.

## The CUDA set

`--set cuda` measures `vsparse._cuda`'s kernels against the alternative they
exist to beat: expanding the value-compressed layout to one float per nonzero
and letting cuSPARSE do the product. It runs on the self-hosted GPU runner, not
the shared ones, and needs `uv sync --all-extras` (the CPU jobs pass
`--no-extra cuda`).

The baseline is deliberately the *best* materialized option rather than the
cheapest to produce, so the comparison is not a strawman: CSR in every case,
including for a VCSC-backed view whose natural materialization is CSC, and
index-sorted. Both matter more than expected -- `B @ csc` measured ~5x
`B @ csr`, and an unsorted CSR ~6x a sorted one -- so `to_cupy_sparse` sorts by
default and the cases pass `format="csr"`.

Three metrics, of which two are gated:

- `cuda_time_ratio_kernel_over_csr` -- steady-state throughput, ours over
cuSPARSE's on the already-built CSR. Gated loosely, for the same reason the
scipy ratios are: it is a ratio measured on the same device in the same
process, which cancels most of the difference between GPU models but not all
of it (the two development GPUs differed by up to 1.8x on this metric).
- `cuda_device_bytes_ratio_vs_csr` -- device memory held, ours over the CSR's.
Deterministic, so gated tightly. This is the number the kernels exist to buy.
- `cuda_materialize_in_matmuls` -- how many of our matmuls the one-time
materialization costs. Recorded for context, not gated.

Timing a CUDA call needs `best_gpu_time`, not `best_time`: launches are
asynchronous, so an unsynchronized timer measures the launch and not the work.

## Adding a case

Write a function returning `{metric: value}` in `cases.py`, decorated with
`@fast` (runs on every PR, keep it under a minute) or `@slow`. Add any new
gated metric to `margins` in `baselines.json`, then `--record`.
`@fast` (runs on every PR, keep it under a minute), `@slow`, or `@cuda`. Add any
new gated metric to `margins` in `baselines.json`, then `--record`.

Record CUDA ceilings on the *slowest* device you expect to run them on, so a
faster one cannot trip a gate it should pass.

## Memory is a test, not a benchmark

Expand Down
36 changes: 35 additions & 1 deletion benchmarks/baselines.json
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,9 @@
"cpu_ratio_1t_parafac2": 1.5,
"cpu_ratio_1t_pearson": 1.5,
"cpu_ratio_1t_raw": 1.5,
"cpu_ratio_1t_scanpy": 1.5
"cpu_ratio_1t_scanpy": 1.5,
"cuda_time_ratio_kernel_over_csr": 2.0,
"cuda_device_bytes_ratio_vs_csr": 1.1
},
"cases": {
"layout_bytes_per_nonzero": {
Expand Down Expand Up @@ -99,6 +101,38 @@
},
"misaligned_rmatmat_vs_scipy": {
"time_ratio_vs_scipy": 1.0817
},
"cuda_cp10k_log1p_matmul_vs_csr": {
"cuda_time_ratio_kernel_over_csr": 1.6947,
"cuda_device_bytes_ratio_vs_csr": 0.6792
},
"cuda_parafac2_matmul_vs_csr": {
"cuda_time_ratio_kernel_over_csr": 1.4385,
"cuda_device_bytes_ratio_vs_csr": 0.6792
},
"cuda_pearson_matmul_vs_csr": {
"cuda_time_ratio_kernel_over_csr": 1.4353,
"cuda_device_bytes_ratio_vs_csr": 0.6792
},
"cuda_raw_matmul_vs_csr": {
"cuda_time_ratio_kernel_over_csr": 1.4316,
"cuda_device_bytes_ratio_vs_csr": 0.6792
},
"cuda_scanpy_matmul_vs_csr": {
"cuda_time_ratio_kernel_over_csr": 1.7048,
"cuda_device_bytes_ratio_vs_csr": 0.6792
},
"cuda_parafac2_rmatmul_vs_csr": {
"cuda_time_ratio_kernel_over_csr": 0.6483,
"cuda_device_bytes_ratio_vs_csr": 0.5575
},
"cuda_parafac2_matmul_misaligned_vs_csr": {
"cuda_time_ratio_kernel_over_csr": 1.6738,
"cuda_device_bytes_ratio_vs_csr": 0.5575
},
"cuda_parafac2_rmatmul_misaligned_vs_csr": {
"cuda_time_ratio_kernel_over_csr": 0.4632,
"cuda_device_bytes_ratio_vs_csr": 0.6792
}
}
}
Loading
Loading