Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
40 changes: 24 additions & 16 deletions benchmarks/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,11 +17,10 @@ Exits nonzero if a gated metric exceeds its ceiling.
Correctness tests do not catch cost regressions, so the suite measures two
things that move silently:

**Layout size and memory allocated by an operation.** Both deterministic and
comparable across machines, so they are gated tightly. Memory uses
`tracemalloc` rather than `ru_maxrss`, which is a process-lifetime high-water
mark and reports zero for an operation staying under the peak set while
building its input.
**Layout size.** `bytes_per_nonzero` and `indices_bytes_per_nonzero` -- what
the compressed format actually costs per stored value. Deterministic and
comparable across machines, so they are gated tightly, and nothing else in
the project guards them.

**Throughput relative to scipy**, never absolute seconds. The same work is
timed through scipy in the same process and the ratio recorded, which cancels
Expand All @@ -38,20 +37,29 @@ leak between them otherwise.
metrics named in `margins` are gated; anything else a case returns is
recorded for context.

The checked-in ceilings were recorded before the memory fixes landed, so the
memory ones are deliberately generous and should be re-recorded as those
merge. On a 4M-nonzero array, for results of length `n_minor`:

| metric | recorded | with the fix |
|---|---|---|
| `minor_sum_peak_mb` | 66 MB | 0.8 MB |
| `minor_extrema_peak_mb` | 66 MB | 1.6 MB |
| `minor_getnnz_peak_mb` | 32 MB | 0.8 MB |
| `minor_selection_peak_mb` | 62 MB | 4.8 MB |
| `misaligned_matmul_peak_mb` | 204 MB | 115 MB |
Memory is no longer measured here; see below.

## Adding a case

Write a function returning `{metric: value}` in `cases.py`, decorated with
`@fast` (runs on every PR, keep it under a minute) or `@slow`. Add any new
gated metric to `margins` in `baselines.json`, then `--record`.

## Memory is a test, not a benchmark

Memory is not measured here. `pytest-memray` ceilings in `tests/` assert it
instead, because the claims are structural: "a minor-axis reduction must not
allocate anything nnz-sized" either holds or it does not. A ceiling fails the
moment it stops being true, where a recorded number only shows it drifting,
against a baseline needing a re-record whenever the bound legitimately moves.

Those tests pin `numba.set_num_threads`, so a ceiling means the same thing on
a 4-core runner and a 96-core one. A recorded figure could not: the
accumulators being bounded are sized by the runner's core count.

Neither tool sees RSS, so neither catches allocator *fragmentation* -- many
variably-sized alloc/free cycles driving the resident set far above the live
set. Both report the live high-water mark, which stays small throughout such a
run. That failure mode is real but not reliably gateable, since RSS moves with
the allocator, the runner and the thread count; diagnose it with
`/proc/self/status` `VmHWM` around a workload when it is suspected.
109 changes: 21 additions & 88 deletions benchmarks/baselines.json
Original file line number Diff line number Diff line change
@@ -1,12 +1,10 @@
{
"comment": "Ceilings a metric must stay under, regenerated with `python -m benchmarks.run --set fast --record`. Only metrics listed in `margins` are gated; the rest are recorded by cases for context. Margins are the slack applied to a measured value when recording: tight for deterministic layout/memory numbers, loose for timing ratios, which vary with the runner.",
"comment": "Ceilings a metric must stay under, regenerated with `python -m benchmarks.run --set fast --record`. Only metrics listed in `margins` are gated; the rest are recorded by cases for context. Margins are the slack applied to a measured value when recording: tight for deterministic layout numbers, loose for timing ratios, which vary with the runner.",
"margins": {
"bytes_per_nonzero": 1.1,
"indices_bytes_per_nonzero": 1.1,
"vs_scipy_ratio": 1.1,
"peak_alloc_mb": 2.0,
"time_ratio_vs_scipy": 4.0,
"peak_alloc_mb_view": 2.0,
"time_ratio_view_over_materialize": 4.0,
"cpu_ratio_1t_cp10k_log1p": 1.5,
"cpu_ratio_1t_parafac2": 1.5,
Expand All @@ -20,136 +18,71 @@
"indices_bytes_per_nonzero": 4.4,
"vs_scipy_ratio": 0.4751
},
"minor_sum_peak_mb": {
"peak_alloc_mb": 132.5131
},
"misaligned_matmul_peak_mb": {
"peak_alloc_mb": 407.2383
},
"matvec_vs_scipy": {
"time_ratio_vs_scipy": 1.0
},
"matmat_vs_scipy": {
"time_ratio_vs_scipy": 1.0
},
"minor_extrema_peak_mb": {
"peak_alloc_mb": 132.5453
},
"minor_getnnz_peak_mb": {
"peak_alloc_mb": 64.0324
},
"minor_selection_peak_mb": {
"peak_alloc_mb": 124.5457
},
"normalized_cp10k_log1p_matmat_vs_materialize": {
"time_ratio_view_over_materialize": 0.3,
"peak_alloc_mb_view": 2.7262
"time_ratio_view_over_materialize": 0.3
},
"normalized_cp10k_log1p_matvec_vs_materialize": {
"time_ratio_view_over_materialize": 0.3,
"peak_alloc_mb_view": 0.3858
"time_ratio_view_over_materialize": 0.3
},
"normalized_cp10k_log1p_rmatmat_vs_materialize": {
"time_ratio_view_over_materialize": 0.3,
"peak_alloc_mb_view": 3.6521
"time_ratio_view_over_materialize": 0.3
},
"normalized_cp10k_log1p_rmatvec_vs_materialize": {
"time_ratio_view_over_materialize": 0.3,
"peak_alloc_mb_view": 0.1321
"time_ratio_view_over_materialize": 0.3
},
"normalized_parafac2_matmat_vs_materialize": {
"time_ratio_view_over_materialize": 0.3,
"peak_alloc_mb_view": 2.7262
"time_ratio_view_over_materialize": 0.3
},
"normalized_parafac2_matvec_vs_materialize": {
"time_ratio_view_over_materialize": 0.3,
"peak_alloc_mb_view": 0.3858
"time_ratio_view_over_materialize": 0.3
},
"normalized_parafac2_rmatmat_vs_materialize": {
"time_ratio_view_over_materialize": 0.3,
"peak_alloc_mb_view": 3.6521
"time_ratio_view_over_materialize": 0.3
},
"normalized_parafac2_rmatvec_vs_materialize": {
"time_ratio_view_over_materialize": 0.3,
"peak_alloc_mb_view": 0.1321
"time_ratio_view_over_materialize": 0.3
},
"normalized_pearson_matmat_vs_materialize": {
"time_ratio_view_over_materialize": 0.3,
"peak_alloc_mb_view": 2.7262
"time_ratio_view_over_materialize": 0.3
},
"normalized_pearson_matvec_vs_materialize": {
"time_ratio_view_over_materialize": 0.3,
"peak_alloc_mb_view": 0.3858
"time_ratio_view_over_materialize": 0.3
},
"normalized_pearson_rmatmat_vs_materialize": {
"time_ratio_view_over_materialize": 0.3,
"peak_alloc_mb_view": 3.6521
"time_ratio_view_over_materialize": 0.3
},
"normalized_pearson_rmatvec_vs_materialize": {
"time_ratio_view_over_materialize": 0.3,
"peak_alloc_mb_view": 0.1321
"time_ratio_view_over_materialize": 0.3
},
"normalized_raw_matmat_vs_materialize": {
"time_ratio_view_over_materialize": 0.3,
"peak_alloc_mb_view": 2.7262
"time_ratio_view_over_materialize": 0.3
},
"normalized_raw_matvec_vs_materialize": {
"time_ratio_view_over_materialize": 0.3,
"peak_alloc_mb_view": 0.3858
"time_ratio_view_over_materialize": 0.3
},
"normalized_raw_rmatmat_vs_materialize": {
"time_ratio_view_over_materialize": 0.3,
"peak_alloc_mb_view": 3.6521
"time_ratio_view_over_materialize": 0.3
},
"normalized_raw_rmatvec_vs_materialize": {
"time_ratio_view_over_materialize": 0.3,
"peak_alloc_mb_view": 0.1321
"time_ratio_view_over_materialize": 0.3
},
"normalized_scanpy_matmat_vs_materialize": {
"time_ratio_view_over_materialize": 0.3,
"peak_alloc_mb_view": 2.7262
"time_ratio_view_over_materialize": 0.3
},
"normalized_scanpy_matvec_vs_materialize": {
"time_ratio_view_over_materialize": 0.3,
"peak_alloc_mb_view": 0.3858
"time_ratio_view_over_materialize": 0.3
},
"normalized_scanpy_rmatmat_vs_materialize": {
"time_ratio_view_over_materialize": 0.3,
"peak_alloc_mb_view": 3.6521
"time_ratio_view_over_materialize": 0.3
},
"normalized_scanpy_rmatvec_vs_materialize": {
"time_ratio_view_over_materialize": 0.3,
"peak_alloc_mb_view": 0.1321
},
"normalized_cp10k_log1p_matmat_vs_sparse": {
"peak_alloc_mb_view": 2.7262
},
"normalized_cp10k_log1p_matvec_vs_sparse": {
"peak_alloc_mb_view": 0.3858
},
"normalized_parafac2_matmat_vs_sparse": {
"peak_alloc_mb_view": 2.7262
},
"normalized_parafac2_matvec_vs_sparse": {
"peak_alloc_mb_view": 0.3858
},
"normalized_pearson_matmat_vs_sparse": {
"peak_alloc_mb_view": 2.7262
},
"normalized_pearson_matvec_vs_sparse": {
"peak_alloc_mb_view": 0.3858
},
"normalized_raw_matmat_vs_sparse": {
"peak_alloc_mb_view": 2.7262
},
"normalized_raw_matvec_vs_sparse": {
"peak_alloc_mb_view": 0.3858
},
"normalized_scanpy_matmat_vs_sparse": {
"peak_alloc_mb_view": 2.7262
},
"normalized_scanpy_matvec_vs_sparse": {
"peak_alloc_mb_view": 0.3858
"time_ratio_view_over_materialize": 0.3
},
"normalized_cpu_vs_sparse_1t": {
"cpu_ratio_1t_cp10k_log1p": 7.815,
Expand Down
89 changes: 6 additions & 83 deletions benchmarks/cases.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,6 @@
best_cpu_time,
best_time,
integer_counts_csr,
peak_alloc_mb,
ratio_vs_scipy,
)

Expand Down Expand Up @@ -46,75 +45,6 @@ def layout_bytes_per_nonzero() -> dict[str, float]:
}


# -- memory ceilings ---------------------------------------------------------


@fast
def minor_sum_peak_mb() -> dict[str, float]:
"""Memory allocated by a minor-axis sum."""
from vsparse import VCSRArray

v = VCSRArray.from_scipy(integer_counts_csr(40_000, 2_000, density=0.05))
nnz_mb = v.nnz * 8 / 1e6
return {
"peak_alloc_mb": peak_alloc_mb(lambda: v.sum(axis=0)),
"expanded_nnz_mb": nnz_mb, # what a per-nonzero temporary would cost
}


@fast
def misaligned_matmul_peak_mb() -> dict[str, float]:
"""Memory allocated by the matmul direction the storage isn't aligned for."""
from vsparse import VCSCArray

v = VCSCArray.from_scipy(integer_counts_csr(40_000, 2_000, density=0.05))
rng = np.random.default_rng(0)
B = rng.normal(size=(v.shape[1], 4))
array_mb = (v.values.nbytes + v.value_ptr.nbytes + v.indices.nbytes) / 1e6

return {
"peak_alloc_mb": peak_alloc_mb(lambda: v.normalized() @ B),
"array_mb": array_mb, # what a full second copy would cost
}


@fast
def minor_extrema_peak_mb() -> dict[str, float]:
"""Memory allocated by a minor-axis max/min."""
from vsparse import VCSRArray

v = VCSRArray.from_scipy(integer_counts_csr(40_000, 2_000, density=0.05))
return {
"peak_alloc_mb": peak_alloc_mb(lambda: v.max(axis=0)),
"expanded_nnz_mb": v.nnz * 8 / 1e6,
}


@fast
def minor_getnnz_peak_mb() -> dict[str, float]:
"""Memory allocated by a per-minor-index stored-element count."""
from vsparse import VCSRArray

v = VCSRArray.from_scipy(integer_counts_csr(40_000, 2_000, density=0.05))
return {
"peak_alloc_mb": peak_alloc_mb(lambda: v.getnnz(axis=0)),
"indices_nnz_mb": v.nnz * 8 / 1e6,
}


@fast
def minor_selection_peak_mb() -> dict[str, float]:
"""Memory allocated by a minor-axis selection."""
from vsparse import VCSRArray

v = VCSRArray.from_scipy(integer_counts_csr(40_000, 2_000, density=0.05))
cols = np.arange(0, v.shape[1], 2)
return {
"peak_alloc_mb": peak_alloc_mb(lambda: v[:, cols]),
"indices_nnz_mb": v.nnz * 8 / 1e6,
}


# -- throughput, relative to scipy -------------------------------------------


Expand Down Expand Up @@ -144,12 +74,12 @@ def matmat_vs_scipy() -> dict[str, float]:

# -- normalized views (issue #40 recipes): view-op vs materialize-then-op ---
#
# For every recipe, the view-based matmul/matvec should cost less, both in
# time and in peak allocation, than fully materializing the (dense,
# implicit-zero-filling) normalized matrix and multiplying that -- the whole
# point of a *view*. ``time_ratio_view_over_materialize`` < 1 and
# ``peak_alloc_mb_view`` < ``peak_alloc_mb_materialize`` are the expectation
# for every case below.
# For every recipe, the view-based matmul/matvec should be faster than fully
# materializing the (dense, implicit-zero-filling) normalized matrix and
# multiplying that -- the whole point of a *view*.
# ``time_ratio_view_over_materialize`` < 1 is the expectation for every case
# below. The matching memory claim is asserted as a ceiling in
# ``tests/test_minor_axis_matmul.py`` rather than recorded here.


def _normalized_bench(recipe: str, *, vector: bool) -> Callable[[], dict[str, float]]:
Expand All @@ -170,8 +100,6 @@ def via_materialize() -> np.ndarray:

return {
"time_ratio_view_over_materialize": best_time(via_view) / best_time(via_materialize),
"peak_alloc_mb_view": peak_alloc_mb(via_view),
"peak_alloc_mb_materialize": peak_alloc_mb(via_materialize),
}

bench.__name__ = f"normalized_{recipe}_{'matvec' if vector else 'matmat'}_vs_materialize"
Expand All @@ -196,8 +124,6 @@ def via_materialize() -> np.ndarray:

return {
"time_ratio_view_over_materialize": best_time(via_view) / best_time(via_materialize),
"peak_alloc_mb_view": peak_alloc_mb(via_view),
"peak_alloc_mb_materialize": peak_alloc_mb(via_materialize),
}

bench.__name__ = f"normalized_{recipe}_{'rmatvec' if vector else 'rmatmat'}_vs_materialize"
Expand Down Expand Up @@ -278,8 +204,6 @@ def via_sparse() -> np.ndarray:
# every case doubled the suite's runtime for a number nothing gates.
return {
"wall_ratio_view_over_sparse": best_time(via_view) / best_time(via_sparse),
"peak_alloc_mb_view": peak_alloc_mb(via_view),
"peak_alloc_mb_sparse_delta": peak_alloc_mb(lambda: _sparse_delta(nv, mat)),
}

bench.__name__ = f"normalized_{recipe}_{'matvec' if vector else 'matmat'}_vs_sparse"
Expand Down Expand Up @@ -368,7 +292,6 @@ def large_layout_and_matmul() -> dict[str, float]:
return {
"bytes_per_nonzero": stored / v.nnz,
"time_ratio_vs_scipy": ratio_vs_scipy(lambda: v @ B, lambda: mat @ B),
"matmul_peak_alloc_mb": peak_alloc_mb(lambda: v @ B),
}


Expand Down
20 changes: 0 additions & 20 deletions benchmarks/harness.py
Original file line number Diff line number Diff line change
@@ -1,33 +1,13 @@
from __future__ import annotations

import time
import tracemalloc
from collections.abc import Callable
from typing import Any

import numpy as np
import scipy.sparse as sp


def peak_alloc_mb(fn: Callable[[], Any]) -> float:
"""Peak memory allocated during ``fn``, in MB.

Not RSS, which is a process-lifetime high-water mark and so reports zero
for anything staying under the peak set while building its input. numpy
allocations are traced, numba's internal ones are not.
"""
fn() # JIT compile / warm caches outside the measurement
tracemalloc.start()
try:
before = tracemalloc.get_traced_memory()[0]
tracemalloc.reset_peak()
fn()
peak = tracemalloc.get_traced_memory()[1]
finally:
tracemalloc.stop()
return max(0.0, (peak - before) / 1e6)


def best_time(fn: Callable[[], Any], repeat: int = 7) -> float:
"""Best wall-clock time over ``repeat`` runs, in seconds. Warms up first."""
fn() # JIT compile / allocate caches outside the measurement
Expand Down
Loading
Loading