Skip to content

perf(rocm): profile gfx1151 decode per kernel and rank the #1814 ports - #2086

Merged
inureyes merged 6 commits into
mainfrom
update/issue-2061-rocm-decode-profile
Sep 30, 2026
Merged

inureyes merged 6 commits into
mainfrom
update/issue-2061-rocm-decode-profile

Conversation

@inureyes

@inureyes inureyes commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Measures where ROCm decode time goes per kernel on gfx1151 and ranks the port issues split from #1814 by it; settles MLX's ROCm gather_mm.

What changed

  • docs/benchmark_results/rocm-decode-profile-gfx1151-2026-09-30.md: decode GPU time and host gap per token, top kernels, share per port unit for Llama-3.1-8B, Qwen3-30B-A3B, granite-4.0-h-tiny and Nemotron-3-Nano (greedy, plus temperature and top-p runs), the attribution method, and the ranked order. Raw rocprofv3 stats CSVs, per-kernel decode tables, summaries and the guard log are under benchmarks/rocm_profiles/gfx1151_929c80ab/.
  • Harness: scripts/rocm_gpu_guard.sh (the perf(bench): ROCm support in the benchmark harness and a gfx1151 baseline #2056 idle-GPU guard as a script), scripts/rocm_decode_profile.sh / .py (profile, decode cut, attribution, report), mlxcel-bench-decode --temperature/--top-p and MLXCEL_BENCH_PHASE_MARKS=1; documented in docs/benchmarks.md.
  • grouped_gemm_numeric_tests.rs gates read gpu_backend_available(); the file left BACKEND_ENUMERATION_TODO.

Result: implied order #2067, #2065, #2064, #2068, #2063 (items 7, 5, 4, 8, 3). Every measured run passed the guard (90 s idle, 1 Hz monitor); one attempt was rejected by the guard, one rerun by hand. Profiler cost scales with dispatch count: within noise on Llama, 5% on Qwen3, 20 to 23% on the hybrids, so the doc reports shares.

Verification (gfx1151)

  • cargo test --release --features rocm -p mlxcel-core --lib grouped_gemm_numeric_tests -- --test-threads=1 --nocapture: 3 passed; each also run by exact name under rocprofv3, trace shows gather_batched_gemm_kernel and a hipBLASLt GEMM; all three fail with a wrong-expert reference
  • python3 scripts/ci/check_kernel_port_dispatch.py: 0 awaiting a predicate; make verify-kernel-port-dispatch passes
  • make verify-versions verify-kernel-dtype-keys verify-kernel-port-dispatch verify-llama-compat verify-fmt: pass
  • cargo clippy -p mlxcel --features rocm --bin mlxcel-bench-decode -- -D warnings and cargo clippy -p mlxcel-core --features rocm --lib --tests -- -D warnings: clean (narrow scope instead of make verify-clippy-rocm, which is workspace-wide)
  • cargo test --features rocm --test dead_doc_pointers: pass
  • python3 -m unittest tests/test_rocm_decode_profile.py: 17 passed; bash -n on all three scripts (shellcheck is not installed on this host)

Changes during review

  • The report now states that per-token figures divide by the 128 generated tokens while the decode window holds 127 forward passes (the first token comes from prefill), gives exact per-step SSD dispatch counts (47 granite, 50 Nemotron), and widens the perf(rocm): port fused_add_rms_norm and fused_rope_qk_append to HIP #2063 opt-in ceiling to 0.83 to 0.89%.
  • Script hardening from the security review: rocm_decode_profile.sh scrubs local paths from guard.log in an EXIT trap with a literal replacement; rocm_gpu_guard.sh parses the parent PID after the last ), avoids glob expansion, validates its integer options, and stops the command and monitor on INT/TERM; Dispatch in rocm_decode_profile.py is slotted. rocm_decode_profile.py report output on the committed data is byte-identical before and after.

Not verified: Metal and CUDA (not available here). The gather_mm gate change means the three tests still run there, unchanged. cargo test --test dead_doc_pointers without --features rocm fails to link on this host (copy_gpu_inplace undefined from kv_inplace_write.cpp), unrelated to this change.

Closes #2061

@inureyes inureyes added status:review Under review type:performance Performance improvements priority:medium Medium priority area:core mlxcel-core: MLX FFI, primitives, KV cache, layers platform:linux Linux (CUDA / packaging) specific labels Sep 30, 2026
@inureyes

Copy link
Copy Markdown
Member Author

Implementation Review Summary

Intent

Profile gfx1151 decode per kernel on four models, attribute the time to the #1814 port units, rank #2063/#2064/#2065/#2067/#2068 by measured share, and settle MLX's ROCm gather_mm.

Findings Addressed

  • Attribution check quoted 46.6 and 49.6 SSD dispatches per Mamba2 layer and "whole multiples per token"; the window holds 127 decode steps while per-token figures divide by 128, so per step the counts are exactly 47 and 50. Method now states the 127/128 relation (MEDIUM)
  • perf(rocm): port fused_add_rms_norm and fused_rope_qk_append to HIP #2063 opt-in ceiling given as 0.83%; the top-p run reads 0.89% (LOW)
  • perf(rocm): port gumbel_max_sample and rejection_sample to HIP #2064 text said the reached share excludes the --ignore-eos logit bias; the summary counts the whole sampler tail, so the figure is slightly high (LOW)
  • A guard-rejected attempt stays in granite-4.0-h-tiny-4bit_greedy_bench.log; the doc now says only the last, accepted attempt is read (LOW)

Checked, no change needed

Every figure in the doc's tables matches rocm_decode_profile.py report over the committed summaries; plain host gap, GEMV bandwidth, SSM-state traffic and the 430-dispatch MoE estimate recompute from them. The #2065-before-#2064 and #2068-before-#2063 deviations from pure share are argued explicitly. Cited source lines (layers.rs:808/:820/:978, llama3.rs:669/:1201, qwen3_moe.rs:223, granitemoehybrid.rs:440, nemotron_h.rs:537/:635) hold. Phase marks are off unless MLXCEL_BENCH_PHASE_MARKS=1.

Verification

  • All stated requirements implemented
  • No placeholder/mock code remaining
  • Integrated into project code flow
  • Project conventions followed
  • Existing modules reused where applicable
  • No unintended structural changes
  • Tests pass (python3 -m unittest tests/test_rocm_decode_profile.py: 13 passed; checker and make verify-kernel-port-dispatch: 0 awaiting a predicate). No cargo or GPU run in this review, to keep the shared host idle for other measurements; the follow-up commit is doc-only.

@inureyes

Copy link
Copy Markdown
Member Author

Security and performance review

No CRITICAL or HIGH findings, so nothing was changed on the branch. The scripts are local operator tooling: every input (model paths, flags, ROCPROFV3, the guarded command) comes from the person running them, so none of the items below crosses a trust boundary. python3 -m unittest tests/test_rocm_decode_profile.py passes (13 tests).

Checked and clean: commands are passed as argv arrays ("$@", "${args[@]}") with no eval or sh -c; the default trace dir comes from mktemp -d; the Python post-processor has no subprocess, eval, or unsafe deserialization; the guard never kills or signals a process (it only observes and reruns); the published data under benchmarks/rocm_profiles/gfx1151_929c80ab/ contains no local paths. With MLXCEL_BENCH_PHASE_MARKS unset, the only added runtime cost is one Stamp::now() (two vDSO clock_gettime calls) per measured pass, taken after the generator has already recorded decode_time_ms, so it does not touch the timed region.

MEDIUM

  • scripts/rocm_decode_profile.sh: under set -e, a guard exit 75 or a failing rocprofv3/cp aborts the loop before the final sed -i that scrubs $ROOT and the trace dir from guard.log, leaving absolute local paths in a file that is meant to be committed. A trap on EXIT that runs the scrub would cover every exit path. The abort also skips the remaining models silently apart from the guard's own lines.
  • scripts/rocm_decode_profile.py read_trace: loads the whole kernel trace into a list of non-slotted dataclass instances. For the traces the header describes (hundreds of MB), that is several GB of RAM. @dataclass(slots=True) or streaming only the rows inside [warmup_start, measured_end] would bound it.

LOW

  • scripts/rocm_gpu_guard.sh descends_from: awk '{print $4}' /proc/$pid/stat returns the wrong field when a process's comm contains a space. Parsing after the last ) (e.g. ${stat##*) } then the second word) is exact.
  • scripts/rocm_gpu_guard.sh: --idle-secs, --max-attempts, --max-wait go straight into (( )), which evaluates arbitrary expressions (including $(...) inside an array subscript). Not exploitable here since the caller already chooses the command to run, but a ^[0-9]+$ check gives a clean error for typos.
  • scripts/rocm_gpu_guard.sh foreign_holders: for entry in $2 is unquoted, so a comm containing * or ? (and the '?' fallback itself) is glob-expanded against the cwd. set -f in that function, or read -ra with IFS=,, avoids it.
  • scripts/rocm_gpu_guard.sh: no signal trap, so a SIGTERM to the guard leaves the command and the monitor running. The monitor's kill -0 "$cmd_pid" loop could in principle follow a reused PID after the command is reaped.
  • scripts/rocm_decode_profile.sh: the sed scrub interpolates $ROOT/$TRACE_DIR into the pattern, so a path containing # breaks it (and . matches any character).

@inureyes inureyes added status:done Completed and removed status:review Under review labels Sep 30, 2026
The three grouped_gemm_numeric_tests gated on Metal or CUDA, so on ROCm they returned before touching the GPU and MLX's ROCm GatherMM was never checked. Run by exact name on gfx1151 with the gate widened, all three pass against the f64 host reference, and a rocprofv3 kernel trace shows the overlay's gather_batched_gemm_kernel (f32, bf16, f16) and a hipBLASLt GEMM for the sorted single-row case, so the pass is not vacuous. With the reference pointed at the wrong expert, all three fail at the value assertion.

The gates now read gpu_backend_available(), the file leaves BACKEND_ENUMERATION_TODO in check_kernel_port_dispatch.py, and the checker reports 0 awaiting a predicate. Metal and CUDA still run the tests (not run here).

Refs #2061
ROCm had end-to-end tok/s only, so the #1814 kernel ports had no measured order. This adds what a per-kernel decode profile needs, as reusable tooling rather than a one-off:

- scripts/rocm_gpu_guard.sh: the idle-GPU guard the #2056 baseline described, as a script (90 s with /sys/class/kfd/kfd/proc empty and no compiler, a 1 Hz monitor that ignores the command's own GPU processes, rerun on contention, every sample logged).
- mlxcel-bench-decode: --temperature and --top-p (default greedy, unchanged), and MLXCEL_BENCH_PHASE_MARKS=1, which prints the warmup, measured, decode-start and end times on CLOCK_MONOTONIC and CLOCK_BOOTTIME so a trace can be cut to the measured decode by timestamp.
- scripts/rocm_decode_profile.sh runs a plain and a rocprofv3 --kernel-trace --hip-graph-trace --stats run per model under the guard; scripts/rocm_decode_profile.py cuts the decode window, reports GPU time and host gap per token, profiler cost, top kernels, and attributes dispatches to mlxcel ops and #1814 port units by dispatch order, with which ports mlxcel actually reaches per model.
- tests/test_rocm_decode_profile.py covers the guard against a fake KFD directory and the cut and attribution rules on synthetic steps; docs/benchmarks.md documents the harness.

Refs #2061
Profiles greedy and sampled decode (pp512/tg128) of Llama-3.1-8B, Qwen3-30B-A3B, granite-4.0-h-tiny and Nemotron-3-Nano-30B-A3B on gfx1151 with rocprofv3, every run under the idle-GPU guard (16 accepted runs; one rejected by the guard, one discarded and rerun by hand), and attributes decode GPU time to the port units split from #1814.

Shipped-settings share of decode GPU time: #2067 SSM update 29.8% (granite) and 20.0% (Nemotron) plus more than half their dispatches; #2065 fused MoE 46.8% of Qwen3, mostly GEMVs already at about 181 GB/s; #2064 samplers 0 greedy, 0.4 to 3.8% sampled; #2063 zero (both fusions ship off) and #2068 zero (no paged path in single-stream decode). Implied order: 7, 5, 4, 8, 3.

The rocprofv3 stats CSVs, per-kernel decode tables, summaries, bench logs and the guard log are under benchmarks/rocm_profiles/gfx1151_929c80ab/. The report also records the gather_mm test outcome.

Refs #2061
The decode window holds 127 forward passes (the first of 128 tokens comes from the prefill), while per-token figures divide by 128 to match the bench's tok/s. The attribution check quoted 46.6 and 49.6 SSD dispatches per Mamba2 layer and claimed whole multiples "per token"; per step they are exactly 47 and 50, which is the check that actually holds. The Method section now states the 127/128 relation.

Also: the #2063 opt-in ceiling is 0.83 to 0.89% (the top-p run reads 0.89), the #2064 reached share counts the whole sampler tail including the --ignore-eos logit bias, so it is slightly high rather than net of it, and a guard-rejected attempt stays in the run's _bench.log (granite greedy) while only the last, accepted attempt is read.

Numbers recomputed from the committed summaries and decode CSVs; no GPU run.

Refs #2061
rocm_decode_profile.sh now scrubs local paths from the published guard.log from an EXIT trap, so an early exit (guard 75, rocprofv3 or cp failure) no longer leaves them in, and the replacement is literal so '#' and regex characters in the paths are harmless. rocm_decode_profile.py makes Dispatch slotted to bound the memory of a full kernel trace; the report output for benchmarks/rocm_profiles/gfx1151_929c80ab is byte-identical.

rocm_gpu_guard.sh reads the parent pid after the last ')' of /proc/<pid>/stat, avoids glob expansion when splitting holders, rejects non-integer --idle-secs, --max-attempts and --max-wait (and a missing option value) with exit 2, and stops the command and monitor on INT or TERM (exit 130 or 143). Tests cover the validation, octal-looking values, the SIGTERM path and slots.

Validation: python3 -m unittest tests/test_rocm_decode_profile.py (17 pass), bash -n on both scripts, make verify-fmt verify-kernel-port-dispatch. No GPU workload was run.

Refs #2061
@inureyes
inureyes force-pushed the update/issue-2061-rocm-decode-profile branch from c9d6811 to 0d5d6db Compare September 30, 2026 14:47
@inureyes
inureyes merged commit 78af3e2 into main Sep 30, 2026
25 checks passed
@inureyes
inureyes deleted the update/issue-2061-rocm-decode-profile branch September 30, 2026 15:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:core mlxcel-core: MLX FFI, primitives, KV cache, layers platform:linux Linux (CUDA / packaging) specific priority:medium Medium priority status:done Completed type:performance Performance improvements

Projects

None yet

Development

Successfully merging this pull request may close these issues.

perf(rocm): profile decode on gfx1151 and settle MLX ROCm gather_mm

1 participant