Skip to content

oscillators: size-specialised kernels for the saw/pulse tables and the FM sine table - #1171

Merged
dpwe merged 3 commits into
shorepine:mainfrom
rt-rtos:perf/sized-lut-kernels
Sep 22, 2026
Merged

dpwe merged 3 commits into
shorepine:mainfrom
rt-rtos:perf/sized-lut-kernels

Conversation

@rt-rtos

@rt-rtos rt-rtos commented Sep 18, 2026

Copy link
Copy Markdown
Contributor

Summary

render_lut_cub is generic over the table size, so its two shift amounts
and the mask occupy registers. On Xtensa that pushes the loop over the
register window: gcc spills the counter, the step, the amp increment and
both shift amounts, and emits the loop as decrement-and-branch, 74
instructions per sample with five stack reloads. With the table size a
compile-time constant the shifts and the mask are immediates and the same
arithmetic compiles to a 60-instruction hardware loop. Output is
bit-identical.

The first commit adds one instantiation per saw table a fundamental under
about 4 kHz picks (2048 down to 64 entries), selected once per block by
log_2_table_size in render_lpf_lut; the three smallest tables fall
back to the generic kernel. The second commit does the same for
render_lut_fm, which has the same spilled-counter loop: the FM
operators only ever read the 256-entry sine table, so it dispatches to a
lut_bits = 8 instantiation in the style of render_lut_256 and keeps
the generic body for any other table. That loop goes from 65
instructions to 32, a hardware loop with no stack reloads.

Context

  • The instantiations are noinline. Inlined into the dispatcher's switch
    the loop is lost again, since the six bodies then share one register
    allocation; checked in the per-TU assembly and in an LTO-linked ELF,
    where noinline holds and all six carry a loop instruction.
  • The dispatcher keys on the table's own log_2_table_size, so the baked
    BITS always matches the table that choose_from_lutset picked; a
    table outside 64..2048 entries takes the generic kernel. Only
    render_lpf_lut calls it; the sine and triangle paths are untouched.
  • Inside the loop the table values stay at their 16-bit scale. This is
    exact, not an approximation: MUL0_SS(L2S(x), f) is
    (x * (f >> 7)) >> 8, so the L2S shift-up and the multiply's
    shift-down cancel. It is where the last three instructions went, and it
    is why the kernel reads the LUTSAMPLE taps directly rather than
    converting each one. Like every >> on a SAMPLE in
    amy_fixedpoint.h, it assumes an arithmetic right shift on negative
    values, which gcc and clang give on every target AMY builds for.
  • The FM variant is the generic RENDER_LUT_GUTS(MOD_PART_MOD, NOTHING, INTERP_LINEAR) body with the table size substituted, not a rewrite, so
    its exactness needs no derivation. The feedback variants
    (render_lut_fb, render_lut_fm_fb) are left alone: the same
    treatment would apply, but they were not measured.
  • Fixed point only. The 16-bit-scale arithmetic has no float form, so the
    sized kernels and their dispatcher sit under #ifdef AMY_USE_FIXEDPOINT;
    a float build (amy.h documents removing the define) keeps calling
    render_lut_cub through the same name.
  • Size cutoff at 64 entries: tables of 32, 16 and 8 serve fundamentals
    above 4 kHz, where the cubic reads three harmonics or fewer and the
    render time is elsewhere. Six instantiations cost about 1.2 KB of
    instruction memory on the S3.

Verification

make test against the reference set: every test's error value
identical to a pristine 1.2.171 run on the same host. Host sweeps of a
single saw and pulse osc, 55 Hz to 3520 Hz, 48 kHz fixed point,
byte-identical WAVs against the generic kernel.

Per-TU Xtensa assembly at -O2 (esp-15.2.0 gcc), the loop of each kernel:

kernel before after
render_lut_cub (generic, still used for 32..8-entry tables) plain loop, 74 insns/sample unchanged
render_lut_cub_11 .. render_lut_cub_6 hardware loop, 60 insns/sample
render_lut_fm (generic, kept for any non-256 table) plain loop, 65 insns/sample unchanged
render_lut_fm_256 hardware loop, 32 insns/sample

On an ESP32-S3 at 48 kHz, LTO, 256-sample blocks, three boots x three
passes, run-to-run noise +/-0.01 %, output CRCs identical on every scene:

scene before after delta
one plain saw osc (8-saw scene, idle subtracted) 27.6k cycles 23.0k -16.6 %
8 plain saws 252,540 215,832 -14.5 %
8 saws through an LPF 289,533 261,998 -9.5 %
Juno patch, 6 voices 811,620 728,562 -10.2 %
sine, triangle, wavetable, FM, KS, PCM scenes within noise

The FM commit was not benched on target; its per-TU numbers are above.

Notes for reviewers

  • Two changes are folded into the kernel body and they are separable.
    Baking in the table size alone, with INTERP_CUBIC used verbatim,
    gives a 63-instruction hardware loop; rewriting the taps at their
    16-bit scale takes it to 60. If you would rather read the original
    macro inside the sized kernel, say so and I will drop the fold: it costs
    three instructions per sample and nothing else.
  • The sized kernels are compiled on every fixed-point target, not only
    Xtensa, so the host suite exercises them. On ARM (RP2040) the immediate
    shifts and mask still remove instructions but there is no hardware loop
    to gain and the code size is six 60-instruction bodies; not measured
    there. Gating the block on __XTENSA__ is a one-line change if you
    would rather keep RP2040 as it was.
  • The linear kernel (render_lut, triangle and wavetable paths) already
    compiles as a hardware loop; the same size specialisation only trims
    it from 32 to 28 instructions per sample. Not included here because
    those paths were not benched on target; easy follow-up if wanted.
  • The two commits are independent; dropping the second costs only the FM
    kernel's loop shape.

render_lut_cub is generic over the table size, so its two shift amounts
and the mask occupy registers. On Xtensa that pushes the loop over the
register window: gcc spills the counter, step, amp increment and both
shift amounts, and the loop is emitted as decrement-and-branch, 74
instructions per sample with five stack reloads. With the table size a
compile-time constant the shifts and the mask are immediates, three
registers free up, and the same arithmetic compiles to a 60-instruction
hardware loop.

One instantiation per saw table a fundamental under about 4 kHz picks
(2048 down to 64 entries), chosen once per block by log_2_table_size;
the three smallest tables fall back to the generic kernel. The
instantiations are noinline: inlined into the dispatcher's switch the
loop is lost again. Inside the loop the table values stay at their
16-bit scale, which is exact (MUL0_SS(L2S(x), f) == (x * (f >> 7)) >> 8),
so the L2S shift-up and the multiply's shift-down cancel.

Output is bit-identical to render_lut_cub (amy.test unchanged; host
sweeps of saw and pulse from 55 Hz to 3.5 kHz byte-identical). On an
ESP32-S3 at 48 kHz a plain saw osc drops from 27.6k to 23.0k cycles per
256-sample block; an 8-saw scene from 252.5k to 215.8k (-14.5 %), a
Juno patch -10 %.
render_lut_fm has the same problem as the generic cubic kernel: the
runtime table size keeps two shift amounts and a mask in registers, the
loop counter spills, and gcc emits decrement-and-branch, 65 instructions
per sample on Xtensa. The FM operators only ever read the 256-entry
sine table, so render_lut_fm now dispatches to render_lut_fm_256, the
same expressions with lut_bits = 8 substituted, and keeps the generic
body for any other table. The sized loop is 32 instructions as a
hardware loop with no stack reloads. Output is bit-identical
(amy.test unchanged against the reference set).
@dpwe

dpwe commented Sep 22, 2026

Copy link
Copy Markdown
Collaborator

I'm concerned about the object size growth, particularly for the RP2040. Is that an issue?

At -O3 gcc clones the render_lut_cub_sized switch into render_lpf_lut
and the saw and pulse wrappers. The switch runs once per block on a
runtime table size, so the copies buy nothing. Object growth over
1.2.171 for oscillators.o on cortex-m0plus at -O3 drops from +3,140 to
+1,848 bytes of .text; -O2 and -Os are unchanged.
@rt-rtos

rt-rtos commented Sep 22, 2026

Copy link
Copy Markdown
Contributor Author

Pushed a third commit that keeps render_lut_cub_sized out of line. At -O3
gcc was cloning its switch into render_lpf_lut and the four saw/pulse
wrappers, so the object carried it five times.

.text growth of src/oscillators.c alone over 1.2.171, cortex-m0plus,
arm-none-eabi-gcc 7.2.1:

before now
-O3 (pico-sdk Release) +3,140 +1,848
-O2 +1,848 +1,848
-Os +1,552 +1,552

plus 24 bytes of .rodata. Of the +1,848 at -O3, the six cubic kernels are
240..252 bytes each, the FM kernel 178 and the dispatcher 160. Against the
Pico's 2 MB flash it is under 0.1 %; against oscillators.o itself it is
+20 %.

That is the cost side. I have not measured render time on an RP2040 (don´t own one); the
closest I have is a static count of the per-sample loop in the same
objects, cortex-m0plus, same flags:

per-sample loop, instructions -O3 -Os
render_lut_cub_11 .. _9 81 81
render_lut_cub_8 .. _6 85 83
render_lut_cub (generic) 94 92

Same loop shape in all of them: four ldrsh taps, one branch for the
running max, no unrolling, no calls. The generic kernel spends its extra
instructions reloading its two shift amounts from the stack each sample,
three lsls #8 scale shifts the sized bodies fold into the arithmetic, and
a running max spilled to a stack slot.

So I do not think the size is an issue on the RP2040: the growth is flash,
not SRAM (AMY_IRAM_ATTR is a no-op off ESP), and the kernels are shorter
per sample there too. If the bytes ever need to come back on the M0 parts,
the __ARM_ARCH_6M__ test that gates AMY_HAS_MUL64 is the switch.

arm-none-eabi-gcc -c -mcpu=cortex-m0plus -mthumb -O3 \
    -DARDUINO_ARCH_RP2040 -DPICO_ON_DEVICE -Isrc src/oscillators.c -o osc.o
arm-none-eabi-size -A osc.o                  # per-section
arm-none-eabi-nm -S -t d --size-sort osc.o   # per-kernel
arm-none-eabi-objdump -d osc.o               # loop bodies

@dpwe
dpwe merged commit 380d103 into shorepine:main Sep 22, 2026
12 checks passed
@bwhitman

Copy link
Copy Markdown
Collaborator

⛓️ tulipcc integration PR opened

This merge was pinned into tulipcc for full-system CI: shorepine/tulipcc#1369

Test it there and merge that PR to move tulipcc onto this AMY.

@dpwe

dpwe commented Sep 25, 2026

Copy link
Copy Markdown
Collaborator

This was a nice speed up for Juno voices etc. thanks a lot!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants