oscillators: size-specialised kernels for the saw/pulse tables and the FM sine table - #1171
Conversation
render_lut_cub is generic over the table size, so its two shift amounts and the mask occupy registers. On Xtensa that pushes the loop over the register window: gcc spills the counter, step, amp increment and both shift amounts, and the loop is emitted as decrement-and-branch, 74 instructions per sample with five stack reloads. With the table size a compile-time constant the shifts and the mask are immediates, three registers free up, and the same arithmetic compiles to a 60-instruction hardware loop. One instantiation per saw table a fundamental under about 4 kHz picks (2048 down to 64 entries), chosen once per block by log_2_table_size; the three smallest tables fall back to the generic kernel. The instantiations are noinline: inlined into the dispatcher's switch the loop is lost again. Inside the loop the table values stay at their 16-bit scale, which is exact (MUL0_SS(L2S(x), f) == (x * (f >> 7)) >> 8), so the L2S shift-up and the multiply's shift-down cancel. Output is bit-identical to render_lut_cub (amy.test unchanged; host sweeps of saw and pulse from 55 Hz to 3.5 kHz byte-identical). On an ESP32-S3 at 48 kHz a plain saw osc drops from 27.6k to 23.0k cycles per 256-sample block; an 8-saw scene from 252.5k to 215.8k (-14.5 %), a Juno patch -10 %.
render_lut_fm has the same problem as the generic cubic kernel: the runtime table size keeps two shift amounts and a mask in registers, the loop counter spills, and gcc emits decrement-and-branch, 65 instructions per sample on Xtensa. The FM operators only ever read the 256-entry sine table, so render_lut_fm now dispatches to render_lut_fm_256, the same expressions with lut_bits = 8 substituted, and keeps the generic body for any other table. The sized loop is 32 instructions as a hardware loop with no stack reloads. Output is bit-identical (amy.test unchanged against the reference set).
|
I'm concerned about the object size growth, particularly for the RP2040. Is that an issue? |
At -O3 gcc clones the render_lut_cub_sized switch into render_lpf_lut and the saw and pulse wrappers. The switch runs once per block on a runtime table size, so the copies buy nothing. Object growth over 1.2.171 for oscillators.o on cortex-m0plus at -O3 drops from +3,140 to +1,848 bytes of .text; -O2 and -Os are unchanged.
|
Pushed a third commit that keeps
plus 24 bytes of That is the cost side. I have not measured render time on an RP2040 (don´t own one); the
Same loop shape in all of them: four So I do not think the size is an issue on the RP2040: the growth is flash, |
⛓️ tulipcc integration PR openedThis merge was pinned into tulipcc for full-system CI: shorepine/tulipcc#1369 Test it there and merge that PR to move tulipcc onto this AMY. |
|
This was a nice speed up for Juno voices etc. thanks a lot! |
Summary
render_lut_cubis generic over the table size, so its two shift amountsand the mask occupy registers. On Xtensa that pushes the loop over the
register window: gcc spills the counter, the step, the amp increment and
both shift amounts, and emits the loop as decrement-and-branch, 74
instructions per sample with five stack reloads. With the table size a
compile-time constant the shifts and the mask are immediates and the same
arithmetic compiles to a 60-instruction hardware loop. Output is
bit-identical.
The first commit adds one instantiation per saw table a fundamental under
about 4 kHz picks (2048 down to 64 entries), selected once per block by
log_2_table_sizeinrender_lpf_lut; the three smallest tables fallback to the generic kernel. The second commit does the same for
render_lut_fm, which has the same spilled-counter loop: the FMoperators only ever read the 256-entry sine table, so it dispatches to a
lut_bits = 8instantiation in the style ofrender_lut_256and keepsthe generic body for any other table. That loop goes from 65
instructions to 32, a hardware loop with no stack reloads.
Context
noinline. Inlined into the dispatcher's switchthe loop is lost again, since the six bodies then share one register
allocation; checked in the per-TU assembly and in an LTO-linked ELF,
where
noinlineholds and all six carry aloopinstruction.log_2_table_size, so the bakedBITSalways matches the table thatchoose_from_lutsetpicked; atable outside 64..2048 entries takes the generic kernel. Only
render_lpf_lutcalls it; the sine and triangle paths are untouched.exact, not an approximation:
MUL0_SS(L2S(x), f)is(x * (f >> 7)) >> 8, so theL2Sshift-up and the multiply'sshift-down cancel. It is where the last three instructions went, and it
is why the kernel reads the
LUTSAMPLEtaps directly rather thanconverting each one. Like every
>>on aSAMPLEinamy_fixedpoint.h, it assumes an arithmetic right shift on negativevalues, which gcc and clang give on every target AMY builds for.
RENDER_LUT_GUTS(MOD_PART_MOD, NOTHING, INTERP_LINEAR)body with the table size substituted, not a rewrite, soits exactness needs no derivation. The feedback variants
(
render_lut_fb,render_lut_fm_fb) are left alone: the sametreatment would apply, but they were not measured.
sized kernels and their dispatcher sit under
#ifdef AMY_USE_FIXEDPOINT;a float build (
amy.hdocuments removing the define) keeps callingrender_lut_cubthrough the same name.above 4 kHz, where the cubic reads three harmonics or fewer and the
render time is elsewhere. Six instantiations cost about 1.2 KB of
instruction memory on the S3.
Verification
make testagainst the reference set: every test's error valueidentical to a pristine 1.2.171 run on the same host. Host sweeps of a
single saw and pulse osc, 55 Hz to 3520 Hz, 48 kHz fixed point,
byte-identical WAVs against the generic kernel.
Per-TU Xtensa assembly at -O2 (esp-15.2.0 gcc), the loop of each kernel:
render_lut_cub(generic, still used for 32..8-entry tables)render_lut_cub_11..render_lut_cub_6render_lut_fm(generic, kept for any non-256 table)render_lut_fm_256On an ESP32-S3 at 48 kHz, LTO, 256-sample blocks, three boots x three
passes, run-to-run noise +/-0.01 %, output CRCs identical on every scene:
The FM commit was not benched on target; its per-TU numbers are above.
Notes for reviewers
Baking in the table size alone, with
INTERP_CUBICused verbatim,gives a 63-instruction hardware loop; rewriting the taps at their
16-bit scale takes it to 60. If you would rather read the original
macro inside the sized kernel, say so and I will drop the fold: it costs
three instructions per sample and nothing else.
Xtensa, so the host suite exercises them. On ARM (RP2040) the immediate
shifts and mask still remove instructions but there is no hardware loop
to gain and the code size is six 60-instruction bodies; not measured
there. Gating the block on
__XTENSA__is a one-line change if youwould rather keep RP2040 as it was.
render_lut, triangle and wavetable paths) alreadycompiles as a hardware loop; the same size specialisation only trims
it from 32 to 28 instructions per sample. Not included here because
those paths were not benched on target; easy follow-up if wanted.
kernel's loop shape.