Hi! We're using Jet (commit c56dfc0) as the rasteriser core of "Vertice", the 3D engine of NucleoOS, an OS for the ESP32-P4 (dual RISC-V @360 MHz, PSRAM). Credit and the MIT notice are kept in the repo. A few findings that might be useful upstream:
- RV32 has no int64->float instruction. Evaluating row-start attributes (z, u, v, 1/z, brightness) from the int64 edge weights costs ~18 soft conversions per scanline (~3800 cycles/row on the P4, more than 2/3 of raster time). Anchoring screen-space plane equations once per triangle (
a0 + ax*dx + ay*dy in float) removes them.
- PSRAM writes are latency-bound on the P4 (~33 MB/s per core for cached CPU writes). Rendering into small row tiles in internal SRAM (colour + depth), then copying each finished tile out with the 2D DMA, works much better than rendering straight into a PSRAM framebuffer. Both cores pull tiles from an atomic counter; the render queue is binned per tile once per frame.
- Soft particles need a fill budget: a puff right in front of the camera became a 97x97 sprite with an integer divide per pixel.
Happy to share details or patches if any of this is interesting. Thanks for Jet!
Hi! We're using Jet (commit c56dfc0) as the rasteriser core of "Vertice", the 3D engine of NucleoOS, an OS for the ESP32-P4 (dual RISC-V @360 MHz, PSRAM). Credit and the MIT notice are kept in the repo. A few findings that might be useful upstream:
a0 + ax*dx + ay*dyin float) removes them.Happy to share details or patches if any of this is interesting. Thanks for Jet!