Deadlock-safe mode - #164
Conversation
- `compute_cfg` register is added after existing registers - Controls `enable`, `op`, `mode`, `tensor_m` and `tensor_n` - Connected to the relevant fields of the ND request struct in the register top-level template - Add required additional frontend IDs to CI deploy job
- Transpose-enabled iDMA would previously always read tile-sized chunks, leading to "overreads" (reading past the source matrix's backing memory) - Also, the transposed output had its row alignment padded to the tile width (which is the bus data width) - The `idma_transpose_req_replay` module prevents over-reading (and over-writing, although those writes are zero-strobed) by redirecting redundant reads and writes to known safe addresses, specifically src/dst addresses of the first read/write bursts - Compact output matrix storage is achieved by changing the destination stride in the transpose midend to the actual row width of the transposed matrix - Compact output mode is controlled by a new field `compact` in `transpose_options_t`, which is also configurable by the register frontend
- Replace the legacy Verilator elaboration target with reusable build, run, clean, per-testbench, and full-suite targets - Generate stable Bender file lists and track RTL, headers, DPI sources, and elaboration parameters so binaries rebuild only when their inputs change. - Add configurable tracing and C/C++ compiler and linker flags - Register the directed transpose, ND midend, register frontend, and runtime midend test matrices across relevant widths and operating modes - Extend transpose tests to cover compact and padded layouts, partial tiles, back-to-back requests, backpressure, and multiple bus widths - Use clocking blocks for race-free ready/valid stimulus and sampling in the clocked directed testbenches - Make the transpose DPI model link cleanly when Verilator compiles it as C++
- To decide if an AR request can be accepted/fired, we want to know how many bytes the transfer will be and on which lane it starts - This will allow us to decide whether all read data FIFOs have space to absorb the bytes that will be pushed into them as a result of the read burst
- Currently the only effect it has is that maximum burst length gets clamped to `BufferDepth` in the legalizer - Added to register and descriptor frontends
- The gate tracks (future) usage of the buffer FIFO lanes and computes how many bytes will be pushed into each lane. - In deadlock-safe mode, an AR request is allowed to propagate only if its read response will fit into the read buffer FIFOs. - In unsafe mode, usage is tracked but all requests are allowed to propagate. - If a safe-mode 1D request follows an unsafe one, the gate waits until sufficient capacity is available in the buffer FIFO lanes.
2d30bdb to
2766d33
Compare
|
Reading On the "theoretical occupancy" question: it is tracked, not measured. Admission compares a modeled per lane counter, So safety rests on one thing the gate does not check directly: that the R data actually pushed into each byte lane FIFO matches |
Deadlock-Safe Read Mode
Summary
from overcommitting the backend dataflow buffers.
byte-lane FIFOs can hold while the write path is backpressured.
disabled.
Backend
protocol-specific read module.
Each accepted read-address request reserves exactly the entries required by
its byte count and starting lane; entries are released as bytes leave the
dataflow buffer.
sufficient capacity for its complete response.
between pipelined transfers.
combinational path from the dataflow output back to read-address valid.
Legalization And Configuration
deadlock_freeto the backend option structure and propagate it throughthe register, descriptor, synthesized, midend, legalizer, backend, and
transport-layer interfaces.
legalized read burst to the per-lane buffer capacity. This guarantees that
an individual request can be reserved without overflowing a lane FIFO.
burst sizing cannot then be guaranteed by the backend.
keeping it disabled by default for compatibility.
BufferDepthand AR bursts can only be submitted when the buffer is fully drained, which introduces bubbles.Integration And Verification
byte count, starting lane, and mode bit required for reservation decisions.
to encode and propagate the new configuration bit.