The render thread, the lock-free frame handoff, and the vsync-paced
frame clock (M2-GL-02; PRD §10.2, §11; AGENTS ARCH-002/ARCH-009,
CONC-002/CONC-005/CONC-006, PERF-008). Public header:
src/laige-render/include/laige/render/frame_pipeline.h;
implementation: src/laige-render/frame_pipeline.cpp. Unit suites:
ctest -R render_thread (tests/laige-render/render_thread_tests.cpp)
— FrameClock and RenderThreadHandoff run without any GL
environment; RenderThreadOffscreen needs a usable OpenGL 3.3
environment (always present on the P0 CI runners, the
GlContextSmoke convention). On success the offscreen runs print
one machine-greppable line each in the ctest output (the
docs/testing.md machine-line convention):
render-thread: kind=<lowrate|3000|slowed> frames=<n> rendered=<n> dropped=<n> pixel=ok status=ok
Rendering runs on its own thread — the frame pipeline (cull/batch →
submit) — with the frame descriptor handed off from the main/sim
thread. This step lands the handoff, the thread, the clock, the
backpressure, and the shutdown; the pipeline's actual stages — the
sprite batcher's cull/batch (M2-SPRITE-01) and the GPU submit
(M2-SPRITE-02) — plug in as the plain function callbacks below. Nothing
in this module depends on GL in the handoff itself (the GlContext
render-thread takeover is the onStart hook — release-then-bind per
docs/api/gl_context.md; the window swap/present lands with the
engine's windowed wiring, M2-SPRITE-02 onward).
One producer (the main/sim thread — the engine's frame loop) and one
consumer (the render thread). The handoff is a single slot holding
one FrameDescriptor plus six atomics — no mutex, no condition
variable, no queue:
FrameDescriptor— a 32-byte POD:frameIndex(the producer's 1-based frame counter),simTick(the last completed sim tick at publish time),renderTimeNs(the frame's presentation time on the engine's monotonic steady time base — the presentation/interpolationrender_time, M1-LOOP-02), andframeData(an opaque non-owning per-frame payload pointer — the sim state read arrives with the sprite stages, M2-SPRITE-02; the handoff itself never touches sim types, so no render→sim module edge is needed yet).StageFn— a plainnoexceptfunction-pointer callback (nostd::function, no virtual dispatch — PERF-006).batchStageruns beforesubmitStage, once per rendered frame, on the render thread;nullptrskips the stage.RenderThreadOptions— the stages, their context, and theonStart/onStophooks.
The full synchronization argument (the atomic-slot-word handoff —
memory orders, the ordering re-check, the no-ABA case, the 2^64 wrap
bound) is in the header preamble; the short form: the producer
publishes the descriptor's four 8-byte words (release) then
seq_.store(s + 1, release); the consumer acquire-loads seq_, loads
the four words (8 bytes is natively single-copy-atomic on every P0
platform — never torn), and re-checks seq_ — an unchanged value
means all four words are from the same publication, otherwise a newer
publication crossed the copy window and the copy is retried. The
ATOMIC words (not a plain-memory sequence lock) are what make the
handoff free of data races under the C++ memory model — a plain
seqlock's concurrent plain read/write of the slot is the exact access
pair the P0 CI TSan lane reports (CONC-007); a single 32-byte atomic
would add nothing but a libatomic dependency on the P0 g++/clang
lanes (verified in CI). The diagnostic atomics (consumedSeq_,
inFlight_, stop_, submitted_, rendered_) carry no correctness —
seq_ and the slot words alone do; they exist for the accounting,
waitIdle, and the ordered shutdown (CONC-002: partitioned ownership
and an immutable snapshot, no shared locks).
Backpressure (PERF-008, "never queue unboundedly"). The single slot
means the consumer can lag by at most one frame. When the producer
publishes while a frame is still pending (seq_ > consumedSeq_), the
older frame is dropped in place — replaced in the slot — and one
rate-limited Warn event is emitted (LOG-004: the facade's per-second
window suppresses repeats):
render_thread/frame_dropped fields: dropped_frame, new_frame
The drop runs no stage (it is the scope's backpressure, not a
silent skip — it is logged, FR-12.3). One documented event race: the
drop event fires at publish time from the producer's snapshot of
consumedSeq_, so it can name a frame the consumer finishes consuming
in the same instant — the event count can exceed framesDropped by
at most one per racing window; stats() is the authoritative count.
The exact accounting invariant. Every published frame is in exactly one bucket at every instant:
framesSubmitted = framesRendered + framesDropped + outstanding
where outstanding = the pending slot (≤ 1) + the in-flight frame (≤
1, a consumed frame whose stages are still running) — 2 at most. A
frame moves pending → in-flight → rendered exactly once; a dropped
frame is overwritten before it is ever consumed. rendered is a real
counter (incremented when a frame's stages complete), not a seq
difference: a drop overwrites the slot without the consumer ever
consuming the older frame, so seq/2 would count the frame number
of the last consumed frame, not the number of frames rendered.
RenderThreadStats::framesDropped is the source of the Profiler's
frame-drop field (M2-SPRITE-04, docs/api/profiler.md).
| Operation | Behavior | Complexity / allocation |
|---|---|---|
RenderThread(options) |
Spawns the consumer thread in the constructor body, after every state member is initialized (thread_ is declared last — see Threading and phase), runs options.onStart on it (the GlContext::makeCurrent takeover hook) and, at shutdown, options.onStop on it after the last frame and before the thread exits (the GlContext::release hand-back hook — the P0 EGL stack cannot rebind a context last held by a dead thread). One Info event, render_thread/thread_started. A thread-spawn failure terminates the process (exceptions disabled, NFR-8.10 — a documented platform boundary, CORE-008: the failure is never silent) |
one-time setup: one thread + one log line |
submitFrame(frame) |
The owner-thread (main/sim) publish — the hot path: a few atomic loads + one plain 32-byte copy + one release store, no allocation, no lock, no log on the healthy path. Single-slot backpressure: a pending frame (the consumer more than one frame behind) is dropped in place — one rate-limited render_thread/frame_dropped warn. Stopped → InvalidArgument (no log — the stopped-state precedent) |
O(1); no allocation; one release store |
waitIdle() |
The owner-thread barrier: blocks until every published frame is fully processed (no pending frame, no in-flight pipeline). Bounded by the single slot plus the stage callbacks' bound (API-005). No-op on a stopped object | O(1) yield-spin; one bounded wait per frame |
shutdown() |
Ordered idempotent shutdown (CONC-006): stop request + join + stopped mark; a second call is a no-op; safe on a stopped object. It does not flush a pending frame — the owner calls waitIdle() first when the last frame must render (the M1-HEAD-01 ordered-shutdown precedent). One Info event per actual stop, render_thread/thread_stopped |
O(1) + the join (bounded by one frame's pipeline work) |
stats() |
The since-construction counters (RenderThreadStats); the invariant above holds at every instant |
a few relaxed atomic reads; no allocation |
running() |
True while the thread is running (false once stopped/moved-from) | O(1) |
FrameClock::create(options) |
Validates the frame rate (outside [kMinFrameRateHz, kMaxFrameRateHz] → InvalidArgument + one rate-limited render_thread/frame_rate_invalid warn) and records the reference (one steady_clock read). Returns Result<FrameClock, ErrorCode> |
O(1), one division, no allocation |
FrameClock::deadlineNs(n) |
Frame n's presentation deadline (1-based; exact integer grid referenceNs + n × periodNs; clamped to INT64_MAX beyond the 2^63 ns frame-time bound — never UB, CPP-004) |
O(1) integer math |
FrameClock::waitFrame(n) |
Paces the owner thread: blocks until frame n's deadline (one bounded sleep_until, no spin), then returns the frame's actual presentation time — on time the deadline, late the real clock reading (ARCH-009 wall-clock fact; the M1-LOOP-02 alpha contract clamps a late frame to 1.0). Stopped clock → 0 |
O(1) + one bounded sleep |
FrameClock accessors |
valid(), frameRateHz(), periodNs(), referenceNs(), nowNs() — a stopped (default/moved-from) clock returns 0 for all |
O(1) |
Move semantics. RenderThread is move-only; the move stops the
source (joins a live thread) and the moved-to object is stopped
(the GlContext moved-out precedent) — construct it in place.
FrameClock is move-only; a moved-from clock is stopped (every
operation returns 0 — the clock has no failure to log).
Ownership and lifetime. submitFrame/waitIdle/shutdown are
owner-thread calls (the one producer); the render thread only runs the
stage callbacks. FrameDescriptor::frameData is a non-owning
per-frame payload pointer (e.g. the engine's PresentationSnapshot
view): the producer owns it and it must outlive the frame's render —
waitIdle() before release (one producer, one consumer, one pending
frame: the handoff carries no reference count). The stage callbacks and
onStart/onStop run on the render thread: the stages once per
rendered frame, onStart once before the first frame, onStop once
after the last frame and before the thread exits (the GL context's
release hand-back — the P0 EGL stack cannot rebind a context last
held by a dead thread); they must be bounded and non-blocking
(API-005: a stage or hook that blocks, waitIdle and the shutdown
join block too).
frameIndex is the producer's 1-based frame counter; the handoff does
not validate it (seq_ is the handoff's own ordering).
Threading and phase. The render thread is spawned by the
constructor and joined by shutdown() (or the destructor) — never
detached (CONC-005). The engine's shutdown calls waitIdle() then
shutdown() when the last frame must render, or shutdown() alone
when it does not (CONC-006: ordered, testable, idempotent). No engine
locks anywhere in the module (CONC-002); the only thread-join is the
shutdown join.
The construction order is load-bearing. C++ initializes members in
declaration order; thread_ is declared last and is started in
the constructor body — never in a member initializer. The
thread-start synchronization edge ([intro.multithread]) publishes only
what happened before the start, so spawning from a member
initializer would let the consumer read the state members before
their in-class initializers ran (a data race caught by
ThreadSanitizer in CI). With the thread started last, the consumer
never observes a partially-initialized object.
Failure behavior (NFR-008 / CORE-008). The handoff has no runtime
failure to report: submitFrame on a stopped object returns
InvalidArgument (a precondition query, no log); the only logged
events are the two lifecycle Info events, the rate-limited Warn
backpressure drop, and the rate-limited Warn invalid clock rate.
FrameClock::create returns InvalidArgument with one rate-limited
Warn (never silent). A thread-spawn failure terminates the process
(see the API table) — with exceptions disabled there is no recoverable
error path (documented platform boundary).
The clock is a frame deadline grid on the engine's monotonic
steady time base — the same epoch the GameLoop and the M1 headless
engine use (the M1-LOOP-01 clock base):
deadlineNs(N) = referenceNs + N × periodNs (N 1-based; exact integer)
periodNs = 10⁹ / frameRateHz (one division at create)
- Windowed: the caller sets
frameRateHzto the display's refresh rate (GlContext::refreshRateHz— 0 when unavailable, in which case the target rate stands in) so the deadline grid coincides with the display refresh period; the window's swap (vsync on, GLFW's default) lands each frame on a refresh boundary. This replaces the M1 headless monotonic clock read (M1-LOOP-01) as the frame pacer: the engine's frame loop callswaitFrame(N)and hands the returned frame time to the presentation/interpolation path (PresentationSnapshot::onRenderFrame, M1-LOOP-02). - Headless:
frameRateHzis the target pace (the CI path — no display, no vsync; the same deadline grid, the samerender_timecontract). - The
render_timeis a wall-clock fact (ARCH-009): non-deterministic by design, never part of replay state or the simulation state hash.
The documented frame-rate range is [kMinFrameRateHz, kMaxFrameRateHz]
= 1–1000 Hz (kDefaultFrameRateHz = 60).
Hot path (submitFrame). A few atomic loads + four 8-byte atomic
release stores + one release store — no allocation, no lock, no
virtual dispatch, no std::function (PERF-006), no logging on the
healthy path (the drop path is cold: one rate-limited Warn). The consumer's
between-frame wait is a yield spin (no busy-burn: the OS reschedules
during the ~16 ms vsync gap). FrameClock: deadlineNs is O(1)
integer math; waitFrame is one bounded sleep_until (no spin — the
producer's wait is cadence, not work, PERF-002).
No hot-path logging, no hot-path allocation. stats() is a few
relaxed atomic reads (cold); the stage callbacks are the frame's real
work (M2-SPRITE-02 budget, not this step's).
Misuse warnings.
- Never publish from a thread other than the owner (a second producer
races the handoff's single-slot protocol —
seq_/slot_are producer-owned) — the sim thread is the producer. - Never release per-frame data the handoff still carries:
waitIdle()shutdown()first.
- A stopped (destroyed/moved-from)
RenderThreadrejectssubmitFramewithInvalidArgument(no log — the stopped-state precedent). - A stage that blocks blocks
waitIdleand the shutdown join (API-005) — the stage callbacks' bound is the caller's responsibility. waitFrameis the owner-thread call: calling it from the render thread paces nothing (the render thread runs the stages, it does not publish).
// Setup (once, e.g. in the render module's start — the M2-GL-02
// windowed path; the headless CI path uses a target rate):
laige::render::GlContext gl = /* createWindowed / createHeadless */;
laige::render::FrameClockOptions co;
co.frameRateHz = gl.refreshRateHz(); // 0 → the target rate stands in
laige::render::FrameClock clock =
std::move(laige::render::FrameClock::create(co)).takeValue();
// Hand off the GL context (release-then-bind — docs/api/gl_context.md):
// the P0 EGL stack rejects a takeover while the context is still
// current on another live thread, so the old owner releases first.
gl.release().ok(); // main/sim thread: no longer current
laige::render::RenderThreadOptions opts;
opts.batchStage = &cullBatch; // M2-SPRITE-01 (render thread)
opts.submitStage = &gpuSubmit; // M2-SPRITE-02 (render thread)
opts.stageContext = &batcher;
opts.onStart = &takeover; // render thread: gl.makeCurrent()
opts.onStartContext = ≷
opts.onStop = &handBack; // render thread: gl.release()
opts.onStopContext = ≷ // after the last frame, before the
// thread exits (a dead thread's
// context cannot be rebound on the
// P0 EGL stack)
laige::render::RenderThread thread(opts);
// Per frame (the main/sim thread — the one producer):
const std::int64_t renderTime = clock.waitFrame(frameIndex);
laige::render::FrameDescriptor d;
d.frameIndex = frameIndex;
d.simTick = loop.currentTick();
d.renderTimeNs = renderTime; // → PresentationSnapshot::onRenderFrame
d.frameData = &snapshot; // non-owning; waitIdle before release
thread.submitFrame(d); // O(1); a late producer drops the
// older frame (logged), never queues
// Engine shutdown (ordered, CONC-006):
thread.waitIdle(); // the last frame renders
thread.shutdown(); // joins the render thread