Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
42 changes: 21 additions & 21 deletions docs/adr/0001-kernel-adapter-boundary.md

Large diffs are not rendered by default.

118 changes: 59 additions & 59 deletions docs/design/native-feature-parity.md

Large diffs are not rendered by default.

8 changes: 4 additions & 4 deletions docs/native-refactor-progress.md
Original file line number Diff line number Diff line change
Expand Up @@ -173,7 +173,7 @@ them for the record only. "Linux aarch64" is the maintainers' Cortex-A725 host.

| Item | What it is | Measured size | Deferred by | Prototype and evidence | Reopens when |
| --- | --- | --- | --- | --- | --- |
| D2-B, the quiet-bar protocol and one receipt read per bar | PERF-D2's lane: a per-bar bar-open request, a bulk fold of quiet driver points, the adapter's quiet-open predicate with lazy open bookkeeping, and a gate before the receipt poll. D2-A landed the hook declarations and the O(1) `has_command_after` (`src/native_execution_consumer.hpp:286`). PERF-L4 gated the read inside `observe_terminal_receipts` (`src/source/pine_adapter.cpp:6307`), which still runs about three times a bar | PERF-D2, Linux aarch64 (Cortex-A725), callgrind and GBench: about ×0.88 on the 14 probes (an estimate from a measured −13 to −14 % floor bound for the bar-open request alone); after PB the receipt poll still costs about 250 instructions a bar on the no-order floor. On this tree the quiet-bar probe asks the receipt gate 40,996 times and each once-per-bar-open gate 13,248 times: 3.09 reads a bar, 82 % of them stopped at the gate (macOS arm64, AppleClang 17, Release) | "R5 PERF-D2 DONE 2026-09-24 19:56" (D2-B after V19-D); "R5 SUPERVISOR RULING 2026-09-25 00:15 Taipei": "D2-B (quiet-bar protocol) waits for PERF-D3's verdict"; TRIAGE1 §D | An unpublished prototype branch and its patches. PERF-D3's S1 (next row) prototyped the request and the lazy opening | A post-release performance epoch. Done when the receipt gate is asked as often as the once-per-bar gates |
| D2-B, the quiet-bar protocol and one receipt read per bar | PERF-D2's lane: a per-bar bar-open request, a bulk fold of quiet driver points, the adapter's quiet-open predicate with lazy open bookkeeping, and a gate before the receipt poll. D2-A landed the hook declarations and the O(1) `has_command_after` (`src/native_execution_consumer.hpp:295`). PERF-L4 gated the read inside `observe_terminal_receipts` (`src/source/pine_adapter.cpp:6307`), which still runs about three times a bar | PERF-D2, Linux aarch64 (Cortex-A725), callgrind and GBench: about ×0.88 on the 14 probes (an estimate from a measured −13 to −14 % floor bound for the bar-open request alone); after PB the receipt poll still costs about 250 instructions a bar on the no-order floor. On this tree the quiet-bar probe asks the receipt gate 40,996 times and each once-per-bar-open gate 13,248 times: 3.09 reads a bar, 82 % of them stopped at the gate (macOS arm64, AppleClang 17, Release) | "R5 PERF-D2 DONE 2026-09-24 19:56" (D2-B after V19-D); "R5 SUPERVISOR RULING 2026-09-25 00:15 Taipei": "D2-B (quiet-bar protocol) waits for PERF-D3's verdict"; TRIAGE1 §D | An unpublished prototype branch and its patches. PERF-D3's S1 (next row) prototyped the request and the lazy opening | A post-release performance epoch. Done when the receipt gate is asked as often as the once-per-bar gates |
| S1a, kernel quiet-bar driving (PERF-D3) | A bar with nothing live, a flat book and a declined opening is driven in a quiet form: the kernel writes the bookkeeping its four price points would leave (ordinals, the last two driver marks, the floor, the point epochs) without walking them or calling the bar-open callback. Additive host cadence declarations; C hosts declare from their callback table; one C symbol under the PF_API checklist | PERF-D3, Linux aarch64 (Cortex-A725), GBench time and `perf stat` instructions, on the L3 tree `b33ea659`, S1a with S1b: no-order floor 0.764 / 0.660 in time (0.718 / 0.656 in instructions), 14 probes 0.958 / 0.921 (0.892 / 0.856), 100 slots 0.972 / 0.956 per slot. A bare kernel host that opts in: idle 0.545× (35 % fewer instructions a bar), market 0.81×, bracket 0.92×. Byte-identical: corpus 312/312, the 132-configuration battery with its hash columns | "FINDING PERF-D3 (filed for a post-release epoch per the owner's cap)" (ledger, 2026-09-24 16:58 UTC); "OWNER DECISION 2026-09-25 00:20 Taipei"; TRIAGE1 §D: "PERF-D3 S1/S2 (post-release epoch)" | An unpublished prototype branch (five commits on L3's tree) and its patches | The post-release epoch, with the owner calls PERF-D3 lists (the API under the existing ruling, the restated gate count) |
| S1b, the Pine lazy opening | The Pine host records what its skipped opening and its uncalled input callback would have written, at the bar's first host callback. Member functions only, no layout change | Measured with S1a. A quiet bar then asks 15 gates, not 30, so the quiet-bar cost check `CHECK(asked >= 20 * o.boundary.size())` (`tests/test_adapter_quiet_bar.cpp:1054`) must be restated | As S1a | As S1a. The in-script hash witnesses caught the prototype's one miss (`awaiting_legacy_script_open_ms_`) | As S1a |
| S1c, the held-position wake band | S1 for bars that hold a position with nothing live, 41 % of the median public slot's bars. The host declares a price band, and the kernel calls the opening only when the bar's range leaves it. The adapter derives the band from its margin model and declines when it cannot bound it | PERF-D3, a measurement-only upper bound without the band, which is not sound in general: 100 slots 0.776 / 0.716 per slot, 14 probes 0.830 / 0.761, cumulative against L3 with S1, S2 and D2-A | As S1a; the owner call "Fund the price band" | An unpublished measurement-only script | As S1a. Size L, parity risk high: TradingView's margin money rules |
Expand All @@ -184,7 +184,7 @@ them for the record only. "Linux aarch64" is the maintainers' Cortex-A725 host.
| A checked-values handoff for executions (L3) | `execution_values` runs twice, in `check_execution` (`src/native_order.cpp:4791`) and again in `apply_execution` (`src/native_order.cpp:4805`); handing the checked values across needs a public type, an API addition with no layout change | L3, as above: 2.21 % of bracket for the two together, about 1.1 % recoverable | L3 ruling 2: "a public checked-values handoff for executions (~1.1% bracket) -> DEFERRED" | As above | An owner decision on the API addition |
| A chunked `trades_` store (L3) | Closed trades in chunks rather than one vector; `trades_` (`include/pineforge/engine.hpp:546`) is protected `BacktestEngine` state, so script ABI | L3, as above: on bracket, vector growth 0.64 %, `Trade` moves 0.74 % and `record_close_trade` 0.43 % (at most about 1.4 %); at most 0.35 % on the 14 probes | L3 ruling 4: "chunked trades_ store (ABI) -> DEFERRED" | As above | An `engine_script_run` epoch that changes the engine's protected layout |
| A zoned chart-day memo (D2-D finding 4) | On a chart with a timezone, `libc_chart_day_key` (`src/source/pine_adapter.cpp:17130`) takes the process timezone lock through `ScopedTimezone` and calls `localtime_r` on every bar, because `on_bar_open` (`src/source/pine_adapter.cpp:22011`) reads `chart_day_key` (`src/source/pine_adapter.cpp:22248`) unconditionally. The non-UTC monthly Sharpe/Sortino walk, `month_key_local` (`src/engine_metrics.cpp:33`), pays one `localtime_r` per equity point. UTC charts use D2-D's memo, `civil_chart_day_key` (`src/source/pine_adapter.cpp:17114`) | Not measured on a zoned chart. The UTC path cost 134 instructions a bar before D2-D's memo and 68 after (D2-D, Linux aarch64, callgrind on the floor). PERF-K1: `ScopedTimezone` copies the zone name, which allocates for names of 16 or more characters under libstdc++ | TRIAGE1 §D: "D2-D finding 4 (the timezone `localtime_r` memo)", under the owner's cap. D2-D's report: an exact zoned memo "would need the zone's transition table, which belongs in `timezone.cpp`" | None; the UTC memo's witness is `tests/test_chart_day_memo.cpp` | A lane that owns `src/timezone.cpp` and builds an exact zone-transition memo (PERF-K1's Sitka case shows why a shortcut is not exact), byte-identical, with the zoned keys of `test_chart_day_memo` still equal to the computed ones |
| Thread-local reads that remain per call | D2-C moved the three scopes' writes onto the pump's block, `pump_ambient` (`src/native_execution_consumer.hpp:461`). A reader still reaches the block through `tl_runtime_ambient` (`src/ta_extremes_volume.cpp:29`): `bar_context` (`src/ta_extremes_volume.cpp:82`), read by every `ExtremeRing::update` (`src/ta_extremes_volume.cpp:117`), which backs `ta.highest`, `ta.lowest`, `ta.highestbars` and `ta.lowestbars` and, through an embedded `Highest` and `Lowest`, `ta.stoch`, `ta.wpr` and `ta.range`; and `matching_day_partition` (`src/timeframe.cpp:628`), read by the default-anchored `ta.vwap` through `session_day_index`, by `crosses_boundary` on a DAY period (`timeframe.change("D")`) and by the session-period helpers. `decompose_ms_local` (`src/session_time.cpp:250`) reads two thread-locals per call on a zoned clock, and the prepared order storage's `free_blocks` (`src/native_order.cpp:605`) one per prepared mutation. `ema_na_warmup_flag` is read once per EMA, at its first compute | One thread-local access per reader call. AUDIT4, macOS arm64, a dlopen'd tutorial-MACD probe, counted by `_tlv_get_addr` breakpoints and by descriptor patching: 1.000 a bar for one `ta::Highest` and for one default `ta::VWAP`, 1.999 for one `pine_hour` on a New York clock, 0.000 on the no-order floor and with orders. On Linux a TLS-descriptor call costs about 20 instructions: D2-C's 13 → 4 calls a bar was 260 → 80 instructions (Linux aarch64, GCC 13, callgrind). INT23: 2.00 a bar on the zoned probe 064, 0.08 on probe 008 | None for the TA and day-partition readers: "RULING: INT23 ACCEPTED" (ledger, 2026-09-24 23:30 UTC) reads "dlopen TLS 13->0/bar", and INT23's report names only the timezone and order-core residuals. This row is their record | Unpublished (AUDIT4's scratch) | A byte-identical lane that hands these readers the pump's block without a thread-local read. `ta::bar_context()` is script ABI, so its public declaration stays as it is |
| Thread-local reads that remain per call | D2-C moved the three scopes' writes onto the pump's block, `pump_ambient` (`src/native_execution_consumer.hpp:470`). A reader still reaches the block through `tl_runtime_ambient` (`src/ta_extremes_volume.cpp:29`): `bar_context` (`src/ta_extremes_volume.cpp:82`), read by every `ExtremeRing::update` (`src/ta_extremes_volume.cpp:117`), which backs `ta.highest`, `ta.lowest`, `ta.highestbars` and `ta.lowestbars` and, through an embedded `Highest` and `Lowest`, `ta.stoch`, `ta.wpr` and `ta.range`; and `matching_day_partition` (`src/timeframe.cpp:628`), read by the default-anchored `ta.vwap` through `session_day_index`, by `crosses_boundary` on a DAY period (`timeframe.change("D")`) and by the session-period helpers. `decompose_ms_local` (`src/session_time.cpp:250`) reads two thread-locals per call on a zoned clock, and the prepared order storage's `free_blocks` (`src/native_order.cpp:605`) one per prepared mutation. `ema_na_warmup_flag` is read once per EMA, at its first compute | One thread-local access per reader call. AUDIT4, macOS arm64, a dlopen'd tutorial-MACD probe, counted by `_tlv_get_addr` breakpoints and by descriptor patching: 1.000 a bar for one `ta::Highest` and for one default `ta::VWAP`, 1.999 for one `pine_hour` on a New York clock, 0.000 on the no-order floor and with orders. On Linux a TLS-descriptor call costs about 20 instructions: D2-C's 13 → 4 calls a bar was 260 → 80 instructions (Linux aarch64, GCC 13, callgrind). INT23: 2.00 a bar on the zoned probe 064, 0.08 on probe 008 | None for the TA and day-partition readers: "RULING: INT23 ACCEPTED" (ledger, 2026-09-24 23:30 UTC) reads "dlopen TLS 13->0/bar", and INT23's report names only the timezone and order-core residuals. This row is their record | Unpublished (AUDIT4's scratch) | A byte-identical lane that hands these readers the pump's block without a thread-local read. `ta::bar_context()` is script ABI, so its public declaration stays as it is |
| "audit-off mode", "LTO", "data-oriented layout" | Named in "OWNER DECISION 2026-09-25 00:20 Taipei" among the post-release structural items: "S1 bulk quiet bars, S2 per-strategy specialization, audit-off mode, LTO+PGO, data-oriented layout". No engine or campaign document defines, designs or measures the three | None | The owner decision | None | Struck from this inventory until a design note defines one of them |

### Recorded witness drift
Expand All @@ -194,9 +194,9 @@ them for the record only. "Linux aarch64" is the maintainers' Cortex-A725 host.
K4, the batch log pre-size. Since V19-B's step 6 (`08e71bae`) a run under the
default `NativeEventRetention::Window` keeps no driver point and retires its
journal at every script-bar boundary, so K4 has nothing to size.
`reserve_driver_log` (`src/native_execution_consumer.cpp:2870-2875`) returns
`reserve_driver_log` (`src/native_execution_consumer.cpp:2884-2889`) returns
unless the retention is Full, and `presize_logs`
(`src/native_execution_consumer.cpp:9655-9666`) sizes no journal under Window.
(`src/native_execution_consumer.cpp:9669-9680`) sizes no journal under Window.
`NativeRunSpec::event_retention` (`include/pineforge/native_run_spec.hpp:714`)
defaults to Window, and the Pine adapter declares
`NativeEventRetention::Window` (`src/source/pine_adapter.cpp:2931`). K4 still
Expand Down
Loading
Loading