Skip to content

refactor(bbapi): migrate bb + bb.js to ipc-codegen/ipc-runtime; delete legacy bb::ipc - #25362

Open
charlielye wants to merge 17 commits into
nextfrom
cl/ipc-bb-bbjs-migrate
Open

charlielye wants to merge 17 commits into
nextfrom
cl/ipc-bb-bbjs-migrate

Conversation

@charlielye

@charlielye charlielye commented Aug 29, 2026 •

Copy link
Copy Markdown
Contributor

Combined redo of #23612 + #23614 off next. Those PRs predate envelope ids (the 8-byte request-id frame prefix), which broke wire compatibility between old and new stacks — bb, bb.js, and bb-rs must move together, so this is one PR. The old PRs were used as the recipe; they'll be closed pointing here.

bb (C++)

  • Schema-first bbapi: bbapi/bb_schema.json (65 commands) and bb_curve_constants.json checked in, in the friendly ipc-codegen dialect (same as avm/cdb/wsdb). CMake runs ipc-codegen at build time (client+server, bb::bbapi namespace); the schema is also embedded so bb msgpack schema prints it.
  • Handlers: bbapi_handlers.* (+ bbapi_chonk_handlers.cpp, bbapi_ultra_honk_handlers.cpp, bbapi_wire_convert.hpp) implement the generated dispatch (Responder-style, exactly-once ok/error) as thin adapters over the existing domain Cmd{...}.execute(ctx) commands.
  • Serve: bb msgpack run serves every transport through ipc-runtime — -/empty → stdio pipe server, .sock → UDS, .shm → shared memory — one generated handler, envelope framing everywhere. It uses ipc-runtime's serial zero-copy run() loop rather than run_reactor(): every bb handler responds inline, so the reactor's per-request copy, closure, lock and wake bought nothing. A plain-file input keeps the offline bare-frame replay for wasm/debug flows.
  • FFI/wasm: typed-cbind bbapi() replaced by the ipc-codegen FFI contract ipc_ffi_entry(input, len, out, out_len) — same msgpack payload as the transports, no envelope, output is free()-compatible.
  • Deletions (consumer-free after the above): legacy barretenberg/ipc/ + ipc_bench; the whole nodejs_module/ (it only hosted the msgpack_client NAPI wrappers — bb.js now uses ipc-runtime's addon) along with its CMake presets targets and the bb-cpp-yarn Makefile step; bbapi_execute.* (Command/CommandResponse named unions + dispatcher); func_traits/schema reflection + CBIND macros + msgpack_schema test.

ipc-runtime

  • New pipe transport (C++): PipeServer/PipeClient over an fd pair or stdio, identical envelope framing to UDS, peer EOF → shutdown; serve_helper maps - to stdio, serving on a private duplicate of stdout and pointing fd 1 at stderr so stray stdout writes can't corrupt frames (next's bb redirected std::cout for the same reason). Exposed through the C ABI (ipc_client_create_pipe, dup'ing the caller's fds) and the Rust crate (IpcClient::from_fds). Used by bb msgpack run on stdio and by bb-rs. The Rust and Zig builds now discover sources instead of listing them.
  • No stream frame cap: UDS and pipe frames are limited only by the u32 length field (the old 256 MiB cap rejected bb proving requests). Receivers grow their buffer as bytes arrive (C++ read_payload, TS FrameReader), so a corrupt prefix can't force a large allocation, and the TS client/server no longer re-concatenate per chunk (receiving a large frame was quadratic). SHM stays bounded by half its ring capacity.
  • SHM conditional wakes made safe: the always-wake that SHM ships with was masking a startup race, not a memory-ordering bug — a client could attach to a segment after ftruncate but before create() initialised it, and that init then wiped the client's consumer_blocked flag, so a later publish skipped the wake and the client slept forever. create() now publishes capacity (rings) / num_slots (doorbell) last with release ordering, connect() refuses a segment until it is set, and wakes are conditional again (~6 futex syscalls per request saved). Each side of the wake handshake stores then loads the other's value, so both carry a seq_cst fence between the two; release/acquire alone would let both miss each other. Regression test ShmTest.ConnectRefusesSizedButUninitialisedSegments. The MPSC consumer spin now reads the clock every 256 iterations rather than every iteration.
  • TS:
    • SpawnedProcessBackend: unref/unrefStdio options (child + idle-socket unref, re-ref while calls are in flight) so bb.js callers that never destroy() can't hang node; stale shm segments/sockets from a SIGKILLed previous occupant are removed before spawn; a transport error against a still-live server kills it so respawn can replace it; the connect is aborted when the child dies first.
    • New SpawnedProcessBackendSync for the SHM sync client.
    • Call fast path: a live process is called directly, with no async frame or await ensureUp() per call. Errors are attributed to the dead process (exit code, log path) by an IpcErrorMapper hook the clients apply on their rejection paths only, instead of a per-call .catch.
    • SHM and UDS clients pair responses through a shared PendingQueue: small-integer ids (V8 Smis) and an issue-order queue with a head-of-queue fast path and a search fallback for out-of-order servers. This replaces Map<bigint>, which allocated and hashed heap BigInts on every call. The NAPI async client passes ids as numbers; the wire id stays u64.
    • A request too large for the length field rejects up front with a non-transport IpcError, so it doesn't make SpawnedProcessBackend retire a healthy server.

ipc-codegen

  • --strip-type-prefix: strips the service prefix from generated type and converter names too (BbCircuitProve → CircuitProve); wire tags always keep the full schema name.
  • TS: fixed-size arrays of length 2 (Fq2 coordinates) are emitted as tuples.
  • Rust: [u8; N] fields get an explicit bytes codec (msgpack bin, not an integer sequence — the C++ side rejects the latter); generated Bin32 keeps the byte-oriented surface callers had (deref to slice, comparisons with Vec/arrays); fq[2] (Fq2) fields get serde_array2_bytes, so G2 coordinates serialize as bin32.

bb.js

  • The in-tree src/cbind generator is deleted. scripts/generate.sh runs ipc-codegen on the checked-in bb_schema.json (plain node --experimental-strip-types, ts-node dropped) with --strip-method-prefix --strip-type-prefix, so the public API surface is unchanged while wire tags carry the Bb prefix.
  • Backends: UDS, SHM async and SHM sync backends are thin wrappers over ipc-runtime's SpawnedProcessBackend / SpawnedProcessBackendSync and NAPI clients (no separate bb.js NAPI addon); wasm backends call the ipc_ffi_entry export. Unused stdio pipe backend deleted, and findNapiBinary removed (it looked for the no-longer-shipped nodejs_module.node).
  • poseidon.bench.test.ts warms each mode with one untimed pass before timing it; previously every mode was timed cold, and the pipelined result in particular swung by several µs between CI runs.
  • barretenberg/ts and bb.js bootstraps include ipc-runtime/ts/package.json in the node-modules cache key, since the workspaces portal into it.

bb-rs

  • The crate ships no transport code of its own: the generated Backend trait has a bridge impl for ipc_runtime::IpcClient (UDS / SHM / pipe), behind the ipc-runtime feature; the FFI backend (binding ipc_ffi_entry) is generated. The in-crate pipe/FFI backends, error and types modules are replaced by the generated backend, bb_client, bb_types, error, ffi_backend.
  • Breaking for external users (7.0): the pre-codegen API is removed rather than shimmed: BarretenbergApi → BbApi, PipeBackend → IpcClient::from_fds, typed scalars (Fr/Fq/… as Bin32) in arguments and responses, no shutdown() (bb exits on EOF), native feature → ipc-runtime, async removed. Older versions stay on crates.io and semver keeps "6" manifests off 7.x; the README has a migration note. The only public dependent, noir-zk-backend, pins an exact 7.0 nightly.
  • Publishing: barretenberg-rs now depends on the ipc-runtime crate, which wasn't on crates.io (cargo package refused the path-only dependency). ipc-runtime's release now also publishes its crate (it runs before barretenberg/rust's), and the bb-rs release pins it exactly (=$version). The published crate carries a copy of ipc-runtime/cpp/ipc_runtime in cpp/, and build.rs falls back to it when ../cpp is absent. ipc-runtime's release also runs for private releases (internal npm only), so the crates.io step is skipped when assert_public_release refuses.
  • Shutdown is gone (pipe EOF replaces it — it was the only sender).

Schema dialect notes (vs the raw dump the old PRs used)

  • Shutdown removed; CircuitKind is a plain u8 on the wire (the alias collided with the domain enum at bb::bbapi scope); Fq2/G2 coords are fq[2]; the shared vk-data struct is VkData (prefix-stripping made the old name collide with CircuitComputeVkResponse).

Drive-by fixes

  • Two latent breakages from the ts/ → ts/bb.js/ restructure, both one directory short: the src/barretenberg_wasm/*.wasm.gz symlinks resolved to barretenberg/ts/cpp/..., and chonk_pinned_inputs.test.ts's findRepoRoot() produced barretenberg/barretenberg/cpp/chonk-pinned-flows.
  • barretenberg/docs recursive example: 1500s lane / 1200s jest budget. It uses the published bb.js, and runs alongside the bb build, which slows it by an order of magnitude when bb rebuilds everything.

Performance

The poseidon bench's "Native Shared Pipelined" (2 fields) was ~19–20 µs/call on this branch vs ~12 µs on next. Causes, each measured in isolation: the always-wake futex traffic, the reactor serve loop, the async hop in SpawnedProcessBackend.call, the per-call .catch, and the Map<bigint> pending map (only visible under jest, whose VM realm runs this JS ~3.5x slower and saturates the main thread in a pipelined burst). With the fixes above, CI reports 12.71 µs/call — level with next — and the other modes match next too (shared sync 21.5, shared 34.7, socket pipelined 29.8, socket 48.3).

Validation

  • ipc-runtime: C++ tests incl. 6 new pipe tests and the SHM init-race regression test, + ASAN; TS tests 19/19 (incl. SpawnedProcessBackend exit attribution, kill-on-broken-connection and respawn).
  • The SHM hang was reproduced before fixing: with a 20 ms test-only delay before ring init, 100/100 runs hung with conditional wakes and no ready marker, 0/100 with always-wake; the regression test pins the ready-marker fix.
  • bb: bbapi_tests green; c_bind exception tests green through ipc_ffi_entry.
  • bb.js full suite, exercising every transport — wasm, native socket, socket pipelined, SHM async, SHM pipelined, SHM sync.
  • Pinned Chonk IVC flows prove and verify through bb.js on both the native and wasm backends.
  • Live UDS smoke against real bb: pipelined poseidon2, correct hashes, clean SIGTERM. The wasm ipc_ffi_entry path returns the identical hash. (In wasm a failing command aborts via the throw_or_abort host import rather than returning an error response — unchanged behavior, since the wasm preset defines BB_NO_EXCEPTIONS and the old cbind guarded its catch the same way.)
  • bb-rs: pipe tests 14/14 against the new bb binary; FFI tests 72 pass (3 ignored) against libbb-external.a. Release dry run in a scratch copy: the packaged ipc-runtime crate builds with no ../cpp present, barretenberg-rs packages against it, and a consumer built only from the two packages hashes through bb over UDS.

🤖 Generated with Claude Code

https://claude.ai/code/session_01NuUzj3qpkpJor6GMpWB4T4

@socket-security

socket-security Bot commented Aug 29, 2026 •

Copy link
Copy Markdown

Review the following changes in direct dependencies. Learn more about Socket for GitHub.

Diff Package Supply Chain
Security
Vulnerability Quality Maintenance License
Addednpm/​@​aztec-foundation/​ipc-runtime@​0.0.0-use.localN/AN/AN/AN/AN/A

View full report

charlielye added a commit that referenced this pull request Sep 24, 2026
Moves the world-state DB service (world_state engine, persistent
content-addressed merkle storage, lmdb tree store, IPC server) out of
barretenberg into native-packages/wsdb, compiling with zero barretenberg
headers: bb is linked only as a prebuilt archive for the poseidon2 c_bind and
for the new bb_wsref_* C ABI over the in-memory reference world state
(world_state_reference), against which this package's conformance test drives
its WorldState and asserts agreement on roots, sibling paths, low-leaf
lookups, preimages and checkpointing.

Constants stay in lockstep via a wsdb-local remake-constants hook on the
protocol constants-codegen (generated header, no longer checked in).

Rebased onto next after the @aztec-foundation scope rename (#25328). wsdb keeps
the name next gave it; the new kvdb package follows the same scope rather than
introducing an @aztec-scoped foundation package. The labs series is re-exported
against the current pin: the wsdb/kvdb consumption patch is rewritten for the new
scope, and the indexed/nullifier tree reference update is dropped because
upstream removed that references frontmatter, leaving only its cspell additions.

Rebased onto next: the labs patch series is re-derived against the current pin
(42d7d24b) and renumbered -- next's own scope patch has been absorbed upstream and
dropped out. scripts/labs_fnd_hashes.sh keeps next's acvm -> noir-execute rename
alongside the native-packages/{wsdb,kvdb} component entries.

A test pins the on-disk lmdb key bytes, which nothing covered: the bb-linked
parity target compiles only field_element.test.cpp, so it checks fr's hash,
msgpack and ordering but not how a key reaches disk -- and lmdblib's concrete
serialise_key(const uint256_t&) became a generic template that memcpys from
&key rather than uint256_t::data, which agree only while data[4] is the sole
member at offset 0.

Rebased onto next after it drained its labs patch queue and moved the labs pin: the series is
re-derived against the new pin (54e3383c) and our three patches are now 0001-0003, the only
ones left. The genesis seeds take next's regenerated values (#25515 changed the protocol
nullifier derivation) in wsdb's types.

The NAPI scaffolding moves rather than being copied. lmdb_store_wrapper was its
only consumer -- msgpack_client includes just ipc_client.hpp and napi.h, and
nothing outside the module referenced barretenberg/messaging -- so extracting the
store orphaned util/{promise,async_op,message_processor} and messaging/{dispatcher,
header} on the bb side. Deleting them there leaves nodejs_module as init_module.cpp
plus msgpack_client (still needed until #25362 removes it), and lets the diff read
as renames instead of ~450 lines of apparently new code. stream_parser.hpp goes
with them: it was already unreferenced and only header.hpp kept it compiling.

Wire and domain types are layered so that the generated wire records are the only
serialisation types, for IPC and for the lmdb store, and nothing in the domain carries
msgpack. Pure-data records (WorldStateRevision, TreeMeta, DBStats, TreeDBStats, the
WorldStateStatus/Meta/DBStats family, SiblingPathAndIndex) simply are the wire record via
an alias; their former static helpers are free functions. Types with behaviour (FieldElement,
the leaf values, IndexedLeaf, LeafUpdateWitnessData and the insertion results) keep their
identity and convert through one Wire<T> trait in merkle_tree/wire.hpp (to_wire(x) /
from_wire<Domain>(w)), written once as templates over the leaf kind; the store packs leaf
preimages through it and FieldElement's msgpack is a non-intrusive adaptor onto wire::Fr.
The only shape conversions left are StateReference (map vs list) and the tree-id enum. A
test pins the persisted bytes of an IndexedLeaf and a TreeMeta, captured before the change,
so the move is proven byte-identical on disk. bb's client uses the same Wire<T> pattern
over its own merkle types (wsdb_wire.hpp). libbb-external.a, the archive barretenberg
publishes for external consumers, now also carries world_state_reference and its bb_wsref_*
C ABI, so an external world-state implementation links the release archive alone for its
conformance tests. Two hand converter files (426 + 284 lines)
become two headers of 231 + 221, and the six duplicated domain structs go.

Rebased onto next: the genesis seeding (#25497, #25503) landed in bb's world_state,
so its generated seed header moves with the rest of world_state into
native-packages/wsdb in wsdb's own types, regenerate_genesis_constants.sh and the
genesis-constants skill point at the new path, GENESIS_NULLIFIER_TREE_ROOT joins
wsdb's constants selection, and the three labs patches renumber to 0021-0023 behind
next's grown series.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012EUKia6wteDk9kZZuT2Gju
@charlielye
charlielye force-pushed the cl/ipc-bb-bbjs-migrate branch from 1c5d177 to 3cbcfe2 Compare September 25, 2026 12:56
@charlielye
charlielye force-pushed the cl/ipc-bb-bbjs-migrate branch from 3cbcfe2 to c79d6db Compare September 29, 2026 17:43
charlielye added a commit that referenced this pull request Sep 30, 2026
A test that binds a port has to run in its own container. Which tests
those were was a list of package prefixes in labs' `test_cmds`:

```bash
# These need isolation due to network stack usage (p2p, anvil, etc).
if [[ "$test" =~ ^(prover-node|p2p|ethereum|aztec|prover-client/src/test|...) ]]; then
  prefix+=":ISOLATE=1:NAME=$test"
```

A list only works if whoever adds a test in a new package knows to
extend it, and an unisolated test that spawns anvil passes until the day
it races another one.

### Two tests were already in that position

`cli/src/cmds/l1/attester_exit`, added on 18 September, fails before
reaching any assertion:

```
Anvil exited with code 0 before listening.
Output: Error: Address already in use (os error 98)
```

three retries, all the same. It failed the `x-full-no-test-cache` run on
#25362, a bb and bb.js migration it has nothing to do with.
`no-test-cache` is what surfaced it: every test executes in the same
window rather than most being served from cache, which is exactly when
two unisolated anvils meet.

The two it lost to were `ethereum/src/test/blob_kzg_warmup`, isolated
because `ethereum` is listed, and
`epoch-cache/src/epoch_cache.integration`, which is not listed and has
been unisolated since April. Neither is at fault; `cli` and
`epoch-cache` were simply never on the list.

### The change

Name a test `*.isolate.test.ts` and it gets `ISOLATE=1`, wherever it
lives. The two above are renamed accordingly.

The package list stays, so nothing loses the isolation it has today.
Prefer the suffix for anything new; an entry can leave the list later by
renaming its tests. The point is that the requirement is stated where
someone writes the test, instead of somewhere else they have to
remember.

Not a default-port change: fixed ports with full isolation is the
established preference here, and ephemeral ports have brought their own
flakiness before.

### Verified

- The condition isolates both renamed tests, leaves ordinary tests
shared, and keeps every currently listed package isolated.
- The glob still enumerates both renamed files, 649 tests in all, and
`.isolate.test.ts` is not caught by the `.bench.test.ts` skip.
- `bootstrap.sh` parses with extglob and globstar, which is how ci3 runs
it. Plain `bash -n` reports a syntax error on the `!(...)` glob both
before and after this change.

### Landing

Rides as a labs patch, `labs-patches/0009`, so CI gets it without
waiting on a pin bump. It should be upstreamed to aztec-node and the
patch dropped once the pin passes it. Worth a look from whoever owns
`cli`, since it is their test failing other people's PRs.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
@charlielye
charlielye force-pushed the cl/ipc-bb-bbjs-migrate branch from c79d6db to 9ff4d14 Compare September 30, 2026 10:04
charlielye added a commit that referenced this pull request Oct 1, 2026
…request (#25548)

Addresses
AztecProtocol/barretenberg-claude#4427 from
the bb.js side. The aztec-node side is
[aztec-node#354](aztec-labs-eng/aztec-node#354).

A pooled bb verifier whose process dies is handed back to the pool and
handed out again on every later borrow. The pool cannot tell, because
the socket backend marks itself permanently unusable and every later
call fails with a bare `Socket not connected`. The node therefore
reports a dead helper process as an invalid transaction proof,
persistently, and books it into the metric that means users are
submitting bad proofs.

This gives an owner the two things it needs, in the shape the AVM
simulator pool already relies on.

### A failed call says whether retrying can help

An environmental failure now carries `retry: true`: the bb process died,
its connection broke, or it could not be started. Only a binary that
cannot be executed stays non-retryable, because retrying cannot fix it.

The bare property is the whole contract, feature-detected rather than
imported:

```ts
if (err instanceof Error && (err as Error & { retry?: unknown }).retry === true) { ... }
```

That is deliberately the same convention `ipc-runtime`'s transports use,
so a caller written against one works unchanged against the other. It is
also what lets a verifier distinguish a bad proof from a dead helper,
which is the misattribution half of the issue.

### The socket backend can replace a dead process, on request

`respawn` is a new `BackendOptions` flag, off by default. With it on,
the bb process and its connection are one unit swapped as a whole: the
next call starts a replacement, while calls already in flight still fail
retryably. Concurrent callers that find the connection down share one
replacement, so a single death costs a single process, and a replacement
that arrives after `destroy()` is not left running.

It is opt-in because a replacement has none of the state a command
sequence establishes: no SRS loaded over the connection, no Chonk
accumulation, no batch-verifier session with its registered keys. An
owner that runs any such sequence must leave it off, so a death fails
loudly instead of the next call quietly running against a process that
has forgotten everything. The verifier pool qualifies because every
verification carries its proof and key in the call.

With this, a pool needs no liveness check and no maintenance loop:
returning an instance unconditionally becomes correct, exactly as it
already is for the AVM pool.

### Also

`BarretenbergSync.initSingleton()` cached a failed initialization for
the life of the process, so one bad spawn was permanent. It now clears
the failure, as the asynchronous singleton already did.

### Testing

`native_socket.test.ts`, against the fake bb the existing tests use: a
killed bb fails the call retryably and keeps failing without the option;
with it on the next call is served by a replacement process, a different
pid; four concurrent calls that find the connection down start exactly
one replacement; `destroy()` during a replacement leaves no process
running; a replacement that cannot start fails retryably too; a
connection that breaks while the process keeps running leaves no bb
behind.

`singleton.test.ts` covers the initialization fix and fails without it.

Checked against a real bb as well: with the option off a killed bb gives
a retryable error and keeps doing so, and with it on the next call
transparently returns the same hash from a fresh process.

### What this does not cover

`BarretenbergSync` runs on the shared-memory backend, which has no
respawn and no way to report a death mid-call: the NAPI receive loop
retries without a deadline, and because the call blocks the event loop
the process exit is never even observed, so the caller wedges rather
than fails. That path is untouched here and is deliberately left alone:
the synchronous bb runs one thread doing hashes and signatures, so it is
the least likely process on the machine to be killed, and the fix would
mean threading a liveness check into the C++ client for a case that may
never happen. #25546 covers the idle-death half of it by replacing a
dead singleton, so the two PRs cover different backends rather than one
superseding the other.

### Note on direction

bb.js's hand-written backends are replaced by `ipc-runtime`'s in the
codegen migration (#25362). Nothing above is lost in that move:
`ipc-runtime`'s spawned backend already carries the same `retry`
contract and the same opt-in respawn, so the callers written against
this keep working and the implementation here is deleted. That is why
this is expressed as the retry contract rather than as a liveness query,
which would have to become part of the generated client's interface.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
charlielye and others added 2 commits October 1, 2026 18:06
…e legacy bb::ipc

Combined redo of #23612 + #23614 + #23613. Those predate envelope ids (the
8-byte request-id frame prefix), which broke wire compatibility between the
old and new stacks, so bb, bb.js and bb-rs have to move together.

bb now serves msgpack through ipc-runtime on every transport: `bb msgpack run`
maps ""/"-" to a stdio pipe server, .sock to UDS and .shm to shared memory,
all dispatching through one generated handler with envelope framing (a plain
file keeps the offline bare-frame replay). The command surface comes from a
checked-in bb_schema.json in the friendly ipc-codegen dialect, with handlers
implemented as thin adapters over the existing domain commands. The typed
cbind entrypoint is replaced by the ipc-codegen FFI contract ipc_ffi_entry,
so wasm and static-linking consumers speak the same payload as the transports.

That makes the legacy machinery consumer-free, so it goes: the in-tree ipc
library and its benchmark, the nodejs_module msgpack client (bb.js now uses
ipc-runtime's NAPI addon, and nodejs_module only exports LMDBStore), the
Command/CommandResponse named unions with their execute dispatcher, and the
func_traits/schema reflection behind the CBIND macros.

Both client libraries now generate from ipc-codegen rather than forks of it.
bb.js/src/cbind (a stale copy of the generator: schema visitor, TS and rust
backends, naming, friendly-schema lowering) is deleted in favour of calling
ipc-codegen, and barretenberg-rs drops its hand-written Backend trait, error
type, Fr/Point types and both backends — codegen emits those, and
ipc_runtime::IpcClient plugs in as the Backend, so the crate is ~89% generated
with a deprecated BarretenbergApi shim keeping the published surface. The
hand-rolled PipeBackend goes with them; its tests now run over ipc-runtime's
UDS transport.

ipc-codegen gains --strip-type-prefix (bb.js and barretenberg-rs publish
unprefixed names while wire tags keep the service prefix) and three rust
serde fixes: fq[2] pairs, [u8; N] and above-cutoff byte arrays were encoded as
sequences of integers rather than msgpack bin, which the C++ side rejects.

ipc-runtime gains the pipe transport (fd pair or stdio, peer EOF requests
shutdown), a sync spawned-process backend for callers that cannot await,
pre-spawn removal of stale shm segments (they are created O_EXCL, so a
killed server's leftovers blocked the next one), and an unref option so a
spawned backend cannot hold the Node event loop open while idle. Its rust and
zig build files now discover the C++ sources instead of keeping hand-written
copies of the CMake list, which had silently drifted.

Also fixes two latent breakages from the ts/ -> ts/bb.js/ restructure that
blocked wasm and pinned-flow tests: the barretenberg_wasm symlinks and the
chonk pinned-inputs repo-root resolution were both one directory short.

Raises MAX_FRAME_SIZE from 256 MiB to 1 GiB. bb.js reaches bb over ipc-runtime's
socket transport now rather than the deleted nodejs_module msgpack client, so
proving requests are newly subject to the frame cap: the rollup circuits exceed
it, and bb_prover_full_rollup failed with the server refusing a 314525552-byte
frame and the client seeing the closed connection as "write EPIPE". The TS client
also checks the outgoing size, so a request over the cap names itself instead of
surfacing as an unexplained EPIPE on the next write.

Fixes two gaps in unref handling that left the Node event loop open. The
shared-memory backends dropped the caller's option -- their new() did not take it
and createBackend never passed it -- so unrefStdio was never set and the child's
stdout/stderr pipes (which exist whenever a logger is set) kept the loop alive;
they thread it through as the socket backend already does. And
BarretenbergSync.initSingleton did not force unref: true, unlike its async
counterpart, so the sync backend it creates held the loop for the life of the
process. Nothing destroys a singleton, so it opts in now too. Together these had
a labs pxe test pass all 7 cases and then hang until the job timed out.

Rebased onto next, which added two things to the surface this replaces. #25487
hardened the msgpack entrypoints: the generated dispatch already answers a
malformed envelope, an unknown command and a throwing handler with an error
frame, but an unpack or convert failure still escapes it, so api_msgpack.cpp and
ipc_ffi_entry catch that rather than lose the response. Poseidon2AbsorbChain was
a new command registered in the NamedUnion this deletes, so it moves to
bb_schema.json with a handler alongside the other crypto ones, and its malformed
-length test is reexpressed against the FFI entrypoint.

barretenberg/acir_tests is its own yarn project and consumes bb.js through a
portal, so it has to resolve bb.js's dependencies itself. bb.js declares
@aztec-foundation/ipc-runtime with the placeholder specifier that only a
resolution gives meaning to, which barretenberg/ts has and acir_tests did not:
`yarn install` there failed with "isn't supported by any available resolver".
It gets the same resolution.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016einpgfthfjLYGwCB3iQqD
…ull bb build

It takes ~35s idle but exceeded its 360s jest timeout (408s) when the lane
ran concurrently with a full bb rebuild. It uses the published bb.js, so the
slowdown is load, not the tree: raise the per-test timeout to 20 minutes and
the lane's to 25.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuUzj3qpkpJor6GMpWB4T4
@charlielye
charlielye force-pushed the cl/ipc-bb-bbjs-migrate branch from b4c4cc3 to 53105d0 Compare October 1, 2026 18:06
charlielye and others added 15 commits October 6, 2026 18:19
…al run() loop

The shared-memory transport woke the other side with a futex syscall on every
publish and release, about six per bb request, which roughly doubled the
pipelined poseidon benchmark against next. Those wakes were made unconditional
to fix an intermittent hang, whose cause turned out to be a startup race rather
than the conditional wake itself: segments are mappable from ftruncate, before
create() initialises them, so a fast client could attach, set its blocked flag
on a response ring, and have the server's initialisation reset it; the next
publish then skipped the wake and the client slept on published data forever.

connect() now refuses a segment until it is fully initialised. create() writes
the ring capacity, and the doorbell's num_slots, last with release ordering,
and a zero value means not ready: SpscShm::connect and MpscProducer::connect
throw and the existing connect-retry loops try again. With that in place the
conditional wake is restored on the rings, the doorbell and the reactor's
notify().

bb also served every request through run_reactor, paying a request copy, a
heap-allocated respond closure, a locked completion queue and a notify() (a
self-pipe write on sockets) per request, although every bb handler responds
before returning. It now uses the serial run() loop, asserting that.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuUzj3qpkpJor6GMpWB4T4
…and stop the MPSC spin reading the clock every iteration

Pipelined bb.js calls over shared memory were still ~50% slower than next on
the poseidon benchmark after the futex fix. Layer by layer, the transport
itself was at parity; the extra cost was main-thread JavaScript per call, which
left the main thread saturated so the API's own encode/decode no longer
overlapped with bb's work. Most of it was SpawnedProcessBackend.call: an async
function awaiting ensureUp() before every call. A live process now takes a
direct path with no async frame; only a missing or dead process goes through
ensureUp(). On the benchmark's pattern (the first pipelined burst after 10k
sequential calls) this takes the fixed build from 17.3 to 13.4 us/call (min of
8, next 14.1).

MpscConsumer::wait_for_data's spin also called clock_gettime on every
iteration, about a tenth of bb's instructions while serving; it now checks the
clock every 256 iterations, as SpscShm's spin does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuUzj3qpkpJor6GMpWB4T4
… with small-int ids

NapiShmAsyncClient kept pending calls in a Map<bigint>. Each call allocated
two heap BigInts (the id and the one napi hands back) and hashed them three
times. Under jest, where the main thread saturates during a pipelined burst,
that overhead was directly visible in bb.js's poseidon bench.

Ids are now numbers below 2^30, so V8 keeps them as small integers. Pending
calls sit in an issue-order queue: in-order servers (bb's serial run() loop)
always match the head, and out-of-order responses fall back to a search.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuUzj3qpkpJor6GMpWB4T4
…a per-call .catch

SpawnedProcessBackend.call chained .catch(attributeCallError) onto every
call, so each successful call paid for an extra promise and microtask. In a
pipelined burst under jest that was ~0.8us/call of main-thread CPU in steady
state and ~2us in the first burst.

The SHM and UDS clients now take an optional mapError hook, applied only on
their rejection paths. SpawnedProcessBackend passes attributeCallError for
the incarnation when it connects, and the fast path returns the client's
promise directly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuUzj3qpkpJor6GMpWB4T4
Every mode was timed on its first run, so the numbers included JIT warm-up
for that call pattern. A pipelined burst exercises different code paths
from the sequential warm-up that preceded it, so it was the most affected.
Its first-burst result swung by several µs between CI runs. Each mode now
runs once untimed first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuUzj3qpkpJor6GMpWB4T4
UdsIpcClient still kept pending calls in a Map<bigint>, with the same
per-call BigInt allocation and hashing the SHM client had. Both clients now
use one PendingQueue: small-int ids, head-of-queue fast path, search
fallback. The UDS wire id stays u64; live ids sit in the low word, and a
nonzero high word never matches a call.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuUzj3qpkpJor6GMpWB4T4
The 7.0 line is a breaking release, and older barretenberg-rs versions stay on
crates.io for anyone who needs the old API (semver keeps `"6"` manifests off
7.x). Removes the legacy BarretenbergApi/PipeBackend/byte-valued types, the
old module paths and the inert `native`/`async` features, so the crate exposes
only the generated BbApi.

The tests move to BbApi with typed scalars; the pipe tests spawn bb and drive
it through ipc_runtime::IpcClient::from_fds. The README gains a 6.x migration
note.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuUzj3qpkpJor6GMpWB4T4
barretenberg-rs now depends on the ipc-runtime crate, which had no version and
was not on crates.io, so `cargo publish` refused the package. The release step
now publishes ipc-runtime at the same version first and pins it exactly.

ipc-runtime's build compiles the C++ sources from ../cpp, which a published
crate does not have; the release copies them into the crate's cpp/ directory
and build.rs falls back to it. Publishing happens from the barretenberg-rs
release rather than ipc-runtime's, because the latter also runs for private
releases, which must not reach crates.io.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuUzj3qpkpJor6GMpWB4T4
…release

ipc-runtime owns its crate, so its release publishes it. The release also runs
for private releases, which publish only to the internal npm registry, so the
crates.io step is skipped when assert_public_release refuses. ipc-runtime is
released before barretenberg/rust, whose release now only pins the version.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuUzj3qpkpJor6GMpWB4T4
The wake is skipped unless the waiter has flagged that it is about to sleep.
Each side stores (head/tail/seq, or the flag) and then loads the other's
value, and release/acquire lets each load run ahead of its own store. Both
sides could then miss each other, leaving the waiter asleep on published data,
with no timeout on the default call path. A seq_cst fence between the store and
the load on both sides guarantees one of them sees the other.

The futex lock already fenced the waiter in practice; the SPSC producer had
no such cover. No measurable cost in a pipelined poseidon A/B on a loaded host.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuUzj3qpkpJor6GMpWB4T4
UDS and pipe frames were capped at 1 GiB (256 MiB until bb proving requests
outgrew it). Nothing in a stream transport needs a limit below the u32 length
field; the cap only guarded against a corrupt prefix driving a huge up-front
allocation. Receivers now grow their buffer as payload bytes arrive (C++
read_payload, TS FrameReader), so that guard is no longer needed and the
limit is the wire format's own. SHM stays bounded by half its ring capacity.

FrameReader also fixes the TS client and server re-concatenating their whole
buffer on every chunk, which made receiving a large frame quadratic.

A request too large for the length field now rejects with a non-transport
IpcError: it is the caller's mistake on a healthy connection, and as a
transport error SpawnedProcessBackend killed the server for it. Senders also
check against the payload limit (frame limit minus the 8-byte id) rather than
the frame limit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuUzj3qpkpJor6GMpWB4T4
Serving "-" made the process's stdout the frame stream, so any other write to
it (std::cout, printf, a library's diagnostics) landed inside a frame and
desynced the peer. next's bb guarded this by redirecting std::cout; that guard
did not carry over to the ipc-runtime pipe path. make_server now serves on a
private duplicate of stdout and points fd 1 at stderr, which also covers C
stdio and raw writes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuUzj3qpkpJor6GMpWB4T4
It looked for nodejs_module.node, which no longer ships now that bb.js uses
ipc-runtime's addon, so it could only ever return null.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuUzj3qpkpJor6GMpWB4T4

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants