Skip to content

runtime/wasip1: add opt-in threaded GC for WAMR - #2669

Merged
cpunion merged 42 commits into
xgo-dev:mainfrom
cpunion:codex/wasi-threaded-gc
Sep 30, 2026
Merged

cpunion merged 42 commits into
xgo-dev:mainfrom
cpunion:codex/wasi-threaded-gc

Conversation

@cpunion

@cpunion cpunion commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

WAMR's WASI pthread backend currently requires -tags nogc. With LLGO_WASI_THREADS=1, this PR enables the linear-memory collector by default while preserving -tags nogc as an explicit override. It supports multiple application pthreads: one collector stops the others and performs serial, non-moving, conservative mark-and-sweep. The current scheduler still uses one goroutine per pthread.

Threads publish compiler root chains and pthread stack bounds before acknowledging a stop. The runtime also roots host-TLS slots used by sync.Pool, uses the correct pthread startup and main-thread TLS, and defaults worker stacks to 1 MiB. Go allocations occupy independent libc arenas so they cannot overwrite pthread stacks or TLS.

Arena capacity includes block-state metadata and alignment padding. The allocator now accounts for both before choosing an arena size, fixing requests at and just below 32 MiB that previously added multiple arenas without finding a sufficiently large contiguous allocation. Overflow is rejected before calling the allocator. Adding a disjoint WASI arena under the allocator lock no longer requests another STW, avoiding a second 500 ms wait after an unsuccessful collection.

An uninstrumented C call can prevent a safe collection. After 500 ms, GC resumes stopped threads without marking or sweeping; runtime.GC() returns without incrementing MemStats.NumGC. Repeated allocations may still attempt further collections and grow until the memory limit is reached. Foreign C threads that explicitly enter Go remain registered until pthread destruction, including while idle after a callback. Empty arenas are not returned to libc. These progress and memory-ownership limits are documented in the threaded GC contract.

The branch also clears retired browser Fiber/Asyncify buffers to avoid retaining finalizable objects, parks the initial WASI execution environment on main Goexit, and excludes the persistent timer service from the user-G deadlock count. The callback-poll hook keeps an explicit normal-return reset because a Go defer is unsafe on the system Fiber without a G.

Validation:

  • The new host regression executes the actual arena-sizing and metadata-layout code across default-size boundaries, all start alignments, 16/32-byte blocks, allocation failure and overflow. It also checks that disjoint growth skips STW while contiguous growth retains its existing handshake. go test ./internal/build -run '^TestWasmGC' -count=1 passes.
  • Five fresh WAMR instances allocate 32 MiB minus 128 KiB, 32 MiB minus one byte, exactly 32 MiB, 32 MiB plus one byte, and 33 MiB. Each verifies collection progress, retained contents and bounded heap growth. All pass. The old 32 MiB artifact did not complete within 120 seconds; the old sizing code also fails the host capacity regression.
  • The expanded threaded-GC fixture passes with the CI eight-thread limit. It covers worker-private roots without passing pointers to the main thread, thread registration/retirement overlapping collection, an idle foreign pthread after a callback, blocked C calls, timers, panic/recover, cross-thread pointer handoff and arena growth. It verifies that collection resumes after the foreign thread exits.
  • Existing acceptance covers public llgo run/test, startup with 100 goroutines, GC/nogc, filesystem access, weak references, select stress and focused W32 standard-library packages. The Node/Chrome worker acceptance and retired-fiber finalizer regressions were validated before this focused WASI follow-up. Fresh CI runs on the updated head.

This contribution is rebased on main 07a0059ce. WASI pthread mode is still opt-in here. Follow-up #2695 addresses broader WAMR exception/thread stability, safepoint fast paths, GC-aware pthread waits and segment lookup cost; #2697 makes the thread profile the default. Those follow-ups are not claimed as changes in this PR.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: WASI threaded GC support

Thorough, carefully-engineered change. The concurrency-sensitive code (STW epoch parity, wake-race snapshotting, GC-root rebuild guards) is well-reasoned and unusually well-commented, downloads are checksum-pinned, and invariant violations trap rather than corrupting memory. Ran four review passes (quality, performance, security, docs).

Highlights of what was checked and found sound: futex-backed waits (not spin loops) in wasmWorkerStopForGC/wasmGCStopTheWorld; overflow guards in AllocRoot/gc_wasm.c stack math; path-traversal protection in the test server.mjs; no-shell subprocess invocation in the python/emulator paths; the STW target/ready registration race (handled via initWasmWorkerGCSystem's increment-then-stop ordering).

The one item I'd treat as blocking-worthy is the llgo run exit-code regression (inline). The rest are performance/robustness suggestions.

Additional notes (no reliable inline anchor)

  • O(n²) goroutine churn in the GC-root registry (runtime/internal/gcroot/gcroot.go): every goroutine spawn calls Register → registeredLocked (O(n) duplicate-check scan) and every exit calls Unregister (O(n) linked-list walk), both under the single global registry lock that also blocks Visit. For workloads spawning many short-lived goroutines this is O(n²) under one lock. The duplicate-check scan could be debug-only, and Unregister could use an intrusive doubly-linked list (as rootAllocation in root_wasm_workers.go already does) for O(1) removal.

  • Binaryen source switched to a project fork (.github/actions/setup-binaryen/action.yml): the toolchain now pulls wasm-opt from github.com/xgo-dev/binaryen (tag llgo-v132.3) instead of upstream WebAssembly/binaryen. Mitigation is solid — per-platform SHA-256 pinned in checksums.txt, malformed-checksum guard, no trust in a server-provided .sha256. Flagging so a maintainer can confirm the fork + pinned hashes are intentional and verified.

Additional findings

  • internal/build/run.go:335: [P1] llgo run of a native binary loses the guest exit code: ModeRun now returns cmd.Run()'s error directly instead of the previous mockable.Exit(s.ExitCode()). The error flows up to runCmdEx (cmd/internal/run/run.go:104-107), which unconditionally calls mockable.Exit(1). So a native program that exits with code 2, 3, ... now makes llgo run exit with 1, whereas before it propagated the child's exact code. The inline comment only justifies the zero-exit case and doesn't acknowledge the non-zero regression.

Comment thread runtime/internal/wasmsync/mutex.go
Comment thread runtime/internal/wasmsync/mutex.go
Comment thread runtime/internal/runtime/scheduler_events_wasm_workers.go
Comment thread runtime/internal/runtime/tinygogc/gc_tinygo.go Outdated
@codecov

codecov Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

c6c3ce4dc326 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 7160 B 0 B / +0.0% 387 B 0 B / +0.0% 527.105 ms +3.673 ms / +0.7% (worse) 1.263 ms -1.259 us / -0.1% (better)
Linux cprintf-lto 6912 B 0 B / +0.0% 368 B 0 B / +0.0% 531.977 ms -48.09 ms / -8.3% (better) 1.278 ms -104.6 us / -7.6% (better)
Linux fmtprintf 1663312 B +1040 B / +0.1% (worse) 498386 B +27 B / +0.005418% (worse) 3.721 s +137.8 ms / +3.8% (worse) 3.212 ms +65.31 us / +2.1% (worse)
Linux fmtprintf-lto 1501104 B +720 B / +0.04799% (worse) 436123 B +10 B / +0.002293% (worse) 10.798 s -100.2 ms / -0.9% (better) 2.913 ms +35.22 us / +1.2% (worse)
Linux println 69232 B +240 B / +0.3% (worse) 16805 B +22 B / +0.1% (worse) 557.013 ms -37.34 ms / -6.3% (better) 1.651 ms +54.78 us / +3.4% (worse)
Linux println-lto 59960 B +112 B / +0.2% (worse) 14209 B +10 B / +0.1% (worse) 776.519 ms -67.6 ms / -8.0% (better) 1.592 ms -220.2 us / -12.2% (better)
macOS cprintf 68064 B 0 B / +0.0% 4429 B 0 B / +0.0% 993.527 ms -106.9 ms / -9.7% (better) 3.373 ms -871.9 us / -20.5% (better)
macOS cprintf-lto 84288 B 0 B / +0.0% 193 B 0 B / +0.0% 905.922 ms -98.11 ms / -9.8% (better) 3.155 ms -610 us / -16.2% (better)
macOS fmtprintf 1504832 B +160 B / +0.01063% (worse) 875320 B +748 B / +0.1% (worse) 3.792 s +17.12 ms / +0.5% (worse) 5.079 ms -1.789 ms / -26.1% (better)
macOS fmtprintf-lto 1192704 B 0 B / +0.0% 848952 B +736 B / +0.1% (worse) 7.543 s -3.9 s / -34.1% (better) 5.351 ms -2.488 ms / -31.7% (better)
macOS println 117216 B +80 B / +0.1% (worse) 37549 B +128 B / +0.3% (worse) 944.485 ms -163.5 ms / -14.8% (better) 3.864 ms -3.029 ms / -43.9% (better)
macOS println-lto 119472 B 0 B / +0.0% 34928 B +104 B / +0.3% (worse) 1.237 s -244.9 ms / -16.5% (better) 6.580 ms +1.552 ms / +30.9% (worse)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.416 s -30.59 ms / -2.1% (better) 3.942 ms -4.6 us / -0.1% (better)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.433 s -27.51 ms / -1.9% (better) 4.011 ms -27.2 us / -0.7% (better)
Windows MinGW fmtprintf 1933312 B +512 B / +0.02649% (worse) 598950 B 0 B / +0.0% 4.334 s -448.4 ms / -9.4% (better) 8.780 ms -1.699 ms / -16.2% (better)
Windows MinGW fmtprintf-lto 1957888 B +512 B / +0.02616% (worse) 547318 B 0 B / +0.0% 10.902 s -354.1 ms / -3.1% (better) 10.602 ms +786 us / +8.0% (worse)
Windows MinGW println 75776 B 0 B / +0.0% 25142 B 0 B / +0.0% 1.420 s -19.63 ms / -1.4% (better) 7.423 ms +77 us / +1.0% (worse)
Windows MinGW println-lto 69120 B 0 B / +0.0% 21990 B 0 B / +0.0% 1.658 s +10.95 ms / +0.7% (worse) 7.843 ms +41.8 us / +0.5% (worse)
Windows MinGW 386 cprintf 43520 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.261 s +27.19 ms / +2.2% (worse) 5.084 ms -19.8 us / -0.4% (better)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.281 s +15.65 ms / +1.2% (worse) 5.284 ms +265.9 us / +5.3% (worse)
Windows MinGW 386 fmtprintf 1896448 B 0 B / +0.0% 472430 B +16 B / +0.003387% (worse) 4.167 s +118.9 ms / +2.9% (worse) 9.922 ms -409.9 us / -4.0% (better)
Windows MinGW 386 fmtprintf-lto 2181120 B +512 B / +0.02348% (worse) 451274 B +16 B / +0.003546% (worse) 9.632 s +46.78 ms / +0.5% (worse) 10.143 ms +56.7 us / +0.6% (worse)
Windows MinGW 386 println 96256 B 0 B / +0.0% 21474 B +16 B / +0.1% (worse) 1.254 s +5.514 ms / +0.4% (worse) 8.452 ms -75.3 us / -0.9% (better)
Windows MinGW 386 println-lto 74240 B 0 B / +0.0% 19334 B +20 B / +0.1% (worse) 1.465 s -5.213 ms / -0.4% (better) 8.351 ms -311.1 us / -3.6% (better)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.567 s +38.82 ms / +2.5% (worse) 6.242 ms -367.9 us / -5.6% (better)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.608 s +43.82 ms / +2.8% (worse) 6.444 ms +40.8 us / +0.6% (worse)
Windows MinGW ARM64 fmtprintf 1819648 B +512 B / +0.02815% (worse) 509616 B +8 B / +0.00157% (worse) 4.352 s +104.7 ms / +2.5% (worse) 13.164 ms +849.8 us / +6.9% (worse)
Windows MinGW ARM64 fmtprintf-lto 1880064 B +1024 B / +0.1% (worse) 476176 B +24 B / +0.00504% (worse) 9.790 s +265.6 ms / +2.8% (worse) 13.364 ms +265.4 us / +2.0% (worse)
Windows MinGW ARM64 println 72192 B 0 B / +0.0% 23884 B +8 B / +0.03351% (worse) 1.563 s +19.67 ms / +1.3% (worse) 10.651 ms -724.3 us / -6.4% (better)
Windows MinGW ARM64 println-lto 68608 B 0 B / +0.0% 21232 B +8 B / +0.03769% (worse) 1.779 s +25.48 ms / +1.5% (worse) 11.553 ms +380 us / +3.4% (worse)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65798 B 0 B / +0.0% 1.020 s +7.526 ms / +0.7% (worse) 3.512 ms -145.5 us / -4.0% (better)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65734 B 0 B / +0.0% 1.200 s +37.92 ms / +3.3% (worse) 4.004 ms -278.7 us / -6.5% (better)
Windows MSVC fmtprintf 1643520 B +512 B / +0.03116% (worse) 694502 B 0 B / +0.0% 3.517 s -139.1 ms / -3.8% (better) 9.848 ms +629.3 us / +6.8% (worse)
Windows MSVC fmtprintf-lto 1635328 B +512 B / +0.03132% (worse) 647078 B +16 B / +0.002473% (worse) 8.638 s -165.6 ms / -1.9% (better) 10.291 ms +624.3 us / +6.5% (worse)
Windows MSVC println 194560 B 0 B / +0.0% 120822 B 0 B / +0.0% 1.025 s -14.07 ms / -1.4% (better) 7.572 ms -50.8 us / -0.7% (better)
Windows MSVC println-lto 192512 B 0 B / +0.0% 118358 B +16 B / +0.01352% (worse) 1.211 s -16.55 ms / -1.3% (better) 7.456 ms +366.1 us / +5.2% (worse)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.288 s -960.3 us / -0.1% (better) 7.539 ms +1.102 ms / +17.1% (worse)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.309 s +101.6 ms / +8.4% (worse) 7.793 ms +1.406 ms / +22.0% (worse)
Windows MSVC 386 fmtprintf 1204736 B +512 B / +0.04252% (worse) 455820 B +16 B / +0.00351% (worse) 4.415 s +237 ms / +5.7% (worse) 14.661 ms +64.2 us / +0.4% (worse)
Windows MSVC 386 fmtprintf-lto 1241600 B +512 B / +0.04125% (worse) 427211 B +16 B / +0.003745% (worse) 10.148 s +623.7 ms / +6.5% (worse) 15.861 ms +174.5 us / +1.1% (worse)
Windows MSVC 386 println 36352 B 0 B / +0.0% 20324 B 0 B / +0.0% 1.279 s -81.84 ms / -6.0% (better) 11.206 ms -886.6 us / -7.3% (better)
Windows MSVC 386 println-lto 35840 B 0 B / +0.0% 18565 B +16 B / +0.1% (worse) 1.479 s +2.179 ms / +0.1% (worse) 10.179 ms -668.1 us / -6.2% (better)
Windows MSVC ARM64 cprintf 11776 B 0 B / +0.0% 4192 B 0 B / +0.0% 1.361 s +17.47 ms / +1.3% (worse) 8.002 ms -173.6 us / -2.1% (better)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.414 s +75.42 ms / +5.6% (worse) 8.239 ms +107.4 us / +1.3% (worse)
Windows MSVC ARM64 fmtprintf 1387008 B +512 B / +0.03693% (worse) 509560 B +16 B / +0.00314% (worse) 4.127 s +130.3 ms / +3.3% (worse) 16.418 ms +18.5 us / +0.1% (worse)
Windows MSVC ARM64 fmtprintf-lto 1405440 B +512 B / +0.03644% (worse) 476836 B +16 B / +0.003356% (worse) 9.646 s +387.7 ms / +4.2% (worse) 16.602 ms +293.4 us / +1.8% (worse)
Windows MSVC ARM64 println 45056 B 0 B / +0.0% 23908 B 0 B / +0.0% 1.364 s +35.65 ms / +2.7% (worse) 14.691 ms +1.285 ms / +9.6% (worse)
Windows MSVC ARM64 println-lto 42496 B 0 B / +0.0% 21380 B 0 B / +0.0% 1.538 s -34.65 ms / -2.2% (better) 13.594 ms -1.391 ms / -9.3% (better)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.570 ns/op +0.15 ns/op / +1.0% (worse)
Linux BenchmarkMergeCompilerFlags 198.500 ns/op -3.9 ns/op / -1.9% (better)
Linux BenchmarkMergeLinkerFlags 126.800 ns/op -0.8 ns/op / -0.6% (better)
Linux BenchmarkChannelBuffered 61.620 ns/op +6.71 ns/op / +12.2% (worse)
Linux BenchmarkChannelHandoff 13600 ns/op +597 ns/op / +4.6% (worse)
Linux BenchmarkDefer 50.510 ns/op -3.14 ns/op / -5.9% (better)
Linux BenchmarkDirectCall 1.601 ns/op +0.039 ns/op / +2.5% (worse)
Linux BenchmarkGlobalRead 1.178 ns/op +0.009 ns/op / +0.8% (worse)
Linux BenchmarkGlobalWrite 7.802 ns/op +0.016 ns/op / +0.2% (worse)
Linux BenchmarkGoroutine 36202 ns/op +11146 ns/op / +44.5% (worse)
Linux BenchmarkInterfaceCall 6.114 ns/op +0.221 ns/op / +3.8% (worse)
Linux BenchmarkRuntimeGetG 2.907 ns/op -0.083 ns/op / -2.8% (better)
macOS BenchmarkLookupPCRandom 16.920 ns/op -1.32 ns/op / -7.2% (better)
macOS BenchmarkMergeCompilerFlags 122.700 ns/op -19 ns/op / -13.4% (better)
macOS BenchmarkMergeLinkerFlags 85.080 ns/op -13.65 ns/op / -13.8% (better)
macOS BenchmarkChannelBuffered 29.280 ns/op -7.19 ns/op / -19.7% (better)
macOS BenchmarkChannelHandoff 8743 ns/op +1020 ns/op / +13.2% (worse)
macOS BenchmarkDefer 52.660 ns/op +4.15 ns/op / +8.6% (worse)
macOS BenchmarkDirectCall 1.239 ns/op -0.05 ns/op / -3.9% (better)
macOS BenchmarkGlobalRead 1.154 ns/op 0 ns/op / +0.0%
macOS BenchmarkGlobalWrite 1.106 ns/op -0.331 ns/op / -23.0% (better)
macOS BenchmarkGoroutine 99128 ns/op +36666 ns/op / +58.7% (worse)
macOS BenchmarkInterfaceCall 5.891 ns/op +1.257 ns/op / +27.1% (worse)
macOS BenchmarkRuntimeGetG 2.309 ns/op -0.33 ns/op / -12.5% (better)
Windows MinGW BenchmarkLookupPCRandom 13.300 ns/op +0.18 ns/op / +1.4% (worse)
Windows MinGW BenchmarkMergeCompilerFlags 623 ns/op -5.6 ns/op / -0.9% (better)
Windows MinGW BenchmarkMergeLinkerFlags 559.500 ns/op +6.8 ns/op / +1.2% (worse)
Windows MinGW BenchmarkChannelBuffered 30.230 ns/op +0.38 ns/op / +1.3% (worse)
Windows MinGW BenchmarkChannelHandoff 945.600 ns/op -6.8 ns/op / -0.7% (better)
Windows MinGW BenchmarkDefer 58.400 ns/op -4.51 ns/op / -7.2% (better)
Windows MinGW BenchmarkDirectCall 1.548 ns/op -0.002 ns/op / -0.1% (better)
Windows MinGW BenchmarkGlobalRead 1.858 ns/op -0.004 ns/op / -0.2% (better)
Windows MinGW BenchmarkGlobalWrite 2.460 ns/op +0.001 ns/op / +0.04067% (worse)
Windows MinGW BenchmarkGoroutine 101161 ns/op +6103 ns/op / +6.4% (worse)
Windows MinGW BenchmarkInterfaceCall 8.371 ns/op -0.042 ns/op / -0.5% (better)
Windows MinGW BenchmarkRuntimeGetG 2.168 ns/op -0.005 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 26.550 ns/op -0.07 ns/op / -0.3% (better)
Windows MinGW 386 BenchmarkMergeCompilerFlags 752 ns/op -10.3 ns/op / -1.4% (better)
Windows MinGW 386 BenchmarkMergeLinkerFlags 700.600 ns/op -1 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkChannelBuffered 38.730 ns/op -0.57 ns/op / -1.5% (better)
Windows MinGW 386 BenchmarkChannelHandoff 883.500 ns/op +43.4 ns/op / +5.2% (worse)
Windows MinGW 386 BenchmarkDefer 43.990 ns/op +2.22 ns/op / +5.3% (worse)
Windows MinGW 386 BenchmarkDirectCall 1.547 ns/op -0.002 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGlobalRead 1.549 ns/op -0.001 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGlobalWrite 7.769 ns/op -0.008 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGoroutine 104683 ns/op -1241 ns/op / -1.2% (better)
Windows MinGW 386 BenchmarkInterfaceCall 8.362 ns/op -0.014 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 2.172 ns/op +0.004 ns/op / +0.2% (worse)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.080 ns/op -0.01 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 564.600 ns/op +2.4 ns/op / +0.4% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 531.200 ns/op -8.5 ns/op / -1.6% (better)
Windows MinGW ARM64 BenchmarkChannelBuffered 37.790 ns/op -0.37 ns/op / -1.0% (better)
Windows MinGW ARM64 BenchmarkChannelHandoff 2701 ns/op +470 ns/op / +21.1% (worse)
Windows MinGW ARM64 BenchmarkDefer 56.590 ns/op +1.15 ns/op / +2.1% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.590 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkGlobalRead 0.663 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkGlobalWrite 0.663 ns/op +0.0001 ns/op / +0.01508% (worse)
Windows MinGW ARM64 BenchmarkGoroutine 61113 ns/op -1185 ns/op / -1.9% (better)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.148 ns/op +0.007 ns/op / +0.2% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.805 ns/op +0.035 ns/op / +2.0% (worse)
Windows MSVC BenchmarkLookupPCRandom 10.950 ns/op +0.01 ns/op / +0.1% (worse)
Windows MSVC BenchmarkMergeCompilerFlags 535 ns/op -7.5 ns/op / -1.4% (better)
Windows MSVC BenchmarkMergeLinkerFlags 481.900 ns/op +0.1 ns/op / +0.02076% (worse)
Windows MSVC BenchmarkChannelBuffered 41.810 ns/op -0.1 ns/op / -0.2% (better)
Windows MSVC BenchmarkChannelHandoff 1425 ns/op +46 ns/op / +3.3% (worse)
Windows MSVC BenchmarkDefer 62.730 ns/op +2.49 ns/op / +4.1% (worse)
Windows MSVC BenchmarkDirectCall 1.147 ns/op +0.032 ns/op / +2.9% (worse)
Windows MSVC BenchmarkGlobalRead 1.029 ns/op -0.019 ns/op / -1.8% (better)
Windows MSVC BenchmarkGlobalWrite 8.203 ns/op +0.044 ns/op / +0.5% (worse)
Windows MSVC BenchmarkGoroutine 72340 ns/op -1106 ns/op / -1.5% (better)
Windows MSVC BenchmarkInterfaceCall 6.125 ns/op +0.272 ns/op / +4.6% (worse)
Windows MSVC BenchmarkRuntimeGetG 1.458 ns/op +0.051 ns/op / +3.6% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 26.940 ns/op +0.3 ns/op / +1.1% (worse)
Windows MSVC 386 BenchmarkMergeCompilerFlags 851.100 ns/op +69 ns/op / +8.8% (worse)
Windows MSVC 386 BenchmarkMergeLinkerFlags 796.400 ns/op +53.3 ns/op / +7.2% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 45.170 ns/op -0.07 ns/op / -0.2% (better)
Windows MSVC 386 BenchmarkChannelHandoff 971.900 ns/op -21.3 ns/op / -2.1% (better)
Windows MSVC 386 BenchmarkDefer 55.380 ns/op +5.01 ns/op / +9.9% (worse)
Windows MSVC 386 BenchmarkDirectCall 1.590 ns/op +0.035 ns/op / +2.3% (worse)
Windows MSVC 386 BenchmarkGlobalRead 1.866 ns/op -0.004 ns/op / -0.2% (better)
Windows MSVC 386 BenchmarkGlobalWrite 7.795 ns/op -0.015 ns/op / -0.2% (better)
Windows MSVC 386 BenchmarkGoroutine 131417 ns/op -2520 ns/op / -1.9% (better)
Windows MSVC 386 BenchmarkInterfaceCall 8.125 ns/op +0.003 ns/op / +0.03694% (worse)
Windows MSVC 386 BenchmarkRuntimeGetG 2.525 ns/op +0.038 ns/op / +1.5% (worse)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.090 ns/op -0.05 ns/op / -0.4% (better)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 579.400 ns/op +2.2 ns/op / +0.4% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 548.600 ns/op +1.6 ns/op / +0.3% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 38.850 ns/op +0.01 ns/op / +0.02575% (worse)
Windows MSVC ARM64 BenchmarkChannelHandoff 3424 ns/op +1152 ns/op / +50.7% (worse)
Windows MSVC ARM64 BenchmarkDefer 64.980 ns/op -0.64 ns/op / -1.0% (better)
Windows MSVC ARM64 BenchmarkDirectCall 0.590 ns/op -0.0001 ns/op / -0.01696% (better)
Windows MSVC ARM64 BenchmarkGlobalRead 0.664 ns/op +0.0011 ns/op / +0.2% (worse)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.764 ns/op +0.004 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkGoroutine 63224 ns/op -1573 ns/op / -2.4% (better)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.140 ns/op -0.003 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.770 ns/op +0.001 ns/op / +0.1% (worse)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 905.600 ns/op +7.7 ns/op / +0.9% (worse)
Linux AfterFuncZeroDelivery/LLGo 46715 ns/op +9647 ns/op / +26.0% (worse)
Linux CreateStop/Go 292.900 ns/op +3.6 ns/op / +1.2% (worse)
Linux CreateStop/LLGo 1793 ns/op -5 ns/op / -0.3% (better)
Linux RearmStopped/Go 116.200 ns/op 0 ns/op / +0.0%
Linux RearmStopped/LLGo 1389 ns/op -50 ns/op / -3.5% (better)
Linux ResetActive/Go 68.550 ns/op +0.08 ns/op / +0.1% (worse)
Linux ResetActive/LLGo 720 ns/op -49.9 ns/op / -6.5% (better)
Linux ResetHeap1024/Go 67.700 ns/op +0.57 ns/op / +0.8% (worse)
Linux ResetHeap1024/LLGo 180.300 ns/op +7.3 ns/op / +4.2% (worse)
macOS AfterFuncZeroDelivery/Go 470.600 ns/op -47.9 ns/op / -9.2% (better)
macOS AfterFuncZeroDelivery/LLGo 120493 ns/op +21214 ns/op / +21.4% (worse)
macOS CreateStop/Go 189 ns/op -23.6 ns/op / -11.1% (better)
macOS CreateStop/LLGo 745 ns/op -148.7 ns/op / -16.6% (better)
macOS RearmStopped/Go 61.670 ns/op -15.93 ns/op / -20.5% (better)
macOS RearmStopped/LLGo 432.100 ns/op -236.8 ns/op / -35.4% (better)
macOS ResetActive/Go 45.770 ns/op -14 ns/op / -23.4% (better)
macOS ResetActive/LLGo 236.200 ns/op -26.5 ns/op / -10.1% (better)
macOS ResetHeap1024/Go 51.520 ns/op -2.31 ns/op / -4.3% (better)
macOS ResetHeap1024/LLGo 89.550 ns/op -4 ns/op / -4.3% (better)
Windows MinGW AfterFuncZeroDelivery/Go 556.600 ns/op -11.8 ns/op / -2.1% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 181019 ns/op -1197 ns/op / -0.7% (better)
Windows MinGW CreateStop/Go 117.300 ns/op -4.2 ns/op / -3.5% (better)
Windows MinGW CreateStop/LLGo 494.600 ns/op +12.1 ns/op / +2.5% (worse)
Windows MinGW RearmStopped/Go 31.680 ns/op +0.41 ns/op / +1.3% (worse)
Windows MinGW RearmStopped/LLGo 281.400 ns/op -10.2 ns/op / -3.5% (better)
Windows MinGW ResetActive/Go 20.160 ns/op +0.13 ns/op / +0.6% (worse)
Windows MinGW ResetActive/LLGo 176.600 ns/op +2.4 ns/op / +1.4% (worse)
Windows MinGW ResetHeap1024/Go 20.530 ns/op +0.01 ns/op / +0.04873% (worse)
Windows MinGW ResetHeap1024/LLGo 125 ns/op -2.3 ns/op / -1.8% (better)
Windows MinGW 386 AfterFuncZeroDelivery/Go 950.900 ns/op -4.1 ns/op / -0.4% (better)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 201497 ns/op +342 ns/op / +0.2% (worse)
Windows MinGW 386 CreateStop/Go 190.600 ns/op +0.6 ns/op / +0.3% (worse)
Windows MinGW 386 CreateStop/LLGo 497.600 ns/op -9.2 ns/op / -1.8% (better)
Windows MinGW 386 RearmStopped/Go 63.380 ns/op -0.02 ns/op / -0.03155% (better)
Windows MinGW 386 RearmStopped/LLGo 361.400 ns/op +17.1 ns/op / +5.0% (worse)
Windows MinGW 386 ResetActive/Go 38.960 ns/op -0.03 ns/op / -0.1% (better)
Windows MinGW 386 ResetActive/LLGo 957.100 ns/op -35.1 ns/op / -3.5% (better)
Windows MinGW 386 ResetHeap1024/Go 39.350 ns/op -0.12 ns/op / -0.3% (better)
Windows MinGW 386 ResetHeap1024/LLGo 186.100 ns/op -2.6 ns/op / -1.4% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 673.800 ns/op +9.7 ns/op / +1.5% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 144611 ns/op +3888 ns/op / +2.8% (worse)
Windows MinGW ARM64 CreateStop/Go 200.200 ns/op +1.1 ns/op / +0.6% (worse)
Windows MinGW ARM64 CreateStop/LLGo 385.500 ns/op +17.3 ns/op / +4.7% (worse)
Windows MinGW ARM64 RearmStopped/Go 70.600 ns/op -0.04 ns/op / -0.1% (better)
Windows MinGW ARM64 RearmStopped/LLGo 253.900 ns/op -1 ns/op / -0.4% (better)
Windows MinGW ARM64 ResetActive/Go 31.120 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 ResetActive/LLGo 123.200 ns/op -4.7 ns/op / -3.7% (better)
Windows MinGW ARM64 ResetHeap1024/Go 30.960 ns/op -0.2 ns/op / -0.6% (better)
Windows MinGW ARM64 ResetHeap1024/LLGo 130 ns/op +3.1 ns/op / +2.4% (worse)
Windows MSVC AfterFuncZeroDelivery/Go 680.100 ns/op +16.6 ns/op / +2.5% (worse)
Windows MSVC AfterFuncZeroDelivery/LLGo 140232 ns/op -145 ns/op / -0.1% (better)
Windows MSVC CreateStop/Go 186 ns/op +4.3 ns/op / +2.4% (worse)
Windows MSVC CreateStop/LLGo 668.400 ns/op -100.7 ns/op / -13.1% (better)
Windows MSVC RearmStopped/Go 66.990 ns/op -0.22 ns/op / -0.3% (better)
Windows MSVC RearmStopped/LLGo 292.700 ns/op -18.6 ns/op / -6.0% (better)
Windows MSVC ResetActive/Go 29.840 ns/op +0.06 ns/op / +0.2% (worse)
Windows MSVC ResetActive/LLGo 185.500 ns/op -8.8 ns/op / -4.5% (better)
Windows MSVC ResetHeap1024/Go 30.010 ns/op -0.13 ns/op / -0.4% (better)
Windows MSVC ResetHeap1024/LLGo 124.400 ns/op -0.4 ns/op / -0.3% (better)
Windows MSVC 386 AfterFuncZeroDelivery/Go 996.500 ns/op +32.2 ns/op / +3.3% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 240591 ns/op +25033 ns/op / +11.6% (worse)
Windows MSVC 386 CreateStop/Go 205.900 ns/op +3 ns/op / +1.5% (worse)
Windows MSVC 386 CreateStop/LLGo 599.200 ns/op +122.3 ns/op / +25.6% (worse)
Windows MSVC 386 RearmStopped/Go 65.110 ns/op +0.32 ns/op / +0.5% (worse)
Windows MSVC 386 RearmStopped/LLGo 323.600 ns/op +1.3 ns/op / +0.4% (worse)
Windows MSVC 386 ResetActive/Go 39.470 ns/op -0.73 ns/op / -1.8% (better)
Windows MSVC 386 ResetActive/LLGo 743.700 ns/op -198.9 ns/op / -21.1% (better)
Windows MSVC 386 ResetHeap1024/Go 40.160 ns/op +0.57 ns/op / +1.4% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 182 ns/op +9.1 ns/op / +5.3% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 665.500 ns/op +1.1 ns/op / +0.2% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 159189 ns/op +859 ns/op / +0.5% (worse)
Windows MSVC ARM64 CreateStop/Go 197.500 ns/op -1.6 ns/op / -0.8% (better)
Windows MSVC ARM64 CreateStop/LLGo 396.500 ns/op +10.2 ns/op / +2.6% (worse)
Windows MSVC ARM64 RearmStopped/Go 70.560 ns/op +0.05 ns/op / +0.1% (worse)
Windows MSVC ARM64 RearmStopped/LLGo 277.600 ns/op +7.3 ns/op / +2.7% (worse)
Windows MSVC ARM64 ResetActive/Go 31.080 ns/op -0.03 ns/op / -0.1% (better)
Windows MSVC ARM64 ResetActive/LLGo 135 ns/op +1.1 ns/op / +0.8% (worse)
Windows MSVC ARM64 ResetHeap1024/Go 31.260 ns/op +0.15 ns/op / +0.5% (worse)
Windows MSVC ARM64 ResetHeap1024/LLGo 136.700 ns/op +1 ns/op / +0.7% (worse)

Compared with 07a0059ce2e6 measured in the same runner job.

@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

c6c3ce4dc326 | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 153508 B +6535 B / +4.4% (worse) 74890 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 150996 B +5589 B / +3.8% (worse) 73165 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 140669 B +6153 B / +4.6% (worse) 78778 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 146696 B +4740 B / +3.3% (worse) 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 147361 B +5657 B / +4.0% (worse) 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3211577 B +8200 B / +0.3% (worse) 118590 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3190036 B +7260 B / +0.2% (worse) 101635 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2948334 B +8015 B / +0.3% (worse) 125433 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2842186 B +6398 B / +0.2% (worse) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2709561 B +7340 B / +0.3% (worse) 0 B 0 B / 0.0%
j32-emscripten/LLGo 152698 B +6493 B / +4.4% (worse) 74890 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 150427 B +5551 B / +3.8% (worse) 73165 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 139949 B +6102 B / +4.6% (worse) 78778 B 0 B / +0.0%
reflectcall/j32-emscripten/LLGo 1542074 B +7989 B / +0.5% (worse) 92056 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1543894 B +7064 B / +0.5% (worse) 90331 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1426872 B +7615 B / +0.5% (worse) 97789 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1546821 B +6153 B / +0.4% (worse) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1473288 B +7077 B / +0.5% (worse) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 146031 B +4861 B / +3.4% (worse) 0 B 0 B / 0.0%
w32-wasi/LLGo 146762 B +5773 B / +4.1% (worse) 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 5.799 s -56.45 ms / -1.0% (better)
j32-goos-js 6.358 s +556.3 ms / +9.6% (worse)
j64-emscripten-memory64 5.169 s -160.1 ms / -3.0% (better)
reflectcall/w32-wasi 24.904 s -1.233 s / -4.7% (better)
w32-goos-wasip1 4.685 s +127.5 ms / +2.8% (worse)
w32-wasi 4.583 s +39.07 ms / +0.9% (worse)

Compared with 07a0059ce2e6 measured in the same runner job.

@cpunion

cpunion commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator Author

Thanks for the review. This branch is now restacked on current main plus the latest #2661 and #2666; GitHub reports no merge conflict. #2661 fixes the blocking llgo run exit-code regression and the mutex wake-one/thundering-herd issue. Its 1 ms wait is documented as necessary for stop-the-world liveness. The GC arena comment is scoped to WASI threads. The Binaryen fork and pinned hashes are intentional.

Correction on the callback-poll flag: I initially added a defer reset, but repeat Chrome runs exposed address-zero corruption because callback polling can execute on a system Fiber without a G. Reverting that defer made the same module pass 15/15 Chrome runs. The explicit reset is safe on normal return; an abnormal unwind from this runtime hook terminates the worker rather than resuming polling. The code now documents why a Go defer is unsafe here.

The GC-root registry's linear registration/removal under churn is a valid performance concern. I am keeping the current duplicate-registration check and list representation in this functional GC PR; an intrusive O(1) registry should be benchmarked and tested separately because it changes shared root-enumeration invariants. The restacked WAMR threaded-GC acceptance passes locally. The full J32 Worker test passed on rerun after one finalizer timeout, and the focused finalizer tests each passed five repetitions. CI will rerun on the callback fix.

@cpunion

cpunion commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Addressed the source review of 884219e in ac4a1fd.

  • Confirmed the arena-capacity bug. The actual sizing/layout regression fails on the previous code, and the previous 32 MiB WAMR artifact does not complete within 120 seconds. The fix computes object + metadata + alignment requirements before comparing with the nominal arena size; changing > to >= would not cover requests just below the boundary.
  • Added five fresh-process WAMR cases from 32 MiB minus 128 KiB through 33 MiB, including the exact boundary and ±1 byte. All pass. Host tests additionally cover every start alignment, 16/32-byte blocks, allocator failure and integer overflow.
  • Disjoint WASI arena growth now uses the already-held allocator lock without a second STW request. Contiguous browser/embedded growth keeps its existing handshake. This removes the duplicate timeout on the allocation/growth path; later allocation attempts can still pay another collection deadline.
  • Added worker-private-root, thread-lifecycle/GC overlap and idle foreign-thread callback coverage. The expanded fixture passes with the CI --max-threads=8 setting. The foreign-thread test confirms the existing limitation: collection is skipped after a callback returns to an idle C thread and resumes after its destructor unregisters it.
  • Documented serial STW collection, the 500 ms skipped-collection behavior, foreign-thread registration lifetime and non-returned arenas in doc/wasm-threaded-gc.md.

The global-mutex safepoint fast path, 20 ms polling and repeated linear segment lookup observations apply to this PR's reviewed head; their optimizations are already in dependent #2695. The dependent PRs are only being rebased onto this fix. No parallel collector or native scheduler redesign is included.

Validation: go test ./internal/build -run '^TestWasmGC' -count=1; five isolated WAMR arena runs; expanded WAMR threaded-GC fixture (including an eight-thread run). Fresh CI is pending.

@cpunion

cpunion commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Fixed the two failure causes observed across the current PR stack in 63eb425 and ba903fd.

  • runtime/wasip1: add opt-in threaded GC for WAMR #2669's wasm-runtime (test-command) and fix(wasm): stabilize WAMR thread GC and EH boundaries #2695's wasm-runtime (runtime) exhausted their 45-minute budgets during Install dependencies, before any Wasm test ran. Both logs show the runner's APT mirror list selecting https:/us.archive.ubuntu.com/ubuntu and then stalling on another mirror. Setup now replaces the two existing Ubuntu mirror lists with the official archive/security endpoints (Ubuntu Ports on ARM), retains source suites/components/signing keys and unrelated repositories, applies IPv4/download timeout/retry settings, and bounds each complete APT update/install transaction to 5/10 minutes. Partial index updates are treated as errors. The workflow job budgets are unchanged.
  • feat(wasm): integrate browser filesystem with bounded workers #2696's Memory64 worker failure was TestStdinReadDoesNotFreezeScheduler: the timer ran after 167.7 ms, before the read finished at 447.7 ms, but exceeded the old absolute 150 ms assertion. Moved the already-reviewed fix from feat(wasm): make WAMR threads the default W32 profile #2697 into this shared ancestor: count forbidden Sync calls explicitly and require the timer to precede operation completion. The delayed async-operation check remains.

Validation: configuration regressions for amd64/arm64, absent mirror lists, preserved source/signature configuration and unrelated repositories; real apt-get update plus dependency resolution in an isolated Ubuntu 24.04 amd64 container seeded with the failing mirror list (27.4 seconds); 40 filesystem/timer test executions across Memory32/Memory64 × one/two workers; 10 additional Memory64/two-worker executions on Node 24 in #2696. All passed.

#2695, #2696 and #2697 are only rebased onto these common fixes. #2697's duplicate timer-test commit is removed by rebase. Its previous head had no failed checks at the last audit. Fresh CI is running/queued for the updated stack; full CI success is not yet confirmed.

@visualfc visualfc left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking follow-up: parkInitialWasiThread never acknowledges STW. Fine to merge as-is; the 500 ms skip path already covers uninstrumented C, and the current Goexit probes do not collect after main returns.

Comment thread runtime/internal/runtime/goexit_initial_wasi_threads.go
@cpunion
cpunion force-pushed the codex/wasi-threaded-gc branch from ba903fd to e702f81 Compare September 30, 2026 02:30
@cpunion
cpunion merged commit 653957f into xgo-dev:main Sep 30, 2026
90 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants