Skip to content

feat(wasm): make WAMR threads the default W32 profile - #2697

Merged
cpunion merged 6 commits into
xgo-dev:mainfrom
cpunion:codex/wasi-default-threads
Sep 30, 2026
Merged

cpunion merged 6 commits into
xgo-dev:mainfrom
cpunion:codex/wasi-default-threads

Conversation

@cpunion

@cpunion cpunion commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

W32 now uses the WAMR pthread runtime by default. llgo run -target W32 and raw GOOS=wasip1 GOARCH=wasm builds use the same shared-memory/threaded profile, without an opt-in environment variable. Explicitly disabling LLGO_WASI_THREADS reports a migration error. The obsolete single-threaded WASI scheduler, context switching and GC implementation are removed.

The target definitions, package test runners, DWARF checks, CI and execution documentation follow that default. Browser single-worker scheduling and native/embedded pthread behavior are unchanged.

Rebased onto main at 4d51d8e, including merged #2669, #2695, and #2696. There are no outstanding PR dependencies. This contribution contains only the W32 default, build, runner, DWARF, and test changes; all six contribution patches are preserved unchanged.

Validation completed locally:

  • Cross-compilation/profile, runner and WASI configuration unit tests.
  • The expanded WAMR threaded GC/Goexit/EH/filesystem/stdlib/GOROOT acceptance suite, with LLGO_WASI_THREADS unset.
  • W32 O0/O2 final DWARF validation and execution, including Go/C++ source coverage.
  • All target-profile checks, including named/raw WASI and browser targets.
  • The complete public test-command suite now passes: 10 checks across EC32, EC64, named WASI, raw GoJS and raw WASI, including compile-only/run paths. The default WAMR adapter explicitly forwards PWD/PATH and the reviewed stress profile, without inheriting unrelated host variables.

The GC/allocator fixes from #2695 now come from main. Its complete WAMR test/go run passes all 242 top-level tests, including the previously timing-out concurrent function-info lookup. Named/raw WASI public commands and W32 O0/O2 DWARF/runtime checks were re-run after integrating that fix and pass. EH validation is consolidated in #2695.

The full standard-library matrix has not been re-audited on every host at this head; see #2695 for the exact audit slices and results. Cross-platform CI remains the merge gate.

Review/CI follow-up: inherited the segment/safepoint/condition-wait fixes from #2695; updated stale Wasmtime assertions and reflect bridge source tags; routed direct LLGO_WASM_RUNTIME=iwasm runs through the same threaded host contract; corrected GOROOT/reference-runner documentation. DWARF source validation now uses each unit's actual filename index and prints the relevant table on failure. Added parser regressions and retained strict source-line resolution checks. After rebase, focused Go configuration/SSA/runner tests and all J32/J64/W32 O0/O2 DWARF/runtime combinations with the build cache enabled pass locally.

Final local verification at a7639a8: complete WAMR test/go passes all 242 top-level tests using the runner's absolute preopens/PWD contract; all 10 public CLI checks pass. Linux amd64 LLVM 22 also passes W32 O0/O2 final DWARF checks with the build cache enabled. The CI-only DWARF failure is now reproduced on both macOS and Linux amd64: when wasm-opt is present on PATH, Clang implicitly runs an unrequested Binaryen pass without debug preservation for W32 O2. The fix disables implicit passes for direct Wasm Clang links, including WASI threads. All six J32/J64/W32 O0/O2 checks pass with Binaryen on PATH after the fix, as does the Linux amd64 W32 O2 regression. The WAMR GC acceptance suite now runs both 20-goroutine startup shapes with separate 300-second execution budgets, preserving every test and repetition; Linux amd64 validation passes. The per-worker callback bridge fix is supplied by main via #2700, with Memory32/Memory64 two-worker output/filesystem regressions passing.

Earlier validation at main 9853dc8: host configuration/runner/SSA tests and Memory32/Memory64 callback checks with one and two workers pass. The previous CI failure in TestStdinReadDoesNotFreezeScheduler classified a 162 ms timer as blocking even though the read took 442 ms. The stdin/fsync checks now assert timer-before-completion ordering and count calls to the forbidden Sync API explicitly. Both checks pass five repetitions in each of the four configurations (40 executions).

Current rebase validation at main 4d51d8e: cross-compilation, target definitions, command runners, wasmstdlib, browser proxy, focused Wasm/WASI build and SSA tests, DWARF parser regressions, and shell syntax checks pass. With LLGo Binaryen llgo-v132.3 on PATH, W32 O0/O2 final DWARF verification, Go/C++ source resolution, and runtime execution pass. Named WASI and raw GOOS=wasip1 GOARCH=wasm both run the threaded-GC fixture successfully with LLGO_WASI_THREADS unset.

@cpunion
cpunion force-pushed the codex/wasi-default-threads branch from 50bab17 to 8652066 Compare September 29, 2026 02:53
@cpunion
cpunion marked this pull request as ready for review September 29, 2026 03:21

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: WASI threads as default for the WASM GC runtime

This is a large, carefully-constructed change. The concurrency design (epoch-based stop-the-world, root publishing before blocking C calls, cooperative safepoints) is coherent, and the build-tag matrix across the many _default/_wasi_threads file pairs is consistent and mutually exclusive. Memory-safety review of the C _wrap coordination layer found no defects, and the new crypto PEM files are legitimate test fixtures (not leaked secrets).

Findings below, ordered by impact. Inline comments carry the specific ones; a few cross-cutting items are summarized here.

Performance (main concern): segmentForBlock/segmentForAddress (segments.go) linearly scan up to heapSegmentCount segments on every call, and they now back the per-block state helpers (gcStateOf, gcSetState, gcMarkFree, ...) used in the allocator, mark, and sweep hot loops. gcStateOf scans twice and gcSetState three times per block. The mark/sweep loops already iterate segment-by-segment yet re-resolve the segment for each block, turning O(blocks) into ~O(blocks × segments). Since the segmented allocator grows by doubling arenas, GC pauses grow with arena count. Consider threading the known *heapSegment into the state helpers within the segment loops, fast-pathing heapSegmentCount == 1, and collapsing the redundant re-scans in gcStateOf/gcSetState. (See inline note on sweep.)

Polling cadence: condWait (inline) and the timer loop (time_heap_llgo.go / timerGCWaitQuantum) both fall back to a fixed 20ms wakeup so threads reach GC safepoints. The C layer already counts genuinely-blocked threads via world_blocked in llgo_wasi_gc_stop, so a thread parked in pthread_cond_timedwait is already counted toward quorum — a longer quantum (or on-demand wake) would cut idle-program CPU churn that scales with the number of blocked goroutines. At minimum, name the 20*1e6 literal and share it with timerGCWaitQuantum.

Maintainability: configureHeap (gc_tinygo.go) still writes heapSegments[0].end/.metadata/.last even though initGC now uses addHeapSegment; the two functions duplicate the same metadata-layout formula. Worth confirming configureHeap still has a live caller and, if so, having it delegate to addHeapSegment.

Docs: test/goroot/README.md:76 still says LLGO_WASI_THREADS=1 selects WAMR and mentions a "Wasmtime adapter" for W32 — but the harness now hardcodes runner: "iwasm" for W32-WASI unconditionally (no env gate). doc/wasm-wasi-threads.md's "official Go reference tests may still use Wasmtime" is likewise stale for the goroot comparison (only the separate wasmstdlib reference profile still uses wasmtime).

Findings without inline locations

  • internal/build/run.go:333: This direct LLGO_WASM_RUNTIME=iwasm run path still emits the old --stack-size=819200000 --heap-size=800000000 with no --max-threads, while the standardized WASIThreadedEmulator invocation (used by tests, targets/wasi.json, the workflow, goroot, and wasmstdlib) is now iwasm --max-threads=128 --stack-size=1048576 --heap-size=0 --dir=. --dir=/tmp. Since threads/shared memory are now the default module shape, this path may fail to spawn workers (no --max-threads). Align it with WASIThreadedEmulator or document the exemption.

Comment thread runtime/internal/runtime/tinygogc/segments.go Outdated
Comment thread runtime/internal/sync/sync_wait_wasi_threads_gc.go Outdated
Comment thread runtime/internal/runtime/tinygogc/gc_tinygo.go Outdated
Comment thread test/goroot/README.md Outdated
Comment thread runtime/internal/runtime/_wrap/wasi_gc_world.c Outdated
@codecov

codecov Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.36842% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
internal/build/build.go 91.66% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@cpunion
cpunion force-pushed the codex/wasi-default-threads branch from 9255796 to a7639a8 Compare September 29, 2026 04:53
@cpunion

cpunion commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Addressed the review and CI configuration findings. Rebased through #2695 onto current main and kept this PR ready for review.

The inherited collector update removes periodic condition/timer polling, reuses segment metadata and names the uncooperative-C timeout. configureHeap remains live for contiguous browser/embedded heap growth: it shares the layout helper, but cannot call addHeapSegment because that would erase live allocation metadata.

The direct LLGO_WASM_RUNTIME=iwasm path now uses the standard threaded runner, including thread/stack/heap limits, preopens and PWD/PATH handling; a command-level regression verifies it. GOROOT docs now consistently describe unconditional WAMR execution for both W32 artifacts. Fixed the obsolete Wasmtime assertion and the missing threads source tag in the reflection bridge test.

DWARF validation now resolves each unit's actual filename index and prints the relevant line table if resolution fails; parser tests retain rejection of missing source rows. All six J32/J64/W32 O0/O2 debug/runtime combinations pass locally with the build cache enabled, as do the rebased target/SSA/runner/profile tests. The earlier Linux O2 failure has not reproduced locally, so fresh Linux CI is still needed to confirm that check. No debug validation or coverage gate was disabled.

@cpunion
cpunion force-pushed the codex/wasi-default-threads branch from e2c3c3a to 49c8721 Compare September 29, 2026 05:39
@cpunion

cpunion commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

The W32 O2 DWARF failure is now reproduced and fixed. Clang implicitly runs wasm-opt when it finds it on PATH, without preserving debug lines. This only showed up in the CI environment because Binaryen was on PATH; setting WASMOPT alone did not reproduce it locally. Disable the implicit pass for direct Wasm Clang links, including the new default WASI-thread profile. The explicit Asyncify and Emscripten pipelines keep their existing behavior.

Verified the old failure and fixed behavior on both macOS and Linux amd64 with Binaryen on PATH. All six J32/J64/W32 O0/O2 final-DWARF/source-line/runtime checks pass, and linker argument regressions cover both Asyncify and non-Asyncify Clang links.

Rebased onto #2695 at fb2a438 to include its GC test-budget separation and per-worker callback bridge fix. Current head: 49c8721. Fresh CI is queued; all existing review threads are resolved and this PR remains ready for review.

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

79fd653245c8 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 7160 B 0 B / +0.0% 387 B 0 B / +0.0% 624.327 ms -32.19 ms / -4.9% (better) 1.371 ms +35.84 us / +2.7% (worse)
Linux cprintf-lto 6912 B 0 B / +0.0% 368 B 0 B / +0.0% 644.982 ms +46.48 ms / +7.8% (worse) 1.331 ms -18.05 us / -1.3% (better)
Linux fmtprintf 1663432 B +8 B / +0.0004809% (worse) 498437 B 0 B / +0.0% 4.318 s -88.24 ms / -2.0% (better) 3.359 ms +75.78 us / +2.3% (worse)
Linux fmtprintf-lto 1501168 B 0 B / +0.0% 436123 B 0 B / +0.0% 11.905 s +165.1 ms / +1.4% (worse) 3.066 ms -181.1 us / -5.6% (better)
Linux println 69248 B 0 B / +0.0% 16805 B 0 B / +0.0% 757.191 ms +146.4 ms / +24.0% (worse) 1.758 ms +82.27 us / +4.9% (worse)
Linux println-lto 59976 B 0 B / +0.0% 14209 B 0 B / +0.0% 925.768 ms +57.23 ms / +6.6% (worse) 1.823 ms +33.86 us / +1.9% (worse)
macOS cprintf 68064 B 0 B / +0.0% 4429 B 0 B / +0.0% 1.076 s +112.4 ms / +11.7% (worse) 4.221 ms +20.75 us / +0.5% (worse)
macOS cprintf-lto 84288 B 0 B / +0.0% 193 B 0 B / +0.0% 969.200 ms -8.204 ms / -0.8% (better) 2.800 ms -1.197 ms / -29.9% (better)
macOS fmtprintf 1504832 B 0 B / +0.0% 875484 B 0 B / +0.0% 4.039 s -166.7 ms / -4.0% (better) 5.221 ms -499.6 us / -8.7% (better)
macOS fmtprintf-lto 1192704 B 0 B / +0.0% 848980 B 0 B / +0.0% 12.506 s +2.746 s / +28.1% (worse) 9.925 ms +3.558 ms / +55.9% (worse)
macOS println 117216 B 0 B / +0.0% 37565 B 0 B / +0.0% 873.298 ms -254.1 ms / -22.5% (better) 4.679 ms -2.17 ms / -31.7% (better)
macOS println-lto 119472 B 0 B / +0.0% 34944 B 0 B / +0.0% 1.100 s -656 ms / -37.4% (better) 4.238 ms -3.921 ms / -48.1% (better)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.093 s -41.36 ms / -3.6% (better) 2.784 ms -514.8 us / -15.6% (better)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.121 s -30.09 ms / -2.6% (better) 2.763 ms -641.4 us / -18.8% (better)
Windows MinGW fmtprintf 1933312 B 0 B / +0.0% 599030 B 0 B / +0.0% 3.461 s -79.85 ms / -2.3% (better) 6.349 ms -885.1 us / -12.2% (better)
Windows MinGW fmtprintf-lto 1957888 B 0 B / +0.0% 547318 B 0 B / +0.0% 8.608 s +48.49 ms / +0.6% (worse) 7.149 ms +402.7 us / +6.0% (worse)
Windows MinGW println 75776 B 0 B / +0.0% 25142 B 0 B / +0.0% 1.100 s -24.86 ms / -2.2% (better) 5.473 ms -739 us / -11.9% (better)
Windows MinGW println-lto 69120 B 0 B / +0.0% 21990 B 0 B / +0.0% 1.285 s -24.48 ms / -1.9% (better) 5.551 ms +36.4 us / +0.7% (worse)
Windows MinGW 386 cprintf 43520 B 0 B / +0.0% 5326 B 0 B / +0.0% 954.757 ms -102 ms / -9.7% (better) 3.489 ms +116.2 us / +3.4% (worse)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.145 s +195.9 ms / +20.7% (worse) 4.623 ms +1.29 ms / +38.7% (worse)
Windows MinGW 386 fmtprintf 1896448 B 0 B / +0.0% 472462 B 0 B / +0.0% 3.327 s +221.5 ms / +7.1% (worse) 7.201 ms +442.2 us / +6.5% (worse)
Windows MinGW 386 fmtprintf-lto 2181120 B 0 B / +0.0% 451274 B 0 B / +0.0% 7.339 s -131.9 ms / -1.8% (better) 6.928 ms +501.2 us / +7.8% (worse)
Windows MinGW 386 println 96256 B 0 B / +0.0% 21474 B 0 B / +0.0% 1.290 s +201.4 ms / +18.5% (worse) 6.818 ms +1.046 ms / +18.1% (worse)
Windows MinGW 386 println-lto 74240 B 0 B / +0.0% 19334 B 0 B / +0.0% 1.517 s +53.6 ms / +3.7% (worse) 6.147 ms +858 us / +16.2% (worse)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.584 s -61.17 ms / -3.7% (better) 6.922 ms -43.3 us / -0.6% (better)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.625 s +20.62 ms / +1.3% (worse) 6.576 ms -277.8 us / -4.1% (better)
Windows MinGW ARM64 fmtprintf 1819648 B 0 B / +0.0% 509716 B 0 B / +0.0% 4.331 s +85.43 ms / +2.0% (worse) 13.215 ms +445.1 us / +3.5% (worse)
Windows MinGW ARM64 fmtprintf-lto 1880064 B 0 B / +0.0% 476176 B 0 B / +0.0% 9.856 s +255.2 ms / +2.7% (worse) 13.824 ms +744.3 us / +5.7% (worse)
Windows MinGW ARM64 println 72192 B 0 B / +0.0% 23884 B 0 B / +0.0% 1.588 s +40.34 ms / +2.6% (worse) 11.643 ms +150 us / +1.3% (worse)
Windows MinGW ARM64 println-lto 68608 B 0 B / +0.0% 21232 B 0 B / +0.0% 1.813 s +33.09 ms / +1.9% (worse) 11.326 ms +298.2 us / +2.7% (worse)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65798 B 0 B / +0.0% 1.130 s +3.117 ms / +0.3% (worse) 3.479 ms +113.5 us / +3.4% (worse)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65734 B 0 B / +0.0% 1.141 s -207.2 ms / -15.4% (better) 3.328 ms -149.8 us / -4.3% (better)
Windows MSVC fmtprintf 1643520 B 0 B / +0.0% 694582 B 0 B / +0.0% 4.030 s +133.5 ms / +3.4% (worse) 9.689 ms -169.2 us / -1.7% (better)
Windows MSVC fmtprintf-lto 1635328 B 0 B / +0.0% 647078 B 0 B / +0.0% 9.386 s -9.231 ms / -0.1% (better) 9.551 ms -347.4 us / -3.5% (better)
Windows MSVC println 194560 B 0 B / +0.0% 120822 B 0 B / +0.0% 1.150 s +32.56 ms / +2.9% (worse) 7.460 ms -766.9 us / -9.3% (better)
Windows MSVC println-lto 192512 B 0 B / +0.0% 118358 B 0 B / +0.0% 1.349 s -29.65 ms / -2.2% (better) 8.676 ms +689.4 us / +8.6% (worse)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.161 s -131.5 ms / -10.2% (better) 6.037 ms -921.6 us / -13.2% (better)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.191 s +51.43 ms / +4.5% (worse) 6.114 ms +393.7 us / +6.9% (worse)
Windows MSVC 386 fmtprintf 1204736 B 0 B / +0.0% 455852 B 0 B / +0.0% 3.958 s +25.92 ms / +0.7% (worse) 13.400 ms +1.423 ms / +11.9% (worse)
Windows MSVC 386 fmtprintf-lto 1241600 B 0 B / +0.0% 427211 B 0 B / +0.0% 8.951 s -314.4 ms / -3.4% (better) 12.549 ms -783.4 us / -5.9% (better)
Windows MSVC 386 println 36352 B 0 B / +0.0% 20324 B 0 B / +0.0% 1.148 s +4.702 ms / +0.4% (worse) 10.156 ms -464.6 us / -4.4% (better)
Windows MSVC 386 println-lto 35840 B 0 B / +0.0% 18565 B 0 B / +0.0% 1.360 s +7.148 ms / +0.5% (worse) 9.687 ms +49.1 us / +0.5% (worse)
Windows MSVC ARM64 cprintf 11776 B 0 B / +0.0% 4192 B 0 B / +0.0% 1.277 s +37.54 ms / +3.0% (worse) 7.176 ms +588.5 us / +8.9% (worse)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.284 s -3.321 ms / -0.3% (better) 7.260 ms -280.4 us / -3.7% (better)
Windows MSVC ARM64 fmtprintf 1387008 B 0 B / +0.0% 509656 B 0 B / +0.0% 3.848 s -50.61 ms / -1.3% (better) 15.051 ms -705.8 us / -4.5% (better)
Windows MSVC ARM64 fmtprintf-lto 1405440 B 0 B / +0.0% 476836 B 0 B / +0.0% 8.987 s -325.9 ms / -3.5% (better) 15.456 ms -935.7 us / -5.7% (better)
Windows MSVC ARM64 println 45056 B 0 B / +0.0% 23908 B 0 B / +0.0% 1.259 s -34.75 ms / -2.7% (better) 13.265 ms -289.1 us / -2.1% (better)
Windows MSVC ARM64 println-lto 42496 B 0 B / +0.0% 21380 B 0 B / +0.0% 1.441 s -37.2 ms / -2.5% (better) 13.598 ms -134.9 us / -1.0% (better)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.750 ns/op +0.02 ns/op / +0.1% (worse)
Linux BenchmarkMergeCompilerFlags 198.500 ns/op -2.5 ns/op / -1.2% (better)
Linux BenchmarkMergeLinkerFlags 136 ns/op +2.1 ns/op / +1.6% (worse)
Linux BenchmarkChannelBuffered 55.630 ns/op +0.26 ns/op / +0.5% (worse)
Linux BenchmarkChannelHandoff 12798 ns/op +358 ns/op / +2.9% (worse)
Linux BenchmarkDefer 49.770 ns/op +2.05 ns/op / +4.3% (worse)
Linux BenchmarkDirectCall 1.567 ns/op -0.021 ns/op / -1.3% (better)
Linux BenchmarkGlobalRead 1.200 ns/op +0.025 ns/op / +2.1% (worse)
Linux BenchmarkGlobalWrite 7.788 ns/op +0.023 ns/op / +0.3% (worse)
Linux BenchmarkGoroutine 25790 ns/op -754 ns/op / -2.8% (better)
Linux BenchmarkInterfaceCall 5.897 ns/op -0.002 ns/op / -0.0339% (better)
Linux BenchmarkRuntimeGetG 2.996 ns/op +0.007 ns/op / +0.2% (worse)
macOS BenchmarkLookupPCRandom 15.510 ns/op +2.74 ns/op / +21.5% (worse)
macOS BenchmarkMergeCompilerFlags 148.700 ns/op +4.7 ns/op / +3.3% (worse)
macOS BenchmarkMergeLinkerFlags 104.600 ns/op +12.96 ns/op / +14.1% (worse)
macOS BenchmarkChannelBuffered 30.950 ns/op -3.41 ns/op / -9.9% (better)
macOS BenchmarkChannelHandoff 7471 ns/op -758 ns/op / -9.2% (better)
macOS BenchmarkDefer 40.510 ns/op -13 ns/op / -24.3% (better)
macOS BenchmarkDirectCall 1.331 ns/op +0.022 ns/op / +1.7% (worse)
macOS BenchmarkGlobalRead 1.196 ns/op -0.052 ns/op / -4.2% (better)
macOS BenchmarkGlobalWrite 1.344 ns/op +0.088 ns/op / +7.0% (worse)
macOS BenchmarkGoroutine 91455 ns/op +51705 ns/op / +130.1% (worse)
macOS BenchmarkInterfaceCall 4.388 ns/op -0.141 ns/op / -3.1% (better)
macOS BenchmarkRuntimeGetG 3.770 ns/op -0.278 ns/op / -6.9% (better)
Windows MinGW BenchmarkLookupPCRandom 9.658 ns/op +0.084 ns/op / +0.9% (worse)
Windows MinGW BenchmarkMergeCompilerFlags 377.300 ns/op -4.8 ns/op / -1.3% (better)
Windows MinGW BenchmarkMergeLinkerFlags 354 ns/op +20 ns/op / +6.0% (worse)
Windows MinGW BenchmarkChannelBuffered 24.460 ns/op -0.54 ns/op / -2.2% (better)
Windows MinGW BenchmarkChannelHandoff 1190 ns/op +218.7 ns/op / +22.5% (worse)
Windows MinGW BenchmarkDefer 43.600 ns/op -0.43 ns/op / -1.0% (better)
Windows MinGW BenchmarkDirectCall 1.356 ns/op 0 ns/op / +0.0%
Windows MinGW BenchmarkGlobalRead 1.366 ns/op -0.108 ns/op / -7.3% (better)
Windows MinGW BenchmarkGlobalWrite 2.168 ns/op -0.001 ns/op / -0.0461% (better)
Windows MinGW BenchmarkGoroutine 63718 ns/op +799 ns/op / +1.3% (worse)
Windows MinGW BenchmarkInterfaceCall 6.580 ns/op -0.01 ns/op / -0.2% (better)
Windows MinGW BenchmarkRuntimeGetG 1.631 ns/op -0.001 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 15.810 ns/op +0.01 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkMergeCompilerFlags 1051 ns/op +59.6 ns/op / +6.0% (worse)
Windows MinGW 386 BenchmarkMergeLinkerFlags 863.300 ns/op -60.9 ns/op / -6.6% (better)
Windows MinGW 386 BenchmarkChannelBuffered 36.240 ns/op +0.24 ns/op / +0.7% (worse)
Windows MinGW 386 BenchmarkChannelHandoff 564.200 ns/op -12.7 ns/op / -2.2% (better)
Windows MinGW 386 BenchmarkDefer 27.800 ns/op +1.23 ns/op / +4.6% (worse)
Windows MinGW 386 BenchmarkDirectCall 0.934 ns/op +0.0144 ns/op / +1.6% (worse)
Windows MinGW 386 BenchmarkGlobalRead 0.928 ns/op +0.0009 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkGlobalWrite 8.238 ns/op -0.19 ns/op / -2.3% (better)
Windows MinGW 386 BenchmarkGoroutine 51364 ns/op -1748 ns/op / -3.3% (better)
Windows MinGW 386 BenchmarkInterfaceCall 4.521 ns/op -0.055 ns/op / -1.2% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 0.981 ns/op +0.0006 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkLookupPCRandom 11.980 ns/op -0.15 ns/op / -1.2% (better)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 577.700 ns/op +8.4 ns/op / +1.5% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 538.800 ns/op -3.3 ns/op / -0.6% (better)
Windows MinGW ARM64 BenchmarkChannelBuffered 39.390 ns/op +1.88 ns/op / +5.0% (worse)
Windows MinGW ARM64 BenchmarkChannelHandoff 2493 ns/op +44 ns/op / +1.8% (worse)
Windows MinGW ARM64 BenchmarkDefer 57.980 ns/op +0.57 ns/op / +1.0% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.589 ns/op -0.0001 ns/op / -0.01696% (better)
Windows MinGW ARM64 BenchmarkGlobalRead 0.663 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkGlobalWrite 0.663 ns/op -0.0004 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkGoroutine 62907 ns/op -1073 ns/op / -1.7% (better)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.141 ns/op -0.003 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.769 ns/op -0.036 ns/op / -2.0% (better)
Windows MSVC BenchmarkLookupPCRandom 13.100 ns/op -0.16 ns/op / -1.2% (better)
Windows MSVC BenchmarkMergeCompilerFlags 675 ns/op +68.1 ns/op / +11.2% (worse)
Windows MSVC BenchmarkMergeLinkerFlags 585.400 ns/op +38.6 ns/op / +7.1% (worse)
Windows MSVC BenchmarkChannelBuffered 27.900 ns/op +0.01 ns/op / +0.03586% (worse)
Windows MSVC BenchmarkChannelHandoff 1094 ns/op +51 ns/op / +4.9% (worse)
Windows MSVC BenchmarkDefer 55.180 ns/op -0.23 ns/op / -0.4% (better)
Windows MSVC BenchmarkDirectCall 1.545 ns/op -0.004 ns/op / -0.3% (better)
Windows MSVC BenchmarkGlobalRead 1.549 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC BenchmarkGlobalWrite 2.470 ns/op -0.003 ns/op / -0.1% (better)
Windows MSVC BenchmarkGoroutine 87941 ns/op -1391 ns/op / -1.6% (better)
Windows MSVC BenchmarkInterfaceCall 8.384 ns/op +0.004 ns/op / +0.04773% (worse)
Windows MSVC BenchmarkRuntimeGetG 2.173 ns/op +0.004 ns/op / +0.2% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 26.440 ns/op -0.11 ns/op / -0.4% (better)
Windows MSVC 386 BenchmarkMergeCompilerFlags 732.100 ns/op -23.2 ns/op / -3.1% (better)
Windows MSVC 386 BenchmarkMergeLinkerFlags 690.400 ns/op +3.1 ns/op / +0.5% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 39.160 ns/op -0.05 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkChannelHandoff 827.700 ns/op -155.2 ns/op / -15.8% (better)
Windows MSVC 386 BenchmarkDefer 45.570 ns/op -0.67 ns/op / -1.4% (better)
Windows MSVC 386 BenchmarkDirectCall 1.547 ns/op -0.002 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGlobalRead 1.862 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGlobalWrite 7.775 ns/op -0.005 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGoroutine 108351 ns/op +1496 ns/op / +1.4% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 8.357 ns/op -0.009 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkRuntimeGetG 1.928 ns/op -0.274 ns/op / -12.4% (better)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.150 ns/op +0.03 ns/op / +0.2% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 582.800 ns/op +6.9 ns/op / +1.2% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 538.900 ns/op -7.6 ns/op / -1.4% (better)
Windows MSVC ARM64 BenchmarkChannelBuffered 38.890 ns/op +0.02 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkChannelHandoff 3038 ns/op -126 ns/op / -4.0% (better)
Windows MSVC ARM64 BenchmarkDefer 62.710 ns/op +0.05 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkDirectCall 0.590 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 BenchmarkGlobalRead 0.664 ns/op +0.0005 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.762 ns/op +0.017 ns/op / +0.5% (worse)
Windows MSVC ARM64 BenchmarkGoroutine 60438 ns/op +937 ns/op / +1.6% (worse)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.139 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.802 ns/op +0.001 ns/op / +0.1% (worse)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 915.800 ns/op -57.4 ns/op / -5.9% (better)
Linux AfterFuncZeroDelivery/LLGo 41492 ns/op +2168 ns/op / +5.5% (worse)
Linux CreateStop/Go 314.800 ns/op +12.9 ns/op / +4.3% (worse)
Linux CreateStop/LLGo 1913 ns/op +105 ns/op / +5.8% (worse)
Linux RearmStopped/Go 116 ns/op +0.1 ns/op / +0.1% (worse)
Linux RearmStopped/LLGo 1330 ns/op -70 ns/op / -5.0% (better)
Linux ResetActive/Go 68.720 ns/op +0.17 ns/op / +0.2% (worse)
Linux ResetActive/LLGo 804.700 ns/op +32.3 ns/op / +4.2% (worse)
Linux ResetHeap1024/Go 67.700 ns/op +0.63 ns/op / +0.9% (worse)
Linux ResetHeap1024/LLGo 179.400 ns/op +1.6 ns/op / +0.9% (worse)
macOS AfterFuncZeroDelivery/Go 504.100 ns/op -89.9 ns/op / -15.1% (better)
macOS AfterFuncZeroDelivery/LLGo 81659 ns/op +9971 ns/op / +13.9% (worse)
macOS CreateStop/Go 186.200 ns/op +29.8 ns/op / +19.1% (worse)
macOS CreateStop/LLGo 512.500 ns/op -185 ns/op / -26.5% (better)
macOS RearmStopped/Go 64.780 ns/op -9.83 ns/op / -13.2% (better)
macOS RearmStopped/LLGo 363.300 ns/op -359.4 ns/op / -49.7% (better)
macOS ResetActive/Go 45.700 ns/op -1.28 ns/op / -2.7% (better)
macOS ResetActive/LLGo 168.100 ns/op -116.6 ns/op / -41.0% (better)
macOS ResetHeap1024/Go 46.130 ns/op -4.47 ns/op / -8.8% (better)
macOS ResetHeap1024/LLGo 90.160 ns/op -47.64 ns/op / -34.6% (better)
Windows MinGW AfterFuncZeroDelivery/Go 379.800 ns/op -9.2 ns/op / -2.4% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 112110 ns/op -450 ns/op / -0.4% (better)
Windows MinGW CreateStop/Go 91.520 ns/op +1.18 ns/op / +1.3% (worse)
Windows MinGW CreateStop/LLGo 340.900 ns/op -11.2 ns/op / -3.2% (better)
Windows MinGW RearmStopped/Go 24.430 ns/op -0.04 ns/op / -0.2% (better)
Windows MinGW RearmStopped/LLGo 221.300 ns/op +1.7 ns/op / +0.8% (worse)
Windows MinGW ResetActive/Go 14.830 ns/op +0.03 ns/op / +0.2% (worse)
Windows MinGW ResetActive/LLGo 128.500 ns/op +9.8 ns/op / +8.3% (worse)
Windows MinGW ResetHeap1024/Go 15.030 ns/op +0.04 ns/op / +0.3% (worse)
Windows MinGW ResetHeap1024/LLGo 108.600 ns/op -1.2 ns/op / -1.1% (better)
Windows MinGW 386 AfterFuncZeroDelivery/Go 595.900 ns/op -15.2 ns/op / -2.5% (better)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 101943 ns/op -3749 ns/op / -3.5% (better)
Windows MinGW 386 CreateStop/Go 160.700 ns/op -0.2 ns/op / -0.1% (better)
Windows MinGW 386 CreateStop/LLGo 418.200 ns/op +21.7 ns/op / +5.5% (worse)
Windows MinGW 386 RearmStopped/Go 61.170 ns/op -0.37 ns/op / -0.6% (better)
Windows MinGW 386 RearmStopped/LLGo 250 ns/op -6.9 ns/op / -2.7% (better)
Windows MinGW 386 ResetActive/Go 28.960 ns/op -0.36 ns/op / -1.2% (better)
Windows MinGW 386 ResetActive/LLGo 725.100 ns/op +145 ns/op / +25.0% (worse)
Windows MinGW 386 ResetHeap1024/Go 29.190 ns/op -0.24 ns/op / -0.8% (better)
Windows MinGW 386 ResetHeap1024/LLGo 120.200 ns/op +1.2 ns/op / +1.0% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 651.700 ns/op -10.3 ns/op / -1.6% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 151020 ns/op +4511 ns/op / +3.1% (worse)
Windows MinGW ARM64 CreateStop/Go 191.600 ns/op +0.4 ns/op / +0.2% (worse)
Windows MinGW ARM64 CreateStop/LLGo 368.600 ns/op +3.7 ns/op / +1.0% (worse)
Windows MinGW ARM64 RearmStopped/Go 70.650 ns/op +0.03 ns/op / +0.04248% (worse)
Windows MinGW ARM64 RearmStopped/LLGo 256.300 ns/op +2.5 ns/op / +1.0% (worse)
Windows MinGW ARM64 ResetActive/Go 31.090 ns/op +0.09 ns/op / +0.3% (worse)
Windows MinGW ARM64 ResetActive/LLGo 125.600 ns/op -2.1 ns/op / -1.6% (better)
Windows MinGW ARM64 ResetHeap1024/Go 31.140 ns/op -0.05 ns/op / -0.2% (better)
Windows MinGW ARM64 ResetHeap1024/LLGo 126.400 ns/op +0.3 ns/op / +0.2% (worse)
Windows MSVC AfterFuncZeroDelivery/Go 557.300 ns/op -5.1 ns/op / -0.9% (better)
Windows MSVC AfterFuncZeroDelivery/LLGo 176592 ns/op +3615 ns/op / +2.1% (worse)
Windows MSVC CreateStop/Go 115.800 ns/op 0 ns/op / +0.0%
Windows MSVC CreateStop/LLGo 418 ns/op -12.6 ns/op / -2.9% (better)
Windows MSVC RearmStopped/Go 31.380 ns/op -0.24 ns/op / -0.8% (better)
Windows MSVC RearmStopped/LLGo 254.100 ns/op +2.1 ns/op / +0.8% (worse)
Windows MSVC ResetActive/Go 20.100 ns/op -0.04 ns/op / -0.2% (better)
Windows MSVC ResetActive/LLGo 156.400 ns/op +5 ns/op / +3.3% (worse)
Windows MSVC ResetHeap1024/Go 20.330 ns/op -0.25 ns/op / -1.2% (better)
Windows MSVC ResetHeap1024/LLGo 125.800 ns/op +0.9 ns/op / +0.7% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/Go 942.900 ns/op -3.3 ns/op / -0.3% (better)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 191650 ns/op -341 ns/op / -0.2% (better)
Windows MSVC 386 CreateStop/Go 193.900 ns/op -1.5 ns/op / -0.8% (better)
Windows MSVC 386 CreateStop/LLGo 443.700 ns/op -26 ns/op / -5.5% (better)
Windows MSVC 386 RearmStopped/Go 63.300 ns/op -0.02 ns/op / -0.03159% (better)
Windows MSVC 386 RearmStopped/LLGo 310.100 ns/op -11.2 ns/op / -3.5% (better)
Windows MSVC 386 ResetActive/Go 38.950 ns/op +0.08 ns/op / +0.2% (worse)
Windows MSVC 386 ResetActive/LLGo 1009 ns/op +117.1 ns/op / +13.1% (worse)
Windows MSVC 386 ResetHeap1024/Go 39.400 ns/op +0.01 ns/op / +0.02539% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 171.100 ns/op +0.3 ns/op / +0.2% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 670 ns/op +1.6 ns/op / +0.2% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 150131 ns/op +902 ns/op / +0.6% (worse)
Windows MSVC ARM64 CreateStop/Go 216.300 ns/op +16.9 ns/op / +8.5% (worse)
Windows MSVC ARM64 CreateStop/LLGo 407.700 ns/op +23.6 ns/op / +6.1% (worse)
Windows MSVC ARM64 RearmStopped/Go 70.540 ns/op -0.04 ns/op / -0.1% (better)
Windows MSVC ARM64 RearmStopped/LLGo 274.200 ns/op +1.9 ns/op / +0.7% (worse)
Windows MSVC ARM64 ResetActive/Go 31.040 ns/op +0.05 ns/op / +0.2% (worse)
Windows MSVC ARM64 ResetActive/LLGo 129.500 ns/op -5.6 ns/op / -4.1% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.120 ns/op +0.01 ns/op / +0.03214% (worse)
Windows MSVC ARM64 ResetHeap1024/LLGo 137.800 ns/op +0.7 ns/op / +0.5% (worse)

Compared with 4d51d8e95de5 measured in the same runner job.

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

79fd653245c8 | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 154193 B 0 B / +0.0% 88742 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 152295 B 0 B / +0.0% 73165 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 141248 B 0 B / +0.0% 92630 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 152848 B +5159 B / +3.5% (worse) 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 153560 B +5763 B / +3.9% (worse) 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3213428 B 0 B / +0.0% 132442 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3192523 B 0 B / +0.0% 101635 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2950033 B 0 B / +0.0% 139285 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2347351 B -496992 B / -17.5% (better) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2344013 B -367118 B / -13.5% (better) 0 B 0 B / 0.0%
j32-emscripten/LLGo 153428 B 0 B / +0.0% 88742 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 151768 B 0 B / +0.0% 73165 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 140582 B 0 B / +0.0% 92630 B 0 B / +0.0%
reflectcall/j32-emscripten/LLGo 1543771 B 0 B / +0.0% 105908 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1546222 B 0 B / +0.0% 90331 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1428426 B 0 B / +0.0% 111641 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1276693 B -272122 B / -17.6% (better) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1274301 B -200398 B / -13.6% (better) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 152495 B +5472 B / +3.7% (worse) 0 B 0 B / 0.0%
w32-wasi/LLGo 153207 B +6010 B / +4.1% (worse) 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 6.052 s -76.18 ms / -1.2% (better)
j32-goos-js 5.943 s +44.21 ms / +0.7% (worse)
j64-emscripten-memory64 5.260 s -294.2 ms / -5.3% (better)
reflectcall/w32-wasi 22.550 s -2.772 s / -10.9% (better)
w32-goos-wasip1 4.176 s -585.3 ms / -12.3% (better)
w32-wasi 3.971 s -839.5 ms / -17.5% (better)

Compared with 4d51d8e95de5 measured in the same runner job.

@cpunion
cpunion force-pushed the codex/wasi-default-threads branch from 49c8721 to 924370c Compare September 29, 2026 09:29
@cpunion
cpunion force-pushed the codex/wasi-default-threads branch 2 times, most recently from 8dc7714 to a46e23d Compare September 29, 2026 11:34
@cpunion
cpunion force-pushed the codex/wasi-default-threads branch from a46e23d to 79fd653 Compare September 30, 2026 11:14
@cpunion
cpunion merged commit f06143b into xgo-dev:main Sep 30, 2026
91 of 93 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants