Skip to content

runtime/wasip1: discover current pthread stack top for GC - #2666

Merged
cpunion merged 5 commits into
xgo-dev:mainfrom
cpunion:codex/wasi-threaded-stack-bounds
Sep 28, 2026
Merged

cpunion merged 5 commits into
xgo-dev:mainfrom
cpunion:codex/wasi-threaded-stack-bounds

Conversation

@cpunion

@cpunion cpunion commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

The linear-memory collector uses __stack_high as the stack top on WASI. That symbol describes the main thread, so collecting from a WASI pthread would scan the wrong stack and could reclaim live objects. Resolve worker stack bounds through WASI libc; retain __stack_high for the main thread, for which pthread_getattr_np returns ENOSYS. Trap if another stack-bound lookup fails.

The WAMR nogc probe compiles and calls the collector helper on both the main and worker threads. It now requires pthread stack discovery on workers, so a worker cannot pass by falling back to the main-thread stack top. The nogc gate remains until thread root registration, stop-the-world coordination, and allocator locking in #2669 are complete. This PR builds on merged #2653.

Validation: focused WASI build tests and runtime compilation pass; the WAMR probe completed on macOS, although repeated runs with the local WAMR 2.4.5 binary were intermittent both before and after this probe change. The earlier Windows MinGW TLS timeout passed on rerun. A later run canceled the Linux Dev LTO demo after its one-hour limit while downloading the default PyTorch CUDA dependencies; the demo needs only CPU tensors, so Linux now installs PyTorch from the official CPU-only index. The CPU-only setup and Dev LTO demos passed on the next CI run. A separate README link check received a transient GitHub 503 while following a file URL written with /tree/; it now links directly to /blob/, and local lychee checked all 96 README links successfully. On the latest head, Dev LTO GlobalDCE and Go Method Drop and doc_verify both pass; the remaining matrix jobs are still running.

@cpunion
cpunion force-pushed the codex/wasi-threaded-stack-bounds branch from 619a781 to c97d9e4 Compare September 24, 2026 06:19

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review summary

Solid, well-scoped experimental WASI-threads enablement. The build-tag partitioning across the nanotime_*/time_*/sema_* files is exhaustive and non-overlapping, the per-thread GC stack-bound helper (gc_wasm.c) has rigorous overflow/NULL guards and traps cleanly on failure, and the -tags nogc gate with its actionable error message is a sensible guardrail until a threaded collector exists. The Realloc change correctly moves the metadata read inside the lock (an improvement) and fixes the zeroSizedAlloc sentinel path so realloc(malloc(0), n) no longer walks block metadata from the sentinel. The new iwasm cache-id keying and the WASI-threads probe test are good hardening. Findings below are minor/non-blocking.

Note on Realloc (gc_tinygo.go:388): the Memcpy/freeObject after unlock remains a copy-after-unlock window, but this PR does not regress it (the unlock actually moved later than before) and the inline comment correctly flags it as future work for the not-yet-existent threaded collector. No action needed now; worth a tracking note when concurrent collection lands.

Findings without inline locations

  • runtime/internal/lib/runtime/time_wait_unix_llgo.go:20: Misleading panic message on the WASI-threads path. This file now compiles for wasip1 && llgo.wasi_threads, where c_timerCondInit routes to llgo_timer_cond_init, which deliberately initializes the condition with the default clock (not CLOCK_MONOTONIC) because WAMR 2.4.5 workers can't read CLOCK_MONOTONIC — llgo_timer_cond_timedwait uses CLOCK_REALTIME there. Calling it a "monotonic timer condition" is inaccurate for this build. Consider "failed to initialize timer condition variable" so the message stays correct across clock domains.

// CLOCK_MONOTONIC there. Use one clock domain on every M so deadlines
// created on one thread can be consumed by the timer thread.
if ct.ClockGettime(ct.CLOCK_REALTIME, tv) != 0 {
return atomic.LoadInt64(&wasiThreadNanoLast)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On the first-ever call, if ClockGettime(CLOCK_REALTIME) fails, wasiThreadNanoLast is still 0 and this returns 0 as a legitimate timestamp — any deadline computed against it would be wildly wrong. This mirrors the existing single-worker fallback (which also returns 0), so it's not a regression, but a CLOCK_REALTIME read failure on a live worker is unexpected enough that trapping (like the stack-bound helper does) may be preferable to silently returning epoch-relative garbage. Non-blocking.

if atomic.CompareAndSwapInt64(&wasiThreadNanoLast, last, now) {
return now
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor: this publishes to the shared wasiThreadNanoLast via CAS on every forward clock tick, making that global a cross-core cache-line hotspot under multi-thread load (nanotime is called from the scheduler, timers, and sema spin paths). The fast path (now <= last) is correctly a load-only, no-store return. If contention shows up in profiling, consider a single unconditional CAS-and-return-now when now > last (the loop's retry is only needed to correct an observed backward jump). Correct as-is; flagging for later.

@codecov

codecov Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

864f6cec2af7 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 7160 B 0 B / +0.0% 387 B 0 B / +0.0% 501.368 ms -85.27 ms / -14.5% (better) 1.248 ms -6.171 ms / -83.2% (better)
Linux cprintf-lto 6912 B 0 B / +0.0% 368 B 0 B / +0.0% 542.064 ms -41.08 ms / -7.0% (better) 1.273 ms -57.39 us / -4.3% (better)
Linux fmtprintf 1662208 B 0 B / +0.0% 498331 B 0 B / +0.0% 3.911 s -225.5 ms / -5.5% (better) 3.117 ms +45.39 us / +1.5% (worse)
Linux fmtprintf-lto 1500368 B 0 B / +0.0% 436107 B 0 B / +0.0% 11.539 s -63.29 ms / -0.5% (better) 3.154 ms +90.19 us / +2.9% (worse)
Linux println 68992 B 0 B / +0.0% 16783 B 0 B / +0.0% 521.804 ms -75.73 ms / -12.7% (better) 1.664 ms -8.874 us / -0.5% (better)
Linux println-lto 59848 B 0 B / +0.0% 14199 B 0 B / +0.0% 777.599 ms -83.52 ms / -9.7% (better) 1.659 ms -9.537 us / -0.6% (better)
macOS cprintf 68064 B 0 B / +0.0% 4429 B 0 B / +0.0% 877.056 ms -40.3 ms / -4.4% (better) 3.453 ms -532.2 us / -13.4% (better)
macOS cprintf-lto 84288 B 0 B / +0.0% 193 B 0 B / +0.0% 983.492 ms -63.49 ms / -6.1% (better) 4.595 ms -78.08 us / -1.7% (better)
macOS fmtprintf 1504672 B 0 B / +0.0% 874360 B 0 B / +0.0% 4.018 s +306.2 ms / +8.2% (worse) 12.053 ms +6.306 ms / +109.7% (worse)
macOS fmtprintf-lto 1192704 B 0 B / +0.0% 848024 B 0 B / +0.0% 9.029 s -340.7 ms / -3.6% (better) 9.221 ms +3.288 ms / +55.4% (worse)
macOS println 117136 B 0 B / +0.0% 37421 B 0 B / +0.0% 980.991 ms -172.6 ms / -15.0% (better) 8.077 ms +230 us / +2.9% (worse)
macOS println-lto 119472 B 0 B / +0.0% 34824 B 0 B / +0.0% 1.320 s +284.2 ms / +27.4% (worse) 4.399 ms +778.3 us / +21.5% (worse)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.316 s -23.68 ms / -1.8% (better) 3.420 ms -299.2 us / -8.0% (better)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.441 s +84.72 ms / +6.2% (worse) 3.812 ms +229.6 us / +6.4% (worse)
Windows MinGW fmtprintf 1932800 B 0 B / +0.0% 598918 B 0 B / +0.0% 4.746 s +313.7 ms / +7.1% (worse) 9.248 ms +251.2 us / +2.8% (worse)
Windows MinGW fmtprintf-lto 1957376 B 0 B / +0.0% 547318 B 0 B / +0.0% 11.411 s +949.5 ms / +9.1% (worse) 10.166 ms +881.3 us / +9.5% (worse)
Windows MinGW println 75776 B 0 B / +0.0% 25142 B 0 B / +0.0% 1.416 s +46.79 ms / +3.4% (worse) 8.215 ms +921 us / +12.6% (worse)
Windows MinGW println-lto 69120 B 0 B / +0.0% 21990 B 0 B / +0.0% 1.686 s +95.89 ms / +6.0% (worse) 7.171 ms -180 us / -2.4% (better)
Windows MinGW 386 cprintf 43520 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.363 s +8.614 ms / +0.6% (worse) 6.664 ms +1.07 ms / +19.1% (worse)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.523 s +129.1 ms / +9.3% (worse) 6.285 ms +560.5 us / +9.8% (worse)
Windows MinGW 386 fmtprintf 1895936 B 0 B / +0.0% 472398 B 0 B / +0.0% 4.793 s +129.6 ms / +2.8% (worse) 12.521 ms +842.8 us / +7.2% (worse)
Windows MinGW 386 fmtprintf-lto 2180608 B 0 B / +0.0% 451258 B 0 B / +0.0% 11.057 s +93.47 ms / +0.9% (worse) 12.685 ms +282.3 us / +2.3% (worse)
Windows MinGW 386 println 96256 B 0 B / +0.0% 21458 B 0 B / +0.0% 1.642 s +251.3 ms / +18.1% (worse) 12.582 ms +2.743 ms / +27.9% (worse)
Windows MinGW 386 println-lto 74240 B 0 B / +0.0% 19314 B 0 B / +0.0% 1.716 s +83.48 ms / +5.1% (worse) 11.767 ms +1.504 ms / +14.7% (worse)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.514 s +18.91 ms / +1.3% (worse) 6.092 ms -28 us / -0.5% (better)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.548 s +24.46 ms / +1.6% (worse) 6.114 ms -335.7 us / -5.2% (better)
Windows MinGW ARM64 fmtprintf 1819136 B 0 B / +0.0% 509584 B 0 B / +0.0% 4.387 s +72.23 ms / +1.7% (worse) 12.743 ms +303.6 us / +2.4% (worse)
Windows MinGW ARM64 fmtprintf-lto 1879040 B 0 B / +0.0% 476152 B 0 B / +0.0% 9.633 s -40.18 ms / -0.4% (better) 12.332 ms -304.8 us / -2.4% (better)
Windows MinGW ARM64 println 72192 B 0 B / +0.0% 23876 B 0 B / +0.0% 1.510 s -29.32 ms / -1.9% (better) 10.711 ms -37.8 us / -0.4% (better)
Windows MinGW ARM64 println-lto 68608 B 0 B / +0.0% 21224 B 0 B / +0.0% 1.730 s -10.98 ms / -0.6% (better) 10.561 ms -498.7 us / -4.5% (better)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65798 B 0 B / +0.0% 1.185 s +73.29 ms / +6.6% (worse) 3.616 ms -549 us / -13.2% (better)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65734 B 0 B / +0.0% 1.103 s -1.37 ms / -0.1% (better) 3.453 ms -230.9 us / -6.3% (better)
Windows MSVC fmtprintf 1643008 B 0 B / +0.0% 694470 B 0 B / +0.0% 3.923 s -109.4 ms / -2.7% (better) 9.503 ms +305.1 us / +3.3% (worse)
Windows MSVC fmtprintf-lto 1634816 B 0 B / +0.0% 647062 B 0 B / +0.0% 9.328 s +96.7 ms / +1.0% (worse) 10.690 ms +1.103 ms / +11.5% (worse)
Windows MSVC println 194560 B 0 B / +0.0% 120822 B 0 B / +0.0% 1.075 s -30.84 ms / -2.8% (better) 7.593 ms -383.1 us / -4.8% (better)
Windows MSVC println-lto 192512 B 0 B / +0.0% 118342 B 0 B / +0.0% 1.267 s -254.6 ms / -16.7% (better) 7.546 ms -226.3 us / -2.9% (better)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 845.886 ms +30.49 ms / +3.7% (worse) 5.159 ms +71 us / +1.4% (worse)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 859.835 ms +11.3 ms / +1.3% (worse) 4.940 ms +90.1 us / +1.9% (worse)
Windows MSVC 386 fmtprintf 1204224 B 0 B / +0.0% 455788 B 0 B / +0.0% 3.297 s +43.42 ms / +1.3% (worse) 9.873 ms -609.5 us / -5.8% (better)
Windows MSVC 386 fmtprintf-lto 1240576 B 0 B / +0.0% 427195 B 0 B / +0.0% 7.216 s -341.8 ms / -4.5% (better) 9.947 ms -85 us / -0.8% (better)
Windows MSVC 386 println 36352 B 0 B / +0.0% 20324 B 0 B / +0.0% 874.438 ms +26 ms / +3.1% (worse) 8.718 ms +685 us / +8.5% (worse)
Windows MSVC 386 println-lto 35840 B 0 B / +0.0% 18549 B 0 B / +0.0% 1.011 s +8.559 ms / +0.9% (worse) 8.164 ms -294.1 us / -3.5% (better)
Windows MSVC ARM64 cprintf 11776 B 0 B / +0.0% 4192 B 0 B / +0.0% 1.332 s -57.76 ms / -4.2% (better) 8.025 ms -409.9 us / -4.9% (better)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.362 s +74.83 ms / +5.8% (worse) 7.884 ms +904.6 us / +13.0% (worse)
Windows MSVC ARM64 fmtprintf 1386496 B 0 B / +0.0% 509528 B 0 B / +0.0% 4.096 s -100.6 ms / -2.4% (better) 17.080 ms +428.3 us / +2.6% (worse)
Windows MSVC ARM64 fmtprintf-lto 1404928 B 0 B / +0.0% 476820 B 0 B / +0.0% 9.538 s +308.3 ms / +3.3% (worse) 18.291 ms +2.253 ms / +14.0% (worse)
Windows MSVC ARM64 println 45056 B 0 B / +0.0% 23908 B 0 B / +0.0% 1.387 s +103.9 ms / +8.1% (worse) 15.102 ms +1.753 ms / +13.1% (worse)
Windows MSVC ARM64 println-lto 42496 B 0 B / +0.0% 21380 B 0 B / +0.0% 1.516 s +1.213 ms / +0.1% (worse) 14.083 ms +455.6 us / +3.3% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.560 ns/op +0.01 ns/op / +0.1% (worse)
Linux BenchmarkMergeCompilerFlags 206.800 ns/op -2.5 ns/op / -1.2% (better)
Linux BenchmarkMergeLinkerFlags 128.300 ns/op +2.3 ns/op / +1.8% (worse)
Linux BenchmarkChannelBuffered 55.020 ns/op +0.06 ns/op / +0.1% (worse)
Linux BenchmarkChannelHandoff 16840 ns/op +2750 ns/op / +19.5% (worse)
Linux BenchmarkDefer 49.020 ns/op -0.8 ns/op / -1.6% (better)
Linux BenchmarkDirectCall 1.172 ns/op -0.007 ns/op / -0.6% (better)
Linux BenchmarkGlobalRead 1.181 ns/op +0.016 ns/op / +1.4% (worse)
Linux BenchmarkGlobalWrite 7.774 ns/op +0.01 ns/op / +0.1% (worse)
Linux BenchmarkGoroutine 33651 ns/op +10207 ns/op / +43.5% (worse)
Linux BenchmarkInterfaceCall 5.841 ns/op -0.085 ns/op / -1.4% (better)
Linux BenchmarkRuntimeGetG 2.353 ns/op -0.013 ns/op / -0.5% (better)
macOS BenchmarkLookupPCRandom 12.450 ns/op -6.49 ns/op / -34.3% (better)
macOS BenchmarkMergeCompilerFlags 114.900 ns/op -30.1 ns/op / -20.8% (better)
macOS BenchmarkMergeLinkerFlags 75.460 ns/op -17.09 ns/op / -18.5% (better)
macOS BenchmarkChannelBuffered 34.730 ns/op -0.06 ns/op / -0.2% (better)
macOS BenchmarkChannelHandoff 5740 ns/op -3357 ns/op / -36.9% (better)
macOS BenchmarkDefer 45.270 ns/op +5.49 ns/op / +13.8% (worse)
macOS BenchmarkDirectCall 1.214 ns/op -0.134 ns/op / -9.9% (better)
macOS BenchmarkGlobalRead 1.314 ns/op -0.054 ns/op / -3.9% (better)
macOS BenchmarkGlobalWrite 1.437 ns/op -0.32 ns/op / -18.2% (better)
macOS BenchmarkGoroutine 39337 ns/op -12161 ns/op / -23.6% (better)
macOS BenchmarkInterfaceCall 4.254 ns/op -0.521 ns/op / -10.9% (better)
macOS BenchmarkRuntimeGetG 2.393 ns/op -0.597 ns/op / -20.0% (better)
Windows MinGW BenchmarkLookupPCRandom 13.010 ns/op -0.17 ns/op / -1.3% (better)
Windows MinGW BenchmarkMergeCompilerFlags 633.800 ns/op +27.4 ns/op / +4.5% (worse)
Windows MinGW BenchmarkMergeLinkerFlags 570.600 ns/op +19.7 ns/op / +3.6% (worse)
Windows MinGW BenchmarkChannelBuffered 31.060 ns/op -0.25 ns/op / -0.8% (better)
Windows MinGW BenchmarkChannelHandoff 930.800 ns/op -35.5 ns/op / -3.7% (better)
Windows MinGW BenchmarkDefer 59.590 ns/op +1 ns/op / +1.7% (worse)
Windows MinGW BenchmarkDirectCall 1.550 ns/op +0.002 ns/op / +0.1% (worse)
Windows MinGW BenchmarkGlobalRead 1.882 ns/op +0.017 ns/op / +0.9% (worse)
Windows MinGW BenchmarkGlobalWrite 2.465 ns/op -0.004 ns/op / -0.2% (better)
Windows MinGW BenchmarkGoroutine 84088 ns/op -6320 ns/op / -7.0% (better)
Windows MinGW BenchmarkInterfaceCall 8.712 ns/op -0.011 ns/op / -0.1% (better)
Windows MinGW BenchmarkRuntimeGetG 1.864 ns/op -0.001 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 26.640 ns/op -0.15 ns/op / -0.6% (better)
Windows MinGW 386 BenchmarkMergeCompilerFlags 775.800 ns/op +8.8 ns/op / +1.1% (worse)
Windows MinGW 386 BenchmarkMergeLinkerFlags 721.700 ns/op -2.1 ns/op / -0.3% (better)
Windows MinGW 386 BenchmarkChannelBuffered 39.200 ns/op +0.06 ns/op / +0.2% (worse)
Windows MinGW 386 BenchmarkChannelHandoff 826.100 ns/op -54.1 ns/op / -6.1% (better)
Windows MinGW 386 BenchmarkDefer 45.560 ns/op +0.61 ns/op / +1.4% (worse)
Windows MinGW 386 BenchmarkDirectCall 1.547 ns/op -0.001 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGlobalRead 1.551 ns/op -0.001 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGlobalWrite 7.776 ns/op -0.005 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGoroutine 107074 ns/op -1418 ns/op / -1.3% (better)
Windows MinGW 386 BenchmarkInterfaceCall 8.357 ns/op -0.002 ns/op / -0.02393% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 2.170 ns/op +0.002 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.050 ns/op -0.01 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 577.800 ns/op -3 ns/op / -0.5% (better)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 545.900 ns/op +6.8 ns/op / +1.3% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 39.010 ns/op +1.7 ns/op / +4.6% (worse)
Windows MinGW ARM64 BenchmarkChannelHandoff 2084 ns/op -164 ns/op / -7.3% (better)
Windows MinGW ARM64 BenchmarkDefer 55.490 ns/op +0.45 ns/op / +0.8% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.589 ns/op -0.0003 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkGlobalRead 0.664 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkGlobalWrite 0.663 ns/op +0.0001 ns/op / +0.01508% (worse)
Windows MinGW ARM64 BenchmarkGoroutine 61964 ns/op +1350 ns/op / +2.2% (worse)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.230 ns/op -0.003 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.771 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC BenchmarkLookupPCRandom 13.120 ns/op +0.16 ns/op / +1.2% (worse)
Windows MSVC BenchmarkMergeCompilerFlags 623.400 ns/op -6.1 ns/op / -1.0% (better)
Windows MSVC BenchmarkMergeLinkerFlags 562.800 ns/op +4.8 ns/op / +0.9% (worse)
Windows MSVC BenchmarkChannelBuffered 30.200 ns/op -0.55 ns/op / -1.8% (better)
Windows MSVC BenchmarkChannelHandoff 1031 ns/op +46.2 ns/op / +4.7% (worse)
Windows MSVC BenchmarkDefer 54.300 ns/op -1.02 ns/op / -1.8% (better)
Windows MSVC BenchmarkDirectCall 1.548 ns/op +0.002 ns/op / +0.1% (worse)
Windows MSVC BenchmarkGlobalRead 1.549 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC BenchmarkGlobalWrite 2.471 ns/op -0.01 ns/op / -0.4% (better)
Windows MSVC BenchmarkGoroutine 82953 ns/op -254 ns/op / -0.3% (better)
Windows MSVC BenchmarkInterfaceCall 8.362 ns/op -0.039 ns/op / -0.5% (better)
Windows MSVC BenchmarkRuntimeGetG 1.864 ns/op +0.006 ns/op / +0.3% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 52.870 ns/op -1.1 ns/op / -2.0% (better)
Windows MSVC 386 BenchmarkMergeCompilerFlags 570.400 ns/op -18.1 ns/op / -3.1% (better)
Windows MSVC 386 BenchmarkMergeLinkerFlags 565.600 ns/op -82.8 ns/op / -12.8% (better)
Windows MSVC 386 BenchmarkChannelBuffered 36.420 ns/op +0.09 ns/op / +0.2% (worse)
Windows MSVC 386 BenchmarkChannelHandoff 1820 ns/op -228 ns/op / -11.1% (better)
Windows MSVC 386 BenchmarkDefer 40.920 ns/op -0.01 ns/op / -0.02443% (better)
Windows MSVC 386 BenchmarkDirectCall 0.362 ns/op -0.0066 ns/op / -1.8% (better)
Windows MSVC 386 BenchmarkGlobalRead 0.712 ns/op +0.0089 ns/op / +1.3% (worse)
Windows MSVC 386 BenchmarkGlobalWrite 12.870 ns/op -0.2 ns/op / -1.5% (better)
Windows MSVC 386 BenchmarkGoroutine 184975 ns/op -6503 ns/op / -3.4% (better)
Windows MSVC 386 BenchmarkInterfaceCall 4.270 ns/op +0.006 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkRuntimeGetG 0.971 ns/op +0.0005 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.070 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 572.600 ns/op +11.7 ns/op / +2.1% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 547.100 ns/op +14.6 ns/op / +2.7% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 37.340 ns/op -1.55 ns/op / -4.0% (better)
Windows MSVC ARM64 BenchmarkChannelHandoff 2660 ns/op +196 ns/op / +8.0% (worse)
Windows MSVC ARM64 BenchmarkDefer 64.240 ns/op +2.33 ns/op / +3.8% (worse)
Windows MSVC ARM64 BenchmarkDirectCall 0.591 ns/op +0.0012 ns/op / +0.2% (worse)
Windows MSVC ARM64 BenchmarkGlobalRead 0.664 ns/op -0.0002 ns/op / -0.03013% (better)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.758 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 BenchmarkGoroutine 57627 ns/op -1785 ns/op / -3.0% (better)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.238 ns/op -0.004 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.799 ns/op -0.004 ns/op / -0.2% (better)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 914.100 ns/op +9.7 ns/op / +1.1% (worse)
Linux AfterFuncZeroDelivery/LLGo 45054 ns/op +8393 ns/op / +22.9% (worse)
Linux CreateStop/Go 290.600 ns/op +1.6 ns/op / +0.6% (worse)
Linux CreateStop/LLGo 1765 ns/op -157 ns/op / -8.2% (better)
Linux RearmStopped/Go 114.700 ns/op -1.4 ns/op / -1.2% (better)
Linux RearmStopped/LLGo 1429 ns/op -34 ns/op / -2.3% (better)
Linux ResetActive/Go 67.440 ns/op -1.33 ns/op / -1.9% (better)
Linux ResetActive/LLGo 715.900 ns/op -42.6 ns/op / -5.6% (better)
Linux ResetHeap1024/Go 67.120 ns/op -0.12 ns/op / -0.2% (better)
Linux ResetHeap1024/LLGo 175.300 ns/op +1.2 ns/op / +0.7% (worse)
macOS AfterFuncZeroDelivery/Go 448.900 ns/op -189.5 ns/op / -29.7% (better)
macOS AfterFuncZeroDelivery/LLGo 67365 ns/op -39056 ns/op / -36.7% (better)
macOS CreateStop/Go 245 ns/op +58.5 ns/op / +31.4% (worse)
macOS CreateStop/LLGo 415.900 ns/op -544.2 ns/op / -56.7% (better)
macOS RearmStopped/Go 75.380 ns/op +8.99 ns/op / +13.5% (worse)
macOS RearmStopped/LLGo 394.800 ns/op -353.3 ns/op / -47.2% (better)
macOS ResetActive/Go 46.420 ns/op -12.88 ns/op / -21.7% (better)
macOS ResetActive/LLGo 222.900 ns/op -21.4 ns/op / -8.8% (better)
macOS ResetHeap1024/Go 44.550 ns/op -2.94 ns/op / -6.2% (better)
macOS ResetHeap1024/LLGo 92.760 ns/op -19.24 ns/op / -17.2% (better)
Windows MinGW AfterFuncZeroDelivery/Go 590.100 ns/op +27.9 ns/op / +5.0% (worse)
Windows MinGW AfterFuncZeroDelivery/LLGo 185112 ns/op +10356 ns/op / +5.9% (worse)
Windows MinGW CreateStop/Go 120.500 ns/op +5 ns/op / +4.3% (worse)
Windows MinGW CreateStop/LLGo 448.600 ns/op +24.4 ns/op / +5.8% (worse)
Windows MinGW RearmStopped/Go 31.640 ns/op +0.06 ns/op / +0.2% (worse)
Windows MinGW RearmStopped/LLGo 429.200 ns/op +162.6 ns/op / +61.0% (worse)
Windows MinGW ResetActive/Go 20.280 ns/op +0.14 ns/op / +0.7% (worse)
Windows MinGW ResetActive/LLGo 249.600 ns/op +72.1 ns/op / +40.6% (worse)
Windows MinGW ResetHeap1024/Go 20.540 ns/op +0.08 ns/op / +0.4% (worse)
Windows MinGW ResetHeap1024/LLGo 133.100 ns/op +7 ns/op / +5.6% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/Go 978.300 ns/op +14.6 ns/op / +1.5% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 204311 ns/op +607 ns/op / +0.3% (worse)
Windows MinGW 386 CreateStop/Go 199.200 ns/op +0.9 ns/op / +0.5% (worse)
Windows MinGW 386 CreateStop/LLGo 529.800 ns/op +24.5 ns/op / +4.8% (worse)
Windows MinGW 386 RearmStopped/Go 69.300 ns/op +5.49 ns/op / +8.6% (worse)
Windows MinGW 386 RearmStopped/LLGo 347.900 ns/op +8.3 ns/op / +2.4% (worse)
Windows MinGW 386 ResetActive/Go 39.230 ns/op +0.05 ns/op / +0.1% (worse)
Windows MinGW 386 ResetActive/LLGo 431.100 ns/op -473.5 ns/op / -52.3% (better)
Windows MinGW 386 ResetHeap1024/Go 39.620 ns/op +0.09 ns/op / +0.2% (worse)
Windows MinGW 386 ResetHeap1024/LLGo 188.800 ns/op +1.4 ns/op / +0.7% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 672.200 ns/op +2.8 ns/op / +0.4% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 151666 ns/op -11631 ns/op / -7.1% (better)
Windows MinGW ARM64 CreateStop/Go 201.300 ns/op +1.7 ns/op / +0.9% (worse)
Windows MinGW ARM64 CreateStop/LLGo 386.500 ns/op +0.5 ns/op / +0.1% (worse)
Windows MinGW ARM64 RearmStopped/Go 70.580 ns/op +0.02 ns/op / +0.02834% (worse)
Windows MinGW ARM64 RearmStopped/LLGo 260 ns/op +2 ns/op / +0.8% (worse)
Windows MinGW ARM64 ResetActive/Go 31.120 ns/op +0.2 ns/op / +0.6% (worse)
Windows MinGW ARM64 ResetActive/LLGo 128 ns/op +0.3 ns/op / +0.2% (worse)
Windows MinGW ARM64 ResetHeap1024/Go 31.160 ns/op +0.07 ns/op / +0.2% (worse)
Windows MinGW ARM64 ResetHeap1024/LLGo 126.800 ns/op +0.1 ns/op / +0.1% (worse)
Windows MSVC AfterFuncZeroDelivery/Go 554.800 ns/op -6.8 ns/op / -1.2% (better)
Windows MSVC AfterFuncZeroDelivery/LLGo 161046 ns/op +917 ns/op / +0.6% (worse)
Windows MSVC CreateStop/Go 115.800 ns/op +0.6 ns/op / +0.5% (worse)
Windows MSVC CreateStop/LLGo 447.700 ns/op -31.5 ns/op / -6.6% (better)
Windows MSVC RearmStopped/Go 31.400 ns/op -0.14 ns/op / -0.4% (better)
Windows MSVC RearmStopped/LLGo 272 ns/op +6.8 ns/op / +2.6% (worse)
Windows MSVC ResetActive/Go 20.330 ns/op +0.24 ns/op / +1.2% (worse)
Windows MSVC ResetActive/LLGo 141.700 ns/op -5.5 ns/op / -3.7% (better)
Windows MSVC ResetHeap1024/Go 20.430 ns/op +0.07 ns/op / +0.3% (worse)
Windows MSVC ResetHeap1024/LLGo 125.300 ns/op +0.4 ns/op / +0.3% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/Go 711.600 ns/op -42.5 ns/op / -5.6% (better)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 193098 ns/op +12568 ns/op / +7.0% (worse)
Windows MSVC 386 CreateStop/Go 178.700 ns/op -17.1 ns/op / -8.7% (better)
Windows MSVC 386 CreateStop/LLGo 793.100 ns/op -17.7 ns/op / -2.2% (better)
Windows MSVC 386 RearmStopped/Go 65.690 ns/op -0.72 ns/op / -1.1% (better)
Windows MSVC 386 RearmStopped/LLGo 338.400 ns/op +36.1 ns/op / +11.9% (worse)
Windows MSVC 386 ResetActive/Go 31.080 ns/op -0.47 ns/op / -1.5% (better)
Windows MSVC 386 ResetActive/LLGo 197.100 ns/op -60.5 ns/op / -23.5% (better)
Windows MSVC 386 ResetHeap1024/Go 31.890 ns/op -0.59 ns/op / -1.8% (better)
Windows MSVC 386 ResetHeap1024/LLGo 121.800 ns/op -3.6 ns/op / -2.9% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 669.500 ns/op -5.7 ns/op / -0.8% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 155885 ns/op -6387 ns/op / -3.9% (better)
Windows MSVC ARM64 CreateStop/Go 202.200 ns/op +0.8 ns/op / +0.4% (worse)
Windows MSVC ARM64 CreateStop/LLGo 397.500 ns/op +13 ns/op / +3.4% (worse)
Windows MSVC ARM64 RearmStopped/Go 70.700 ns/op -0.18 ns/op / -0.3% (better)
Windows MSVC ARM64 RearmStopped/LLGo 270.200 ns/op +0.8 ns/op / +0.3% (worse)
Windows MSVC ARM64 ResetActive/Go 31.130 ns/op -0.43 ns/op / -1.4% (better)
Windows MSVC ARM64 ResetActive/LLGo 126.200 ns/op -10.5 ns/op / -7.7% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.130 ns/op -0.09 ns/op / -0.3% (better)
Windows MSVC ARM64 ResetHeap1024/LLGo 139.600 ns/op +1 ns/op / +0.7% (worse)

Compared with f59bac1ea985 measured in the same runner job.

@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

864f6cec2af7 | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 146103 B 0 B / +0.0% 74832 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 144479 B 0 B / +0.0% 73147 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 134241 B 0 B / +0.0% 78720 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 141233 B 0 B / +0.0% 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 140983 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3187934 B 0 B / +0.0% 117962 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3179953 B 0 B / +0.0% 101617 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2925417 B 0 B / +0.0% 125341 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2832890 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2699410 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-emscripten/LLGo 145338 B 0 B / +0.0% 74832 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 143950 B 0 B / +0.0% 73147 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 133574 B 0 B / +0.0% 78720 B 0 B / +0.0%
reflectcall/j32-emscripten/LLGo 1531287 B 0 B / +0.0% 91998 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1534006 B 0 B / +0.0% 90313 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1417273 B 0 B / +0.0% 97731 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1538681 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1464240 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 140446 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-wasi/LLGo 140267 B 0 B / +0.0% 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 5.773 s -844.7 ms / -12.8% (better)
j32-goos-js 5.974 s -361.6 ms / -5.7% (better)
j64-emscripten-memory64 5.087 s -745 ms / -12.8% (better)
reflectcall/w32-wasi 25.980 s -603.6 ms / -2.3% (better)
w32-goos-wasip1 4.725 s -531.6 ms / -10.1% (better)
w32-wasi 4.522 s -784.5 ms / -14.8% (better)

Compared with f59bac1ea985 measured in the same runner job.

@cpunion
cpunion force-pushed the codex/wasi-threaded-stack-bounds branch from c4a6fbe to 8db4345 Compare September 27, 2026 15:06
@cpunion

cpunion commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator Author

Thanks for the review. Commit 419970c35 corrects the timer-condition error message for both clock domains. The two inline nanotime comments refer to the older implementation: merged #2653 replaced it before this rebase. The current WASI-threads nanotime path traps on a failed clock read and has no shared CAS/cache-line hotspot. Focused WASI build and runtime compilation checks pass, and the wasm-runtime CI jobs on this head are green.

@cpunion
cpunion merged commit 827d21a into xgo-dev:main Sep 28, 2026
88 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants