Skip to content

ssa: box constant iface data in read-only globals - #2704

Open
visualfc wants to merge 2 commits into
xgo-dev:mainfrom
visualfc:fix/static-iface-box
Open

visualfc wants to merge 2 commits into
xgo-dev:mainfrom
visualfc:fix/static-iface-box

Conversation

@visualfc

Copy link
Copy Markdown
Collaborator

Summary

Non-direct interface values (integers, bools, small aggregates) stay IfaceIndir, matching Go: data is a pointer to a copy. Compile-time constants go in a WeakODR _llgo_ifacebox$ global instead of AllocU, like cmd/compile's static temps / staticuint64s.

Any LLVM-constant payload is eligible. There is no matching on constructor function names.

Together with static itabs, a constant T2I is {itab, box} with no runtime call or heap allocation.

Depends on: #2703 (static itabs). Merge after #2703. This branch contains #2703 plus this commit; only the latest commit is in scope.

Test plan

  • cl/_testrt/staticifacebox (any(int64(42)) has no AllocU)
  • cl/_testrt/eface

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: static iface box / static itab for known T2I

This is a well-targeted change. MakeInterface now boxes compile-time-constant non-direct interface data into deduplicated read-only globals (_llgo_ifacebox$...) instead of AllocU+Store, and staticItab emits a runtime static itab global for known type-to-interface conversions even without LTO plugin markers. The runtime interequal fix correctly restores the invariant that a T2I value (static itab global) and an I2I value (NewItab) of the same interface compare equal, and it is memory-safe (both tabs are nil-guarded before dereference). The verified behaviors:

  • The new comments are accurate — staticItab's hash matches abiCommonFields (sha256(typeName) → little-endian uint32), and integers/bools do stay IfaceIndir.
  • interequal's fallback short-circuits on _type mismatch, so the extra ifaceeq only runs when the type pair actually matches.
  • Test renames/assertions in ssa_test.go are consistent with the new emit-always behavior.

Findings below are refinements, not blockers. The main items are two compile-time hot-path inefficiencies (expensive dedup keys computed before the cache lookup) and a couple of correctness invariants worth documenting.

Semantics note (not inline-placeable)

Boxing all constant non-direct types (structs, strings, slices, arrays, big ints) into shared GlobalConstant globals is broader than cmd/compile, which only shares addresses for the small-integer staticuint64s range and read-only static temps. Two consequences worth confirming are intended:

  • Identical (type, constant) pairs now resolve to a single shared address. Code that extracts the iface data word (via reflect internals, //go:linkname, or unsafe) and relies on distinct addresses will observe shared identity.
  • Because the box is GlobalConstant, mutating the pointee through the interface via unsafe would now fault on a read-only page, whereas the previous AllocU+Store path produced writable storage.

Neither is guaranteed by the Go spec, so this is likely acceptable — flagging so it's a conscious decision.

Comment thread ssa/interface.go
Comment on lines +134 to +136
sum := sha256.Sum256([]byte(typeName + "\x00" + x.impl.String()))
name := "_llgo_ifacebox$" + base64.RawURLEncoding.EncodeToString(sum[:])
if g := b.Pkg.VarOf(name); g != nil {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] staticIfaceBox: compute the expensive dedup key after the cache check

The dedup lookup b.Pkg.VarOf(name) runs after the key is built, so every constant interface conversion — including repeated conversions of a constant that is already boxed — pays x.impl.String() (serializes the full LLVM constant to textual IR, O(constant size)) plus a sha256 and base64 on every call. For programs that box the same large aggregate constant in N places, N-1 of those are cache hits that still do the full IR serialization + hash.

Consider a cheaper first-level key (e.g. the llvm.Value pointer or a small structural key) and only fall back to String()+sha256 on a miss, or memoize the computed name on the constant.

Comment thread ssa/interface.go
Comment on lines +49 to +52
if rawIntf.NumMethods() == 0 || concrete == nil {
return Expr{}, false
}
rawIntf = rawIntf.Complete()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] staticItab: move the VarOf cache check before NewMethodSet/AssignableTo

rawIntf.Complete(), types.AssignableTo, types.NewMethodSet(concrete) and the per-method Lookup loop all run before the VarOf(name) cache check further down. The global name derives only from intfName/typeName (via abi.TypeName), so repeated T2I conversions of the same (interface, concrete) pair rebuild the full method set on every call only to find the itab already exists. Computing the name and checking VarOf first would make repeated conversions O(1).

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two additional inline notes (continuation of the prior review).

Comment thread ssa/interface.go
}
prog := b.Prog
typeName, _ := prog.abi.TypeName(typ.raw.Type)
sum := sha256.Sum256([]byte(typeName + "\x00" + x.impl.String()))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] Fragile constant identity key

The box name hashes x.impl.String() (the LLVM IR text). For constants that reference other globals — notably ConstString, whose data pointer points at an anonymous private global — the printed form can depend on module-local numbering (@0, @1, …). Distinct byte payloads printing identically is extremely unlikely today (referenced string globals are content-addressed via p.strs), but the correctness of the ODR-merged box relies on that implicit invariant. A short comment noting the assumption, or keying on the constant's actual bytes, would make this robust against future changes to anonymous/private global naming.

Comment thread ssa/interface.go
if rawIntf.NumMethods() == 0 || concrete == nil {
return Expr{}, false
}
rawIntf = rawIntf.Complete()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] Dedup trusts abi.TypeName injectivity

Both the itab and the iface box are deduplicated by sha256(abi.TypeName(...)). If two distinct concrete types.Type values can ever map to the same abi.TypeName string, the first-emitted itab/box would be silently reused for the second, wrong type — its fun slots would point at the wrong methods, or wrong data would be shared. This is pre-existing key style, but since the itab global is now emitted in all builds (not just LTO), the blast radius is larger. Worth confirming abi.TypeName is injective for every concrete type reaching here, or folding in a stronger discriminator.

@codecov

codecov Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 82.52427% with 18 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
internal/build/main_module.go 65.38% 9 Missing ⚠️
ssa/interface.go 87.30% 8 Missing ⚠️
ssa/abitype.go 92.30% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

Known concrete-to-interface conversions emit a read-only _llgo_itab$
global (Inter, Type, Hash, Fun) with weak_odr linkage, matching
cmd/compile's go:itab symbols.

Ordinary builds use that global as the runtime vtable. After function
bodies compile, path.init$itabs registers those itabs with
RegisterStaticItab, after runtime.init and before package init (gc
itabsinit). LTO and deadcode-drop keep calling NewItab so unused
interface methods can be dropped; Fun[] would pin them. LTO still
emits the static template for de-virt.

Hash is copied from the type descriptor when present. interequal
compares the (inter, _type) pair so a static itab and a dynamically
allocated itab for the same conversion compare equal.
@visualfc
visualfc force-pushed the fix/static-iface-box branch from 48eba88 to 4b3a724 Compare September 30, 2026 05:18
@github-actions

Copy link
Copy Markdown

LLGo baseline benchmarks

4b3a724c217a | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 7280 B +120 B / +1.7% (worse) 393 B +6 B / +1.6% (worse) 780.762 ms +100.5 ms / +14.8% (worse) 1.435 ms +39.49 us / +2.8% (worse)
Linux cprintf-lto 7112 B +200 B / +2.9% (worse) 374 B +6 B / +1.6% (worse) 780.464 ms +116.7 ms / +17.6% (worse) 1.471 ms +121.3 us / +9.0% (worse)
Linux fmtprintf 1786352 B +124080 B / +7.5% (worse) 547213 B +48854 B / +9.8% (worse) 5.325 s +933.8 ms / +21.3% (worse) 3.757 ms +469.3 us / +14.3% (worse)
Linux fmtprintf-lto 1604056 B +103672 B / +6.9% (worse) 475101 B +38988 B / +8.9% (worse) 14.886 s +2.656 s / +21.7% (worse) 3.611 ms +420.3 us / +13.2% (worse)
Linux println 69824 B +832 B / +1.2% (worse) 16583 B -200 B / -1.2% (better) 767.729 ms +82.53 ms / +12.0% (worse) 1.659 ms -99.82 us / -5.7% (better)
Linux println-lto 60224 B +376 B / +0.6% (worse) 14021 B -178 B / -1.3% (better) 966.102 ms -5.244 ms / -0.5% (better) 1.772 ms -186.5 us / -9.5% (better)
macOS cprintf 68096 B +32 B / +0.04701% (worse) 4437 B +8 B / +0.2% (worse) 1.014 s +253.9 ms / +33.4% (worse) 5.357 ms +1.925 ms / +56.1% (worse)
macOS cprintf-lto 84288 B 0 B / +0.0% 193 B 0 B / +0.0% 1.528 s +750.8 ms / +96.6% (worse) 8.129 ms +4.081 ms / +100.8% (worse)
macOS fmtprintf 1624560 B +119888 B / +8.0% (worse) 929696 B +55124 B / +6.3% (worse) 3.449 s -148.9 ms / -4.1% (better) 4.782 ms -401.7 us / -7.7% (better)
macOS fmtprintf-lto 1258704 B +66000 B / +5.5% (worse) 897464 B +49248 B / +5.8% (worse) 9.391 s +1.383 s / +17.3% (worse) 3.862 ms -565.2 us / -12.8% (better)
macOS println 117776 B +640 B / +0.5% (worse) 37227 B -194 B / -0.5% (better) 1.271 s +365.9 ms / +40.4% (worse) 7.840 ms +241.7 us / +3.2% (worse)
macOS println-lto 119472 B 0 B / +0.0% 34655 B -169 B / -0.5% (better) 1.293 s +90.98 ms / +7.6% (worse) 5.062 ms -775.5 us / -13.3% (better)
Windows MinGW cprintf 19968 B +512 B / +2.6% (worse) 4566 B +16 B / +0.4% (worse) 981.549 ms -29.51 ms / -2.9% (better) 2.741 ms -67 us / -2.4% (better)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.013 s -40.26 ms / -3.8% (better) 2.741 ms -89.2 us / -3.2% (better)
Windows MinGW fmtprintf 2108416 B +175616 B / +9.1% (worse) 687302 B +88352 B / +14.8% (worse) 3.238 s -50.31 ms / -1.5% (better) 6.174 ms -19.6 us / -0.3% (better)
Windows MinGW fmtprintf-lto 2121728 B +164352 B / +8.4% (worse) 620630 B +73312 B / +13.4% (worse) 8.665 s +489 ms / +6.0% (worse) 6.276 ms +143.8 us / +2.3% (worse)
Windows MinGW println 76288 B +512 B / +0.7% (worse) 24950 B -192 B / -0.8% (better) 981.373 ms -38.67 ms / -3.8% (better) 4.988 ms -152.8 us / -3.0% (better)
Windows MinGW println-lto 70144 B +1024 B / +1.5% (worse) 21782 B -208 B / -0.9% (better) 1.179 s -30.24 ms / -2.5% (better) 4.995 ms -802.5 us / -13.8% (better)
Windows MinGW 386 cprintf 43520 B 0 B / +0.0% 5342 B +16 B / +0.3% (worse) 1.276 s -17.09 ms / -1.3% (better) 5.018 ms -81.6 us / -1.6% (better)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.461 s +179.6 ms / +14.0% (worse) 6.185 ms +1.074 ms / +21.0% (worse)
Windows MinGW 386 fmtprintf 2028032 B +131584 B / +6.9% (worse) 527502 B +55088 B / +11.7% (worse) 4.090 s -91.49 ms / -2.2% (better) 10.189 ms -247.9 us / -2.4% (better)
Windows MinGW 386 fmtprintf-lto 2398208 B +217600 B / +10.0% (worse) 500650 B +49392 B / +10.9% (worse) 10.562 s +435.5 ms / +4.3% (worse) 10.687 ms -472.2 us / -4.2% (better)
Windows MinGW 386 println 97792 B +1536 B / +1.6% (worse) 21250 B -208 B / -1.0% (better) 1.431 s +175.4 ms / +14.0% (worse) 9.229 ms +742.8 us / +8.8% (worse)
Windows MinGW 386 println-lto 75264 B +1024 B / +1.4% (worse) 19122 B -192 B / -1.0% (better) 1.526 s +17.25 ms / +1.1% (worse) 8.974 ms +4.3 us / +0.04794% (worse)
Windows MinGW ARM64 cprintf 19456 B +512 B / +2.7% (worse) 4424 B +16 B / +0.4% (worse) 1.562 s -34.8 ms / -2.2% (better) 6.048 ms -686.1 us / -10.2% (better)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.557 s +30.14 ms / +2.0% (worse) 6.272 ms +195.8 us / +3.2% (worse)
Windows MinGW ARM64 fmtprintf 1967104 B +147968 B / +8.1% (worse) 573552 B +63944 B / +12.5% (worse) 4.221 s +37.29 ms / +0.9% (worse) 12.406 ms +185.7 us / +1.5% (worse)
Windows MinGW ARM64 fmtprintf-lto 2035712 B +156672 B / +8.3% (worse) 533168 B +57016 B / +12.0% (worse) 10.236 s +624.3 ms / +6.5% (worse) 12.535 ms -493.5 us / -3.8% (better)
Windows MinGW ARM64 println 73216 B +1024 B / +1.4% (worse) 23644 B -232 B / -1.0% (better) 1.540 s +48.57 ms / +3.3% (worse) 10.732 ms +207.9 us / +2.0% (worse)
Windows MinGW ARM64 println-lto 69120 B +512 B / +0.7% (worse) 21028 B -196 B / -0.9% (better) 1.722 s +2.21 ms / +0.1% (worse) 11.164 ms +378.8 us / +3.5% (worse)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65814 B +16 B / +0.02432% (worse) 1.106 s +3.84 ms / +0.3% (worse) 3.345 ms -1.476 ms / -30.6% (better)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65734 B 0 B / +0.0% 1.095 s -7.206 ms / -0.7% (better) 3.420 ms +134.1 us / +4.1% (worse)
Windows MSVC fmtprintf 1773056 B +130048 B / +7.9% (worse) 782838 B +88336 B / +12.7% (worse) 3.813 s -93.39 ms / -2.4% (better) 9.784 ms +454.2 us / +4.9% (worse)
Windows MSVC fmtprintf-lto 1760256 B +125440 B / +7.7% (worse) 721014 B +73952 B / +11.4% (worse) 9.827 s +651.2 ms / +7.1% (worse) 9.226 ms -24.8 us / -0.3% (better)
Windows MSVC println 195072 B +512 B / +0.3% (worse) 120630 B -192 B / -0.2% (better) 1.093 s +8.774 ms / +0.8% (worse) 8.138 ms +1.04 ms / +14.6% (worse)
Windows MSVC println-lto 192512 B 0 B / +0.0% 118134 B -208 B / -0.2% (better) 1.284 s -23.06 ms / -1.8% (better) 6.740 ms -683.5 us / -9.2% (better)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.074 s -196.8 ms / -15.5% (better) 5.852 ms +480.6 us / +8.9% (worse)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.276 s +185.1 ms / +17.0% (worse) 5.886 ms +616.4 us / +11.7% (worse)
Windows MSVC 386 fmtprintf 1283072 B +78848 B / +6.5% (worse) 510892 B +55088 B / +12.1% (worse) 3.892 s +165.2 ms / +4.4% (worse) 11.431 ms -220.4 us / -1.9% (better)
Windows MSVC 386 fmtprintf-lto 1334272 B +93184 B / +7.5% (worse) 474651 B +47456 B / +11.1% (worse) 9.428 s +704.9 ms / +8.1% (worse) 11.256 ms +117.2 us / +1.1% (worse)
Windows MSVC 386 println 36352 B 0 B / +0.0% 20100 B -224 B / -1.1% (better) 1.082 s +431.8 us / +0.03991% (worse) 9.964 ms +476.1 us / +5.0% (worse)
Windows MSVC 386 println-lto 35328 B -512 B / -1.4% (better) 18341 B -208 B / -1.1% (better) 1.526 s +255.5 ms / +20.1% (worse) 12.505 ms +3.14 ms / +33.5% (worse)
Windows MSVC ARM64 cprintf 11776 B 0 B / +0.0% 4208 B +16 B / +0.4% (worse) 1.319 s +17.65 ms / +1.4% (worse) 7.566 ms +68.1 us / +0.9% (worse)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.305 s -5.104 ms / -0.4% (better) 7.538 ms -118.3 us / -1.5% (better)
Windows MSVC ARM64 fmtprintf 1488896 B +102400 B / +7.4% (worse) 573512 B +63968 B / +12.6% (worse) 3.936 s -75.07 ms / -1.9% (better) 14.362 ms -1.081 ms / -7.0% (better)
Windows MSVC ARM64 fmtprintf-lto 1520640 B +115712 B / +8.2% (worse) 533892 B +57072 B / +12.0% (worse) 9.505 s +435.2 ms / +4.8% (worse) 14.678 ms +800 ns / +0.005451% (worse)
Windows MSVC ARM64 println 45056 B 0 B / +0.0% 23668 B -240 B / -1.0% (better) 1.301 s +8.869 ms / +0.7% (worse) 13.169 ms -1.138 ms / -8.0% (better)
Windows MSVC ARM64 println-lto 42496 B 0 B / +0.0% 21188 B -192 B / -0.9% (better) 1.485 s +16.69 ms / +1.1% (worse) 13.151 ms +292.9 us / +2.3% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.690 ns/op -0.09 ns/op / -0.6% (better)
Linux BenchmarkMergeCompilerFlags 222.100 ns/op -67.1 ns/op / -23.2% (better)
Linux BenchmarkMergeLinkerFlags 155.800 ns/op -68 ns/op / -30.4% (better)
Linux BenchmarkChannelBuffered 55.630 ns/op -0.25 ns/op / -0.4% (better)
Linux BenchmarkChannelHandoff 13100 ns/op -452 ns/op / -3.3% (better)
Linux BenchmarkDefer 58.910 ns/op +6.78 ns/op / +13.0% (worse)
Linux BenchmarkDirectCall 1.167 ns/op -0.402 ns/op / -25.6% (better)
Linux BenchmarkGlobalRead 1.212 ns/op +0.04 ns/op / +3.4% (worse)
Linux BenchmarkGlobalWrite 7.816 ns/op +0.028 ns/op / +0.4% (worse)
Linux BenchmarkGoroutine 25947 ns/op -969 ns/op / -3.6% (better)
Linux BenchmarkInterfaceCall 5.998 ns/op +0.059 ns/op / +1.0% (worse)
Linux BenchmarkRuntimeGetG 2.915 ns/op -0.28 ns/op / -8.8% (better)
macOS BenchmarkLookupPCRandom 12.950 ns/op -1.2 ns/op / -8.5% (better)
macOS BenchmarkMergeCompilerFlags 112 ns/op +1.7 ns/op / +1.5% (worse)
macOS BenchmarkMergeLinkerFlags 73.660 ns/op +5.23 ns/op / +7.6% (worse)
macOS BenchmarkChannelBuffered 30.550 ns/op -7.19 ns/op / -19.1% (better)
macOS BenchmarkChannelHandoff 11191 ns/op -680 ns/op / -5.7% (better)
macOS BenchmarkDefer 43.300 ns/op -2.58 ns/op / -5.6% (better)
macOS BenchmarkDirectCall 1.275 ns/op -0.514 ns/op / -28.7% (better)
macOS BenchmarkGlobalRead 1.278 ns/op -0.137 ns/op / -9.7% (better)
macOS BenchmarkGlobalWrite 1.567 ns/op -0.097 ns/op / -5.8% (better)
macOS BenchmarkGoroutine 74332 ns/op +8891 ns/op / +13.6% (worse)
macOS BenchmarkInterfaceCall 4.572 ns/op -1.137 ns/op / -19.9% (better)
macOS BenchmarkRuntimeGetG 3.041 ns/op +0.384 ns/op / +14.5% (worse)
Windows MinGW BenchmarkLookupPCRandom 9.615 ns/op -0.005 ns/op / -0.1% (better)
Windows MinGW BenchmarkMergeCompilerFlags 371.500 ns/op -18.6 ns/op / -4.8% (better)
Windows MinGW BenchmarkMergeLinkerFlags 338.500 ns/op +2.8 ns/op / +0.8% (worse)
Windows MinGW BenchmarkChannelBuffered 23.650 ns/op +0.36 ns/op / +1.5% (worse)
Windows MinGW BenchmarkChannelHandoff 1077 ns/op -29 ns/op / -2.6% (better)
Windows MinGW BenchmarkDefer 41.130 ns/op -2.34 ns/op / -5.4% (better)
Windows MinGW BenchmarkDirectCall 1.369 ns/op +0.013 ns/op / +1.0% (worse)
Windows MinGW BenchmarkGlobalRead 1.478 ns/op +0.119 ns/op / +8.8% (worse)
Windows MinGW BenchmarkGlobalWrite 2.165 ns/op -0.002 ns/op / -0.1% (better)
Windows MinGW BenchmarkGoroutine 57000 ns/op +60 ns/op / +0.1% (worse)
Windows MinGW BenchmarkInterfaceCall 6.601 ns/op -0.192 ns/op / -2.8% (better)
Windows MinGW BenchmarkRuntimeGetG 1.406 ns/op -0.225 ns/op / -13.8% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 26.470 ns/op +0.01 ns/op / +0.03779% (worse)
Windows MinGW 386 BenchmarkMergeCompilerFlags 749.300 ns/op +50 ns/op / +7.2% (worse)
Windows MinGW 386 BenchmarkMergeLinkerFlags 668.300 ns/op +21.7 ns/op / +3.4% (worse)
Windows MinGW 386 BenchmarkChannelBuffered 39.250 ns/op +0.13 ns/op / +0.3% (worse)
Windows MinGW 386 BenchmarkChannelHandoff 977 ns/op +27.9 ns/op / +2.9% (worse)
Windows MinGW 386 BenchmarkDefer 43.590 ns/op +1.12 ns/op / +2.6% (worse)
Windows MinGW 386 BenchmarkDirectCall 1.546 ns/op -0.002 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGlobalRead 1.548 ns/op -0.001 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGlobalWrite 7.774 ns/op +0.001 ns/op / +0.01287% (worse)
Windows MinGW 386 BenchmarkGoroutine 102021 ns/op -823 ns/op / -0.8% (better)
Windows MinGW 386 BenchmarkInterfaceCall 8.091 ns/op -0.309 ns/op / -3.7% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 2.168 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.130 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 570.200 ns/op -8.2 ns/op / -1.4% (better)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 536.200 ns/op -4.3 ns/op / -0.8% (better)
Windows MinGW ARM64 BenchmarkChannelBuffered 37.880 ns/op +0.5 ns/op / +1.3% (worse)
Windows MinGW ARM64 BenchmarkChannelHandoff 2369 ns/op +316 ns/op / +15.4% (worse)
Windows MinGW ARM64 BenchmarkDefer 56.760 ns/op +2.17 ns/op / +4.0% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.590 ns/op -0.0001 ns/op / -0.01696% (better)
Windows MinGW ARM64 BenchmarkGlobalRead 0.590 ns/op -0.0736 ns/op / -11.1% (better)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.884 ns/op +0.2209 ns/op / +33.3% (worse)
Windows MinGW ARM64 BenchmarkGoroutine 63440 ns/op +204 ns/op / +0.3% (worse)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.203 ns/op +0.059 ns/op / +1.4% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.771 ns/op -0.039 ns/op / -2.2% (better)
Windows MSVC BenchmarkLookupPCRandom 13.200 ns/op +0.38 ns/op / +3.0% (worse)
Windows MSVC BenchmarkMergeCompilerFlags 624.400 ns/op +13.5 ns/op / +2.2% (worse)
Windows MSVC BenchmarkMergeLinkerFlags 550.100 ns/op +8.3 ns/op / +1.5% (worse)
Windows MSVC BenchmarkChannelBuffered 28.690 ns/op -3.37 ns/op / -10.5% (better)
Windows MSVC BenchmarkChannelHandoff 1155 ns/op +77 ns/op / +7.1% (worse)
Windows MSVC BenchmarkDefer 68.810 ns/op +15.11 ns/op / +28.1% (worse)
Windows MSVC BenchmarkDirectCall 1.551 ns/op +0.004 ns/op / +0.3% (worse)
Windows MSVC BenchmarkGlobalRead 1.548 ns/op -0.002 ns/op / -0.1% (better)
Windows MSVC BenchmarkGlobalWrite 2.462 ns/op -0.008 ns/op / -0.3% (better)
Windows MSVC BenchmarkGoroutine 91866 ns/op +3540 ns/op / +4.0% (worse)
Windows MSVC BenchmarkInterfaceCall 8.719 ns/op +0.344 ns/op / +4.1% (worse)
Windows MSVC BenchmarkRuntimeGetG 2.481 ns/op +0.618 ns/op / +33.2% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 26.460 ns/op -0.09 ns/op / -0.3% (better)
Windows MSVC 386 BenchmarkMergeCompilerFlags 730 ns/op +21.4 ns/op / +3.0% (worse)
Windows MSVC 386 BenchmarkMergeLinkerFlags 694.500 ns/op +23.6 ns/op / +3.5% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 38.900 ns/op -5.98 ns/op / -13.3% (better)
Windows MSVC 386 BenchmarkChannelHandoff 912.200 ns/op -47.4 ns/op / -4.9% (better)
Windows MSVC 386 BenchmarkDefer 46.550 ns/op -1.84 ns/op / -3.8% (better)
Windows MSVC 386 BenchmarkDirectCall 1.549 ns/op -0.018 ns/op / -1.1% (better)
Windows MSVC 386 BenchmarkGlobalRead 1.862 ns/op -0.016 ns/op / -0.9% (better)
Windows MSVC 386 BenchmarkGlobalWrite 7.805 ns/op +0.001 ns/op / +0.01281% (worse)
Windows MSVC 386 BenchmarkGoroutine 106993 ns/op +3093 ns/op / +3.0% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 8.120 ns/op -0.049 ns/op / -0.6% (better)
Windows MSVC 386 BenchmarkRuntimeGetG 1.927 ns/op -0.55 ns/op / -22.2% (better)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.190 ns/op +0.14 ns/op / +1.2% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 567.800 ns/op -11.4 ns/op / -2.0% (better)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 539 ns/op -15.5 ns/op / -2.8% (better)
Windows MSVC ARM64 BenchmarkChannelBuffered 38.640 ns/op -0.93 ns/op / -2.4% (better)
Windows MSVC ARM64 BenchmarkChannelHandoff 1981 ns/op +210 ns/op / +11.9% (worse)
Windows MSVC ARM64 BenchmarkDefer 62.420 ns/op +0.04 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkDirectCall 0.590 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 BenchmarkGlobalRead 0.885 ns/op +0.221 ns/op / +33.3% (worse)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.799 ns/op +0.047 ns/op / +1.3% (worse)
Windows MSVC ARM64 BenchmarkGoroutine 58771 ns/op -693 ns/op / -1.2% (better)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.194 ns/op +0.059 ns/op / +1.4% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.805 ns/op +0.014 ns/op / +0.8% (worse)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 1025 ns/op +87.5 ns/op / +9.3% (worse)
Linux AfterFuncZeroDelivery/LLGo 40362 ns/op -5607 ns/op / -12.2% (better)
Linux CreateStop/Go 319.600 ns/op -63 ns/op / -16.5% (better)
Linux CreateStop/LLGo 1803 ns/op -158 ns/op / -8.1% (better)
Linux RearmStopped/Go 115.900 ns/op -0.5 ns/op / -0.4% (better)
Linux RearmStopped/LLGo 1094 ns/op -279 ns/op / -20.3% (better)
Linux ResetActive/Go 68.580 ns/op -0.18 ns/op / -0.3% (better)
Linux ResetActive/LLGo 725.100 ns/op -113.7 ns/op / -13.6% (better)
Linux ResetHeap1024/Go 67.170 ns/op +0.11 ns/op / +0.2% (worse)
Linux ResetHeap1024/LLGo 180.700 ns/op -0.7 ns/op / -0.4% (better)
macOS AfterFuncZeroDelivery/Go 612.800 ns/op +31.4 ns/op / +5.4% (worse)
macOS AfterFuncZeroDelivery/LLGo 124670 ns/op +24192 ns/op / +24.1% (worse)
macOS CreateStop/Go 177.100 ns/op -28.2 ns/op / -13.7% (better)
macOS CreateStop/LLGo 737.600 ns/op +259.9 ns/op / +54.4% (worse)
macOS RearmStopped/Go 75.310 ns/op -0.77 ns/op / -1.0% (better)
macOS RearmStopped/LLGo 605.700 ns/op -99.4 ns/op / -14.1% (better)
macOS ResetActive/Go 55.760 ns/op +1.49 ns/op / +2.7% (worse)
macOS ResetActive/LLGo 306.800 ns/op +1.2 ns/op / +0.4% (worse)
macOS ResetHeap1024/Go 54.330 ns/op -6.49 ns/op / -10.7% (better)
macOS ResetHeap1024/LLGo 128.600 ns/op +37.22 ns/op / +40.7% (worse)
Windows MinGW AfterFuncZeroDelivery/Go 377.500 ns/op +4.1 ns/op / +1.1% (worse)
Windows MinGW AfterFuncZeroDelivery/LLGo 109618 ns/op -400 ns/op / -0.4% (better)
Windows MinGW CreateStop/Go 89.990 ns/op +0.57 ns/op / +0.6% (worse)
Windows MinGW CreateStop/LLGo 334 ns/op -8.5 ns/op / -2.5% (better)
Windows MinGW RearmStopped/Go 24.430 ns/op -0.05 ns/op / -0.2% (better)
Windows MinGW RearmStopped/LLGo 223 ns/op +1.5 ns/op / +0.7% (worse)
Windows MinGW ResetActive/Go 14.840 ns/op +0.05 ns/op / +0.3% (worse)
Windows MinGW ResetActive/LLGo 125.300 ns/op +8.1 ns/op / +6.9% (worse)
Windows MinGW ResetHeap1024/Go 14.820 ns/op -0.04 ns/op / -0.3% (better)
Windows MinGW ResetHeap1024/LLGo 108 ns/op +0.2 ns/op / +0.2% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/Go 943.100 ns/op +8.2 ns/op / +0.9% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 190004 ns/op +356 ns/op / +0.2% (worse)
Windows MinGW 386 CreateStop/Go 192.300 ns/op +1.9 ns/op / +1.0% (worse)
Windows MinGW 386 CreateStop/LLGo 519.500 ns/op +28.8 ns/op / +5.9% (worse)
Windows MinGW 386 RearmStopped/Go 63.650 ns/op +0.31 ns/op / +0.5% (worse)
Windows MinGW 386 RearmStopped/LLGo 352.400 ns/op +2 ns/op / +0.6% (worse)
Windows MinGW 386 ResetActive/Go 38.990 ns/op -0.03 ns/op / -0.1% (better)
Windows MinGW 386 ResetActive/LLGo 373.400 ns/op -626.6 ns/op / -62.7% (better)
Windows MinGW 386 ResetHeap1024/Go 39.430 ns/op -0.03 ns/op / -0.1% (better)
Windows MinGW 386 ResetHeap1024/LLGo 186.400 ns/op -1.5 ns/op / -0.8% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 666 ns/op -8.8 ns/op / -1.3% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 143124 ns/op -1960 ns/op / -1.4% (better)
Windows MinGW ARM64 CreateStop/Go 208.400 ns/op +3.4 ns/op / +1.7% (worse)
Windows MinGW ARM64 CreateStop/LLGo 355.800 ns/op -0.4 ns/op / -0.1% (better)
Windows MinGW ARM64 RearmStopped/Go 70.590 ns/op +0.02 ns/op / +0.02834% (worse)
Windows MinGW ARM64 RearmStopped/LLGo 253.400 ns/op +1 ns/op / +0.4% (worse)
Windows MinGW ARM64 ResetActive/Go 31.020 ns/op -0.1 ns/op / -0.3% (better)
Windows MinGW ARM64 ResetActive/LLGo 127.500 ns/op -1.8 ns/op / -1.4% (better)
Windows MinGW ARM64 ResetHeap1024/Go 31.040 ns/op -0.1 ns/op / -0.3% (better)
Windows MinGW ARM64 ResetHeap1024/LLGo 126.900 ns/op -2.6 ns/op / -2.0% (better)
Windows MSVC AfterFuncZeroDelivery/Go 563.700 ns/op +12.1 ns/op / +2.2% (worse)
Windows MSVC AfterFuncZeroDelivery/LLGo 171395 ns/op -467 ns/op / -0.3% (better)
Windows MSVC CreateStop/Go 117.600 ns/op +3 ns/op / +2.6% (worse)
Windows MSVC CreateStop/LLGo 442 ns/op +42.7 ns/op / +10.7% (worse)
Windows MSVC RearmStopped/Go 31.250 ns/op -0.43 ns/op / -1.4% (better)
Windows MSVC RearmStopped/LLGo 254.100 ns/op -4 ns/op / -1.5% (better)
Windows MSVC ResetActive/Go 20.020 ns/op -0.16 ns/op / -0.8% (better)
Windows MSVC ResetActive/LLGo 207.400 ns/op +58.2 ns/op / +39.0% (worse)
Windows MSVC ResetHeap1024/Go 20.490 ns/op -0.06 ns/op / -0.3% (better)
Windows MSVC ResetHeap1024/LLGo 126.800 ns/op +0.1 ns/op / +0.1% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/Go 937.500 ns/op -10.7 ns/op / -1.1% (better)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 190015 ns/op +1638 ns/op / +0.9% (worse)
Windows MSVC 386 CreateStop/Go 191.800 ns/op +1.6 ns/op / +0.8% (worse)
Windows MSVC 386 CreateStop/LLGo 462.600 ns/op +7.2 ns/op / +1.6% (worse)
Windows MSVC 386 RearmStopped/Go 63.230 ns/op +0.03 ns/op / +0.04747% (worse)
Windows MSVC 386 RearmStopped/LLGo 318.600 ns/op -3.5 ns/op / -1.1% (better)
Windows MSVC 386 ResetActive/Go 38.990 ns/op -0.04 ns/op / -0.1% (better)
Windows MSVC 386 ResetActive/LLGo 878.700 ns/op -71.2 ns/op / -7.5% (better)
Windows MSVC 386 ResetHeap1024/Go 39.410 ns/op +0.1 ns/op / +0.3% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 175.300 ns/op +0.5 ns/op / +0.3% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 665.200 ns/op +3.3 ns/op / +0.5% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 144377 ns/op -6900 ns/op / -4.6% (better)
Windows MSVC ARM64 CreateStop/Go 197.500 ns/op +1.9 ns/op / +1.0% (worse)
Windows MSVC ARM64 CreateStop/LLGo 445.600 ns/op +9.9 ns/op / +2.3% (worse)
Windows MSVC ARM64 RearmStopped/Go 70.670 ns/op -0.2 ns/op / -0.3% (better)
Windows MSVC ARM64 RearmStopped/LLGo 279.600 ns/op +3.3 ns/op / +1.2% (worse)
Windows MSVC ARM64 ResetActive/Go 31.070 ns/op -0.09 ns/op / -0.3% (better)
Windows MSVC ARM64 ResetActive/LLGo 157.500 ns/op +0.5 ns/op / +0.3% (worse)
Windows MSVC ARM64 ResetHeap1024/Go 31.100 ns/op -0.05 ns/op / -0.2% (better)
Windows MSVC ARM64 ResetHeap1024/LLGo 137.400 ns/op -2.5 ns/op / -1.8% (better)

Compared with 653957f840b5 measured in the same runner job.

Non-direct interface values (integers, bools, small aggregates) still
use IfaceIndir, matching Go: data is a pointer to a copy. Compile-time
constants go in a WeakODR _llgo_ifacebox$ global instead of AllocU,
like cmd/compile's static temps.

Depends on static itabs so a constant T2I can be {itab, box} with no
runtime call.
@github-actions

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

4b3a724c217a | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 146278 B -695 B / -0.5% (better) 74890 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 144726 B -681 B / -0.5% (better) 73165 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 134426 B -90 B / -0.1% (better) 78778 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 141539 B -417 B / -0.3% (better) 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 141255 B -449 B / -0.3% (better) 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3445580 B +242203 B / +7.6% (worse) 118590 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3424303 B +241527 B / +7.6% (worse) 101818 B +183 B / +0.2% (worse)
fmtprintf/j64-emscripten-memory64/LLGo 3153559 B +213240 B / +7.3% (worse) 125433 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 3178938 B +343150 B / +12.1% (worse) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 3042048 B +339827 B / +12.6% (worse) 0 B 0 B / 0.0%
j32-emscripten/LLGo 145505 B -700 B / -0.5% (better) 74890 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 144189 B -687 B / -0.5% (better) 73165 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 133758 B -89 B / -0.1% (better) 78778 B 0 B / +0.0%
reflectcall/j32-emscripten/LLGo 1570301 B +36216 B / +2.4% (worse) 92056 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1572736 B +35906 B / +2.3% (worse) 90331 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1454894 B +35637 B / +2.5% (worse) 98119 B +330 B / +0.3% (worse)
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1578914 B +38246 B / +2.5% (worse) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1502668 B +36457 B / +2.5% (worse) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 140756 B -414 B / -0.3% (better) 0 B 0 B / 0.0%
w32-wasi/LLGo 140543 B -446 B / -0.3% (better) 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 5.934 s +8.924 ms / +0.2% (worse)
j32-goos-js 6.242 s +401.1 ms / +6.9% (worse)
j64-emscripten-memory64 5.048 s -137.4 ms / -2.7% (better)
reflectcall/w32-wasi 26.758 s -452.8 ms / -1.7% (better)
w32-goos-wasip1 4.702 s -114.5 ms / -2.4% (better)
w32-wasi 4.937 s +313.5 ms / +6.8% (worse)

Compared with 15732a0d63d9 measured in the same runner job.

@visualfc
visualfc force-pushed the fix/static-iface-box branch from 4b3a724 to 6a8531a Compare September 30, 2026 07:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant