Skip to content

cl: fold trivial constant iface constructors - #2705

Open
visualfc wants to merge 3 commits into
xgo-dev:mainfrom
visualfc:fix/trivial-iface-fold
Open

visualfc wants to merge 3 commits into
xgo-dev:mainfrom
visualfc:fix/trivial-iface-fold

Conversation

@visualfc

Copy link
Copy Markdown
Collaborator

Summary

A call whose SSA body is MakeInterface of a Convert/ChangeType of a constant parameter is lowered at the call site to MakeInterface of that concrete value.

The callee is recognized by that SSA shape (single block, no //go:noinline), not by package path or function name. constant.MakeInt64 / MakeBool match because they are return int64Val(x) / return boolVal(b). User helpers such as func boxMyInt(x int64) any { return myInt(x) } use the same path. Functions with extra control flow (MakeString) are left as real calls.

Depends on: #2703 (static itabs) and #2704 (iface boxes). Without those, this fold would emit AllocU+NewItab per call. Merge after #2704. This branch contains #2703–#2704 plus this commit; only the latest commit is in scope.

Test plan

  • cl/_testrt/constiface (MakeInt64, MakeBool, and a local boxMyInt all become {itab, box})

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: fold trivial constant interface boxing

Solid optimization overall — the static itab / static ifacebox globals replace runtime NewItab and AllocU for compile-time constants, and the interequal fix correctly keeps T2I (static itab) and I2I (NewItab) representations of the same iface equal by comparing the (inter, _type) pair with proper nil guards. Verified: the removed LTO/compiler.used template comments are no longer stale, the staticItab hash matches the concrete type descriptor, and boxed constants are safely immutable under Go semantics (assertions copy out; no addressable path reaches the read-only box).

Two correctness concerns worth addressing before merge, plus one minor doc nit — inline.

Comment thread cl/constiface.go Outdated
case *ssa.ChangeType:
chain[t] = true
v = t.X
case *ssa.Convert:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Convert in fold chain reinterprets the raw constant under the wrong type

analyzeTrivialIfaceBox accepts *ssa.Convert in the chain and returns concrete = mi.X.Type() (the type after the conversion), while foldConstantMakeValue then applies the callee argument's raw constant via b.Const(c.Value, concrete) at cl/constiface.go:50 — skipping the conversion arithmetic.

This is correct only for representation-preserving conversions. For a value-changing Convert it is unsound. Example: func box(x int32) any { return any(string(x)) } called as box(65). Here concrete is string but c.Value is the integer 65; b.Const dispatches on the target kind (ssa/expr.go), hits the types.String branch, and calls constant.StringVal on an integer constant — producing a wrong result or a panic, where real Go yields any("A"). Numeric narrowing/widening and int↔float Convert have the same hazard.

*ssa.Convert in go/ssa is exactly the set of value-changing conversions, so including it here is unsafe. Suggest dropping *ssa.Convert from the recognized chain (keeping only ChangeType/ChangeInterface, which are representation-identical), or skip folding whenever a Convert is present.

Comment thread ssa/interface.go
if g.impl.IsNil() {
return llvm.Value{}, false
}
g.impl.SetGlobalConstant(true)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] staticIfaceBox mutates the shared zero-sized sentinel for zero-sized types

For a zero-sized value type (e.g. struct{}, [0]T, or a named type over them — none are directIfaceType, so MakeInterface routes them here), NewVarEx → doNewVarEx does not create a dedicated global. It returns the shared module zero-sized sentinel (__llgo.moduleZeroSizedAlloc$) as an isZeroSizedAlias Global (ssa/decl.go:209-219).

g.Init(x) is a no-op for that alias, but the following lines are not guarded and run on the shared sentinel:

g.impl.SetGlobalConstant(true)
g.impl.SetUnnamedAddr(true)
b.Pkg.setODRLinkage(g.impl, llvm.WeakODRLinkage)

Boxing any zero-sized constant thus flips the shared sentinel to GlobalConstant and rewrites its linkage (from LinkOnceODRLinkage, or re-COMDATs it on Windows) — a cross-cutting side effect on the address handed out for all zero-sized allocations in the module. It also pollutes the _llgo_ifacebox$ dedup cache with an entry pointing at the sentinel.

Suggest bailing out (return false) before the constant/linkage mutations when the resulting Global is a zero-sized alias — the isZeroSizedAlias field already exists on aGlobal, or guard on TypeAllocSize(storageType)==0. Zero-sized boxes don't benefit from this optimization anyway.

@codecov

codecov Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

d0f53831e9fe | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 7160 B 0 B / +0.0% 387 B 0 B / +0.0% 556.715 ms -7.867 ms / -1.4% (better) 1.318 ms -6.29 us / -0.5% (better)
Linux cprintf-lto 6912 B 0 B / +0.0% 368 B 0 B / +0.0% 570.021 ms -17.8 ms / -3.0% (better) 1.499 ms +159.9 us / +11.9% (worse)
Linux fmtprintf 1673184 B +9752 B / +0.6% (worse) 489761 B -8676 B / -1.7% (better) 3.820 s +133.1 ms / +3.6% (worse) 3.228 ms +52.58 us / +1.7% (worse)
Linux fmtprintf-lto 1502368 B +1200 B / +0.1% (worse) 426027 B -10096 B / -2.3% (better) 11.146 s -325.7 ms / -2.8% (better) 2.877 ms -241.9 us / -7.8% (better)
Linux println 69952 B +704 B / +1.0% (worse) 16599 B -206 B / -1.2% (better) 615.704 ms +33.49 ms / +5.8% (worse) 1.737 ms +71.11 us / +4.3% (worse)
Linux println-lto 60200 B +224 B / +0.4% (worse) 14025 B -184 B / -1.3% (better) 873.670 ms +46.06 ms / +5.6% (worse) 1.694 ms +87.85 us / +5.5% (worse)
macOS cprintf 68064 B 0 B / +0.0% 4429 B 0 B / +0.0% 673.052 ms -245.4 ms / -26.7% (better) 2.540 ms -717.9 us / -22.0% (better)
macOS cprintf-lto 84288 B 0 B / +0.0% 193 B 0 B / +0.0% 725.699 ms -130.6 ms / -15.3% (better) 2.376 ms -1.04 ms / -30.4% (better)
macOS fmtprintf 1534320 B +29488 B / +2.0% (worse) 868632 B -6852 B / -0.8% (better) 2.888 s -927.2 ms / -24.3% (better) 4.296 ms -376.5 us / -8.1% (better)
macOS fmtprintf-lto 1192672 B -32 B / -0.002683% (better) 842380 B -6600 B / -0.8% (better) 6.704 s -1.962 s / -22.6% (better) 3.880 ms -937.3 us / -19.5% (better)
macOS println 117824 B +608 B / +0.5% (worse) 37363 B -202 B / -0.5% (better) 697.733 ms -99.08 ms / -12.4% (better) 3.912 ms +312.6 us / +8.7% (worse)
macOS println-lto 119472 B 0 B / +0.0% 34775 B -169 B / -0.5% (better) 986.075 ms -498.4 ms / -33.6% (better) 3.494 ms -3.195 ms / -47.8% (better)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.311 s +13.72 ms / +1.1% (worse) 3.482 ms -118.4 us / -3.3% (better)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.337 s -31.89 ms / -2.3% (better) 3.514 ms -294.4 us / -7.7% (better)
Windows MinGW fmtprintf 1946624 B +13312 B / +0.7% (worse) 594566 B -4464 B / -0.7% (better) 4.277 s +215.7 ms / +5.3% (worse) 9.165 ms +1.194 ms / +15.0% (worse)
Windows MinGW fmtprintf-lto 1963520 B +5632 B / +0.3% (worse) 540630 B -6688 B / -1.2% (better) 10.589 s -140.4 ms / -1.3% (better) 8.444 ms -626.1 us / -6.9% (better)
Windows MinGW println 76288 B +512 B / +0.7% (worse) 24950 B -192 B / -0.8% (better) 1.302 s -24.84 ms / -1.9% (better) 6.724 ms -717.2 us / -9.6% (better)
Windows MinGW println-lto 70144 B +1024 B / +1.5% (worse) 21782 B -208 B / -0.9% (better) 1.542 s -8.009 ms / -0.5% (better) 6.816 ms -210.8 us / -3.0% (better)
Windows MinGW 386 cprintf 43520 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.301 s -19.98 ms / -1.5% (better) 5.766 ms +51.9 us / +0.9% (worse)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.352 s +40.25 ms / +3.1% (worse) 5.868 ms +420.5 us / +7.7% (worse)
Windows MinGW 386 fmtprintf 1905664 B +9216 B / +0.5% (worse) 466190 B -6272 B / -1.3% (better) 4.290 s +127.6 ms / +3.1% (worse) 11.817 ms +116.6 us / +1.0% (worse)
Windows MinGW 386 fmtprintf-lto 2196992 B +15872 B / +0.7% (worse) 445790 B -5484 B / -1.2% (better) 9.823 s -263.9 ms / -2.6% (better) 10.748 ms -197.1 us / -1.8% (better)
Windows MinGW 386 println 97792 B +1536 B / +1.6% (worse) 21266 B -208 B / -1.0% (better) 1.323 s +27.75 ms / +2.1% (worse) 9.571 ms -322.6 us / -3.3% (better)
Windows MinGW 386 println-lto 75264 B +1024 B / +1.4% (worse) 19138 B -196 B / -1.0% (better) 1.551 s +37.48 ms / +2.5% (worse) 9.744 ms +88.3 us / +0.9% (worse)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.587 s +21.42 ms / +1.4% (worse) 6.529 ms -346.1 us / -5.0% (better)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.628 s +31.48 ms / +2.0% (worse) 6.853 ms -169.7 us / -2.4% (better)
Windows MinGW ARM64 fmtprintf 1830912 B +11264 B / +0.6% (worse) 503620 B -6096 B / -1.2% (better) 4.259 s -40.89 ms / -1.0% (better) 13.194 ms +99.8 us / +0.8% (worse)
Windows MinGW ARM64 fmtprintf-lto 1889792 B +9728 B / +0.5% (worse) 470608 B -5568 B / -1.2% (better) 9.822 s +110.3 ms / +1.1% (worse) 13.616 ms -217.2 us / -1.6% (better)
Windows MinGW ARM64 println 73216 B +1024 B / +1.4% (worse) 23644 B -240 B / -1.0% (better) 1.586 s -20.28 ms / -1.3% (better) 11.799 ms -394.2 us / -3.2% (better)
Windows MinGW ARM64 println-lto 69120 B +512 B / +0.7% (worse) 21036 B -196 B / -0.9% (better) 1.799 s -22.94 ms / -1.3% (better) 12.122 ms +329.1 us / +2.8% (worse)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65798 B 0 B / +0.0% 1.073 s +975.3 us / +0.1% (worse) 3.378 ms +63.1 us / +1.9% (worse)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65734 B 0 B / +0.0% 1.112 s +10.98 ms / +1.0% (worse) 3.450 ms +77 us / +2.3% (worse)
Windows MSVC fmtprintf 1642496 B -1024 B / -0.1% (better) 690134 B -4448 B / -0.6% (better) 3.824 s +8.258 ms / +0.2% (worse) 9.062 ms -608.5 us / -6.3% (better)
Windows MSVC fmtprintf-lto 1631232 B -4096 B / -0.3% (better) 640438 B -6640 B / -1.0% (better) 8.914 s -106.5 ms / -1.2% (better) 9.525 ms +404.8 us / +4.4% (worse)
Windows MSVC println 195072 B +512 B / +0.3% (worse) 120630 B -192 B / -0.2% (better) 1.084 s -23.77 ms / -2.1% (better) 6.796 ms -80 us / -1.2% (better)
Windows MSVC println-lto 192512 B 0 B / +0.0% 118150 B -208 B / -0.2% (better) 1.295 s -144.4 ms / -10.0% (better) 6.811 ms -123.8 us / -1.8% (better)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.099 s +25.67 ms / +2.4% (worse) 5.517 ms -310.1 us / -5.3% (better)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.145 s -147.4 ms / -11.4% (better) 6.062 ms -172.6 us / -2.8% (better)
Windows MSVC 386 fmtprintf 1200640 B -4096 B / -0.3% (better) 449580 B -6272 B / -1.4% (better) 3.861 s -14.77 ms / -0.4% (better) 11.414 ms +362.4 us / +3.3% (worse)
Windows MSVC 386 fmtprintf-lto 1239040 B -2560 B / -0.2% (better) 421179 B -6032 B / -1.4% (better) 8.741 s +45.16 ms / +0.5% (worse) 11.328 ms +192.2 us / +1.7% (worse)
Windows MSVC 386 println 36352 B 0 B / +0.0% 20116 B -208 B / -1.0% (better) 1.104 s -5.197 ms / -0.5% (better) 9.578 ms -785.2 us / -7.6% (better)
Windows MSVC 386 println-lto 35328 B -512 B / -1.4% (better) 18373 B -192 B / -1.0% (better) 1.467 s +173.1 ms / +13.4% (worse) 11.659 ms +1.744 ms / +17.6% (worse)
Windows MSVC ARM64 cprintf 11776 B 0 B / +0.0% 4192 B 0 B / +0.0% 1.208 s -53.9 ms / -4.3% (better) 6.919 ms -118.9 us / -1.7% (better)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.240 s -49.14 ms / -3.8% (better) 6.854 ms -232.8 us / -3.3% (better)
Windows MSVC ARM64 fmtprintf 1384448 B -2560 B / -0.2% (better) 503560 B -6096 B / -1.2% (better) 3.797 s -99.02 ms / -2.5% (better) 13.348 ms -1.126 ms / -7.8% (better)
Windows MSVC ARM64 fmtprintf-lto 1404416 B -1024 B / -0.1% (better) 471284 B -5552 B / -1.2% (better) 8.765 s -176.6 ms / -2.0% (better) 13.934 ms +102.7 us / +0.7% (worse)
Windows MSVC ARM64 println 45056 B 0 B / +0.0% 23668 B -240 B / -1.0% (better) 1.234 s -33.54 ms / -2.6% (better) 11.821 ms -869.2 us / -6.8% (better)
Windows MSVC ARM64 println-lto 42496 B 0 B / +0.0% 21188 B -192 B / -0.9% (better) 1.391 s -43.61 ms / -3.0% (better) 11.979 ms -920.8 us / -7.1% (better)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.610 ns/op -0.4 ns/op / -2.7% (better)
Linux BenchmarkMergeCompilerFlags 193.800 ns/op -9.7 ns/op / -4.8% (better)
Linux BenchmarkMergeLinkerFlags 132.600 ns/op -3.1 ns/op / -2.3% (better)
Linux BenchmarkChannelBuffered 55.070 ns/op -0.37 ns/op / -0.7% (better)
Linux BenchmarkChannelHandoff 12858 ns/op -1047 ns/op / -7.5% (better)
Linux BenchmarkDefer 51.980 ns/op -1.91 ns/op / -3.5% (better)
Linux BenchmarkDirectCall 1.169 ns/op -0.385 ns/op / -24.8% (better)
Linux BenchmarkGlobalRead 1.165 ns/op -0.05 ns/op / -4.1% (better)
Linux BenchmarkGlobalWrite 7.793 ns/op +0.026 ns/op / +0.3% (worse)
Linux BenchmarkGoroutine 25656 ns/op -4413 ns/op / -14.7% (better)
Linux BenchmarkInterfaceCall 6.243 ns/op +0.2 ns/op / +3.3% (worse)
Linux BenchmarkRuntimeGetG 2.716 ns/op -0.183 ns/op / -6.3% (better)
macOS BenchmarkLookupPCRandom 13.050 ns/op +0.42 ns/op / +3.3% (worse)
macOS BenchmarkMergeCompilerFlags 101.400 ns/op -3.1 ns/op / -3.0% (better)
macOS BenchmarkMergeLinkerFlags 67.660 ns/op -0.32 ns/op / -0.5% (better)
macOS BenchmarkChannelBuffered 26.350 ns/op -0.02 ns/op / -0.1% (better)
macOS BenchmarkChannelHandoff 5382 ns/op -2007 ns/op / -27.2% (better)
macOS BenchmarkDefer 33.940 ns/op -4.85 ns/op / -12.5% (better)
macOS BenchmarkDirectCall 1.069 ns/op -0.02 ns/op / -1.8% (better)
macOS BenchmarkGlobalRead 1.073 ns/op -0.02 ns/op / -1.8% (better)
macOS BenchmarkGlobalWrite 1.163 ns/op +0.092 ns/op / +8.6% (worse)
macOS BenchmarkGoroutine 34229 ns/op -221 ns/op / -0.6% (better)
macOS BenchmarkInterfaceCall 3.925 ns/op -0.262 ns/op / -6.3% (better)
macOS BenchmarkRuntimeGetG 2.488 ns/op -1.093 ns/op / -30.5% (better)
Windows MinGW BenchmarkLookupPCRandom 12.310 ns/op -0.05 ns/op / -0.4% (better)
Windows MinGW BenchmarkMergeCompilerFlags 533.100 ns/op -11 ns/op / -2.0% (better)
Windows MinGW BenchmarkMergeLinkerFlags 472.300 ns/op +7.1 ns/op / +1.5% (worse)
Windows MinGW BenchmarkChannelBuffered 31.130 ns/op -0.86 ns/op / -2.7% (better)
Windows MinGW BenchmarkChannelHandoff 1348 ns/op +1 ns/op / +0.1% (worse)
Windows MinGW BenchmarkDefer 55.150 ns/op -0.17 ns/op / -0.3% (better)
Windows MinGW BenchmarkDirectCall 1.749 ns/op +0.002 ns/op / +0.1% (worse)
Windows MinGW BenchmarkGlobalRead 1.785 ns/op +0.034 ns/op / +1.9% (worse)
Windows MinGW BenchmarkGlobalWrite 2.790 ns/op -0.001 ns/op / -0.03583% (better)
Windows MinGW BenchmarkGoroutine 76833 ns/op -6457 ns/op / -7.8% (better)
Windows MinGW BenchmarkInterfaceCall 9.679 ns/op +1.221 ns/op / +14.4% (worse)
Windows MinGW BenchmarkRuntimeGetG 2.098 ns/op -0.007 ns/op / -0.3% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 26.530 ns/op -0.03 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkMergeCompilerFlags 746.400 ns/op -3.1 ns/op / -0.4% (better)
Windows MinGW 386 BenchmarkMergeLinkerFlags 688.300 ns/op +5.8 ns/op / +0.8% (worse)
Windows MinGW 386 BenchmarkChannelBuffered 39.080 ns/op +0.3 ns/op / +0.8% (worse)
Windows MinGW 386 BenchmarkChannelHandoff 838 ns/op +18.1 ns/op / +2.2% (worse)
Windows MinGW 386 BenchmarkDefer 44.170 ns/op +1.89 ns/op / +4.5% (worse)
Windows MinGW 386 BenchmarkDirectCall 1.548 ns/op 0 ns/op / +0.0%
Windows MinGW 386 BenchmarkGlobalRead 1.550 ns/op -0.001 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGlobalWrite 7.776 ns/op -0.004 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGoroutine 108897 ns/op +3152 ns/op / +3.0% (worse)
Windows MinGW 386 BenchmarkInterfaceCall 8.007 ns/op -0.35 ns/op / -4.2% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 2.170 ns/op +0.243 ns/op / +12.6% (worse)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.110 ns/op +0.1 ns/op / +0.8% (worse)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 632 ns/op +65.6 ns/op / +11.6% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 581.600 ns/op +42.7 ns/op / +7.9% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 37.490 ns/op -1.68 ns/op / -4.3% (better)
Windows MinGW ARM64 BenchmarkChannelHandoff 2563 ns/op -150 ns/op / -5.5% (better)
Windows MinGW ARM64 BenchmarkDefer 57.920 ns/op -0.07 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkDirectCall 0.884 ns/op +0.2948 ns/op / +50.0% (worse)
Windows MinGW ARM64 BenchmarkGlobalRead 0.590 ns/op -0.0731 ns/op / -11.0% (better)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.884 ns/op +0.2211 ns/op / +33.3% (worse)
Windows MinGW ARM64 BenchmarkGoroutine 65196 ns/op -928 ns/op / -1.4% (better)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.143 ns/op +0.001 ns/op / +0.02414% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.772 ns/op +0.003 ns/op / +0.2% (worse)
Windows MSVC BenchmarkLookupPCRandom 13.230 ns/op +0.1 ns/op / +0.8% (worse)
Windows MSVC BenchmarkMergeCompilerFlags 618.900 ns/op +12.1 ns/op / +2.0% (worse)
Windows MSVC BenchmarkMergeLinkerFlags 545 ns/op -5.1 ns/op / -0.9% (better)
Windows MSVC BenchmarkChannelBuffered 28.610 ns/op +0.67 ns/op / +2.4% (worse)
Windows MSVC BenchmarkChannelHandoff 1079 ns/op -7 ns/op / -0.6% (better)
Windows MSVC BenchmarkDefer 53.380 ns/op -1.24 ns/op / -2.3% (better)
Windows MSVC BenchmarkDirectCall 1.857 ns/op +0.311 ns/op / +20.1% (worse)
Windows MSVC BenchmarkGlobalRead 1.547 ns/op 0 ns/op / +0.0%
Windows MSVC BenchmarkGlobalWrite 2.471 ns/op +0.001 ns/op / +0.04049% (worse)
Windows MSVC BenchmarkGoroutine 89114 ns/op +780 ns/op / +0.9% (worse)
Windows MSVC BenchmarkInterfaceCall 9.006 ns/op +0.63 ns/op / +7.5% (worse)
Windows MSVC BenchmarkRuntimeGetG 1.860 ns/op -0.314 ns/op / -14.4% (better)
Windows MSVC 386 BenchmarkLookupPCRandom 26.560 ns/op +0.06 ns/op / +0.2% (worse)
Windows MSVC 386 BenchmarkMergeCompilerFlags 746.900 ns/op -19.4 ns/op / -2.5% (better)
Windows MSVC 386 BenchmarkMergeLinkerFlags 688.900 ns/op +5.9 ns/op / +0.9% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 39.010 ns/op -0.24 ns/op / -0.6% (better)
Windows MSVC 386 BenchmarkChannelHandoff 828.700 ns/op -59.5 ns/op / -6.7% (better)
Windows MSVC 386 BenchmarkDefer 48.270 ns/op +4 ns/op / +9.0% (worse)
Windows MSVC 386 BenchmarkDirectCall 1.548 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGlobalRead 1.550 ns/op -0.309 ns/op / -16.6% (better)
Windows MSVC 386 BenchmarkGlobalWrite 7.788 ns/op +0.006 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGoroutine 110290 ns/op +4276 ns/op / +4.0% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 8.355 ns/op -0.011 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkRuntimeGetG 2.170 ns/op +0.244 ns/op / +12.7% (worse)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.100 ns/op -0.04 ns/op / -0.3% (better)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 583.900 ns/op +10.7 ns/op / +1.9% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 545.400 ns/op +3.8 ns/op / +0.7% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 39.020 ns/op +1.27 ns/op / +3.4% (worse)
Windows MSVC ARM64 BenchmarkChannelHandoff 2172 ns/op +169 ns/op / +8.4% (worse)
Windows MSVC ARM64 BenchmarkDefer 62.470 ns/op -0.6 ns/op / -1.0% (better)
Windows MSVC ARM64 BenchmarkDirectCall 0.590 ns/op +0.0003 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkGlobalRead 0.884 ns/op +0.2201 ns/op / +33.1% (worse)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.760 ns/op +0.008 ns/op / +0.2% (worse)
Windows MSVC ARM64 BenchmarkGoroutine 58859 ns/op +534 ns/op / +0.9% (worse)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.234 ns/op +0.074 ns/op / +1.8% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.803 ns/op +0.032 ns/op / +1.8% (worse)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 908.800 ns/op -0.2 ns/op / -0.022% (better)
Linux AfterFuncZeroDelivery/LLGo 38982 ns/op +738 ns/op / +1.9% (worse)
Linux CreateStop/Go 295.300 ns/op +3.1 ns/op / +1.1% (worse)
Linux CreateStop/LLGo 1788 ns/op +31 ns/op / +1.8% (worse)
Linux RearmStopped/Go 115.900 ns/op +1.2 ns/op / +1.0% (worse)
Linux RearmStopped/LLGo 1189 ns/op -254 ns/op / -17.6% (better)
Linux ResetActive/Go 68.540 ns/op +1.09 ns/op / +1.6% (worse)
Linux ResetActive/LLGo 732.700 ns/op +11.7 ns/op / +1.6% (worse)
Linux ResetHeap1024/Go 67.110 ns/op -0.02 ns/op / -0.02979% (better)
Linux ResetHeap1024/LLGo 179.700 ns/op +4.5 ns/op / +2.6% (worse)
macOS AfterFuncZeroDelivery/Go 447.600 ns/op -17.5 ns/op / -3.8% (better)
macOS AfterFuncZeroDelivery/LLGo 71164 ns/op -721 ns/op / -1.0% (better)
macOS CreateStop/Go 139.700 ns/op -1.9 ns/op / -1.3% (better)
macOS CreateStop/LLGo 449.800 ns/op -19.3 ns/op / -4.1% (better)
macOS RearmStopped/Go 58.390 ns/op -0.62 ns/op / -1.1% (better)
macOS RearmStopped/LLGo 368.700 ns/op +33.8 ns/op / +10.1% (worse)
macOS ResetActive/Go 45.090 ns/op +2.26 ns/op / +5.3% (worse)
macOS ResetActive/LLGo 170.400 ns/op +20.4 ns/op / +13.6% (worse)
macOS ResetHeap1024/Go 43.460 ns/op -1.49 ns/op / -3.3% (better)
macOS ResetHeap1024/LLGo 86.180 ns/op -7.58 ns/op / -8.1% (better)
Windows MinGW AfterFuncZeroDelivery/Go 484.700 ns/op -13.4 ns/op / -2.7% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 155423 ns/op +1880 ns/op / +1.2% (worse)
Windows MinGW CreateStop/Go 116.800 ns/op +0.1 ns/op / +0.1% (worse)
Windows MinGW CreateStop/LLGo 462.300 ns/op -13.9 ns/op / -2.9% (better)
Windows MinGW RearmStopped/Go 31.520 ns/op -0.12 ns/op / -0.4% (better)
Windows MinGW RearmStopped/LLGo 289.900 ns/op +6.3 ns/op / +2.2% (worse)
Windows MinGW ResetActive/Go 19.090 ns/op -0.02 ns/op / -0.1% (better)
Windows MinGW ResetActive/LLGo 160.600 ns/op -4.5 ns/op / -2.7% (better)
Windows MinGW ResetHeap1024/Go 19.170 ns/op +0.05 ns/op / +0.3% (worse)
Windows MinGW ResetHeap1024/LLGo 140.200 ns/op +1.4 ns/op / +1.0% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/Go 955.800 ns/op +17.9 ns/op / +1.9% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 204877 ns/op -1049 ns/op / -0.5% (better)
Windows MinGW 386 CreateStop/Go 191.800 ns/op -0.7 ns/op / -0.4% (better)
Windows MinGW 386 CreateStop/LLGo 489.700 ns/op -6.1 ns/op / -1.2% (better)
Windows MinGW 386 RearmStopped/Go 63.390 ns/op +0.15 ns/op / +0.2% (worse)
Windows MinGW 386 RearmStopped/LLGo 340.500 ns/op +4.6 ns/op / +1.4% (worse)
Windows MinGW 386 ResetActive/Go 39.010 ns/op +0.18 ns/op / +0.5% (worse)
Windows MinGW 386 ResetActive/LLGo 960.700 ns/op -30.5 ns/op / -3.1% (better)
Windows MinGW 386 ResetHeap1024/Go 39.530 ns/op +0.26 ns/op / +0.7% (worse)
Windows MinGW 386 ResetHeap1024/LLGo 187.500 ns/op -0.6 ns/op / -0.3% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 652.800 ns/op -13.2 ns/op / -2.0% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 155116 ns/op +121 ns/op / +0.1% (worse)
Windows MinGW ARM64 CreateStop/Go 194.500 ns/op -1.7 ns/op / -0.9% (better)
Windows MinGW ARM64 CreateStop/LLGo 362.300 ns/op +4.4 ns/op / +1.2% (worse)
Windows MinGW ARM64 RearmStopped/Go 70.650 ns/op +0.05 ns/op / +0.1% (worse)
Windows MinGW ARM64 RearmStopped/LLGo 252.500 ns/op -0.5 ns/op / -0.2% (better)
Windows MinGW ARM64 ResetActive/Go 31.040 ns/op +0.13 ns/op / +0.4% (worse)
Windows MinGW ARM64 ResetActive/LLGo 118.600 ns/op -9.5 ns/op / -7.4% (better)
Windows MinGW ARM64 ResetHeap1024/Go 31.070 ns/op +0.06 ns/op / +0.2% (worse)
Windows MinGW ARM64 ResetHeap1024/LLGo 128.800 ns/op +2 ns/op / +1.6% (worse)
Windows MSVC AfterFuncZeroDelivery/Go 556.800 ns/op -0.4 ns/op / -0.1% (better)
Windows MSVC AfterFuncZeroDelivery/LLGo 179192 ns/op +385 ns/op / +0.2% (worse)
Windows MSVC CreateStop/Go 117.600 ns/op +0.9 ns/op / +0.8% (worse)
Windows MSVC CreateStop/LLGo 426.400 ns/op -110.3 ns/op / -20.6% (better)
Windows MSVC RearmStopped/Go 31.610 ns/op +0.33 ns/op / +1.1% (worse)
Windows MSVC RearmStopped/LLGo 268.500 ns/op +15.7 ns/op / +6.2% (worse)
Windows MSVC ResetActive/Go 20.410 ns/op +0.28 ns/op / +1.4% (worse)
Windows MSVC ResetActive/LLGo 132.800 ns/op -7.7 ns/op / -5.5% (better)
Windows MSVC ResetHeap1024/Go 20.590 ns/op +0.19 ns/op / +0.9% (worse)
Windows MSVC ResetHeap1024/LLGo 125.700 ns/op +0.7 ns/op / +0.6% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/Go 961.900 ns/op -3.5 ns/op / -0.4% (better)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 199116 ns/op +1187 ns/op / +0.6% (worse)
Windows MSVC 386 CreateStop/Go 198.500 ns/op +5.9 ns/op / +3.1% (worse)
Windows MSVC 386 CreateStop/LLGo 482.400 ns/op +6.4 ns/op / +1.3% (worse)
Windows MSVC 386 RearmStopped/Go 63.360 ns/op -0.15 ns/op / -0.2% (better)
Windows MSVC 386 RearmStopped/LLGo 324.400 ns/op +11.6 ns/op / +3.7% (worse)
Windows MSVC 386 ResetActive/Go 38.980 ns/op +0.01 ns/op / +0.02566% (worse)
Windows MSVC 386 ResetActive/LLGo 963.800 ns/op -4.4 ns/op / -0.5% (better)
Windows MSVC 386 ResetHeap1024/Go 39.320 ns/op -0.03 ns/op / -0.1% (better)
Windows MSVC 386 ResetHeap1024/LLGo 173.900 ns/op +1.3 ns/op / +0.8% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 662.300 ns/op -5.9 ns/op / -0.9% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 148947 ns/op -1481 ns/op / -1.0% (better)
Windows MSVC ARM64 CreateStop/Go 203.100 ns/op -2.5 ns/op / -1.2% (better)
Windows MSVC ARM64 CreateStop/LLGo 441.600 ns/op -7.3 ns/op / -1.6% (better)
Windows MSVC ARM64 RearmStopped/Go 70.580 ns/op +0.03 ns/op / +0.04252% (worse)
Windows MSVC ARM64 RearmStopped/LLGo 281.200 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 ResetActive/Go 31.090 ns/op -0.02 ns/op / -0.1% (better)
Windows MSVC ARM64 ResetActive/LLGo 138.800 ns/op -7.2 ns/op / -4.9% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.120 ns/op -0.03 ns/op / -0.1% (better)
Windows MSVC ARM64 ResetHeap1024/LLGo 135.100 ns/op -0.5 ns/op / -0.4% (better)

Compared with f06143ba2aac measured in the same runner job.

@visualfc
visualfc force-pushed the fix/trivial-iface-fold branch from b0f449e to aaa3c6e Compare September 30, 2026 05:18
@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

d0f53831e9fe | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 153510 B -683 B / -0.4% (better) 88742 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 151500 B -795 B / -0.5% (better) 73165 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 141143 B -105 B / -0.1% (better) 92630 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 152668 B -180 B / -0.1% (better) 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 153380 B -180 B / -0.1% (better) 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3189496 B -23932 B / -0.7% (better) 132442 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3168726 B -23797 B / -0.7% (better) 101635 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2929079 B -20954 B / -0.7% (better) 139285 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2333963 B -13388 B / -0.6% (better) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2330582 B -13431 B / -0.6% (better) 0 B 0 B / 0.0%
j32-emscripten/LLGo 152744 B -684 B / -0.4% (better) 88742 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 150972 B -796 B / -0.5% (better) 73165 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 140468 B -114 B / -0.1% (better) 92629 B -1 B / -0.00108% (better)
reflectcall/j32-emscripten/LLGo 1537919 B -5852 B / -0.4% (better) 105908 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1540185 B -6037 B / -0.4% (better) 90331 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1424269 B -4157 B / -0.3% (better) 111641 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1275225 B -1468 B / -0.1% (better) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1272820 B -1481 B / -0.1% (better) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 152315 B -180 B / -0.1% (better) 0 B 0 B / 0.0%
w32-wasi/LLGo 153027 B -180 B / -0.1% (better) 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 6.220 s -515.3 ms / -7.7% (better)
j32-goos-js 5.937 s -548.2 ms / -8.5% (better)
j64-emscripten-memory64 5.382 s -554.7 ms / -9.3% (better)
reflectcall/w32-wasi 24.223 s -149.8 ms / -0.6% (better)
w32-goos-wasip1 4.193 s -456.1 ms / -9.8% (better)
w32-wasi 4.038 s -489 ms / -10.8% (better)

Compared with f06143ba2aac measured in the same runner job.

@visualfc
visualfc force-pushed the fix/trivial-iface-fold branch 2 times, most recently from 2780773 to 347be55 Compare September 30, 2026 12:54
Known concrete-to-interface conversions emit a read-only _llgo_itab$
global (Inter, Type, Hash, Fun) with weak_odr linkage, matching
cmd/compile's go:itab symbols.

Ordinary builds use that global as the runtime vtable. The T2I site is
the only reference, so --gc-sections/-dead_strip can drop itabs that
belong to dead functions. LTO and deadcode-drop keep calling NewItab so
unused interface methods can be dropped; Fun[] would pin them. LTO still
emits the static template for de-virt and erases unused templates after
the plugin runs.

Hash is copied from the type descriptor when present. interequal
compares the (inter, _type) pair so a static itab and a dynamically
allocated itab for the same conversion compare equal.
Non-direct interface values (integers, bools, small aggregates) still
use IfaceIndir, matching Go: data is a pointer to a copy. Compile-time
constants go in a WeakODR _llgo_ifacebox$ global instead of AllocU,
like cmd/compile's static temps.

Box names hash the constant's bits, not LLVM's printed form, so
same-length strings in different packages do not collide under
Windows COMDAT. Repeated boxing of the same LLVM value reuses a
per-package pointer cache.

Depends on static itabs so a constant T2I can be {itab, box} with no
runtime call.
A call whose SSA body is MakeInterface of a Convert/ChangeType of a
constant parameter is lowered at the call site to MakeInterface of
that concrete value. No function-name matching: constant.MakeInt64
and user helpers such as boxMyInt(x int64) any { return myInt(x) }
use the same path.

Depends on static itabs and iface boxes so the folded value is a
compile-time {itab, box} pair.
@visualfc
visualfc force-pushed the fix/trivial-iface-fold branch from 347be55 to d0f5383 Compare September 30, 2026 14:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant