Skip to content

cl: lower large constant map literals to a loop - #2706

Open
visualfc wants to merge 4 commits into
xgo-dev:mainfrom
visualfc:fix/static-maplit
Open

visualfc wants to merge 4 commits into
xgo-dev:mainfrom
visualfc:fix/static-maplit

Conversation

@visualfc

Copy link
Copy Markdown
Collaborator

Summary

Map literals with more than 25 constant keys become a counted mapassign loop over private constant key/value arrays, matching cmd/compile's maplit.

Composite values whose fields are constants or trivial iface constructors (the SSA shape from #2705) are rebuilt as LLVM constants so the value array is a ConstArray. Mixed MakeInt64/MakeBool-style boxes in one literal are allowed. Identification is by SSA (MakeMap, constant keys, complit fields), not by function name.

Depends on: #2703, #2704, and #2705. The large {string, constant.Value} literals need folded constant ifaces. Merge after #2705. This branch contains #2703–#2705 plus this commit; only the latest commit is in scope.

Test plan

  • go test ./cl -run TestCollect
  • cl/_testrt/maplit (map[string]int and mixed int/bool complits)

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: static maplit + static itab/interface boxing

This is a well-structured change. The interface-representation split (T2I now uses static itab globals while I2I still goes through NewItab) is handled carefully, and the interequal fix correctly restores cross-representation equality. interhash hashes by tab._type, and both itab representations point _type/hash at the same ODR-merged type descriptor, so map insert/lookup stays consistent. The maplit-to-loop lowering is safe by construction (keys required constant, arrays synthesized at known length).

Findings below are mostly correctness-hardening and readability; none are blocking.

Verified sound: full 32-byte sha256 in box/itab names (no truncation collision), mapLitIndexAddr only indexes fixed-length synthesized arrays, static globals are SetGlobalConstant(true) and never mutated after Init, and the staticItab.hash computation matches abiCommonFields.

Comment thread cl/maplit.go Outdated
if classifyStaticComplit(updates, plan) {
plan.kind = mapLitStaticComplit
} else {
plan.skip = nil

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] classifyStaticComplit mutates plan on failure; plan.skip=nil is misleading

classifyStaticComplit mutates the passed plan (sets plan.structTy at the top, appends to skip) before it may return false on a later entry. After a false return, plan.structTy is left populated while plan.kind stays mapLitGeneric. The else { plan.skip = nil } here only resets one of the two mutated fields, and plan.skip is already nil on a freshly-allocated plan, so the line is both dead and misleading. This works today only because the generic branch of compileMapLitUpdate never reads plan.structTy — a fragile invariant. Suggest making classifyStaticComplit compute structTy/skip in locals and commit to plan only on success (return them, assign in the caller), removing the need for plan.skip = nil.

Comment thread cl/maplit.go
return false
}
if block != afterBlock {
return block.Index > afterBlock.Index

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] instrIsAfter treats BasicBlock.Index as execution order for all functions

instrIsAfter uses block.Index > afterBlock.Index as "executes after". Block index is the position in fn.Blocks (roughly RPO for x/tools SSA) and does not guarantee execution/dominance order across branches. collectLargeMapLits runs for every function (compile.go:759), not just the synthetic package initializer. For a large map literal inside a normal function with branching, a makeMap referrer in a higher-indexed block that is actually on a parallel/earlier path could be wrongly judged "after" the last update, permitting an unsafe delay. The sibling static-map-init pass gates on initFn.Synthetic == "package initializer". Consider restricting this optimization to synthetic init functions, or strengthening the ordering check (e.g. real dominance) and documenting the assumption.

Comment thread cl/constiface.go
if !ok || c.Value == nil {
return llssa.Expr{}, false
}
x := b.Const(c.Value, p.type_(concrete, llssa.InGo))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Constant re-typing in foldConstantMakeValue may skip Convert semantics

analyzeTrivialIfaceBox accepts a *ssa.Convert in the unwrap chain and returns mi.X.Type() as the concrete type. Here the source constant c has the parameter's type, but the value is materialized directly at concrete (= mi.X.Type()). If the chain contains a real numeric Convert (e.g. param int32 widened/narrowed to int64, or a representation-changing conversion), building b.Const(c.Value, concrete) reinterprets the untyped constant instead of applying the conversion, which can diverge from actually calling the constructor. ChangeType/ChangeInterface are representation-preserving and safe. Consider rejecting *ssa.Convert in the accepted chain (or applying the conversion to the constant). Low likelihood in practice, but worth hardening.

Comment thread ssa/interface.go
rawIntf.NumMethods() == 0 || concrete == nil {
if rawIntf.NumMethods() == 0 || concrete == nil {
return Expr{}, false
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] staticItab does method-set work before the dedup cache check

staticItab computes types.NewMethodSet(concrete) plus per-method Lookup and AssignableTo before it builds the cache key name and checks b.Pkg.VarOf(name). Since the itab is deduplicated per (interface, concrete) pair and the cache key does not need the method set, repeated boxings of the same pair recompute the full method set only to discard it on the cache hit. Hoisting the VarOf(name) early-return above the method-set work would make repeated boxings O(1).

Comment thread ssa/datastruct.go
return !v.impl.IsNil() && !v.impl.IsAConstant().IsNil()
}

var mapLitSeq atomic.Uint64

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] mapLitArray globals use a non-dedup, order-dependent counter

mapLitArray names emitted constant arrays via a process-global atomic.Uint64 (mapLitSeq), unlike the other two globals in this PR (staticItab, staticIfaceBox) which are content-addressed and thus deduplicated + deterministic. Consequences: identical constant map-literal arrays never merge (code/data-size bloat), and names depend on compilation order, hurting build reproducibility under parallel/reordered compilation. Content-hashing the array would fix both.

Comment thread runtime/internal/runtime/alg.go Outdated
if x.tab == y.tab {
return ifaceeq(x.tab, x.data, y.data)
}
// T2I uses a static itab global; I2I still goes through NewItab.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] interequal/foldConstantMakeValue comments slightly overstate the mapping

Two minor doc nits: (1) alg.go's // T2I uses a static itab global; I2I still goes through NewItab — T2I falls back to NewItab too when staticItab returns false (empty interface, unassignable, no-interface-method), so it is not a strict one-to-one mapping; the equality logic is still correct. (2) foldConstantMakeValue's doc lists Convert/ChangeType but the recognizer (analyzeTrivialIfaceBox) also accepts ChangeInterface; align the two comments.

Comment thread ssa/interface.go
}
g := b.Pkg.NewVarEx(name, prog.Pointer(typ))
g.Init(x)
if g.impl.IsNil() {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] staticIfaceBox: potential leftover global if Init fails

If g.Init(x) leaves g.impl nil, the function returns (zero, false), but the named global was already created via NewVarEx. A later call with the same name would hit the VarOf(name) cache and return that leftover, uninitialized global with true. Confirm NewVarEx does not register the var when init fails, or clean up on the failure path. Also worth a one-line note that x.impl.String() is used purely as a dedup key so correctness does not depend on LLVM's textual-print stability.

@codecov

codecov Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 81.36646% with 30 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
ssa/datastruct.go 79.31% 12 Missing ⚠️
internal/build/main_module.go 65.38% 9 Missing ⚠️
ssa/interface.go 87.30% 8 Missing ⚠️
ssa/abitype.go 92.30% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

ab860f11ba65 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 7280 B +120 B / +1.7% (worse) 393 B +6 B / +1.6% (worse) 682.955 ms +124.3 ms / +22.2% (worse) 1.656 ms +387.6 us / +30.6% (worse)
Linux cprintf-lto 7112 B +200 B / +2.9% (worse) 374 B +6 B / +1.6% (worse) 697.446 ms +53.99 ms / +8.4% (worse) 1.292 ms -15.46 us / -1.2% (better)
Linux fmtprintf 1786208 B +123936 B / +7.5% (worse) 545733 B +47374 B / +9.5% (worse) 4.078 s +271.7 ms / +7.1% (worse) 3.441 ms +231.4 us / +7.2% (worse)
Linux fmtprintf-lto 1603960 B +103576 B / +6.9% (worse) 473719 B +37606 B / +8.6% (worse) 12.941 s +1.679 s / +14.9% (worse) 3.114 ms +71.42 us / +2.3% (worse)
Linux println 69824 B +832 B / +1.2% (worse) 16583 B -200 B / -1.2% (better) 625.895 ms +37.07 ms / +6.3% (worse) 1.824 ms +149.1 us / +8.9% (worse)
Linux println-lto 60224 B +376 B / +0.6% (worse) 14021 B -178 B / -1.3% (better) 880.125 ms +40.67 ms / +4.8% (worse) 1.642 ms -124.3 us / -7.0% (better)
macOS cprintf 68096 B +32 B / +0.04701% (worse) 4437 B +8 B / +0.2% (worse) 791.816 ms -53.39 ms / -6.3% (better) 3.173 ms +67 us / +2.2% (worse)
macOS cprintf-lto 84288 B 0 B / +0.0% 193 B 0 B / +0.0% 1.231 s +305.3 ms / +33.0% (worse) 7.107 ms +3.58 ms / +101.5% (worse)
macOS fmtprintf 1624560 B +119888 B / +8.0% (worse) 928272 B +53700 B / +6.1% (worse) 3.798 s +942.7 ms / +33.0% (worse) 8.020 ms +3.026 ms / +60.6% (worse)
macOS fmtprintf-lto 1275120 B +82416 B / +6.9% (worse) 896044 B +47828 B / +5.6% (worse) 10.184 s +1.183 s / +13.1% (worse) 7.417 ms +2.358 ms / +46.6% (worse)
macOS println 117776 B +640 B / +0.5% (worse) 37227 B -194 B / -0.5% (better) 1.010 s +63.62 ms / +6.7% (worse) 5.212 ms +459.2 us / +9.7% (worse)
macOS println-lto 119472 B 0 B / +0.0% 34655 B -169 B / -0.5% (better) 946.542 ms +96.93 ms / +11.4% (worse) 3.670 ms +110.2 us / +3.1% (worse)
Windows MinGW cprintf 19968 B +512 B / +2.6% (worse) 4566 B +16 B / +0.4% (worse) 991.390 ms +4.183 ms / +0.4% (worse) 2.822 ms -38.6 us / -1.3% (better)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.025 s -41.06 ms / -3.9% (better) 2.842 ms +43.2 us / +1.5% (worse)
Windows MinGW fmtprintf 2108416 B +175616 B / +9.1% (worse) 677382 B +78432 B / +13.1% (worse) 3.405 s +195.4 ms / +6.1% (worse) 6.593 ms +384.4 us / +6.2% (worse)
Windows MinGW fmtprintf-lto 2122752 B +165376 B / +8.4% (worse) 612822 B +65504 B / +12.0% (worse) 9.436 s +1.31 s / +16.1% (worse) 6.233 ms +150.8 us / +2.5% (worse)
Windows MinGW println 76288 B +512 B / +0.7% (worse) 24950 B -192 B / -0.8% (better) 1.029 s -19.59 ms / -1.9% (better) 5.067 ms -371 us / -6.8% (better)
Windows MinGW println-lto 70144 B +1024 B / +1.5% (worse) 21782 B -208 B / -0.9% (better) 1.172 s -31.54 ms / -2.6% (better) 5.213 ms -532.2 us / -9.3% (better)
Windows MinGW 386 cprintf 43520 B 0 B / +0.0% 5342 B +16 B / +0.3% (worse) 1.314 s +43.88 ms / +3.5% (worse) 5.044 ms -861.2 us / -14.6% (better)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.336 s -40.6 ms / -2.9% (better) 5.022 ms -576.5 us / -10.3% (better)
Windows MinGW 386 fmtprintf 2025984 B +129536 B / +6.8% (worse) 520718 B +48304 B / +10.2% (worse) 4.321 s +101.1 ms / +2.4% (worse) 10.672 ms -317.7 us / -2.9% (better)
Windows MinGW 386 fmtprintf-lto 2394624 B +214016 B / +9.8% (worse) 493286 B +42028 B / +9.3% (worse) 10.934 s +1.041 s / +10.5% (worse) 10.553 ms +14.1 us / +0.1% (worse)
Windows MinGW 386 println 97792 B +1536 B / +1.6% (worse) 21250 B -208 B / -1.0% (better) 1.293 s -78.56 ms / -5.7% (better) 9.523 ms -147.9 us / -1.5% (better)
Windows MinGW 386 println-lto 75264 B +1024 B / +1.4% (worse) 19122 B -192 B / -1.0% (better) 1.570 s +41.36 ms / +2.7% (worse) 9.260 ms +777.1 us / +9.2% (worse)
Windows MinGW ARM64 cprintf 19456 B +512 B / +2.7% (worse) 4424 B +16 B / +0.4% (worse) 1.517 s +13.34 ms / +0.9% (worse) 5.928 ms -283.2 us / -4.6% (better)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.536 s +6.753 ms / +0.4% (worse) 6.089 ms -35.5 us / -0.6% (better)
Windows MinGW ARM64 fmtprintf 1969664 B +150528 B / +8.3% (worse) 566900 B +57292 B / +11.2% (worse) 4.323 s +72.79 ms / +1.7% (worse) 12.681 ms -86.5 us / -0.7% (better)
Windows MinGW ARM64 fmtprintf-lto 2037760 B +158720 B / +8.4% (worse) 526512 B +50360 B / +10.6% (worse) 10.409 s +624.9 ms / +6.4% (worse) 13.304 ms +94 us / +0.7% (worse)
Windows MinGW ARM64 println 73216 B +1024 B / +1.4% (worse) 23644 B -232 B / -1.0% (better) 1.548 s +37.94 ms / +2.5% (worse) 10.991 ms +676.3 us / +6.6% (worse)
Windows MinGW ARM64 println-lto 69120 B +512 B / +0.7% (worse) 21028 B -196 B / -0.9% (better) 1.740 s +22.78 ms / +1.3% (worse) 10.750 ms +26.6 us / +0.2% (worse)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65814 B +16 B / +0.02432% (worse) 970.845 ms -22.99 ms / -2.3% (better) 3.321 ms -37.8 us / -1.1% (better)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65734 B 0 B / +0.0% 995.744 ms -37.62 ms / -3.6% (better) 2.844 ms -605.4 us / -17.6% (better)
Windows MSVC fmtprintf 1772032 B +129024 B / +7.9% (worse) 772902 B +78400 B / +11.3% (worse) 3.957 s +291.9 ms / +8.0% (worse) 9.860 ms +938.4 us / +10.5% (worse)
Windows MSVC fmtprintf-lto 1761792 B +126976 B / +7.8% (worse) 713238 B +66176 B / +10.2% (worse) 9.647 s +1.677 s / +21.0% (worse) 8.989 ms +1.218 ms / +15.7% (worse)
Windows MSVC println 195072 B +512 B / +0.3% (worse) 120630 B -192 B / -0.2% (better) 988.234 ms -27.18 ms / -2.7% (better) 7.278 ms +59.8 us / +0.8% (worse)
Windows MSVC println-lto 192512 B 0 B / +0.0% 118134 B -208 B / -0.2% (better) 1.303 s +121.9 ms / +10.3% (worse) 7.436 ms +621.6 us / +9.1% (worse)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.132 s -165.4 ms / -12.7% (better) 5.812 ms +670.9 us / +13.0% (worse)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.165 s +74.83 ms / +6.9% (worse) 5.864 ms +350.5 us / +6.4% (worse)
Windows MSVC 386 fmtprintf 1280512 B +76288 B / +6.3% (worse) 504108 B +48304 B / +10.6% (worse) 4.011 s +163 ms / +4.2% (worse) 11.371 ms +324.8 us / +2.9% (worse)
Windows MSVC 386 fmtprintf-lto 1330176 B +89088 B / +7.2% (worse) 467275 B +40080 B / +9.4% (worse) 9.792 s +977.2 ms / +11.1% (worse) 13.203 ms +1.797 ms / +15.8% (worse)
Windows MSVC 386 println 36352 B 0 B / +0.0% 20100 B -224 B / -1.1% (better) 1.129 s +48.77 ms / +4.5% (worse) 9.806 ms +601.5 us / +6.5% (worse)
Windows MSVC 386 println-lto 35328 B -512 B / -1.4% (better) 18341 B -208 B / -1.1% (better) 1.331 s +46.25 ms / +3.6% (worse) 9.589 ms +589.4 us / +6.5% (worse)
Windows MSVC ARM64 cprintf 11776 B 0 B / +0.0% 4208 B +16 B / +0.4% (worse) 1.230 s +22.38 ms / +1.9% (worse) 7.203 ms +399.7 us / +5.9% (worse)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.230 s +1.109 ms / +0.1% (worse) 6.939 ms +334.5 us / +5.1% (worse)
Windows MSVC ARM64 fmtprintf 1491456 B +104960 B / +7.6% (worse) 566856 B +57312 B / +11.2% (worse) 3.797 s -171.9 ms / -4.3% (better) 13.925 ms -1.835 ms / -11.6% (better)
Windows MSVC ARM64 fmtprintf-lto 1523200 B +118272 B / +8.4% (worse) 527236 B +50416 B / +10.6% (worse) 9.387 s +461.7 ms / +5.2% (worse) 15.151 ms +1.109 ms / +7.9% (worse)
Windows MSVC ARM64 println 45056 B 0 B / +0.0% 23668 B -240 B / -1.0% (better) 1.209 s -53.29 ms / -4.2% (better) 12.049 ms -1.64 ms / -12.0% (better)
Windows MSVC ARM64 println-lto 42496 B 0 B / +0.0% 21188 B -192 B / -0.9% (better) 1.396 s -56.01 ms / -3.9% (better) 11.917 ms -671.8 us / -5.3% (better)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.520 ns/op -0.04 ns/op / -0.3% (better)
Linux BenchmarkMergeCompilerFlags 197.300 ns/op -1.1 ns/op / -0.6% (better)
Linux BenchmarkMergeLinkerFlags 130.700 ns/op +3.7 ns/op / +2.9% (worse)
Linux BenchmarkChannelBuffered 54.980 ns/op -0.92 ns/op / -1.6% (better)
Linux BenchmarkChannelHandoff 18269 ns/op +4163 ns/op / +29.5% (worse)
Linux BenchmarkDefer 55.810 ns/op -2.93 ns/op / -5.0% (better)
Linux BenchmarkDirectCall 1.171 ns/op -0.435 ns/op / -27.1% (better)
Linux BenchmarkGlobalRead 1.168 ns/op 0 ns/op / +0.0%
Linux BenchmarkGlobalWrite 7.764 ns/op -0.035 ns/op / -0.4% (better)
Linux BenchmarkGoroutine 51893 ns/op +25020 ns/op / +93.1% (worse)
Linux BenchmarkInterfaceCall 6.210 ns/op -0.418 ns/op / -6.3% (better)
Linux BenchmarkRuntimeGetG 2.846 ns/op -0.294 ns/op / -9.4% (better)
macOS BenchmarkLookupPCRandom 16.180 ns/op +1.42 ns/op / +9.6% (worse)
macOS BenchmarkMergeCompilerFlags 180.700 ns/op +65.8 ns/op / +57.3% (worse)
macOS BenchmarkMergeLinkerFlags 86.630 ns/op +13.09 ns/op / +17.8% (worse)
macOS BenchmarkChannelBuffered 33.460 ns/op +1.74 ns/op / +5.5% (worse)
macOS BenchmarkChannelHandoff 8346 ns/op -2127 ns/op / -20.3% (better)
macOS BenchmarkDefer 47.960 ns/op +2.87 ns/op / +6.4% (worse)
macOS BenchmarkDirectCall 1.391 ns/op +0.008 ns/op / +0.6% (worse)
macOS BenchmarkGlobalRead 1.230 ns/op -0.097 ns/op / -7.3% (better)
macOS BenchmarkGlobalWrite 1.386 ns/op +0.092 ns/op / +7.1% (worse)
macOS BenchmarkGoroutine 84211 ns/op +20391 ns/op / +32.0% (worse)
macOS BenchmarkInterfaceCall 5.075 ns/op +0.582 ns/op / +13.0% (worse)
macOS BenchmarkRuntimeGetG 2.618 ns/op -0.019 ns/op / -0.7% (better)
Windows MinGW BenchmarkLookupPCRandom 9.504 ns/op -0.234 ns/op / -2.4% (better)
Windows MinGW BenchmarkMergeCompilerFlags 416.900 ns/op +44.8 ns/op / +12.0% (worse)
Windows MinGW BenchmarkMergeLinkerFlags 338.800 ns/op +4.1 ns/op / +1.2% (worse)
Windows MinGW BenchmarkChannelBuffered 24.380 ns/op +1.09 ns/op / +4.7% (worse)
Windows MinGW BenchmarkChannelHandoff 982 ns/op -43 ns/op / -4.2% (better)
Windows MinGW BenchmarkDefer 41.970 ns/op -2.74 ns/op / -6.1% (better)
Windows MinGW BenchmarkDirectCall 1.358 ns/op +0.001 ns/op / +0.1% (worse)
Windows MinGW BenchmarkGlobalRead 1.360 ns/op -0.016 ns/op / -1.2% (better)
Windows MinGW BenchmarkGlobalWrite 2.165 ns/op -0.004 ns/op / -0.2% (better)
Windows MinGW BenchmarkGoroutine 58693 ns/op -258 ns/op / -0.4% (better)
Windows MinGW BenchmarkInterfaceCall 6.790 ns/op -0.009 ns/op / -0.1% (better)
Windows MinGW BenchmarkRuntimeGetG 1.406 ns/op -0.226 ns/op / -13.8% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 26.490 ns/op -0.14 ns/op / -0.5% (better)
Windows MinGW 386 BenchmarkMergeCompilerFlags 746.500 ns/op -26.5 ns/op / -3.4% (better)
Windows MinGW 386 BenchmarkMergeLinkerFlags 667.400 ns/op -31.2 ns/op / -4.5% (better)
Windows MinGW 386 BenchmarkChannelBuffered 39.630 ns/op +0.35 ns/op / +0.9% (worse)
Windows MinGW 386 BenchmarkChannelHandoff 886.900 ns/op -3.4 ns/op / -0.4% (better)
Windows MinGW 386 BenchmarkDefer 43.540 ns/op +0.15 ns/op / +0.3% (worse)
Windows MinGW 386 BenchmarkDirectCall 1.549 ns/op +0.003 ns/op / +0.2% (worse)
Windows MinGW 386 BenchmarkGlobalRead 1.550 ns/op -0.001 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGlobalWrite 7.773 ns/op 0 ns/op / +0.0%
Windows MinGW 386 BenchmarkGoroutine 108839 ns/op +2554 ns/op / +2.4% (worse)
Windows MinGW 386 BenchmarkInterfaceCall 8.081 ns/op -0.287 ns/op / -3.4% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 2.168 ns/op -0.002 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.100 ns/op +0.03 ns/op / +0.2% (worse)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 659.600 ns/op +71.8 ns/op / +12.2% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 589.400 ns/op +15.2 ns/op / +2.6% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 38.990 ns/op +0.84 ns/op / +2.2% (worse)
Windows MinGW ARM64 BenchmarkChannelHandoff 2644 ns/op +30 ns/op / +1.1% (worse)
Windows MinGW ARM64 BenchmarkDefer 57.540 ns/op +1.59 ns/op / +2.8% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.589 ns/op -0.0001 ns/op / -0.01696% (better)
Windows MinGW ARM64 BenchmarkGlobalRead 0.590 ns/op -0.0738 ns/op / -11.1% (better)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.884 ns/op +0.2209 ns/op / +33.3% (worse)
Windows MinGW ARM64 BenchmarkGoroutine 66922 ns/op -174 ns/op / -0.3% (better)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.170 ns/op +0.025 ns/op / +0.6% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.770 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC BenchmarkLookupPCRandom 11.500 ns/op +0.27 ns/op / +2.4% (worse)
Windows MSVC BenchmarkMergeCompilerFlags 455.700 ns/op -25.6 ns/op / -5.3% (better)
Windows MSVC BenchmarkMergeLinkerFlags 410.800 ns/op -13.6 ns/op / -3.2% (better)
Windows MSVC BenchmarkChannelBuffered 27.990 ns/op +0.85 ns/op / +3.1% (worse)
Windows MSVC BenchmarkChannelHandoff 1225 ns/op +147 ns/op / +13.6% (worse)
Windows MSVC BenchmarkDefer 55.250 ns/op +3.64 ns/op / +7.1% (worse)
Windows MSVC BenchmarkDirectCall 1.652 ns/op +0.318 ns/op / +23.8% (worse)
Windows MSVC BenchmarkGlobalRead 1.661 ns/op +0.112 ns/op / +7.2% (worse)
Windows MSVC BenchmarkGlobalWrite 2.652 ns/op +0.131 ns/op / +5.2% (worse)
Windows MSVC BenchmarkGoroutine 70020 ns/op +6539 ns/op / +10.3% (worse)
Windows MSVC BenchmarkInterfaceCall 8.634 ns/op +1.049 ns/op / +13.8% (worse)
Windows MSVC BenchmarkRuntimeGetG 2.626 ns/op +1.031 ns/op / +64.6% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 26.590 ns/op -0.02 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkMergeCompilerFlags 750.700 ns/op -6 ns/op / -0.8% (better)
Windows MSVC 386 BenchmarkMergeLinkerFlags 716.700 ns/op +44.8 ns/op / +6.7% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 39.060 ns/op -5.83 ns/op / -13.0% (better)
Windows MSVC 386 BenchmarkChannelHandoff 908.800 ns/op -26.8 ns/op / -2.9% (better)
Windows MSVC 386 BenchmarkDefer 46.940 ns/op -1.12 ns/op / -2.3% (better)
Windows MSVC 386 BenchmarkDirectCall 1.548 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGlobalRead 1.861 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGlobalWrite 7.792 ns/op +0.025 ns/op / +0.3% (worse)
Windows MSVC 386 BenchmarkGoroutine 110174 ns/op +856 ns/op / +0.8% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 7.981 ns/op -0.133 ns/op / -1.6% (better)
Windows MSVC 386 BenchmarkRuntimeGetG 1.931 ns/op -0.548 ns/op / -22.1% (better)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.130 ns/op +0.1 ns/op / +0.8% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 566.900 ns/op -1.8 ns/op / -0.3% (better)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 541.200 ns/op +4 ns/op / +0.7% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 38.630 ns/op -0.25 ns/op / -0.6% (better)
Windows MSVC ARM64 BenchmarkChannelHandoff 2940 ns/op -908 ns/op / -23.6% (better)
Windows MSVC ARM64 BenchmarkDefer 62.890 ns/op +0.66 ns/op / +1.1% (worse)
Windows MSVC ARM64 BenchmarkDirectCall 0.591 ns/op +0.0008 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkGlobalRead 0.885 ns/op +0.2219 ns/op / +33.4% (worse)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.797 ns/op +0.05 ns/op / +1.3% (worse)
Windows MSVC ARM64 BenchmarkGoroutine 57866 ns/op +789 ns/op / +1.4% (worse)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.287 ns/op +0.148 ns/op / +3.6% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.768 ns/op -0.001 ns/op / -0.1% (better)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 922.200 ns/op +12.4 ns/op / +1.4% (worse)
Linux AfterFuncZeroDelivery/LLGo 39590 ns/op -10590 ns/op / -21.1% (better)
Linux CreateStop/Go 294.100 ns/op +3.7 ns/op / +1.3% (worse)
Linux CreateStop/LLGo 1784 ns/op -4 ns/op / -0.2% (better)
Linux RearmStopped/Go 116.900 ns/op +1.9 ns/op / +1.7% (worse)
Linux RearmStopped/LLGo 1257 ns/op +63 ns/op / +5.3% (worse)
Linux ResetActive/Go 68.680 ns/op +1.18 ns/op / +1.7% (worse)
Linux ResetActive/LLGo 750.300 ns/op -1.6 ns/op / -0.2% (better)
Linux ResetHeap1024/Go 67.230 ns/op -0.06 ns/op / -0.1% (better)
Linux ResetHeap1024/LLGo 177.600 ns/op +1 ns/op / +0.6% (worse)
macOS AfterFuncZeroDelivery/Go 579.500 ns/op -15.4 ns/op / -2.6% (better)
macOS AfterFuncZeroDelivery/LLGo 116505 ns/op -4980 ns/op / -4.1% (better)
macOS CreateStop/Go 179.200 ns/op -26.2 ns/op / -12.8% (better)
macOS CreateStop/LLGo 730.900 ns/op +47.8 ns/op / +7.0% (worse)
macOS RearmStopped/Go 75.180 ns/op -7.84 ns/op / -9.4% (better)
macOS RearmStopped/LLGo 484.700 ns/op -125.8 ns/op / -20.6% (better)
macOS ResetActive/Go 52.100 ns/op -4.01 ns/op / -7.1% (better)
macOS ResetActive/LLGo 254.800 ns/op -49.8 ns/op / -16.3% (better)
macOS ResetHeap1024/Go 63.300 ns/op +6.68 ns/op / +11.8% (worse)
macOS ResetHeap1024/LLGo 100.800 ns/op -85.1 ns/op / -45.8% (better)
Windows MinGW AfterFuncZeroDelivery/Go 364.200 ns/op -6 ns/op / -1.6% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 112292 ns/op +585 ns/op / +0.5% (worse)
Windows MinGW CreateStop/Go 90.940 ns/op +1.41 ns/op / +1.6% (worse)
Windows MinGW CreateStop/LLGo 331 ns/op -61.9 ns/op / -15.8% (better)
Windows MinGW RearmStopped/Go 24.410 ns/op -0.1 ns/op / -0.4% (better)
Windows MinGW RearmStopped/LLGo 220.200 ns/op -3.1 ns/op / -1.4% (better)
Windows MinGW ResetActive/Go 14.760 ns/op -0.01 ns/op / -0.1% (better)
Windows MinGW ResetActive/LLGo 120.500 ns/op +1.5 ns/op / +1.3% (worse)
Windows MinGW ResetHeap1024/Go 14.910 ns/op -0.14 ns/op / -0.9% (better)
Windows MinGW ResetHeap1024/LLGo 107.400 ns/op +0.6 ns/op / +0.6% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/Go 956.200 ns/op -5.9 ns/op / -0.6% (better)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 196082 ns/op -196 ns/op / -0.1% (better)
Windows MinGW 386 CreateStop/Go 190.500 ns/op -4 ns/op / -2.1% (better)
Windows MinGW 386 CreateStop/LLGo 507 ns/op +16.2 ns/op / +3.3% (worse)
Windows MinGW 386 RearmStopped/Go 63.390 ns/op +0.06 ns/op / +0.1% (worse)
Windows MinGW 386 RearmStopped/LLGo 351.100 ns/op -1.7 ns/op / -0.5% (better)
Windows MinGW 386 ResetActive/Go 39.090 ns/op +0.06 ns/op / +0.2% (worse)
Windows MinGW 386 ResetActive/LLGo 1021 ns/op +49.5 ns/op / +5.1% (worse)
Windows MinGW 386 ResetHeap1024/Go 39.310 ns/op 0 ns/op / +0.0%
Windows MinGW 386 ResetHeap1024/LLGo 185.700 ns/op -3 ns/op / -1.6% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 674.900 ns/op +6 ns/op / +0.9% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 156490 ns/op +1782 ns/op / +1.2% (worse)
Windows MinGW ARM64 CreateStop/Go 201.300 ns/op +0.3 ns/op / +0.1% (worse)
Windows MinGW ARM64 CreateStop/LLGo 354.500 ns/op -11.9 ns/op / -3.2% (better)
Windows MinGW ARM64 RearmStopped/Go 70.540 ns/op -0.01 ns/op / -0.01417% (better)
Windows MinGW ARM64 RearmStopped/LLGo 255.200 ns/op +1.8 ns/op / +0.7% (worse)
Windows MinGW ARM64 ResetActive/Go 31.030 ns/op -0.11 ns/op / -0.4% (better)
Windows MinGW ARM64 ResetActive/LLGo 119.400 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 ResetHeap1024/Go 31.100 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 ResetHeap1024/LLGo 126.900 ns/op +0.2 ns/op / +0.2% (worse)
Windows MSVC AfterFuncZeroDelivery/Go 453.800 ns/op +3.3 ns/op / +0.7% (worse)
Windows MSVC AfterFuncZeroDelivery/LLGo 130097 ns/op +4138 ns/op / +3.3% (worse)
Windows MSVC CreateStop/Go 108.400 ns/op +3.7 ns/op / +3.5% (worse)
Windows MSVC CreateStop/LLGo 421.800 ns/op +31.7 ns/op / +8.1% (worse)
Windows MSVC RearmStopped/Go 29.280 ns/op +0.8 ns/op / +2.8% (worse)
Windows MSVC RearmStopped/LLGo 260.100 ns/op +10.4 ns/op / +4.2% (worse)
Windows MSVC ResetActive/Go 17.880 ns/op +0.73 ns/op / +4.3% (worse)
Windows MSVC ResetActive/LLGo 140.500 ns/op +5.4 ns/op / +4.0% (worse)
Windows MSVC ResetHeap1024/Go 17.910 ns/op +0.66 ns/op / +3.8% (worse)
Windows MSVC ResetHeap1024/LLGo 124.500 ns/op +4.9 ns/op / +4.1% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/Go 959.600 ns/op +2.2 ns/op / +0.2% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 203212 ns/op +2544 ns/op / +1.3% (worse)
Windows MSVC 386 CreateStop/Go 192.900 ns/op +1.6 ns/op / +0.8% (worse)
Windows MSVC 386 CreateStop/LLGo 464.900 ns/op -18.3 ns/op / -3.8% (better)
Windows MSVC 386 RearmStopped/Go 63.490 ns/op +0.14 ns/op / +0.2% (worse)
Windows MSVC 386 RearmStopped/LLGo 321.800 ns/op +2 ns/op / +0.6% (worse)
Windows MSVC 386 ResetActive/Go 39.050 ns/op +0.17 ns/op / +0.4% (worse)
Windows MSVC 386 ResetActive/LLGo 971.900 ns/op -12.9 ns/op / -1.3% (better)
Windows MSVC 386 ResetHeap1024/Go 39.330 ns/op -0.3 ns/op / -0.8% (better)
Windows MSVC 386 ResetHeap1024/LLGo 170.200 ns/op -0.8 ns/op / -0.5% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 676.500 ns/op +17.2 ns/op / +2.6% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 144526 ns/op +3561 ns/op / +2.5% (worse)
Windows MSVC ARM64 CreateStop/Go 212.100 ns/op +17.2 ns/op / +8.8% (worse)
Windows MSVC ARM64 CreateStop/LLGo 385.100 ns/op -29.9 ns/op / -7.2% (better)
Windows MSVC ARM64 RearmStopped/Go 70.570 ns/op -0.02 ns/op / -0.02833% (better)
Windows MSVC ARM64 RearmStopped/LLGo 274.100 ns/op +0.4 ns/op / +0.1% (worse)
Windows MSVC ARM64 ResetActive/Go 31.090 ns/op +0.01 ns/op / +0.03218% (worse)
Windows MSVC ARM64 ResetActive/LLGo 131 ns/op -3.7 ns/op / -2.7% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.110 ns/op -0.04 ns/op / -0.1% (better)
Windows MSVC ARM64 ResetHeap1024/LLGo 137.100 ns/op +0.5 ns/op / +0.4% (worse)

Compared with 653957f840b5 measured in the same runner job.

Known concrete-to-interface conversions emit a read-only _llgo_itab$
global (Inter, Type, Hash, Fun) with weak_odr linkage, matching
cmd/compile's go:itab symbols.

Ordinary builds use that global as the runtime vtable. After function
bodies compile, path.init$itabs registers those itabs with
RegisterStaticItab, after runtime.init and before package init (gc
itabsinit). LTO and deadcode-drop keep calling NewItab so unused
interface methods can be dropped; Fun[] would pin them. LTO still
emits the static template for de-virt.

Hash is copied from the type descriptor when present. interequal
compares the (inter, _type) pair so a static itab and a dynamically
allocated itab for the same conversion compare equal.
@github-actions

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

ab860f11ba65 | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 146278 B -695 B / -0.5% (better) 74890 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 144726 B -681 B / -0.5% (better) 73165 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 134426 B -90 B / -0.1% (better) 78778 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 141539 B -417 B / -0.3% (better) 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 141255 B -449 B / -0.3% (better) 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3423342 B +219965 B / +6.9% (worse) 118590 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3402019 B +219243 B / +6.9% (worse) 101818 B +183 B / +0.2% (worse)
fmtprintf/j64-emscripten-memory64/LLGo 3133396 B +193077 B / +6.6% (worse) 125433 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 3172938 B +337150 B / +11.9% (worse) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 3036055 B +333834 B / +12.4% (worse) 0 B 0 B / 0.0%
j32-emscripten/LLGo 145505 B -700 B / -0.5% (better) 74890 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 144189 B -687 B / -0.5% (better) 73165 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 133758 B -89 B / -0.1% (better) 78778 B 0 B / +0.0%
reflectcall/j32-emscripten/LLGo 1564350 B +30265 B / +2.0% (worse) 92056 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1566777 B +29947 B / +1.9% (worse) 90331 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1449809 B +30552 B / +2.2% (worse) 98119 B +330 B / +0.3% (worse)
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1572899 B +32231 B / +2.1% (worse) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1496687 B +30476 B / +2.1% (worse) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 140756 B -414 B / -0.3% (better) 0 B 0 B / 0.0%
w32-wasi/LLGo 140543 B -446 B / -0.3% (better) 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 5.924 s +91.1 ms / +1.6% (worse)
j32-goos-js 5.966 s -26.47 ms / -0.4% (better)
j64-emscripten-memory64 5.439 s +155.9 ms / +3.0% (worse)
reflectcall/w32-wasi 26.346 s +44.34 ms / +0.2% (worse)
w32-goos-wasip1 4.698 s -96.27 ms / -2.0% (better)
w32-wasi 4.935 s +200 ms / +4.2% (worse)

Compared with 15732a0d63d9 measured in the same runner job.

Non-direct interface values (integers, bools, small aggregates) still
use IfaceIndir, matching Go: data is a pointer to a copy. Compile-time
constants go in a WeakODR _llgo_ifacebox$ global instead of AllocU,
like cmd/compile's static temps.

Depends on static itabs so a constant T2I can be {itab, box} with no
runtime call.
A call whose SSA body is MakeInterface of a Convert/ChangeType of a
constant parameter is lowered at the call site to MakeInterface of
that concrete value. No function-name matching: constant.MakeInt64
and user helpers such as boxMyInt(x int64) any { return myInt(x) }
use the same path.

Depends on static itabs and iface boxes so the folded value is a
compile-time {itab, box} pair.
Map literals with more than 25 constant keys become a counted
mapassign loop over private constant key/value arrays, matching
cmd/compile's maplit. Composite values whose fields are constants or
trivial iface constructors (PR 3) are rebuilt as LLVM constants so
the value array is a ConstArray.

Runtime values stay unrolled. Spilling them into alloca [N x T] makes
LLVM default<Os> SLP scalarize the array.

Depends on trivial iface folding so {string, MakeInt64/MakeBool}
entries are compile-time structs.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant