You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Compounding both: 6 of the 8 startup sites do go func() { _ = runner.Run(ctx) }(), swallowing Run's error — so a bind/startup failure is invisible and surfaces only as the readiness timeout (this is why the #529 flake "looked like a generic timeout" until the error was temporarily surfaced during diagnosis). Two files already use the good pattern (errCh <- runner.Run(ctx) / runErrCh <-).
Durable fix
Stop guessing the port — have the runner accept a pre-bound net.Listener (or expose its resolved port after Start), so tests bind once and use exactly that address. Eliminates both the wrong-port race and the scan's identity ambiguity, and removes the exhaustion-from-guessing.
Surface Run's error in every startup test — send it to a channel (template already in runner_test.go / tracing_runner_test.go) and have the readiness wait select on it, so a startup failure fails fast with the real cause instead of the 20s ceiling.
Scope
Test-harness only; no production behavior change (unless option 1 adds a small, opt-in WithListener/resolved-port accessor to the runner). Follow-up to #529 (which shipped the interim scan fix).
Problem
Runner-startup tests guess a port via
findFreePort(bind:0, close, return the port), then poll it. Two failure modes stem from guessing:server.Startauto-increments toport+1..+9and the test polled the wrong port → opaque "server did not start within 5s". test(runtime): fix flaky server-readiness wait (port auto-increment race) #529 mitigated this by scanning the auto-increment window inwaitForServer, but the scan has no server-identity check (under heavy parallel overlap it could match a neighbor's/healthz).Run's bind fails outright. Not seen in single-pass CI, deliberately scoped out of test(runtime): fix flaky server-readiness wait (port auto-increment race) #529.Compounding both: 6 of the 8 startup sites do
go func() { _ = runner.Run(ctx) }(), swallowingRun's error — so a bind/startup failure is invisible and surfaces only as the readiness timeout (this is why the #529 flake "looked like a generic timeout" until the error was temporarily surfaced during diagnosis). Two files already use the good pattern (errCh <- runner.Run(ctx)/runErrCh <-).Durable fix
net.Listener(or expose its resolved port afterStart), so tests bind once and use exactly that address. Eliminates both the wrong-port race and the scan's identity ambiguity, and removes the exhaustion-from-guessing.Run's error in every startup test — send it to a channel (template already inrunner_test.go/tracing_runner_test.go) and have the readiness waitselecton it, so a startup failure fails fast with the real cause instead of the 20s ceiling.Scope
Test-harness only; no production behavior change (unless option 1 adds a small, opt-in
WithListener/resolved-port accessor to the runner). Follow-up to #529 (which shipped the interim scan fix).