CI flakes: remove each timing cause - #268
Conversation
LLMDB loads ReqLLM's model catalog on first use under a :global.trans lock, and a process that finds the lock taken sleeps a random backoff that doubles up to 8 s. Async tests making their first model lookup together at the start of the suite waited several times the 1.4 s load, and on a loaded runner past a 5 s Task.await in ReqLLMStreamEndTest.
The first request is held until the test ends rather than for 2 s, so the first call times out however slow the machine is, and the second call keeps ReqLLM's own receive timeout instead of one the test picks. A second request on the stuck connection is still never answered.
The walk is found for as long as the caller is alive, since the caller parses the whole file before it starts it. The walk is then suspended, so it cannot finish and exit on its own before the caller is killed, and the test asserts that it ended killed.
A 150 ms timeout fired on a loaded runner before Python had started the tree, so the output held no "tree ready". The three tree tests run with no timer, wait for the tree, and send the port owner the message its timer sends; a separate test checks that the timeout option ends a command. Waiting for the process ids has no attempt limit.
|
CI on 31eda09, run 36621743611, re-run with
Attempts 3, 4 and 5 are three consecutive green runs. An intermittent failure this PR does not explain. In attempt 2, "a tool that traps exits still ends when its caller is killed" ( |
|
Review (fresh-eyes reviewer, checked by me), at 31eda09. I read the failing runs' logs (36501812974, 36514520298, 36488635687, 36353581023, 36423475872), and the table's causes match them.
Not blocking:
Changed to "Part of #267". |
Part of #267
Every failed CI run since 2026-09-27 was read with
gh run view <id> --log-failed(no run in that window had a re-run attempt except 36361631807, whose two attempts failed the same way). Five test failures were intermittent. Each had a timing cause, listed below with its fix. Everything else that failed did so on every run of that commit, and is listed after the table.Intermittent failures
req_llm_stream_end_test.exs:191Task.awaitexited after 5000 ms:global.translock. A process that finds the lock taken sleeps a random backoff that doubles up to 8 s before it tries again. The async tests at the start of the suite all make their first lookup at the same moment. Both failures came within 6 s ofRunning ExUnit. Measured locally: a 1.4 s load, and waits of up to 4.6 s among 16 simultaneous first lookups on an idle machine. With the catalog already loaded, each lookup took 9 ms.test_helper.exsloads the catalog (LLMDB.load()) before any test runs. The slowest stream-end test dropped from 1509 ms to 35 ms.req_llm_stream_end_test.exs:217Task.awaitexited after 5000 mstimed_out_connection_test.exs:32{:error, %Imp.LMError{message: "API request failed: timeout"}}parity_sidecar_test.exs:79output.textwas"", not"tree ready":97,:112) have the same race.:command_timeout, the message its timer sends."tree ready"is printed before the pid file is written, so it is always in the pipe by then. A new test checks the real timer on a command that sleeps.await_pidshas no attempt limit.datasets_contract_test.exs:203{:DOWN, …, :killed}within 1000 ms;walk = #PID<0.53.0>:killed.Failures that were not intermittent
property_invariants_test.exs:46(36432281651): a real defect that the property test found by chance, a signature field namednil. Fixed in Signature fields named nil, true or false keep their names #248.hover_papillon_calibration_pilot_test.exs(5 tests, differential.check: 36350561245, 36352551929, 36359159692, 36361631807 both attempts):Imp candidate tracked tree is dirty. The branch's examplemix.lockfiles lackednimble_csv, sodeps.getrewrote them. Fixed on that branch in cc42aa9. This is also where theassert_raisecame from (:654).deployment_reference_test.exs:453(36430185896, 36434502372, 36437298490): the version assertion failed on the release branch until the version was set.reasoning_continuity,silent_failure_regressionsP03,local_mlx_campaign_test.exs:26,avatar_persistence_test.exs:24andreq_llm_batch_test.exs:273/:780. None of these failed in any CI run from 2026-09-24 on. They show up in the logs only as stack frames printed under ReqLLM'sUsing unverified modelwarning (IO.warn), not as failures.Evidence
The tests still fail when the behaviour they check is broken. Each code change below was made briefly, the test was run, and the code was restored:
deps/finch/lib/finch/http1/conn.ex:260, returning the connection unclosed, as 0.23 did).timed_out_connection_testthen fails at:76after ReqLLM's 30 s receive timeout.Imp.Datasets.walk_csv's guard was changed to not kill the walk. The CSV walk test then fails: it times out waiting for the DOWN.:command_timeoutbranch inImp.ExternalCommand.Lifecyclewas changed to skipterminate_group. All three tree tests then fail. When it returned an empty capture instead, the two"tree ready"assertions fail.Imp.Core.responsewas changed to drop the reported cost.req_llm_stream_end_test:191and:217then fail.Local repeats (
MIX_ENV=test ELIXIR_ERL_OPTIONS="+S 4" mix test <file> --repeat-until-failure 50, one at a time):(51 = the first run plus 50 repeats.)
CI: three consecutive runs of the PR are linked in a comment below.