You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Why. Imp's CI fails on tests that pass on a re-run, and a different one each time. A check that fails at random stops being read, and then a real failure gets waved through as "the flaky one". Seen in the last two days:
test/timed_out_connection_test.exs:32: the call after a timed-out call is answered on a new connection (runs 36488635687, 36514520298).
test/req_llm_stream_end_test.exs:191: Task.await times out after 5000 ms (run 36501812974).
reasoning_continuity, silent_failure_regressions P03, and an assert_raise (reported from stranger.check runs; find the run IDs).
Also seen in failing runs, possibly caused by another test: local_mlx_campaign_test.exs:26, avatar_persistence_test.exs:24, and req_llm_batch_test.exs:273 and :780.
Done means (each can fail):
Every intermittent failure in CI runs since 2026-09-27 is listed here, with its run ID and its cause. A cause is something like shared global state (application env, a named process, a registered name, the file system or ports), a fixed wall-clock timeout on a loaded runner, or ordering between async tests. Not "flaky".
Each cause is fixed at the cause. No longer timeouts, no retries, no @tag :skip, and a test isn't switched to async: false unless shared state really can't be removed. If one is, the reason is written beside it.
Evidence: each fixed test runs 50 times with mix test <file> --repeat-until-failure 50 and doesn't fail, and the full suite passes on three consecutive CI runs of the PR. Both are linked or pasted in the PR. Local runs stay light, with no machine-wide CPU load on the owner's Mac. CI is where load is tested.
A test that only fails because another test leaks state is fixed in the leaking test.
Out of scope. Changing which CI jobs exist. The duplicate stranger.check job is a separate PR.
Why. Imp's CI fails on tests that pass on a re-run, and a different one each time. A check that fails at random stops being read, and then a real failure gets waved through as "the flaky one". Seen in the last two days:
test/timed_out_connection_test.exs:32: the call after a timed-out call is answered on a new connection (runs 36488635687, 36514520298).test/req_llm_stream_end_test.exs:191:Task.awaittimes out after 5000 ms (run 36501812974).silent_failure_regressionsP03, and anassert_raise(reported fromstranger.checkruns; find the run IDs).local_mlx_campaign_test.exs:26,avatar_persistence_test.exs:24, andreq_llm_batch_test.exs:273and:780.Done means (each can fail):
@tag :skip, and a test isn't switched toasync: falseunless shared state really can't be removed. If one is, the reason is written beside it.mix test <file> --repeat-until-failure 50and doesn't fail, and the full suite passes on three consecutive CI runs of the PR. Both are linked or pasted in the PR. Local runs stay light, with no machine-wide CPU load on the owner's Mac. CI is where load is tested.Out of scope. Changing which CI jobs exist. The duplicate
stranger.checkjob is a separate PR.