Skip to content

fix(studio): the timeline scroll gate catches dropped frames in the virtualized view - #4855

Draft
miguel-heygen wants to merge 3 commits into
mainfrom
fix/studio-viewport-gate-virtualized-limit
Draft

miguel-heygen wants to merge 3 commits into
mainfrom
fix/studio-viewport-gate-virtualized-limit

Conversation

@miguel-heygen

@miguel-heygen miguel-heygen commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

The timeline viewport gate now catches a scroll that drops frames in the virtualized timeline on CI. A 45 ms stall on every scroll step used to pass it.

What changes

  • The CI virtualized arm gets its own pair of limits. Pooled over all 315 measured scroll steps, the interaction p95 must be ≤58.3 ms (3.5 frames) and the frame-interval p95 ≤25 ms (1.5 frames). In plain terms, at most 5% of scroll steps may drop a frame, and at most 5% may take 4 or more frames.
  • The unvirtualized arm keeps 75/75, because its normal interaction p95 is already 49-60 ms. The primary tier keeps 50/33.3, and the throttled and high-DPR tiers keep 75/75.
  • Which pair applies is decided in one place, responsivenessLimits in the verdict module. The gate script reads it from there.
  • One rerun per arm. When an arm's timing attempt fails (3 warm-up runs dropped, then 5 measured runs), the gate measures that arm once more with the same protocol and fails only if both attempts fail. The loop is inside the gate script, not a CI retry. Direct-scroll and memory checks run once. The step summary prints every attempt's p95s, limits and verdict.
  • Per-step samples in the artifact. The timeline-viewport-gate-evidence artifact keeps every measured step's interaction and frame interval, per arm and per attempt (attempts[].runs[].interactions / frameIntervals), so a red can be recomputed later.

Why these numbers

A pooled p95 over 315 steps fails once 16 of them are over the limit.

False fails on normal builds. Every CI run today that carries per-step evidence: 68 normal virtualized attempts from 67 PR runs, judged single-shot.

Virtualized limits (interaction / frame) Attempts that fail Most steps over the interaction limit
75 / 75 (main) 0 of 68 3
41.7 / 25 (first pick) 16 of 68 (23.5%) 39
58.3 / 25 (this PR) 0 of 68 6

At 41.7 ms the fails were all on interactions: 17-39 steps at 3 frames, spread across all five runs (a slow runner, not a slow start), and frame intervals stayed under the limit (at most 9 steps over 25 ms). The earlier analysis predicted 0% at 41.7 from too little per-step data, so the fallback rule fired before merge and the interaction limit stepped to 58.3 ms. With the rerun a false fail needs two failed attempts in a row; at 58.3 none of the 68 had a single one.

Catching the 45 ms stall, with the rerun. A throwaway build injected a 45 ms stall into every timeline scroll step and ran the virtualized arm 10 times on CI.

  • 20 of 20 attempts failed, so 10 of 10 gate runs failed.
  • Every attempt had 270 of 315 frame intervals over 25 ms (every step that scrolls stalls), against 16 needed.
  • Interactions over the limit: 142-193 of 315 over 41.7 ms, and 45-47 over 58.3 ms. Either half alone fails the stall.
  • Single-attempt detection d: 20 of 20 observed, with a margin of 270 against 16. With the rerun the gate fails the stall at a rate of d², which is also about 1, above the 95% bar. Counting attempts alone, 20 of 20 bounds d at ≥0.86 with 95% confidence; the step margin is what carries it.

CI time per arm. A normal pass runs one attempt, unchanged: about 40 s for the virtualized arm and 35 s for the unvirtualized one. A failed attempt adds one more, about 17 s on a normal build and 30 s on the stall build (a stall arm takes 67-68 s end to end). The steps per run are unchanged.

False fails are counted per arm, each against that arm's own limit: the unvirtualized arm's against its 75 ms, never toward the virtualized limit. For example, #4836's unvirtualized arm read 82, 82, 39, 49 and 47 ms interaction p95 per run on a busy runner; that counts against 75 ms.

This PR's own CI. The virtualized arm applied 58.3 / 25 and passed on its first attempt: interaction p95 33.2 ms with 0 of 315 steps over 58.3, and frame p95 16.8 ms with 1 of 315 over 25. The unvirtualized arm passed at 75 / 75 (61.8 / 33.4 ms).

Tests

timeline-viewport-verdict.test.mjs:

  • The CI virtualized arm: 16 of 315 steps two frames late fail and 15 pass; the same for frame intervals that drop a frame. Both fail with the virtualized limits set back to 75.
  • The rerun: a failed attempt followed by a passing one passes; two failed attempts fail; a third attempt is never counted.
  • The other arms and tiers stay at their old limits.

Before

Main's limits on recorded CI samples: the 45 ms stall run passes.

Main's limits pass the 45 ms stall run

After

This branch's limits and rerun on the same samples: both stall attempts fail, and a normal run still passes.

This branch's limits fail the stall on both attempts and pass the normal run

@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Edit accuracy: 835 passing here, 835 on the base branch

The gate passes.
Smoothness is reported in the artifact, not gated. A case fails only if it fails 2 of 3 runs.

Quarantined, measured but not gated (9)

@miguel-heygen
miguel-heygen force-pushed the fix/studio-viewport-gate-virtualized-limit branch from e945e4f to e3338bf Compare October 1, 2026 14:56
@miguel-heygen
miguel-heygen force-pushed the fix/studio-viewport-gate-virtualized-limit branch from e3338bf to f8f419d Compare October 1, 2026 15:58
@miguel-heygen
miguel-heygen force-pushed the fix/studio-viewport-gate-virtualized-limit branch from f8f419d to 8d8b4b3 Compare October 1, 2026 16:52

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant