Skip to content

perf: improve first frame and progressive-mount speed, optimize large documents - #54

Merged
denolfe merged 6 commits into
mainfrom
perf/faster-first-frame
Sep 6, 2026
Merged

denolfe merged 6 commits into
mainfrom
perf/faster-first-frame

Conversation

@denolfe

@denolfe denolfe commented Sep 6, 2026 •

Copy link
Copy Markdown
Owner

Profiles each stage of the viewer (parse, first paint, progressive mount, scroll) on synthetic 5k- and 26k-line documents and fixes the cost centers found. README-sized documents are unchanged. Method and fixtures are in bench/BASELINE.md.

Cumulative, baseline f596175 vs head (Bun 1.4.1, M2 Max, headless 120x40, React dev build). large.md = 2202 nodes / 401 headings, huge.md = 11002 nodes / 2001 headings.

Metric Before After Change
first-frame README 96ms 96ms none
first-frame large.md 157ms 120ms −24%
first-frame huge.md 320ms 176ms −45%
buildDocument huge.md 69ms 39ms −44%
full progressive mount large.md 11222ms 4286ms −62%
scroll step p50 large.md 13.9ms 9.6ms −31%
scroll step under search large.md 15.3ms 9.5ms −38%

Root Cause

Four independent defects, each proportional to document length:

  • useProgressiveMount built a fresh nodes.slice() on every Viewer render. NodeList is memoized on that array, so every chunk (and every shell re-render mid-mount) re-reconciled every mounted block and relaid out the whole tree.
  • The TOC sidebar mounted a row for every heading before first paint. Once the Viewer mounted progressively, this was the entire remaining first-frame scaling.
  • computeHeadingLines ran its own full marked.lexer pass on top of the one in buildTree, doubling parse time.
  • The spacer estimate re-walked the unmounted tail on every chunk, and keyboard scrolls resolved the current heading twice (once explicitly, once from the watched scroll setter).

The build scripts also set no NODE_ENV, so both the npm bundle and the compiled binary shipped the development react-reconciler with its per-commit profiling instrumentation.

Key Changes

  • Progressive mount no longer re-renders mounted blocks
    • The mounted slice is memoized on [nodes, mountedCount]. NodeRenderer is memoized on node identity plus path value, since NodeList builds a new path array per element.
  • TOC rows mount progressively
    • The chunked-growth counter is extracted into useProgressiveCount, shared by the Viewer and the TOC. The TOC paints two screenfuls of rows first and grows behind a spacer that keeps the sidebar scrollbar honest.
  • One lex per document
    • Mermaid and DOT preprocessing moved from a text regex to a token pass (renderDiagramBlocks). It rewrites a code token's text and lang and leaves raw intact, so heading line numbers and the tree both come from the same token list. A side effect: a mermaid fence quoted inside a wider fence is no longer rendered.
  • Cheaper per-chunk and per-scroll work
    • Row estimates are a prefix sum built once per document. scroll() relies on the watched scroll setter instead of resolving headings a second time.
  • Production React in shipped artifacts
    • Both build scripts define NODE_ENV=production. Verified by rebuilding: zero react-reconciler.development references in dist/npm/main.js and dist/bin/viewmd, smoke test passes.
  • Bench fixes
    • bench/scroll.tsx drains until the renderable count is stable (60 fixed iterations left large docs half-mounted) and times every loop with the same keypress-plus-render protocol. bench/first-frame.tsx runs the real buildDocument pipeline. New: bench/stages.ts (per-stage parse timing) and bench/gen-fixture.ts (sized fixtures).

Per-change A/B, each taken when the commit landed:

Change Metric Before After Change
memoized slice + memo(NodeRenderer) (a71fc69) large.md full mount 11222ms 5104ms −55%
large.md step p50 13.9ms 10.9ms −22%
large.md step under search p50 15.3ms 13.5ms −12%
NODE_ENV=production (a71fc69), measured from source via env var large.md full mount 5104ms 4382ms −14%
large.md step p50 / p95 10.9 / 20.8ms 10.5 / 17.2ms −4% / −17%
prefix-sum estimate + drop duplicate resolve (69588bd), measured together large.md full mount 5104ms 4286ms −16%
large.md step p50 10.9ms 9.6ms −12%
large.md step under search p50 13.5ms 9.5ms −30%
lex once (1a2689e) huge.md buildDocument 69.4ms 38.6ms −44%
huge.md computeHeadingLines 34.4ms 1.3ms −96%
README buildDocument 0.44ms 0.24ms −45%
TOC progressive mount (253e866), hyperfine 10 runs first-frame README 96.0ms 96.5ms noise
first-frame large.md 157.0ms 120.5ms −23%
first-frame huge.md 320.4ms 175.9ms −45%

Caveats: the prefix-sum and duplicate-resolve changes were measured together. The production-React row is a from-source proxy; the shipped binary was verified by string inspection and smoke test, not benched. The cumulative mount and scroll numbers exclude the production-React gain because the benches run from source in dev mode.

Design Decisions

TOC: progressive mount over windowing. Windowing would also shrink the per-frame layout walk, but TOC rows wrap when the sidebar hits its 40% width clamp, so fixed-height windowing needs truncation (a visible UI change) or post-layout height measurement. Progressive mount gets the first-frame win with no UI change and a small diff.

Diagram preprocessing at token level rather than correcting line offsets after a text replace. The lexer already handles fence nesting and indentation inside lists and blockquotes; the regex did not.

Tried and dropped: deferring projectionMap past first paint. The scroll-marks path builds it before first paint regardless, so the change only moved the cost (huge.md first frame 313 ± 17ms to 347 ± 21ms).

Not addressed: the remaining floor of about 7ms per scroll step on an 18k-renderable tree is OpenTUI's per-frame layout walk. Only virtualizing the Viewer moves it.

… first-frame

Match ticks paint over the thumb as soon as a pattern is typed, so the
scroll bench could not locate the scrollbar column on docs where every
track row has a match. Locate the column before the search bar opens and
poll until ticks land instead of assuming one task is enough.

first-frame now goes through buildDocument so DOT, frontmatter, and
heading-line work is measured, and reports the parse slice.

Add bench/stages.ts (per-stage pipeline timing) and bench/gen-fixture.ts
(parameterized large fixture generator).
useProgressiveMount minted a new nodes slice on every Viewer render, so
NodeList's memo never held and every chunk re-reconciled every mounted
block and relaid out the whole tree. Memoize the slice and memo
NodeRenderer on node identity plus path value. Full mount of a 2200-node
doc drops from 11.2s to 5.1s.

Pin NODE_ENV=production in both build scripts: React picked the
development reconciler at bundle time, shipping its per-commit profiling
and prop-diff instrumentation.

bench/scroll.tsx now drains until the renderable count is stable (60
fixed iterations left large docs half-mounted) and times every step loop
with the same keypress-plus-render protocol.
…per scroll

estimateTotalRows re-walked the unmounted tail on every chunk commit,
O(n^2) over a full mount. Build a row-estimate prefix sum once per
document and difference it.

scroll() resolved headings after every keyboard scroll, but the watched
scrollPosition setter already resolves synchronously on any change, so
each keypress walked the renderable tree twice.

large.md: full mount 5.1s -> 4.3s, scroll step p50 10.9ms -> 9.6ms.
computeHeadingLines ran its own full marked.lexer pass over the body,
doubling parse time. The mermaid/DOT preprocess is now a token pass
(renderDiagramBlocks) that rewrites code tokens' text/lang and leaves
raw intact, so heading line numbers and the tree are both derived from
one lex. Handling diagrams at token level also stops a mermaid fence
quoted inside a wider fence from being rendered.

countNewlines walks with indexOf instead of per-character iteration.

huge.md (26k lines): buildDocument 69ms -> 39ms; README 0.44 -> 0.24ms.
The sidebar mounted every heading row before first paint, which was the
whole of first-frame's scaling with document length once the Viewer
mounted progressively. Extract the chunked-growth counter the Viewer
uses into useProgressiveCount and drive the TOC from it: two screenfuls
of rows on first paint, the rest one chunk per task behind a spacer
that keeps the sidebar scrollbar honest.

First frame: 2200-node doc 157ms -> 120ms; 11000-node doc 320ms -> 176ms;
README unchanged.
@github-actions

github-actions Bot commented Sep 6, 2026

Copy link
Copy Markdown

Startup benchmark (--render test/exhaustive.md, linux-x64)

build mean ratio verdict
baseline (main@f596175) 394.0ms ± 14.8ms —
PR 373.6ms ± 10.6ms 0.95× ✅ ok

Thresholds: warn ≥ 1.1×, fail ≥ 1.25×. Baseline built from main.

@denolfe denolfe changed the title perf: cut first-frame and progressive-mount cost on large documents perf: improve first frame and progressive-mount speed, optimize large documents Sep 6, 2026
@denolfe
denolfe merged commit 3b6d9f2 into main Sep 6, 2026
8 checks passed
@denolfe
denolfe deleted the perf/faster-first-frame branch September 6, 2026 02:24
denolfe added a commit that referenced this pull request Sep 9, 2026
- feat: copy selected text (#58) (a9f836e)
- fix(hr): stop adding a resize listener per horizontal rule (#57) (beedc6d)
- ci: add first-frame and scroll benchmarks to the PR bench job (#56) (b7ab07d)
- perf: improve first frame and progressive-mount speed, optimize large documents (#54) (3b6d9f2)
- chore(ci): pin bun 1.4.1 via mise (#53) (f596175)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant