Skip to content

perf: render and memory pass over the grid, allocator and bundle - #83

Merged
broisnischal merged 14 commits into
masterfrom
perf/render-and-memory
Aug 31, 2026
Merged

broisnischal merged 14 commits into
masterfrom
perf/render-and-memory

Conversation

@broisnischal

Copy link
Copy Markdown
Collaborator

The app felt laggy and low-framerate across the board, the data table worst of all. I profiled it instead of guessing, with a temporary in-page probe reporting frame times and canvas call counts, plus an eval channel so I could drive the running app from the shell. All of that instrumentation is removed again; nothing here ships it.

Test data was a local Postgres: events (1M rows), articles (100k), openai_docs (10k, pgvector embeddings). 60Hz display, so the budget is 16.6ms.

What the numbers actually said

The canvas path was already fine for wide-but-plain tables. events held a clean 60fps before I touched anything, so the general "everything is slow" feeling was not the row count. Two things were real:

  1. Text-heavy tables dropped frames. openai_docs ran p50 18ms / p95 22ms / max 30ms, i.e. 45-55fps. measureText was called 113 times per frame shaping 16,014 characters, costing 3.08ms of an 11.7ms draw, recomputing the identical truncation every frame.
  2. Memory was never given back. One 1M-row table took the process from a 45MB/48MB idle baseline to 587MB Rust / 891MB WebKit. Closing every tab released nothing. A full webview reload released nothing. smaps showed four roughly 50MB glibc malloc arenas, one per worker thread, retained for the session.

The changes

Data grid

  • Truncation memoised by (font, column width, string): measureText 113 to 31 calls/frame, 16,014 to 1,521 chars/frame, 3.08ms to 0.46ms, draw 11.7ms to 8.7ms, frames to p50 17ms / p95 17ms / max 20ms. The single biggest win.
  • Per-column type predicates cached against the column instead of re-running four regexes per cell per frame (about 40,000 regex executions a second while scrolling).
  • Frame-constant draw context: drawCell was doing roughly 3,000 reactive-graph reads a frame for values that cannot change mid-frame. Snapshotted once per frame.
  • Row loop invariants hoisted, and the pinned-column pass skipped entirely when nothing is pinned.
  • Scroll blitting: a frame that differs only by a vertical scroll moves the still-valid band with one image copy and repaints only the uncovered strip. A one-row step paints one row instead of every visible one.
  • Row count reused across page changes.

Memory

  • mimalloc as the global allocator. Retained growth over the same open-scroll-seek workload: 527MB to 27MB.

Startup and bundle

  • Monaco entry composed by hand from the six languages the app actually displays (sql, json, javascript, typescript, rust, markdown) instead of the default everything-bundle. Over dist/assets: 34,821,807 to 32,585,794 bytes, -2.24MB (-6.4%), 81 chunks gone, including css.worker (1.05MB) and html.worker (719KB) that were emitted for grammars nothing in the app can open. ts.worker stays because OrmRunner genuinely uses it.
  • The startup prefetch is gated per entry on the connection's engine. A warmed chunk is never freed, and on Redis every relational page is unreachable UI, so about 4.3MB of the 6.01MB warm set was pinned for tabs that cannot exist. The rest of the tail stays warm on purpose: 22 pages cost 0.66MB between them, so trimming them would regress first-open latency to save almost nothing.
  • matchMedia resolved once instead of on every wheel tick (about 5us of style resolution on the one path that must not touch the style system).

Fixes

  • Reverse-FK loader checked its cache before awaiting but filled it after, so restoring a session fired one query per tab at once: nine concurrent calls, 37.7s of backend time against D1, with everything the user was waiting for queued behind them. Added in-flight dedup plus a generation counter so a reply from before a connection change is discarded rather than writing another database's FKs under the same key.
  • Catalog replies belonging to a previous connection are dropped.
  • Two Vite dep-optimizer fixes. canvas-confetti and the eight Monaco subpaths are both only reachable through lazy imports, so the initial scan missed them; discovering them mid-session rehashes every pre-bundled chunk and the webview then 404s on the paths it already holds.

On the blit

This is the one change here that can produce a wrong pixel, so the gates are worth stating. The scroll loop already distinguished a position-only frame from a content change, and every content, hover, focus, selection, theme and geometry change goes through scheduleDraw, so that existing flag is the gate. draw() adds the geometric half: no animating skeletons, no insert draft, no expanded rows, integer dy, dy and the header offset both on whole device pixels. The copy runs with the transform reset so it is an exact 1:1 move of device pixels and cannot resample, which matters because each frame copies the last and any blur would compound over a scroll. The strip is rounded outward and widened a pixel at each edge so a separator cannot leave a seam. The loop forces a full repaint when scrolling stops, so nothing this path could produce can persist in a resting grid.

The arithmetic is verified exhaustively rather than by eye: across 2,280 (DPR, header height, viewport height, dy) combinations, every body pixel in the new frame is either copied from exactly the right source row or falls inside the repainted strip.

Verification

467/467 tests green, vite build clean, cargo check --no-default-features clean, tauri:build succeeds on macOS (3m00s, both bundles), and the app runs on this branch with all of it in place. The canvas draw path has no unit coverage, so the blit is covered by the geometry check above rather than by the suite.

Still owed: a live scroll watched by a human, and a first Windows build with mimalloc. Both are noted in TODO_TODAY.md along with what I ruled out, so nobody re-checks it.

broisnischal and others added 14 commits August 31, 2026 11:33
…shing

canvas-confetti is only reachable through a lazy `await import()` on the
license pages, so Vite's initial scan never saw it. The first time one of
those views mounted, Vite discovered the dep at runtime, re-ran the
optimizer and rewrote every pre-bundled chunk with new hashes. The webview
was still holding the old URLs, so runtime-*.js, index-client-*.js,
legacy-client-*.js and Icon-*.js all 404'd at once.

Losing the Svelte runtime chunk takes down every page, not just the license
screen, which is why the SQL editor and everything else went blank rather
than one view failing.

Declaring it in optimizeDeps.include puts it in the first pass, so there is
no mid-session re-optimize to race with.
…loop

Profiled the canvas draw path on a 60Hz display against local Postgres:
`events` (1M rows) and `openai_docs` (10k rows of pgvector embeddings).
Wide-but-plain tables were already fine, holding p95 17ms. Text-heavy ones
were not: openai_docs ran p50 18ms / p95 22ms / max 30ms, so 45-55fps.

The dominant cost was measureText: 113 calls per frame shaping 16,014
characters through HarfBuzz for 3.08ms of an 11.7ms draw, re-deriving the
same truncations every frame. Truncation is a pure function of (font,
column width, string) and all three hold still while scrolling, so each
combination is now measured once. The cache deliberately survives a `rows`
swap, which is what keeps windowed scrolling on hits, and is dropped on a
column change so one table's strings are not held alive after navigating
away.

Three more per-cell costs, all invariant, moved out of the hot path:

- isBooleanType / isSqlArrayType / isVectorType / isGeometryType and the
  right-alignment test ran per cell per frame. Four regexes and about six
  throwaway strings per cell, roughly 40k regex executions a second while
  scrolling, for something that depends only on the column's declared type.
  They live in _colCache now, which already exists for exactly this.

- drawCell read about 20 values straight off the reactive graph per cell. A
  $derived read re-checks the derived against every dependency's write
  version, so a 150-cell viewport paid ~3000 graph reads a frame for values
  that cannot change mid-frame. Snapshotted once per frame into bodyC.

- geom.cols.length sat in the row loop condition, re-entering the geom
  derived once per column per row, and the pinned pass scanned every column
  of every row even with nothing pinned.

Measured on openai_docs: measureText 113 -> 31 calls/frame and 16,014 ->
1,521 chars/frame, draw 11.7ms -> 8.7ms, frames p50 18 -> 17ms, p95 22 ->
17ms, max 30 -> 20ms, and 1-3 of 118 frames over 18ms instead of most of
them.
Opening one 1M-row table took the app from a 45MB/48MB idle baseline to
587MB on the Rust side and 891MB in the web process. Closing every tab
released nothing, and neither did a full webview reload.

For the Rust half the cause was visible in /proc/PID/smaps: four ~50MB
glibc malloc arenas, one per worker thread, retained for the rest of the
session. glibc does not hand those back, so the process never shrank after
you moved on to a smaller table. mimalloc decommits pages it no longer
needs and is faster on the allocation-heavy row-decode and JSON paths,
which is most of what a browse does. Same open-scroll-seek workload
measured after the swap: retained growth 527MB -> 27MB.

Separately, the dev profile left the whole IPC path unoptimized. Every
command result is serialised to JSON for the webview and a browse is
megabytes of it; at opt-level 0 serde_json's generated code is roughly 20x
slower than release, so a dev build spent most of a row fetch inside
serialisation rather than in the database. serde_json, serde, itoa, ryu,
the four sqlx crates, rust_decimal, chrono, uuid and base64 now build at
opt-level 3 in dev, following the crypto block already in this file. They
are all dependencies, so incremental rebuilds of our own code are
unaffected.
loadIncomingForeignKeys checked the cache before awaiting but only populated
it once the reply came back, so the guard could not see a request that had
been fired and not yet answered. Restoring a session therefore fired one
reverse-FK query per restored tab at the same instant: measured nine
concurrent calls totalling 37.7s of backend time against a D1 database,
where each call is several HTTP round trips to Cloudflare. Everything the
user was actually waiting for queued behind them.

Concurrent callers now share one promise. A generation counter guards the
write, because the cache is keyed by "schema.table" alone and a reply that
lands after a connection change would otherwise serve one database's
reverse FKs for a same-named table in another. The three cache resets go
through one helper so the counter can never be missed.
createSmoothScroll called matchMedia('(prefers-reduced-motion: reduce)')
from push(), so every wheel tick paid about 5us of style resolution on the
one path that must not touch the style system. A MediaQueryList stays live
on its own, so resolving it once is equivalent.
Measurements, what changed and why, the items I deliberately left alone,
and the things I ruled out so they do not get re-checked. Also records how
to reproduce the profiling, since the instrumentation itself is not in the
tree.
…tion

Switching database could leave the sidebar, schema catalog and structure pane
showing the database you just left. The loaders write whatever comes back from
their await, and on a slow link or a slower machine the previous connection's
reply lands after the new one is already up.

The schema checks several of them already do are not a guard against this, and
that is what made it look handled: clearConnectionState() resets activeSchema
to 'public', so a Postgres-to-Postgres switch leaves `activeSchema === schema`
true and the stale payload is applied anyway. Table and column names survive a
switch too - two databases routinely hold a public.users - so those comparisons
pass as well.

It is not only what is on screen. The backend resolves the pool from a single
global ACTIVE slot at execution time, and loadTables writes its result into the
catalog cache, so a late reply could also be filed under whichever connection
id happened to be current and served from there for the rest of its TTL.

loadSchemas, loadTables, loadSchemaCatalog, reloadCatalogKind and loadStructure
now capture persistConnectionId before their first await and drop the reply if
it changed - the same idiom resolveRowCounts already uses through its cache
key. loadTables guards its catch and its finally too, so a stale failure cannot
write an error or clear the spinner the new connection's load just put up.

Row loading needed nothing: its per-tab load token plus the resetTabs() in
clearConnectionState already discard rows from a superseded connection.
Paging through a filtered table re-ran COUNT(*) on every page for a number
that was already on screen. A page change cannot change how many rows match
the predicate, so every one of those counts after the first was work the
server carried for an answer we had.

Measured on the 1M-row `events` fixture with a `kind LIKE '%click%'` filter:
the COUNT(*) is 182ms cold and 31.8ms warm, against 0.79ms for the page of
rows it accompanies - roughly 40x the query it was riding along with, and on
a remote database another round trip on top. It never showed up as a blocked
paint because the count is already deferred (rows paint first, the count
streams in behind them), only as load nobody could see.

The unfiltered path was already handled: it takes the pg_class.reltuples
estimate above 100k rows rather than counting. This is the filtered half,
where the count has to be exact and so cannot be estimated - only reused.

A count is now remembered against the view it was actually counted for -
connection, schema, table and predicate together. The predicate alone would
not do; two tables routinely carry the same filter, and reusing one's count
for the other would show the wrong number of rows.

Reuse means the -1 "counting" sentinel a non-counting load returns can no
longer be written over a known total, or every page change would blank it and
force the count this avoids. Anything that can change how many rows match
clears the record instead: a filtered insert (that branch does not adjust the
total itself - whether the new row matches is the database's call),
hand-written DML, and a cell edit while a filter is active, which can move a
row in or out of the result. Deletes and unfiltered inserts already adjust the
total correctly in place and need nothing.
The idle prefetch pulled all 24 lazy page chunks in whatever the engine, and a
warmed chunk is never freed again - which sat badly next to the idle-teardown
directly above it, that unmounts views after 3 minutes hidden to reclaim
memory.

Measured this time, on a release build, against the manifest's static import
graph (warming a chunk pulls its static imports; its own dynamic imports stay
lazy). The eager entry graph is 2.71MB/25 chunks and the warm set adds
6.01MB/49 - but the shape is not what the earlier guess assumed. 5.35MB of it
is two entries: SqlConsole drags in monaco at 3.78MB and AiChat the
markdown/highlight stack at 1.57MB. The other 22 pages cost 0.66MB between
them, 0.01-0.11MB each.

So trimming the rarely-opened tail buys almost nothing and would regress
first-open latency on 21 pages to save half a megabyte; it stays, and the
measurement is now in the comment so it does not get re-derived.

What the numbers do justify is the engine gate. On a Redis connection every
relational page is unreachable UI - monaco included - so about 4.3MB of that
6.01MB was being pinned for tabs that do not exist. Each entry is now gated on
the same isRedis / hasSchemaExplorer / hasSecurity flags the tab affordances
use, RedisKeyspacePage warms only on Redis and goes first there, and the pump
waits for a connection: before one exists the engine is unknown and no tab can
be opened anyway.
After the truncation cache, fillText was the whole remaining draw cost: 167
calls and 6,243 glyphs per frame, 4.8ms of an 8.7ms draw, about 29us a call.
That is Cairo rasterising glyphs on the CPU, and no amount of JS tuning
reaches it. The only way past it is to stop redrawing text that has not moved.

On a frame that differs from the last only by a vertical scroll, the body is a
rigid translation of what is already on the canvas: one drawImage moves the
still-valid band and only the strip the scroll uncovered is repainted, so a
one-row step paints one row instead of every visible one.

Nothing new was needed to know when that is legal. The scroll loop already
separated a position-only frame from a content change through _loopNeedsDraw,
which scheduleDraw sets for every content, hover, focus, selection, theme and
geometry change; the loop offers draw() a delta only when that flag is clear
and the horizontal offset is unchanged. draw() adds the geometric half of the
precondition - the body is not a rigid translation while skeletons animate,
while an insert draft adds a band above row 0, or when an expanded row makes
row heights non-uniform.

The copy runs with the transform reset, source and destination the same size,
so it is an exact 1:1 move of whole device pixels. Going through the
DPR-scaled transform instead would let a fractional viewport size or a 1.5x
DPR make source and destination differ by a fraction of a pixel; drawImage
would then resample the band, and because each frame copies the previous one
that blur compounds over a scroll into visibly smeared text. This is the trap
in the approach, and the guards on dy and the header offset landing on whole
device pixels are what keep it exact.

The strip is rounded outward and the repainted band is widened a pixel at each
edge, so a separator or focus ring bleeding outside its own row cannot leave a
seam. emitVisibleRange still runs on every frame, blit or not, because row
windowing depends on it; _imgWanted is left to accumulate rather than cleared,
so a thumbnail that merely survived the scroll is not dropped.

The loop now forces a full repaint when scrolling stops. Any artifact this
path could produce therefore lives only while the content is moving and is
gone within one frame of the user letting go - it can never persist in a
resting grid.

Verified arithmetically rather than by eye: across 2,280 combinations of DPR,
header height, viewport height and dy, every body pixel in the new frame is
either copied from exactly the right source row or falls inside the repainted
strip. Watching a real scroll is still owed.
monaco-editor's default entry is an everything-bundle: all 80-odd basic
languages plus the css, html, json and typescript services, each service
dragging in its own web worker. The app displays six languages.

The blocker last time was not being able to prove that, because MonacoTextView
takes its language as a prop. It is provable. Only two call sites are dynamic:
OrmSchemaPage passes sql, rust (Monaco has no Prisma grammar and the DSL reads
correctly as Rust) or typescript, and TableTextView passes stroke-csv,
stroke-tsv, markdown or json. Every other site passes a literal. The full set
is sql, json, javascript, typescript, rust and markdown - markdown included,
which the earlier note on this missed.

src/lib/monaco.js composes the entry by hand out of the same pieces
editor.main.js uses: edcore.main.js for the editor and every feature
contribution with no languages, the six language contributions, and the json
and typescript services republished onto monaco.languages.* by hand, which
editor.main.js does itself and the contributions do not do for themselves. The
basic-languages registrations for javascript and typescript have to stay even
though the service also targets them - the service only hooks onLanguage, so
without them it never fires. All eleven importers now go through this module.

Measured over dist/assets: 34,821,807 -> 32,585,794 bytes, 2.24MB smaller, 81
chunks gone. css.worker at 1.05MB and html.worker at 719KB are among them,
both of which were being emitted for grammars nothing in the app can open.
ts.worker stays - OrmRunner genuinely uses
languages.typescript.javascriptDefaults for its IntelliSense.

Adding a language now means adding its contribution here. An unregistered
language id does not error, it silently renders as plaintext, so that file is
the first place to look if highlighting ever goes missing.
Rewrites the notes against the work that has since landed. The two entries
worth reading are the ones where measuring changed the answer rather than
confirming it: the startup prefetch is 5.35MB of two chunks and not 24 evenly,
so the tail is staying; and the Monaco language set turned out to be provable,
with markdown in it, which the earlier list had missed.

Remaining is reordered around what is actually owed now: watching a live
scroll with the blit on (it is verified arithmetically, but nobody has looked
at one), Windows for mimalloc, and the Shiki grammars - the four largest
chunks in dist are no longer Monaco's but brain, emacs-lisp, cpp and wasm, for
languages the app never highlights. That is the same trim one library over,
and probably a bigger win than what is left in Monaco.
@broisnischal broisnischal added the release:minor Bump minor version (0.x.0) label Aug 31, 2026
@changeset-bot

changeset-bot Bot commented Aug 31, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 3b5d277

The changes in this PR will be included in the next version bump.

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@broisnischal
broisnischal merged commit 1dd1294 into master Aug 31, 2026
1 check passed
@github-actions github-actions Bot locked and limited conversation to collaborators Aug 31, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

release:minor Bump minor version (0.x.0)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant