skip chunk quota checks for superusers - #12
Merged
Merged
Conversation
alexandrusavin
marked this pull request as ready for review
August 3, 2026 16:53
christophwitzko
approved these changes
Aug 4, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
limits_overrides["max_chunks"]behavior for normal R2R users.list_chunks, while normal-user ingestion still does.Exact issue
Monoloom does not use R2R ownership for file permissions. It filters searches with the UUIDs of files and notes that the caller may access. Because Monoloom does not authenticate individual users to R2R, every stored chunk belongs to R2R's default admin user.
The Ruinart database confirmed:
owner_id.store_embeddings, which calledlist_chunks(limit=1, filters={"owner_id": ...})to enforcedefault_max_chunks_per_user.That call generated query fingerprint
-2001702420117276229:LIMIT 1does not make this cheap: PostgreSQL must process all chunks for that owner to calculateCOUNT(*) OVER(). The query also carries the widetextandmetadatacolumns through that operation. Since all Monoloom chunks have the same owner, every ingestion counted the entire chunk corpus.Increasing
default_max_chunks_per_userwould not help because the count query runs before R2R reads and compares the configured limit.Incident evidence
This was observed during the Ruinart 504 incident.
Azure Query Store for 2026-08-03 15:30:14–15:45:14 UTC, which overlaps the incident report at 15:38 UTC, recorded:
-2001702420117276229.At PostgreSQL's 8 KiB block size, that is approximately 2.54 GiB read and 2.54 GiB written per execution on average, or roughly 5.08 GiB of temporary I/O per quota check.
A nearby 15-minute window recorded 52 executions averaging 147.1 seconds, with a maximum of 239.1 seconds.
The browser polling storm investigated in the Slack thread amplified overall load, but it did not create this SQL. This query came from ingestion-side quota bookkeeping and can recur under concurrent indexing even after polling is reduced.
Why this fix
R2R's document API already skips its route-level quota preflight for superusers. The ingestion service did not apply the same exemption and still counted every chunk before storing embeddings.
This change makes the ingestion path consistent with the API path:
This does not change Monoloom permissions or R2R access filtering.
Validation
ruff check,ruff format --check, andgit diff --checkalso pass.Deployment follow-up
After merging this PR:
interloom.azurecr.io/r2rimage.-2001702420117276229no longer appears during file or note ingestion.A separate defense-in-depth improvement can replace the general
COUNT(*) OVER()implementation with a narrow count query for deployments that use normal R2R user quotas.