Repository navigation
Reserve pure execution for lifespan loops and native workers - #22
Merged
Merged
Conversation
added 3 commits
October 8, 2026 09:51
The recovery and schedule traversal logged only "Tenant recovery scan failed" and "Tenant schedule scan failed", so a capacity refusal could not be told apart from any other error. Both lines now name the error class and, for a CatalogError, its code, never the message or a value. A unit test drives the real RecoveryLoop cycle through the recovery quanta into transition_async while two in-flight requests hold every work slot. The cycle fails with WV-OPERATION-CAPACITY and skips that tenant's schedule scan; once the requests finish, it reaches the kernel. The test pins today's admission design, in which lifespan loops borrow the request work pool.
In-flight requests hold a work or control slot for their whole response, and lifespan loops borrowed those same no-queue pools. Two Studio mutations, or four reads, refused recovery quanta, deadline settlements, schedule starts, provider and email dispatch, and Kafka consumer turns. Native connector executions held a request work slot for the whole connector call, so two of them took every mutation slot from Studio and a third failed at entry. ReservedSlots gives each owner its own pool. Inside its execution() context, every pure call takes one slot from that pool instead of the request pools, and an inherited request lease is set aside. A call that outlived its cancelled caller keeps its slot until it ends, so the size bounds the owner's pure threads, and a spare slot absorbs one such call as the shared work pool did. Recovery and provider cycles get 2 slots, Kafka turns 8 (at most max_clients - 1 turns, max_clients <= 8), and the native dispatcher its configured capacity plus one. Native claims, heartbeats and settlements keep request admission and their replay policy. Request admission itself is unchanged. The recovery catalog failure log now names the error class and code as the tenant logs do.
The ReservedSlots tests now share the request saturation helper in test_background_admission.py instead of duplicating it at the end of test_execution.py, which the compatibility rescan change also extends.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
A local acceptance run logged "Tenant recovery scan failed; traversal will continue" at the same time Studio requests were getting 429s.
Pure CPU work runs under no-queue leases:
WORK_SLOTS = 2for mutations andCONTROL_SLOTS = 4for reads and terminal control. Every HTTP request holds one lease for its whole response (BodyBoundary._serve).execute_pure()without an inherited lease takes a fresh one from those same pools and raisesCatalogError(429, "WV-OPERATION-CAPACITY")when none is free. The lifespan loops have no lease of their own, so they borrowed request capacity:wait_elapsedandsignal_receivedsettlementstimed_outdeadlines, terminal capacityruntime.startWV-SCHEDULE-READINESS)runtime.start,signals.deliver,settleruntime.start,signals.deliverper recordConnectorExecutionService.executeheld a request work slot for the whole connector call, network I/O includedHANDLER_FAILEDThe failure logs also gave no cause, so a capacity refusal could not be told apart from any other error.
What changes
Logs.
runtime/scheduler.pynow names the error class and, for aCatalogError, its code in "Tenant recovery scan failed", "Tenant schedule scan failed" and "Recovery catalog scan failed". It never logs the message or a value. Example:Tenant recovery scan failed (CatalogError WV-OPERATION-CAPACITY); traversal will continue.ReservedSlots(operations/execution.py). Inside itsexecution()context, everyexecute_purecall takes one slot from the owner's own semaphore instead of the request pools, and any inherited request lease is set aside.execute_purechanges by one line.dispatch_onetimeout. That is why slots are now taken per call.Owners:
RECOVERY_EXECUTION = ReservedSlots(2)PROVIDER_EXECUTION = ReservedSlots(2)TURN_EXECUTION = ReservedSlots(8)max_clients - 1turns,max_clients <= 8)ReservedSlots(sum of executor capacities + 1)execute(..., reservation=)Unchanged:
WORK_SLOTS,CONTROL_SLOTS) and the publicProcessCapabilitiescontract.ServiceTransport._replay, in parity with remote workers.ConnectorExecutionService.executewithout a reservation are still admitted like a request.Pure thread bound. It rises by at most 2 + 2 + 8 + (native capacity + 1). Native executions used to be capped at 2 by the work pool, with the third failing; the operator's configured capacity is now the ceiling.
Related
contextvars.Context(). Without it, a loop opened byPOST .../operations/compatibility/checkinherits the request's lease and keeps failing after the response. It touches only thecreate_tasklines of the same files.inventory_execution()for the compatibility rescan. That lease deliberately refuses an overlapping rescan, so it keeps its one-lease semantics and its own slot.Residuals
HANDLER_FAILEDwith an unknown outcome, although nothing ran. That only happens when orphaned calls hold every slot, the spare included.WV-OPERATION-CAPACITYis shared with database admission (uow.py, sqlstatesWQ001and55P03). After this change, one left in a loop's log comes from the database or from orphaned calls that hold a whole reservation.Verification
make checkon the branch head: all stages,All requested checks passed.tests/unit/operations/test_execution.py:tests/unit/runtime/test_recovery_admission.py, through the realRecoveryLoop.cycle→_quanta→_recover→transition_async→execute_purepath:timed_outdeadline while they hold every control slot;tests/unit/operations/test_background_admission.py:BrokerPolicy.max_clientscap.tests/unit/connectors/test_local_build.py: the handlersNativeDispatcher.open()registers execute through its reservation.execute_purechange fails three reservation tests;reservation=self.reservationfrom the dispatcher fails the wiring test.