Repository navigation
2026.10.6: load-balance weights, upstream concurrency, model specs, key usage limits (core v0.63.0) - #287
Merged
Conversation
The next round of Lite work (load-balance weights and balance_by, switching on a slow stream start, per-upstream concurrency caps, hand-set model specs, per-key usage limits) needs the protocol from core's integ/routing-2026-10, which has no tag yet. The six core crates are pinned to that commit by `rev` until core is released; the lock file changes only in their source lines, since the crates' versions and dependencies did not move. Types: tw-api.ts is regenerated. The new required fields (ClientView.limits, GroupView.weights/balance_by, FailoverView.next_on_slow_start/slot_wait_secs) are filled in test fixtures and in the notices snapshot with the values core now reports for an unchanged configuration (no limits, weight 1, balance by weights, switching off, 30 s slot wait), so every test still describes today's behaviour. The failover settings form keeps its seven fields: the two new ones are excluded from its field type until that section is designed. The new `busy` skip reason gets a label so the exhaustive switch keeps compiling. Control plane: `SetModelSpec` (PUT /provider-model-spec) is added to the endpoints the interface may call, in call.rs, control.ts and the screenshot mock, so the upcoming model-spec UI does not have to touch the allowlist. Messages: core adds 30 codes; each has a Chinese sentence now, so the Chinese interface does not fall back to English for slow-start switches, busy upstreams, key limit refusals or the new validation errors. Core fills `per` and `measure` with English words (day, tokens) and expects the interface to look them up, so three small word tables are added; new template cases run the table lookups through both the TypeScript and the Rust implementation. Screenshot data: the stored core answers are regenerated with oracle.sh, which now also reads a `rev` pin, so the shots mock still type-checks against the core types. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Core's routing round lets a load-balance group split new conversations by
member weights (1 to 100) and, on top of that ratio, by speed, reliability
or both. Without UI the only way to use either was editing config.yaml,
and nothing on the routing page told a 7 : 3 group from an even one.
Group dialog: a load-balance group gets a "Distribute by" segmented
control (by ratio / by speed / by reliability / by speed and reliability)
with one line of what it means, and a Weight column for the chosen
members, like the Preferred column of a select group. Weights default to
an explicit 1, are checked on the spot (whole numbers 1 to 100, the same
range core enforces) and are sent with balance_by through the existing
group save; other strategies send neither, so their configs stay as they
are. The load-balance description now says new conversations, which is
what the strategy distributes.
Summaries: the group table shows the ratio under the strategy when the
weights are not all 1, and the distribute-by mode when it is not by
ratio, each on its own line so the narrow column does not break a phrase.
The routing map's tooltip, the rule target descriptions and the command
palette use the same wording, e.g. "Round robin (7 : 3 · By speed)".
Dry run: each candidate of a load-balance group shows its weight; in an
automatic mode also its TTFB, success rate ("no data yet" when core has
too few samples) and its share of new conversations, computed as core
does (weight × factor, members with an open circuit sit out unless all
do). The strategy badge names the mode.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Core now records three new things in a request's routing: an attempt abandoned because no content arrived in time (outcome slow_start, with the input the upstream may have billed), a candidate skipped because it was at its concurrency limit (skipped: busy), and the time an attempt waited for a free slot (queued_ms). It also refuses requests with two new kinds of error: a gateway key's usage limit (gw.key_limit.*) and every upstream being full (gw.busy_all). Attempt chain: a slow start reads "Start timed out" and a busy skip "At its concurrency limit", each with core's own sentence on hover (how long it waited, how many requests the upstream already had). A busy skip was never sent, so it shows no duration, like a hop a rule denied. A wait for a slot shows as "Queued 1.2 s" next to the hop, since core keeps it out of the hop's duration. Under an abandoned hop one line gives the input tokens (the upstream's own numbers when it reported them, otherwise "about N" from the gateway's estimate) and says the upstream may have billed them; the hover adds that they are not part of this request's cost. The note under the chain no longer says "the first N upstreams failed" when some of those hops were busy skips or slow starts: neither is the upstream's failure, so it says the first N attempts did not take the request. Refusals: until now a request with an empty upstream and an error was either a rule denial or "No upstream". A key-limit refusal also has an empty upstream and would have been shown as "No upstream", which sends the user to the upstreams page; it now reads "Usage limit" with its own icon. A busy refusal is recorded on the last upstream looked at, which received nothing, so the upstream column, the drawer and the session upstream lists would have blamed that upstream; it now reads "At capacity" there and is left out of a session's upstreams. Both keep core's translated sentence as the reason and stay under "Failed only". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… notifications at 80% and at the limit
Core can now cap a gateway key by requests, tokens or cost per minute,
hour, day, week or month (clients[].limits), and reports each limit with
what it has used. This gives that a place in Lite.
- Key dialog: a "Usage limits" section. Each limit is one row that reads
as a sentence (Per [day] at most [5.00] [USD]); token limits add
"Count cache reads". Rows are checked on the spot with core's rules
(a positive amount, no two limits with the same period, measure and
cache-read setting, and a monthly limit needs records kept for 31
days), so a wrong limit is caught before Save. A blank amount is only
flagged once the field is left, so a newly added row is not red. Cost
is typed in dollars and sent as micros.
- For an existing key each row shows what is used ("Today $1.23 /
$5.00 · resets at 00:00", "Last minute 12 / 30") and marks a reached
limit. With a cost limit, the models this key can use that have no
price are listed (folded after six): they count as $0, so the limit
does not hold them back.
- Keys table: a "Limit reached" chip next to the disabled state; hover
names the limit and when it resets. The list is fetched again on
key_limit_alert, at the next period reset and after the clock jumps,
so the chip does not linger after midnight on an idle key.
- Disabling or enabling a key from the table now sends its limits back:
the key input replaces the whole list, so leaving them out would have
deleted every limit.
- Notifications: KeyLimitAlert becomes a system notification through the
notice bus, naming the key, the limit and when it resets. One notice
per limit: reaching the limit replaces the 80% notice, and the next
period's 80% notice replaces yesterday's "reached". Each period counts
as a new event, so a notice read yesterday does not silence today's.
Clicking it opens the Keys page.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…art switch and slot wait Core's routing round adds three things an upstream owner sets by hand and two failover settings; this makes all of them reachable without editing config.yaml. Upstream dialog: an optional "Concurrency limit" (max_concurrent, 1-1000, blank shows "No limit"). It sits under the proxy fields in the Connection section and in the Account section of ChatGPT accounts, since accounts limit concurrent requests too. It is checked on the spot and blocks saving like the other required fields. The form now carries max_concurrent both ways, so the table's enable/disable toggle, which saves the upstream from its view, keeps the limit instead of dropping it. The row's in-flight tooltip names the limit when one is set. Models popover: hovering a model offers "Specs..." next to "Add alias..."; it opens a small dialog for the context window and max output of that model on that upstream, saved through PUT /provider-model-spec with the version the dialog opened on. A blank field shows the price table's value (or that the table has none) as its placeholder, so blank never means something hidden; clearing both removes the hand-set entry, and the footer says so. Models with a hand-set value carry a "Manual specs" mark next to their name (kept on the name side so it stays visible while the row's right side gives way to the buttons); its tooltip lists which values are manual. The context window column of the edit dialog's Models section now prefers the hand-set value too, so the two places never show different numbers for one model. Settings > Failover: "Move to the next upstream when the start times out" is a switch hung under the stream-start wait, since it says what happens when that wait runs out; "Wait for a free slot at most" (slot_wait_secs, 0-300, default 30 shown) follows. With the switch on the wait must be 5 to 120 seconds, checked on the spot like the section's other ranges, including when only the switch changed; core's config.slow_start_too_short still surfaces in the section's error banner if the file changed underneath. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…equests Core measures the speed sample behind url-test and load-balance by speed from sending a hop to the first content of the answer, not to the response headers. The dry run and the url-test description called it TTFB / first byte, the word the request detail keeps for the headers; they now say first token, as the request detail, traffic and upstream table already do for that measurement. The terminology table records both terms. Weights share requests, not new conversations: a turn that stays on an upstream is charged to it. The load-balance and distribute-by descriptions now talk about requests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Core's routing round is final at adadbf5 (integ/routing-2026-10, protocol 39). The six core crates move to that rev; the lock file changes only in their source lines. tw-api.ts is regenerated: the protocol constant and doc comments changed, no field did. Core's checks changed, so the UI follows them: - Key limits: a cost limit has to be at least $0.01, and day, week and month limits need request records kept for 1, 7 and 31 days (config.key_limit_cost_too_small, config.key_limit_retention, which replaces the month-only code). The key dialog checks both on the spot, in core's order, and names the period in the retention line. - New codes translated: config.key_limit_retention, config.key_limit_cost_too_small, gw.ws.upstream_closed; the retired config.key_limit_month_retention is removed. - failover.slot_wait_secs is now one total wait per request, for a full upstream and for a key's minute or hour limit alike, so the setting no longer says it is only about concurrency. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The description that now covers both waits (a full upstream, a key's minute or hour limit) wrapped to two lines where every other failover setting takes one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ion commit) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The UI for core v0.63.0 (control-plane protocol 39), released as 2026.10.6.
Strategy groups — load-balance groups take a weight per member (1–100) and a "Distribute by" choice: by ratio, by speed, by reliability, or both. The group table shows the ratio and mode; the route test shows each candidate's weight, time to first token, success rate and share.
Upstreams — an optional concurrency limit in the upstream dialog; "Specs…" in the models popover sets a model's context window and max output by hand (
PUT /provider-model-spec), marked "Manual specs" and used instead of the price table. The edit dialog's context-window column prefers the manual value.Failover settings — "Move to the next upstream when the start times out" with the existing stream-start wait, and "Wait for a free slot at most" (default 30 s, the total wait for full upstreams and a key's minute/hour limits).
Keys — a "Usage limits" section: each limit a row (per minute/hour/day/week/month, at most N requests/tokens/USD; token rows can count cache reads), validated as core does, with usage and reset time per row, and a warning listing the reachable models without a price (counted as $0). Keys at a limit get a "Limit reached" chip;
KeyLimitAlertbecomes a notification (80% and reached). Toggling a key from the table now sends its limits back, since core replaces the list on save.Traffic — the attempt chain shows "Start timed out" (with the possibly billed input tokens), "At its concurrency limit" skips and time queued for a slot; refusals by key usage limits and by all-busy upstreams are labelled in the list, drawer and sessions.
Core v0.63.0 — six core crates pinned to the tag, regenerated types,
SetModelSpecallowlisted, Chinese translations for the 30 new message codes plus thelimit_per/limit_measureword tables.Also included from dev: check-up findings on upstream rows (#285) and the menu bar menu following the app's appearance (#286).
2026.10.6 — version in the four files and
release-notes/2026.10.6.md. A 0.62.0 initial configuration and the current local configuration both load on 0.63.0; request history is kept (schema unchanged).Checked locally against twcore 0.63.0: typecheck, 869 unit tests, build,
cargo fmt, clippy and 859 Rust tests.🤖 Generated with Claude Code