Skip to content

2026.10.6: load-balance weights, upstream concurrency, model specs, key usage limits (core v0.63.0) - #287

Merged
fylorn merged 15 commits into
devfrom
ui-routing-integ
Oct 5, 2026
Merged

fylorn merged 15 commits into
devfrom
ui-routing-integ

Conversation

@fylorn

@fylorn fylorn commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

The UI for core v0.63.0 (control-plane protocol 39), released as 2026.10.6.

Strategy groups — load-balance groups take a weight per member (1–100) and a "Distribute by" choice: by ratio, by speed, by reliability, or both. The group table shows the ratio and mode; the route test shows each candidate's weight, time to first token, success rate and share.

Upstreams — an optional concurrency limit in the upstream dialog; "Specs…" in the models popover sets a model's context window and max output by hand (PUT /provider-model-spec), marked "Manual specs" and used instead of the price table. The edit dialog's context-window column prefers the manual value.

Failover settings — "Move to the next upstream when the start times out" with the existing stream-start wait, and "Wait for a free slot at most" (default 30 s, the total wait for full upstreams and a key's minute/hour limits).

Keys — a "Usage limits" section: each limit a row (per minute/hour/day/week/month, at most N requests/tokens/USD; token rows can count cache reads), validated as core does, with usage and reset time per row, and a warning listing the reachable models without a price (counted as $0). Keys at a limit get a "Limit reached" chip; KeyLimitAlert becomes a notification (80% and reached). Toggling a key from the table now sends its limits back, since core replaces the list on save.

Traffic — the attempt chain shows "Start timed out" (with the possibly billed input tokens), "At its concurrency limit" skips and time queued for a slot; refusals by key usage limits and by all-busy upstreams are labelled in the list, drawer and sessions.

Core v0.63.0 — six core crates pinned to the tag, regenerated types, SetModelSpec allowlisted, Chinese translations for the 30 new message codes plus the limit_per / limit_measure word tables.

Also included from dev: check-up findings on upstream rows (#285) and the menu bar menu following the app's appearance (#286).

2026.10.6 — version in the four files and release-notes/2026.10.6.md. A 0.62.0 initial configuration and the current local configuration both load on 0.63.0; request history is kept (schema unchanged).

Checked locally against twcore 0.63.0: typecheck, 869 unit tests, build, cargo fmt, clippy and 859 Rust tests.

🤖 Generated with Claude Code

fylorn and others added 15 commits October 5, 2026 17:20
The next round of Lite work (load-balance weights and balance_by,
switching on a slow stream start, per-upstream concurrency caps,
hand-set model specs, per-key usage limits) needs the protocol from
core's integ/routing-2026-10, which has no tag yet. The six core crates
are pinned to that commit by `rev` until core is released; the lock file
changes only in their source lines, since the crates' versions and
dependencies did not move.

Types: tw-api.ts is regenerated. The new required fields (ClientView.limits,
GroupView.weights/balance_by, FailoverView.next_on_slow_start/slot_wait_secs)
are filled in test fixtures and in the notices snapshot with the values
core now reports for an unchanged configuration (no limits, weight 1,
balance by weights, switching off, 30 s slot wait), so every test still
describes today's behaviour. The failover settings form keeps its seven
fields: the two new ones are excluded from its field type until that
section is designed. The new `busy` skip reason gets a label so the
exhaustive switch keeps compiling.

Control plane: `SetModelSpec` (PUT /provider-model-spec) is added to the
endpoints the interface may call, in call.rs, control.ts and the
screenshot mock, so the upcoming model-spec UI does not have to touch
the allowlist.

Messages: core adds 30 codes; each has a Chinese sentence now, so the
Chinese interface does not fall back to English for slow-start switches,
busy upstreams, key limit refusals or the new validation errors. Core
fills `per` and `measure` with English words (day, tokens) and expects
the interface to look them up, so three small word tables are added; new
template cases run the table lookups through both the TypeScript and the
Rust implementation.

Screenshot data: the stored core answers are regenerated with oracle.sh,
which now also reads a `rev` pin, so the shots mock still type-checks
against the core types.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Core's routing round lets a load-balance group split new conversations by
member weights (1 to 100) and, on top of that ratio, by speed, reliability
or both. Without UI the only way to use either was editing config.yaml,
and nothing on the routing page told a 7 : 3 group from an even one.

Group dialog: a load-balance group gets a "Distribute by" segmented
control (by ratio / by speed / by reliability / by speed and reliability)
with one line of what it means, and a Weight column for the chosen
members, like the Preferred column of a select group. Weights default to
an explicit 1, are checked on the spot (whole numbers 1 to 100, the same
range core enforces) and are sent with balance_by through the existing
group save; other strategies send neither, so their configs stay as they
are. The load-balance description now says new conversations, which is
what the strategy distributes.

Summaries: the group table shows the ratio under the strategy when the
weights are not all 1, and the distribute-by mode when it is not by
ratio, each on its own line so the narrow column does not break a phrase.
The routing map's tooltip, the rule target descriptions and the command
palette use the same wording, e.g. "Round robin (7 : 3 · By speed)".

Dry run: each candidate of a load-balance group shows its weight; in an
automatic mode also its TTFB, success rate ("no data yet" when core has
too few samples) and its share of new conversations, computed as core
does (weight × factor, members with an open circuit sit out unless all
do). The strategy badge names the mode.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Core now records three new things in a request's routing: an attempt
abandoned because no content arrived in time (outcome slow_start, with
the input the upstream may have billed), a candidate skipped because it
was at its concurrency limit (skipped: busy), and the time an attempt
waited for a free slot (queued_ms). It also refuses requests with two
new kinds of error: a gateway key's usage limit (gw.key_limit.*) and
every upstream being full (gw.busy_all).

Attempt chain: a slow start reads "Start timed out" and a busy skip
"At its concurrency limit", each with core's own sentence on hover
(how long it waited, how many requests the upstream already had). A
busy skip was never sent, so it shows no duration, like a hop a rule
denied. A wait for a slot shows as "Queued 1.2 s" next to the hop,
since core keeps it out of the hop's duration. Under an abandoned hop
one line gives the input tokens (the upstream's own numbers when it
reported them, otherwise "about N" from the gateway's estimate) and
says the upstream may have billed them; the hover adds that they are
not part of this request's cost.

The note under the chain no longer says "the first N upstreams failed"
when some of those hops were busy skips or slow starts: neither is the
upstream's failure, so it says the first N attempts did not take the
request.

Refusals: until now a request with an empty upstream and an error was
either a rule denial or "No upstream". A key-limit refusal also has an
empty upstream and would have been shown as "No upstream", which sends
the user to the upstreams page; it now reads "Usage limit" with its own
icon. A busy refusal is recorded on the last upstream looked at, which
received nothing, so the upstream column, the drawer and the session
upstream lists would have blamed that upstream; it now reads
"At capacity" there and is left out of a session's upstreams. Both keep
core's translated sentence as the reason and stay under "Failed only".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… notifications at 80% and at the limit

Core can now cap a gateway key by requests, tokens or cost per minute,
hour, day, week or month (clients[].limits), and reports each limit with
what it has used. This gives that a place in Lite.

- Key dialog: a "Usage limits" section. Each limit is one row that reads
  as a sentence (Per [day] at most [5.00] [USD]); token limits add
  "Count cache reads". Rows are checked on the spot with core's rules
  (a positive amount, no two limits with the same period, measure and
  cache-read setting, and a monthly limit needs records kept for 31
  days), so a wrong limit is caught before Save. A blank amount is only
  flagged once the field is left, so a newly added row is not red. Cost
  is typed in dollars and sent as micros.
- For an existing key each row shows what is used ("Today $1.23 /
  $5.00 · resets at 00:00", "Last minute 12 / 30") and marks a reached
  limit. With a cost limit, the models this key can use that have no
  price are listed (folded after six): they count as $0, so the limit
  does not hold them back.
- Keys table: a "Limit reached" chip next to the disabled state; hover
  names the limit and when it resets. The list is fetched again on
  key_limit_alert, at the next period reset and after the clock jumps,
  so the chip does not linger after midnight on an idle key.
- Disabling or enabling a key from the table now sends its limits back:
  the key input replaces the whole list, so leaving them out would have
  deleted every limit.
- Notifications: KeyLimitAlert becomes a system notification through the
  notice bus, naming the key, the limit and when it resets. One notice
  per limit: reaching the limit replaces the 80% notice, and the next
  period's 80% notice replaces yesterday's "reached". Each period counts
  as a new event, so a notice read yesterday does not silence today's.
  Clicking it opens the Keys page.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…art switch and slot wait

Core's routing round adds three things an upstream owner sets by hand and
two failover settings; this makes all of them reachable without editing
config.yaml.

Upstream dialog: an optional "Concurrency limit" (max_concurrent, 1-1000,
blank shows "No limit"). It sits under the proxy fields in the Connection
section and in the Account section of ChatGPT accounts, since accounts limit
concurrent requests too. It is checked on the spot and blocks saving like the
other required fields. The form now carries max_concurrent both ways, so the
table's enable/disable toggle, which saves the upstream from its view, keeps
the limit instead of dropping it. The row's in-flight tooltip names the
limit when one is set.

Models popover: hovering a model offers "Specs..." next to "Add alias...";
it opens a small dialog for the context window and max output of that model
on that upstream, saved through PUT /provider-model-spec with the version
the dialog opened on. A blank field shows the price table's value (or that
the table has none) as its placeholder, so blank never means something
hidden; clearing both removes the hand-set entry, and the footer says so.
Models with a hand-set value carry a "Manual specs" mark next to their name
(kept on the name side so it stays visible while the row's right side gives
way to the buttons); its tooltip lists which values are manual. The context
window column of the edit dialog's Models section now prefers the hand-set
value too, so the two places never show different numbers for one model.

Settings > Failover: "Move to the next upstream when the start times out"
is a switch hung under the stream-start wait, since it says what happens
when that wait runs out; "Wait for a free slot at most" (slot_wait_secs,
0-300, default 30 shown) follows. With the switch on the wait must be 5 to
120 seconds, checked on the spot like the section's other ranges, including
when only the switch changed; core's config.slow_start_too_short still
surfaces in the section's error banner if the file changed underneath.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…equests

Core measures the speed sample behind url-test and load-balance by speed from sending a
hop to the first content of the answer, not to the response headers. The dry run and the
url-test description called it TTFB / first byte, the word the request detail keeps for
the headers; they now say first token, as the request detail, traffic and upstream table
already do for that measurement. The terminology table records both terms.

Weights share requests, not new conversations: a turn that stays on an upstream is
charged to it. The load-balance and distribute-by descriptions now talk about requests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Core's routing round is final at adadbf5 (integ/routing-2026-10, protocol 39). The six
core crates move to that rev; the lock file changes only in their source lines.
tw-api.ts is regenerated: the protocol constant and doc comments changed, no field did.

Core's checks changed, so the UI follows them:

- Key limits: a cost limit has to be at least $0.01, and day, week and month limits need
  request records kept for 1, 7 and 31 days (config.key_limit_cost_too_small,
  config.key_limit_retention, which replaces the month-only code). The key dialog checks
  both on the spot, in core's order, and names the period in the retention line.
- New codes translated: config.key_limit_retention, config.key_limit_cost_too_small,
  gw.ws.upstream_closed; the retired config.key_limit_month_retention is removed.
- failover.slot_wait_secs is now one total wait per request, for a full upstream and for
  a key's minute or hour limit alike, so the setting no longer says it is only about
  concurrency.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The description that now covers both waits (a full upstream, a key's minute or hour
limit) wrapped to two lines where every other failover setting takes one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ion commit)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@fylorn
fylorn merged commit effd9bc into dev Oct 5, 2026
4 checks passed
@fylorn
fylorn deleted the ui-routing-integ branch October 5, 2026 14:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant