Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs/gentle-shell.md
Original file line number Diff line number Diff line change
Expand Up @@ -116,7 +116,7 @@ The panel rows a provider reports its windows with:
glm5.3-flash ▰▰▱▱▱▱▱▱▱▱▱▱▱▱▱▱ 11% · resets in 12d 17h
```

- For Codex, usage comes from the same account usage endpoint the Codex CLI reads, using the OAuth token pi already holds. It is fetched at session start, at most every 5 minutes after a turn, and on `r` in the panel. Rate-limit headers on SSE responses are picked up too. The same refresh covers every provider the session targets: the active model's own provider plus every provider the active profile's subagent routing names — a repository pin decides which profile that is, falling back to the global active profile when no pin applies. Each provider keeps its own 5-minute window, its own stale-source guard, and its own last good snapshot. Providers refresh concurrently, each inside its own bounded window (default 10s, `GENTLE_PI_SHELL_USAGE_TIMEOUT_MS`): the abort signal reaches the underlying fetch — composed with the caller's own signal when it carries one, never replacing it — a provider that outlives its window wears the generic failure note, and whatever it answers afterwards is discarded — a late answer never replaces what the timeout settled, exactly like the stale-source guard above.
- For Codex, usage comes from the same account usage endpoint the Codex CLI reads, using the OAuth token pi already holds. It is fetched at session start, at most every 5 minutes after a turn, and on `r` in the panel. Rate-limit headers on SSE responses are picked up too. Codex windows require a finite positive duration reported by the provider; missing or invalid durations are ignored, never shown as `0m` or replaced with an assumed `5h` quota. A real `0%` remains visible when its window is valid, and a response with no valid header windows leaves the previous snapshot untouched. The same refresh covers every provider the session targets: the active model's own provider plus every provider the active profile's subagent routing names — a repository pin decides which profile that is, falling back to the global active profile when no pin applies. Each provider keeps its own 5-minute window, its own stale-source guard, and its own last good snapshot. Providers refresh concurrently, each inside its own bounded window (default 10s, `GENTLE_PI_SHELL_USAGE_TIMEOUT_MS`): the abort signal reaches the underlying fetch — composed with the caller's own signal when it carries one, never replacing it — a provider that outlives its window wears the generic failure note, and whatever it answers afterwards is discarded — a late answer never replaces what the timeout settled, exactly like the stale-source guard above.
- A routing entry names its provider with a qualified ref (`provider/model`); a bare model id is resolved through the model registry only when exactly one provider carries that id, and is left untargeted rather than guessed when none or several do. The targeted scope is resolved when a refresh runs — at session start, on each turn's throttled refresh, on `r` or reopening the panel, and when a usage source registers — so a profile switch is picked up by the next refresh rather than live per render.
- For Claude Pro/Max, usage arrives in the rate-limit headers of every response, so the 5h and weekly windows appear after the first turn.
- For NaN Cloud, usage comes from the quota endpoint the official dashboard reads, with the same API key pi already holds. Each metered model reports one allowance for the billing period, and that window carries no label: the model id names it in the bar and the reset text says what it is in the panel. A model that also reports a rolling window shows that one labeled next to it (`4h`), which today's payload does not send; percentages are tokens used over the allowance, exactly as the dashboard draws them, and the allowance is the full-period cap (`fullCap`) whenever the model reports a positive one, because `cap` alone is the prorated allowance of the period in progress. It is fetched under the same 5-minute rule as Codex, counted per provider so a switch fetches the provider it switched to, refuses redirects so the bearer cannot be replayed to another origin, and keeps no cached copy. The endpoint sits outside NaN's published OpenAPI, so the parser reads it defensively: a model that reports no allowance is skipped, as the dashboard skips it, while a metered model whose usage cannot be read fails the whole read, so a partial payload never replaces a complete snapshot with a cheaper-looking one. A session that already has a snapshot keeps the last valid one through a malformed payload or a failed fetch, and the pending note appears only while there is nothing to draw.
Expand Down
7 changes: 4 additions & 3 deletions lib/shell-usage.ts
Original file line number Diff line number Diff line change
Expand Up @@ -269,7 +269,7 @@ export function formatReset(resetAt: number | null, now: number): string {
}

function parseWindow(raw: RawWindow | null | undefined, now: number): UsageWindow | undefined {
if (!raw || typeof raw.used_percent !== "number" || typeof raw.limit_window_seconds !== "number") return undefined;
if (!raw || !isFiniteNumber(raw.used_percent) || !isFiniteNumber(raw.limit_window_seconds) || raw.limit_window_seconds <= 0) return undefined;
const resetAt =
typeof raw.reset_at === "number" ? raw.reset_at * 1000 : typeof raw.reset_after_seconds === "number" ? now + raw.reset_after_seconds * 1000 : null;
return { label: windowLabel(raw.limit_window_seconds), usedPercent: raw.used_percent, windowSeconds: raw.limit_window_seconds, resetAt };
Expand Down Expand Up @@ -297,9 +297,10 @@ export function parseCodexUsage(payload: unknown, now: number): ProviderUsage {
function headerWindow(headers: Record<string, string>, kind: "primary" | "secondary"): UsageWindow | undefined {
const used = Number.parseFloat(headers[`${HEADER_PREFIX}${kind}-used-percent`] ?? "");
if (!Number.isFinite(used)) return undefined;
const minutes = Number.parseInt(headers[`${HEADER_PREFIX}${kind}-window-minutes`] ?? "", 10);
// Missing or invalid duration is not a zero-minute quota or an assumed 5h window.
const seconds = Number(headers[`${HEADER_PREFIX}${kind}-window-minutes`]) * MINUTE;
if (!Number.isFinite(seconds) || seconds <= 0) return undefined;
const resetAt = Number.parseInt(headers[`${HEADER_PREFIX}${kind}-reset-at`] ?? "", 10);
const seconds = Number.isFinite(minutes) ? minutes * MINUTE : 0;
return { label: windowLabel(seconds), usedPercent: used, windowSeconds: seconds, resetAt: Number.isFinite(resetAt) ? resetAt * 1000 : null };
}

Expand Down
22 changes: 22 additions & 0 deletions tests/gentle-shell.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -5192,6 +5192,28 @@ test("gentleShell records SSE rate-limit headers from provider responses", async
assert.doesNotMatch(renderFooter(ui), /codex/);
});

test("gentleShell ignores incomplete Codex headers and retains the last real quota without assuming 5h", async () => {
const { pi, handlers } = fakePi();
gentleShell(pi, { GENTLE_PI_SHELL_CHANGES_WATCH_MS: "off" }, { fetch: fakeFetch({}, false).fetchFn, now: () => 0 });
const { ctx, ui } = fakeContext();
await fire(handlers, "session_start", ctx);
const respond = (headers: Record<string, string>) => {
for (const handler of handlers.get("after_provider_response") ?? []) handler({ status: 200, headers }, ctx);
};
respond({ "x-codex-primary-used-percent": "0" });
assert.doesNotMatch(renderFooter(ui), /codex|0m/, "missing duration is not a real quota");
respond({ "x-codex-secondary-used-percent": "31", "x-codex-secondary-window-minutes": "10080" });
const valid = renderFooter(ui);
assert.match(valid, /codex week .*31%/);
assert.doesNotMatch(valid, /5h|0m/);
for (const minutes of [undefined, "0", "invalid"]) {
const headers: Record<string, string> = { "x-codex-primary-used-percent": "0" };
if (minutes !== undefined) headers["x-codex-primary-window-minutes"] = minutes;
respond(headers);
assert.equal(renderFooter(ui), valid, "a malformed response must not replace the known weekly quota");
}
});

test("gentleShell registers /gentle:usage and opens the subscriptions overlay", async () => {
const { pi, handlers, commands } = fakePi();
gentleShell(pi, { GENTLE_PI_SHELL_CHANGES_WATCH_MS: "off" }, { fetch: fakeFetch().fetchFn, now: () => 1_788_600_000_000 });
Expand Down
35 changes: 35 additions & 0 deletions tests/shell-usage.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -122,6 +122,41 @@ test("parseCodexHeaders reads the SSE rate-limit headers when a provider sends t
assert.equal(parseCodexHeaders({ "content-type": "text/event-stream" }, NOW), undefined);
});

test("parseCodexHeaders ignores windows without a valid duration instead of inventing 0m", () => {
for (const minutes of [undefined, "", " ", "0", "-1", "invalid", "300oops", "Infinity", "1e309"]) {
const headers: Record<string, string> = { "x-codex-primary-used-percent": "0" };
if (minutes !== undefined) headers["x-codex-primary-window-minutes"] = minutes;
assert.equal(parseCodexHeaders(headers, NOW), undefined, `minutes: ${String(minutes)}`);
}
});

test("parseCodexHeaders keeps genuine zero usage and only the valid provider windows", () => {
const usage = parseCodexHeaders({
"x-codex-primary-used-percent": "0",
"x-codex-primary-window-minutes": "0",
"x-codex-secondary-used-percent": "0",
"x-codex-secondary-window-minutes": "10080",
}, NOW);
assert.ok(usage);
assert.deepEqual(usage.limits[0].windows.map((w) => `${w.label}:${w.usedPercent}`), ["week:0"]);
assert.equal(renderUsageBar(usage, plainTheme), "codex week ▱▱▱▱▱▱▱▱ 0%");
assert.doesNotMatch(renderUsageBar(usage, plainTheme)!, /0m|5h/);
});

test("parseCodexUsage filters invalid numeric windows without inventing a five-hour quota", () => {
for (const seconds of [undefined, 0, -1, Number.NaN, Number.POSITIVE_INFINITY, "18000"]) {
const usage = parseCodexUsage({ rate_limit: {
primary_window: { used_percent: 0, limit_window_seconds: seconds },
secondary_window: { used_percent: 0, limit_window_seconds: 604_800 },
} }, NOW);
assert.deepEqual(usage.limits[0].windows.map((w) => `${w.label}:${w.usedPercent}`), ["week:0"], `seconds: ${String(seconds)}`);
}
for (const used of [undefined, Number.NaN, Number.POSITIVE_INFINITY]) {
assert.deepEqual(parseCodexUsage({ rate_limit: { primary_window: { used_percent: used, limit_window_seconds: 604_800 } } }, NOW).limits, []);
}
assert.deepEqual(parseCodexUsage({ rate_limit: { primary_window: { used_percent: 0, limit_window_seconds: 0 } } }, NOW).limits, []);
});

test("accountIdFromToken decodes the chatgpt account claim from an OAuth JWT", () => {
const claims = Buffer.from(JSON.stringify({ "https://api.openai.com/auth": { chatgpt_account_id: "acct-123" } })).toString("base64url");
assert.equal(accountIdFromToken(`header.${claims}.sig`), "acct-123");
Expand Down
Loading