Skip to content

[Live API] gemini-3.8-live-extended-thinking: announces tool call then says "I apologize, but a system error occurred." with no functionCall frame - 61% of tool runs (n=64) #3015

Description

@jauharymanshura9-max

Summary

On models/gemini-3.8-live-extended-thinking, tool calls fail server-side in a way that is undetectable from the client: the model first speaks a normal turn committing to the call (e.g. "Let me update that file…"), then 1–3 seconds later says "I apologize, but a system error occurred." — and no functionCall frame is ever sent. The turn closes with a normal turnComplete and no APIError/close frame, so exception-based detection is impossible.

We have measured this over n=64 tool-scenario runs: 25 pass / 39 fail = 61% failure rate, across all thinking levels. A fully independent raw-WebSocket reproduction (no application code, minimal client) reports the identical signature: masudl-hub/theoremai#25

Environment

  • google-genai==2.25.0, Python 3.14, Live API over websocket
  • Model models/gemini-3.8-live-extended-thinking, thinking_level tried at high, low, medium
  • Tools declared as functionDeclarations with behavior: NON_BLOCKING (13 tools in our production session; failures reproduce with a single trivial tool as well)
  • Session also uses session_resumption, context_window_compression, media_resolution: MEDIUM — but the minimal reproduction in the thread above uses none of these and still fails

Reproduction steps

  1. Open a Live session on gemini-3.8-live-extended-thinking with one tool declared (behavior: NON_BLOCKING).
  2. Send a text/audio prompt that requires the tool (in our harness: "edit the file … replace the word done with finished").
  3. Observed: audio/content turn announces the action, then the apology text; toolCall/functionCall frame never arrives; turnComplete fires normally; no client-side error object.
  4. Expected: a functionCall frame so the client can dispatch and respond.

Repeat several times — the failure is intermittent within the same session config (some runs fully clean, others fail 3–4 times in a row).

Measured rates (our E2E harness, one row per run)

Config Pass / Fail
extended + tool, thinking_level=high 16 / 29
extended + tool, thinking_level=low 4 / 4
extended + tool, thinking_level=medium 1 / 2
extended + tool, unlabelled legacy runs 4 / 4
extended + tool, total 25 / 39 = 61% fail (n=64)
extended, chat only (no tools) 2 / 0
base gemini-3.8-live + same tools 2 / 2 (small-n; earlier A/B: 0/3 errors vs 2/3 for extended)

A fresh 10-run batch on 2026-09-28 was worse (3/10 pass), confirming high run-to-run variance rather than a fixed rate.

Additional contract observations

  • behavior: BLOCKING is rejected at session setup with close code 1007: "BLOCKING function calls are not supported for this model." (reported in the thread above), so NON_BLOCKING is the only path for this model.
  • Because there is no error signal, the only practical detection is text-signature based (the apology phrase). We mitigated by injecting a short text nudge after the model goes IDLE ("verify the actual result, then re-call the tool"): this recovered 4 of 6 detected failures in one batch, but it is a workaround, not a fix, and it cannot help clients that don't watch the transcript.
  • This looks related to Live API (gemini-3.1-flash-live-preview): model composes a function call in thinking but never emits it, speaks fabricated data instead #2827 ("model composes a function call in thinking but never emits it"), possibly the same underlying class surfaced differently (apology text vs. fabricated answer).

Ask

  1. Investigate why the extended-thinking Live model aborts tool-call emission server-side while the base gemini-3.8-live model on the same setup succeeds more often.
  2. Consider whether the failure should surface as an explicit error/event instead of plain assistant text, so clients can retry or fail over.

Happy to share redacted session logs (full websocket transcripts + per-run timing).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

priority: p2Moderately-important priority. Fix may not be included in next release.type: bugError or flaw in code with unintended results or allowing sub-optimal usage patterns.

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions