Skip to content

[Feature] Custom base URL for the realtime provider + server-side VAD for the turn-based pipeline (factory firmware never sends listen:stop) #14

Description

@woshiyig

[Feature] Add openai_base_url and server-side VAD for the openai_compatible pipeline

Chinese abstract:出厂固件以 auto(服务端 VAD)模式工作,从不发送 listen:stop,导致 openai_compatible 回合制管线永远收不到音频、会话在 2–15 秒后以 WS 1006 异常关闭。请求:①给 openai 实时通道加 openai_base_url(可接 GLM-Realtime 等 OpenAI 兼容实时端点);②给 openai_compatible 管线加服务端 VAD(修好后 Hermes / MiniMax / DeepSeek 等任意 LLM 均可用)。


Summary

Two related capability gaps prevent using self-hosted / third-party LLMs with the M5Stack StackChan factory firmware:

  1. openai realtime provider has no configurable base URL. It is hard-wired to api.openai.com, so OpenAI-Realtime-compatible providers (e.g. Zhipu GLM-Realtime, which has a first-party realtime voice API) cannot be used.
  2. openai_compatible (turn-based) pipeline never receives audio from the factory firmware, because the firmware never emits listen:stop. The pipeline therefore never triggers STT.

Net effect: the only working configuration today is provider=gemini (Gemini Live). Any text-only LLM the user already pays for (Hermes agent, MiniMax-M3, DeepSeek, …) is unreachable.


Environment

Item Value
Server StackChan AI Server macOS standalone preview 0.1.1 (source rev 2b5b09de6e38daa7417b31866ba7437609fe58a8), run directly from the app bundle binary
Host macOS, STACKCHAN_LOCAL_HOST=192.168.1.36, device port 12800
Device M5Stack StackChan (SKU:K151), CoreS3, factory firmware 1.5.1 (app_init: App version: 1.5.1, Launcher banner 2.3.3, build Jul 31 2026)
Device MAC 44:1b:f6:e5:5c:20
Onboarding OTA URL injected into NVS (wifi namespace, ota_url = http://192.168.1.36:12800/xiaozhi/ota/) — device successfully fetches it and receives {"websocket":{"url":"ws://192.168.1.36:12800/xiaozhi/ws","version":3}}

What works

  • Device discovers the local server via the injected OTA URL ✅
  • WebSocket handshake succeeds every time: hello → hello OK (session id assigned) ✅
  • provider=gemini (gemini-2.5-flash-native-audio-latest) — full voice conversation works ✅
  • Custom STT/TTS adapters (local whisper + macOS say) both serve /v1/models, /v1/audio/transcriptions, /v1/audio/speech and are reachable from the server ✅

What fails

provider=openai_compatible, with LLM pointing at either a local Hermes agent (http://127.0.0.1:8642/v1, model hermes-agent) or MiniMax (https://api.minimaxi.com/v1, model MiniMax-M3), and STT/TTS pointing at the local adapters.

The LLM is never reached — in fact STT is never reached. Across 14 consecutive attempts the server log shows the same pattern:

[WS] device=44:1b:f6:e5:5c:20 realtime session ready provider=openai_compatible
[WS] device=44:1b:f6:e5:5c:20 recv type=hello
[WS] device=44:1b:f6:e5:5c:20 session=6d7810d4-… hello OK
[WS] device=44:1b:f6:e5:5c:20 recv type=listen state=start
[WS] device=44:1b:f6:e5:5c:20 listening started
      ↓ nothing else ↓
[WS] device=44:1b:f6:e5:5c:20 read error: websocket: close 1006 (abnormal closure): unexpected EOF
[WS] device=44:1b:f6:e5:5c:20 session closed

Event tally over all attempts (log level debug):

Event Count
recv type=hello 12
recv type=listen state=start 14
recv type=listen state=stop 0
audio frames observed 0 (opus/audio/frame/vad keyword hits = 0)
STT adapter invocations 0
TTS adapter invocations 0
Server-side errors none

Time from listening started to the 1006 close varied between ~2 s and ~14 s (also one clean server-side timeout: conversation idle; returning device to wake-word standby after ~45 s).

No error/warn/fail entries appear in the server log.

Not the cause (ruled out)

Hypothesis Verdict
Network / reachability ❌ ruled out — same device, same network, Gemini path works
Power supply ❌ ruled out — the device holds normal conversations with the stock Xiaozhi cloud on the same PSU
Bad API key / upstream model ❌ ruled out — MiniMax-M3 returns 200 for the exact same credentials/base URL used in the settings; llm_api_key reads back at full length
Missing /v1/models on my adapters ❌ ruled out — added it (now both return 200); failure unchanged
Wake word / microphone ❌ ruled out — green LED lights, wake word is recognized

Root-cause analysis

Xiaozhi protocol v3 has two listening modes:

Mode Who decides "user finished speaking"
manual Device — sends {"type":"listen","state":"stop"}
auto Server — server-side VAD detects end of speech

The factory firmware appears to run in auto mode: it opens with listen:start and then waits for the server to drive the turn. It never sends listen:stop.

But the openai_compatible pipeline is turn-based — per the README it is a STT → Chat Completions → TTS path, and it waits for the device's listen:stop before invoking STT.

Both sides wait for the other → deadlock → the device eventually tears down the socket (1006).

This is consistent with the fact that provider=gemini (and presumably provider=openai) work: those are realtime sessions with server-side VAD, so they never need listen:stop.

Note this is independent of which LLM is configured — swapping MiniMax-M3 for Hermes, DeepSeek, or anything else changes nothing, because no audio is ever handed to STT.


Requests

1. Add openai_base_url to the openai (realtime) provider

Currently:

$ curl -X PUT .../api/settings -d '{"openai_base_url":"https://open.bigmodel.cn/api/paas/v4"}'
HTTP 400  {"error":"unsupported setting"}   # (exact body wording may differ)

The realtime provider is hard-wired to api.openai.com. Making the base URL configurable would immediately unlock any OpenAI-Realtime-protocol-compatible realtime endpoint — for example Zhipu GLM-Realtime (https://docs.bigmodel.cn/cn/guide/models/sound-and-video/glm-realtime, SDK: https://github.com/MetaGLM/glm-realtime-sdk), which is a genuine realtime voice model and is substantially cheaper than OpenAI Realtime for CN users.

This is a small, low-risk change: one settings field, one base-URL constant.

2. Add server-side VAD to the openai_compatible pipeline

If the compatible pipeline performed (or accepted) server-side VAD — i.e. treated "N ms of silence after detected speech" as end-of-turn instead of requiring listen:stop — then any text LLM would work: local Hermes agents, MiniMax-M3, DeepSeek, OpenRouter models, etc.

This greatly widens the project's usefulness: users can keep their existing (already paid for) LLM subscriptions instead of buying realtime-audio tokens.

Possible implementation notes:

  • Reuse whatever VAD already exists for the openai realtime frontend.
  • If the device advertises its mode in hello (features / listen mode), branch on it: auto → server VAD; manual → current behaviour.
  • Expose an option such as compatible_server_vad: true / compatible_vad_silence_ms: 800 so users can opt in.
  • Consider logging received audio frame counts at debug level — right now it is impossible to tell from the log whether audio is arriving at all. (This alone cost me hours; see "Debugging note" below.)

Debugging note (suggestion)

Please consider adding a debug log line for binary/incoming audio frames (e.g. count + bytes per second). The current debug output shows only JSON control messages, so a "no audio at all" failure is indistinguishable from "audio arriving but never triggering a turn". A single counter line would have made this report much shorter.


Reproduce

# 1. Device points at the local server (NVS: wifi.ota_url)
curl -s http://192.168.1.36:12800/xiaozhi/ota/
# → {"websocket":{"token":"","url":"ws://192.168.1.36:12800/xiaozhi/ws","version":3}}

# 2. Configure the compatible pipeline
curl -X PUT http://127.0.0.1:8099/api/settings \
  -H "Authorization: Bearer $SETTINGS_TOKEN" -H 'Content-Type: application/json' \
  -d '{"provider":"openai_compatible",
       "llm_base_url":"https://api.minimaxi.com/v1","llm_api_key":"...","llm_model":"MiniMax-M3",
       "stt_base_url":"http://127.0.0.1:8801/v1","stt_api_key":"sk-local","stt_model":"whisper-1",
       "tts_base_url":"http://127.0.0.1:8802/v1","tts_api_key":"sk-local","tts_model":"tts-1","tts_voice":"Tingting"}'

# 3. Wake the device ("Hi, StackChan") and speak
# 4. Observe: listen state=start, then WS 1006 within 2–15 s; STT adapter never called
tail -f "~/Library/Application Support/StackChan AI Server/logs/$(date +%F).log"

Impact

Without either change, M5Stack StackChan factory-firmware users are limited to provider=gemini (or OpenAI Realtime). Every already-paid-for text LLM subscription — including locally hosted agents — is unusable with this server, which is arguably the project's main value proposition.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions