[Feature] Add openai_base_url and server-side VAD for the openai_compatible pipeline
Chinese abstract:出厂固件以 auto(服务端 VAD)模式工作,从不发送 listen:stop,导致 openai_compatible 回合制管线永远收不到音频、会话在 2–15 秒后以 WS 1006 异常关闭。请求:①给 openai 实时通道加 openai_base_url(可接 GLM-Realtime 等 OpenAI 兼容实时端点);②给 openai_compatible 管线加服务端 VAD(修好后 Hermes / MiniMax / DeepSeek 等任意 LLM 均可用)。
Summary
Two related capability gaps prevent using self-hosted / third-party LLMs with the M5Stack StackChan factory firmware:
openai realtime provider has no configurable base URL. It is hard-wired to api.openai.com, so OpenAI-Realtime-compatible providers (e.g. Zhipu GLM-Realtime, which has a first-party realtime voice API) cannot be used.
openai_compatible (turn-based) pipeline never receives audio from the factory firmware, because the firmware never emits listen:stop. The pipeline therefore never triggers STT.
Net effect: the only working configuration today is provider=gemini (Gemini Live). Any text-only LLM the user already pays for (Hermes agent, MiniMax-M3, DeepSeek, …) is unreachable.
Environment
| Item |
Value |
| Server |
StackChan AI Server macOS standalone preview 0.1.1 (source rev 2b5b09de6e38daa7417b31866ba7437609fe58a8), run directly from the app bundle binary |
| Host |
macOS, STACKCHAN_LOCAL_HOST=192.168.1.36, device port 12800 |
| Device |
M5Stack StackChan (SKU:K151), CoreS3, factory firmware 1.5.1 (app_init: App version: 1.5.1, Launcher banner 2.3.3, build Jul 31 2026) |
| Device MAC |
44:1b:f6:e5:5c:20 |
| Onboarding |
OTA URL injected into NVS (wifi namespace, ota_url = http://192.168.1.36:12800/xiaozhi/ota/) — device successfully fetches it and receives {"websocket":{"url":"ws://192.168.1.36:12800/xiaozhi/ws","version":3}} |
What works
- Device discovers the local server via the injected OTA URL ✅
- WebSocket handshake succeeds every time:
hello → hello OK (session id assigned) ✅
provider=gemini (gemini-2.5-flash-native-audio-latest) — full voice conversation works ✅
- Custom STT/TTS adapters (local
whisper + macOS say) both serve /v1/models, /v1/audio/transcriptions, /v1/audio/speech and are reachable from the server ✅
What fails
provider=openai_compatible, with LLM pointing at either a local Hermes agent (http://127.0.0.1:8642/v1, model hermes-agent) or MiniMax (https://api.minimaxi.com/v1, model MiniMax-M3), and STT/TTS pointing at the local adapters.
The LLM is never reached — in fact STT is never reached. Across 14 consecutive attempts the server log shows the same pattern:
[WS] device=44:1b:f6:e5:5c:20 realtime session ready provider=openai_compatible
[WS] device=44:1b:f6:e5:5c:20 recv type=hello
[WS] device=44:1b:f6:e5:5c:20 session=6d7810d4-… hello OK
[WS] device=44:1b:f6:e5:5c:20 recv type=listen state=start
[WS] device=44:1b:f6:e5:5c:20 listening started
↓ nothing else ↓
[WS] device=44:1b:f6:e5:5c:20 read error: websocket: close 1006 (abnormal closure): unexpected EOF
[WS] device=44:1b:f6:e5:5c:20 session closed
Event tally over all attempts (log level debug):
| Event |
Count |
recv type=hello |
12 |
recv type=listen state=start |
14 |
recv type=listen state=stop |
0 |
| audio frames observed |
0 (opus/audio/frame/vad keyword hits = 0) |
| STT adapter invocations |
0 |
| TTS adapter invocations |
0 |
| Server-side errors |
none |
Time from listening started to the 1006 close varied between ~2 s and ~14 s (also one clean server-side timeout: conversation idle; returning device to wake-word standby after ~45 s).
No error/warn/fail entries appear in the server log.
Not the cause (ruled out)
| Hypothesis |
Verdict |
| Network / reachability |
❌ ruled out — same device, same network, Gemini path works |
| Power supply |
❌ ruled out — the device holds normal conversations with the stock Xiaozhi cloud on the same PSU |
| Bad API key / upstream model |
❌ ruled out — MiniMax-M3 returns 200 for the exact same credentials/base URL used in the settings; llm_api_key reads back at full length |
Missing /v1/models on my adapters |
❌ ruled out — added it (now both return 200); failure unchanged |
| Wake word / microphone |
❌ ruled out — green LED lights, wake word is recognized |
Root-cause analysis
Xiaozhi protocol v3 has two listening modes:
| Mode |
Who decides "user finished speaking" |
manual |
Device — sends {"type":"listen","state":"stop"} |
auto |
Server — server-side VAD detects end of speech |
The factory firmware appears to run in auto mode: it opens with listen:start and then waits for the server to drive the turn. It never sends listen:stop.
But the openai_compatible pipeline is turn-based — per the README it is a STT → Chat Completions → TTS path, and it waits for the device's listen:stop before invoking STT.
Both sides wait for the other → deadlock → the device eventually tears down the socket (1006).
This is consistent with the fact that provider=gemini (and presumably provider=openai) work: those are realtime sessions with server-side VAD, so they never need listen:stop.
Note this is independent of which LLM is configured — swapping MiniMax-M3 for Hermes, DeepSeek, or anything else changes nothing, because no audio is ever handed to STT.
Requests
1. Add openai_base_url to the openai (realtime) provider
Currently:
$ curl -X PUT .../api/settings -d '{"openai_base_url":"https://open.bigmodel.cn/api/paas/v4"}'
HTTP 400 {"error":"unsupported setting"} # (exact body wording may differ)
The realtime provider is hard-wired to api.openai.com. Making the base URL configurable would immediately unlock any OpenAI-Realtime-protocol-compatible realtime endpoint — for example Zhipu GLM-Realtime (https://docs.bigmodel.cn/cn/guide/models/sound-and-video/glm-realtime, SDK: https://github.com/MetaGLM/glm-realtime-sdk), which is a genuine realtime voice model and is substantially cheaper than OpenAI Realtime for CN users.
This is a small, low-risk change: one settings field, one base-URL constant.
2. Add server-side VAD to the openai_compatible pipeline
If the compatible pipeline performed (or accepted) server-side VAD — i.e. treated "N ms of silence after detected speech" as end-of-turn instead of requiring listen:stop — then any text LLM would work: local Hermes agents, MiniMax-M3, DeepSeek, OpenRouter models, etc.
This greatly widens the project's usefulness: users can keep their existing (already paid for) LLM subscriptions instead of buying realtime-audio tokens.
Possible implementation notes:
- Reuse whatever VAD already exists for the
openai realtime frontend.
- If the device advertises its mode in
hello (features / listen mode), branch on it: auto → server VAD; manual → current behaviour.
- Expose an option such as
compatible_server_vad: true / compatible_vad_silence_ms: 800 so users can opt in.
- Consider logging received audio frame counts at
debug level — right now it is impossible to tell from the log whether audio is arriving at all. (This alone cost me hours; see "Debugging note" below.)
Debugging note (suggestion)
Please consider adding a debug log line for binary/incoming audio frames (e.g. count + bytes per second). The current debug output shows only JSON control messages, so a "no audio at all" failure is indistinguishable from "audio arriving but never triggering a turn". A single counter line would have made this report much shorter.
Reproduce
# 1. Device points at the local server (NVS: wifi.ota_url)
curl -s http://192.168.1.36:12800/xiaozhi/ota/
# → {"websocket":{"token":"","url":"ws://192.168.1.36:12800/xiaozhi/ws","version":3}}
# 2. Configure the compatible pipeline
curl -X PUT http://127.0.0.1:8099/api/settings \
-H "Authorization: Bearer $SETTINGS_TOKEN" -H 'Content-Type: application/json' \
-d '{"provider":"openai_compatible",
"llm_base_url":"https://api.minimaxi.com/v1","llm_api_key":"...","llm_model":"MiniMax-M3",
"stt_base_url":"http://127.0.0.1:8801/v1","stt_api_key":"sk-local","stt_model":"whisper-1",
"tts_base_url":"http://127.0.0.1:8802/v1","tts_api_key":"sk-local","tts_model":"tts-1","tts_voice":"Tingting"}'
# 3. Wake the device ("Hi, StackChan") and speak
# 4. Observe: listen state=start, then WS 1006 within 2–15 s; STT adapter never called
tail -f "~/Library/Application Support/StackChan AI Server/logs/$(date +%F).log"
Impact
Without either change, M5Stack StackChan factory-firmware users are limited to provider=gemini (or OpenAI Realtime). Every already-paid-for text LLM subscription — including locally hosted agents — is unusable with this server, which is arguably the project's main value proposition.
[Feature] Add
openai_base_urland server-side VAD for theopenai_compatiblepipelineSummary
Two related capability gaps prevent using self-hosted / third-party LLMs with the M5Stack StackChan factory firmware:
openairealtime provider has no configurable base URL. It is hard-wired toapi.openai.com, so OpenAI-Realtime-compatible providers (e.g. Zhipu GLM-Realtime, which has a first-party realtime voice API) cannot be used.openai_compatible(turn-based) pipeline never receives audio from the factory firmware, because the firmware never emitslisten:stop. The pipeline therefore never triggers STT.Net effect: the only working configuration today is
provider=gemini(Gemini Live). Any text-only LLM the user already pays for (Hermes agent, MiniMax-M3, DeepSeek, …) is unreachable.Environment
2b5b09de6e38daa7417b31866ba7437609fe58a8), run directly from the app bundle binarySTACKCHAN_LOCAL_HOST=192.168.1.36, device port128001.5.1(app_init:App version: 1.5.1, Launcher banner2.3.3, buildJul 31 2026)44:1b:f6:e5:5c:20wifinamespace,ota_url = http://192.168.1.36:12800/xiaozhi/ota/) — device successfully fetches it and receives{"websocket":{"url":"ws://192.168.1.36:12800/xiaozhi/ws","version":3}}What works
hello→hello OK(session id assigned) ✅provider=gemini(gemini-2.5-flash-native-audio-latest) — full voice conversation works ✅whisper+ macOSsay) both serve/v1/models,/v1/audio/transcriptions,/v1/audio/speechand are reachable from the server ✅What fails
provider=openai_compatible, with LLM pointing at either a local Hermes agent (http://127.0.0.1:8642/v1, modelhermes-agent) or MiniMax (https://api.minimaxi.com/v1, modelMiniMax-M3), and STT/TTS pointing at the local adapters.The LLM is never reached — in fact STT is never reached. Across 14 consecutive attempts the server log shows the same pattern:
Event tally over all attempts (log level
debug):recv type=hellorecv type=listen state=startrecv type=listen state=stopopus/audio/frame/vadkeyword hits = 0)Time from
listening startedto the 1006 close varied between ~2 s and ~14 s (also one clean server-side timeout:conversation idle; returning device to wake-word standbyafter ~45 s).No
error/warn/failentries appear in the server log.Not the cause (ruled out)
llm_api_keyreads back at full length/v1/modelson my adaptersRoot-cause analysis
Xiaozhi protocol v3 has two listening modes:
manual{"type":"listen","state":"stop"}autoThe factory firmware appears to run in
automode: it opens withlisten:startand then waits for the server to drive the turn. It never sendslisten:stop.But the
openai_compatiblepipeline is turn-based — per the README it is aSTT → Chat Completions → TTSpath, and it waits for the device'slisten:stopbefore invoking STT.Both sides wait for the other → deadlock → the device eventually tears down the socket (1006).
This is consistent with the fact that
provider=gemini(and presumablyprovider=openai) work: those are realtime sessions with server-side VAD, so they never needlisten:stop.Note this is independent of which LLM is configured — swapping MiniMax-M3 for Hermes, DeepSeek, or anything else changes nothing, because no audio is ever handed to STT.
Requests
1. Add
openai_base_urlto theopenai(realtime) providerCurrently:
The realtime provider is hard-wired to
api.openai.com. Making the base URL configurable would immediately unlock any OpenAI-Realtime-protocol-compatible realtime endpoint — for example Zhipu GLM-Realtime (https://docs.bigmodel.cn/cn/guide/models/sound-and-video/glm-realtime, SDK: https://github.com/MetaGLM/glm-realtime-sdk), which is a genuine realtime voice model and is substantially cheaper than OpenAI Realtime for CN users.This is a small, low-risk change: one settings field, one base-URL constant.
2. Add server-side VAD to the
openai_compatiblepipelineIf the compatible pipeline performed (or accepted) server-side VAD — i.e. treated "N ms of silence after detected speech" as end-of-turn instead of requiring
listen:stop— then any text LLM would work: local Hermes agents, MiniMax-M3, DeepSeek, OpenRouter models, etc.This greatly widens the project's usefulness: users can keep their existing (already paid for) LLM subscriptions instead of buying realtime-audio tokens.
Possible implementation notes:
openairealtime frontend.hello(features/listen mode), branch on it:auto→ server VAD;manual→ current behaviour.compatible_server_vad: true/compatible_vad_silence_ms: 800so users can opt in.debuglevel — right now it is impossible to tell from the log whether audio is arriving at all. (This alone cost me hours; see "Debugging note" below.)Debugging note (suggestion)
Please consider adding a
debuglog line for binary/incoming audio frames (e.g. count + bytes per second). The currentdebugoutput shows only JSON control messages, so a "no audio at all" failure is indistinguishable from "audio arriving but never triggering a turn". A single counter line would have made this report much shorter.Reproduce
Impact
Without either change, M5Stack StackChan factory-firmware users are limited to
provider=gemini(or OpenAI Realtime). Every already-paid-for text LLM subscription — including locally hosted agents — is unusable with this server, which is arguably the project's main value proposition.