Skip to content

feat: local multimodal agent support via OpenAI-compatible providers (Gemma 4 E4B) - #176

Open
banggaoo wants to merge 23 commits into
google:mainfrom
banggaoo:feat/gemma4-upstream
Open

banggaoo wants to merge 23 commits into
google:mainfrom
banggaoo:feat/gemma4-upstream

Conversation

@banggaoo

@banggaoo banggaoo commented Oct 9, 2026 •

Copy link
Copy Markdown

Design & architecture: #177

Summary

Adds support for running ARTEMIS agents against a local OpenAI-compatible model server (tested with Gemma 4 E4B via MLX), enabling fully on-device multimodal agent execution.

  • Managed model lifecycle: artemis model pull/serve/status/stop/list — catalog aliases (gemma4-e4b) resolve to pinned HF repos; serve launches a managed mlx_vlm.server and applies the catalog's recommended vision soft-token budget (1120, ~1056×2352 effective detail vs the 280 default).
  • LLM routing: route agent LLM endpoints to local OpenAI-compatible providers; custom provider path with function_calling structured output; ready-made config/artemis.gemma4.jsonc.
  • Model readiness: local_model_probe gates runs on reachability + served-alias checks for custom endpoints; catalog aliases get artemis model pull/serve guidance, BYO servers get neutral tooling hints. Self-passes for cloud-only configs.
  • Console-guided setup: Local Model Endpoint card in the setup wizard with one-click Run buttons backed by loopback-guarded, alias-validated pull/serve/stop subprocess endpoints (status polling + log tail); copy-to-terminal fallback when the CLI isn't installed.
  • Image-first ordering: remaining text-first multimodal call sites normalized — screenshot precedes instruction text in planner, checker, explorer, flash runner, image processor, committee, outputter, history, and pixel-validation messages.
  • Completion affordance: mark_done tool gives the Pro/operator loop a structural completion claim (latched flag → convergence gate → settlement/final verification); false claims bounce back to execution. Plans whose top-level items are all resolved ([x] or [!] blocked) now route to settlement instead of looping.
  • Image wire format: image_data_uri sniffs PNG/JPEG/WebP magic bytes instead of hardcoding image/jpeg; strict local gateways reject mislabeled payloads.
  • Small-model prompt discipline: focus reminders keep tool-call protocol rules inside the sliding attention window; bounded injected context.
  • Flash reliability: prune intermediate screenshots before invocation to respect provider image-count limits; LLM_FLASH_HARD_TIMEOUT_SECONDS for slow local providers.
  • Flash tool-call hardening: tool_choice="required" on custom providers holds the one-tool-call-per-turn contract (eliminates narration-only loops on small local models); call:name{args} envelope repair recovers calls that mlx_vlm wraps into unparseable invalid_tool_calls or leaks into text content; a consecutive no-tool-call cap (8) bounds the loop when a provider still refuses to call tools.
  • OCR: Apple Vision OCR utility on macOS; optional/skipped elsewhere.
  • Object detection: coordinate handling and yx_1000 mode for Gemma-style outputs.

Test plan

  • pytest focused suites: object detector coords, local endpoint routing, optional OCR, operator focus reminder — pass
  • Full tests/unit/{agents,memory,tools} sweep: no new failures vs. base branch
  • Settlement/gate tests incl. mark_done claim routing + blocked-plan termination — pass
  • Local-model lifecycle endpoint tests (alias validation, loopback guard, task state/logs, port-scoped stop) — pass
  • Live end-to-end on Android emulator (API 36) with gemma4-e4b on a direct mlx_vlm.server: Flash tasks verified on-device — "Open Settings" x2 (click Settings icon → verified → report_task_status(completed)); complex multi-step "open Settings → Display & touch → Dark theme" completed in 7 turns with ui_night_mode=2 confirmed
  • Maintainer review

🤖 Generated with Devin

Wire api_base/max_tokens/coordinate_format through the LLM config
schema and endpoint resolver so per-role nodes can target local
gateways (e.g. ondevice-agent-platform serving Qwen models) while
Gemini defaults stay byte-identical.

- LLM schema gains api_base, max_tokens, coordinate_format (yx/xy x
  1000/px/norm); _resolve_endpoint forwards all three to ModelEndpoint
- Object detector honors the declared point convention instead of
  hardcoding Gemini ER's [y,x] swap; batch-level rescale guards
  responses that clearly arrive in 0-1 or raw pixels, including the
  Qwen smart-resize grid for pixel-space output
- get_llm(object_detector) now resolves the utils scope, so the
  configured detector endpoint is actually used (was silently falling
  back to the operator model)
- strip_json_comments preserves // and /* */ inside string literals,
  which made URL api_base values unparseable
- Empty exported OPENAI_API_KEY falls back to the "EMPTY" placeholder
  instead of reaching ChatOpenAI as a blank credential
- resources config: document the ondevice-qwen preset and the
  coordinate_format contract for local vision models

Verified live against ondevice-agent-platform (:8080): qwen3.8-9b
serves the planner/text route and qwen-vl grounds UI elements on a
real iOS screenshot within ~3% of the 0-1000 grid.
Accept Gemma's trained box_2d output ([ymin,xmin,ymax,xmax] on the
0-1000 grid) in the object detector: normalize it to the existing
[x1,y1,x2,y2] contract, synthesize the point from the box center when
no usable point exists, and include box coordinates in batch-scale
inference. Add a spike config routing every node to gemma-4-e4b on the
local OpenAI-compatible endpoint, alongside the untouched qwen config.

Also add an on-device Apple Vision OCR provider behind perform_ocr:
a standalone worker subprocess (Vision calls fail intermittently once
cv2 is resident) with a per-call timeout, one retry, and a circuit
breaker that disables the provider after consecutive failures. Apple
Vision is preferred on macOS; Google Cloud Vision remains the keyed
fallback and ARTEMIS_OCR_PROVIDER pins a provider.
Adds a LocalModelEndpointProbe (artemis doctor / console wizard /
mobile_diagnose): walks the active LLM config for custom-provider
endpoints, verifies each required model alias is served via
GET {api_base}/models, and surfaces guided remediation actions -
`ondevice-agent-platform model pull --alias <model>` when an alias is
missing, `ondevice-agent-platform serve` when the endpoint is down.

LLMCredentialsProbe now treats config-bound custom endpoints as a
satisfied credential instead of reporting a missing API key.
… OCR

- default -> gemma4-e4b (custom provider, :8080/v1, 300s timeout),
  fallback -> qwen3.8-9b on the same endpoint
- operator/object_detector/video_analyzer/explorer primaries ->
  vision-hybrid: Apple Vision answers confident OCR turns natively;
  everything else escalates to qwen-vl (kept as each node's fallback)

Includes previously-uncommitted local-provider bindings from the
gemma4 default switch; no behavior change for semantic turns.
operator/object_detector/video_analyzer/explorer primaries now use
gemma4-e4b (semantic/grounding); qwen-vl stays each node's fallback.
vision-hybrid is the OCR-only route per user direction. Note: gemma's
0-1000 grid output verified only on synthetic input - qwen-vl remains
the verified grounding fallback.
…R=platform)

Adds a 'platform' OCR provider calling the local platform's vision-hybrid
route; structured observations arrive via the additive oap_ocr field
(text + pixel vertices matching the boundingPoly contract). In 'auto' it
sits between the native Apple Vision worker and Google Cloud Vision.
Verified live: TOTAL extraction with correct normalized coordinates.
The standalone PyObjC Apple Vision worker is deleted - the platform's
vision-hybrid route (oap-vision-bridge) is now the primary OCR provider
in 'auto', with keyed Google Vision as the fallback. The 'apple'
provider pin is a deprecated alias for 'platform'. Structured
observations arrive via the additive oap_ocr field (pixel vertices,
boundingPoly order). Verified live: TOTAL extraction at [446,475]
through run_ocr_core under default auto.
The platform exposes the vision-hybrid route (oap_ocr structured
observations); whether ARTEMIS calls it is ARTEMIS's own choice. The
original apple_vision_ocr.py worker and provider chain are restored.
Gemma's native box_2d is [ymin, xmin, ymax, xmax]/1000 — the detector's
yx_1000 contract. The previous xy_1000 (correct for the qwen-vl
fallback) axis-flipped every box into garbage; verified with a real
model response normalized through _normalize_detected_box.
- Sniff image MIME from bytes (image_data_uri) instead of hardcoding
  image/jpeg; strict local gateways reject mislabeled PNG screenshots.
- with_structured_output uses function_calling for the custom provider.
- Flash: prune intermediate screenshots before invoke (platform caps
  images-per-request).
- Planner/operator/checker focus reminders keep protocol rules inside
  the sliding attention window for small local models.
- LLM_FLASH_HARD_TIMEOUT_SECONDS for slow local providers.
@google-cla

google-cla Bot commented Oct 9, 2026

Copy link
Copy Markdown

Thanks for your pull request! It looks like this may be your first contribution to a Google open source project. Before we can look at your pull request, you'll need to sign a Contributor License Agreement (CLA).

View this failed invocation of the CLA check for more information.

For the most up to date status, view the checks section at the bottom of the pull request.

@banggaoo
banggaoo force-pushed the feat/gemma4-upstream branch from fa22b22 to 37017be Compare October 10, 2026 03:50
- artemis model pull/serve/stop/status/list manages mlx_vlm.server as a
  background process (state under ~/.artemis/local_models/)
- config/local_models.py: alias catalog (gemma4-e4b -> HF repo)
- readiness probe now suggests artemis model pull/serve <alias> for
  catalog models, generic pull/serve hints otherwise
- pip extra local (mlx-vlm, huggingface-hub; Apple Silicon gated)
- docs/local_models.md leads with the self-operated path; BYO server
  remains for other platforms
@banggaoo
banggaoo force-pushed the feat/gemma4-upstream branch from 37017be to 2c39050 Compare October 10, 2026 04:31
Gemma 4 image processor defaults to 280 soft tokens per image, which
downsamples a phone screenshot to ~528x1152 effective and produced
measured grounding failures (wrong thumbnail, echoed coordinates,
lookalike-anchored detections of absent elements). At 1120 tokens the
same suite goes 8/8: dead-on taps and correct empty-list refusals.

- LocalModelSpec.vision_soft_tokens + catalog default 1120 for e4b
- artemis model serve --vision-tokens {70,140,280,560,1120} patches the
  pulled snapshot processor_config.json before spawning mlx_vlm.server
  (idempotent; graceful noop when the snapshot or key is absent)
- docs/local_models.md: vision-detail tuning note
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant