Repository navigation
Conversation
Wire api_base/max_tokens/coordinate_format through the LLM config schema and endpoint resolver so per-role nodes can target local gateways (e.g. ondevice-agent-platform serving Qwen models) while Gemini defaults stay byte-identical. - LLM schema gains api_base, max_tokens, coordinate_format (yx/xy x 1000/px/norm); _resolve_endpoint forwards all three to ModelEndpoint - Object detector honors the declared point convention instead of hardcoding Gemini ER's [y,x] swap; batch-level rescale guards responses that clearly arrive in 0-1 or raw pixels, including the Qwen smart-resize grid for pixel-space output - get_llm(object_detector) now resolves the utils scope, so the configured detector endpoint is actually used (was silently falling back to the operator model) - strip_json_comments preserves // and /* */ inside string literals, which made URL api_base values unparseable - Empty exported OPENAI_API_KEY falls back to the "EMPTY" placeholder instead of reaching ChatOpenAI as a blank credential - resources config: document the ondevice-qwen preset and the coordinate_format contract for local vision models Verified live against ondevice-agent-platform (:8080): qwen3.8-9b serves the planner/text route and qwen-vl grounds UI elements on a real iOS screenshot within ~3% of the 0-1000 grid.
Accept Gemma's trained box_2d output ([ymin,xmin,ymax,xmax] on the 0-1000 grid) in the object detector: normalize it to the existing [x1,y1,x2,y2] contract, synthesize the point from the box center when no usable point exists, and include box coordinates in batch-scale inference. Add a spike config routing every node to gemma-4-e4b on the local OpenAI-compatible endpoint, alongside the untouched qwen config. Also add an on-device Apple Vision OCR provider behind perform_ocr: a standalone worker subprocess (Vision calls fail intermittently once cv2 is resident) with a per-call timeout, one retry, and a circuit breaker that disables the provider after consecutive failures. Apple Vision is preferred on macOS; Google Cloud Vision remains the keyed fallback and ARTEMIS_OCR_PROVIDER pins a provider.
Adds a LocalModelEndpointProbe (artemis doctor / console wizard /
mobile_diagnose): walks the active LLM config for custom-provider
endpoints, verifies each required model alias is served via
GET {api_base}/models, and surfaces guided remediation actions -
`ondevice-agent-platform model pull --alias <model>` when an alias is
missing, `ondevice-agent-platform serve` when the endpoint is down.
LLMCredentialsProbe now treats config-bound custom endpoints as a
satisfied credential instead of reporting a missing API key.
… OCR - default -> gemma4-e4b (custom provider, :8080/v1, 300s timeout), fallback -> qwen3.8-9b on the same endpoint - operator/object_detector/video_analyzer/explorer primaries -> vision-hybrid: Apple Vision answers confident OCR turns natively; everything else escalates to qwen-vl (kept as each node's fallback) Includes previously-uncommitted local-provider bindings from the gemma4 default switch; no behavior change for semantic turns.
operator/object_detector/video_analyzer/explorer primaries now use gemma4-e4b (semantic/grounding); qwen-vl stays each node's fallback. vision-hybrid is the OCR-only route per user direction. Note: gemma's 0-1000 grid output verified only on synthetic input - qwen-vl remains the verified grounding fallback.
…R=platform) Adds a 'platform' OCR provider calling the local platform's vision-hybrid route; structured observations arrive via the additive oap_ocr field (text + pixel vertices matching the boundingPoly contract). In 'auto' it sits between the native Apple Vision worker and Google Cloud Vision. Verified live: TOTAL extraction with correct normalized coordinates.
The standalone PyObjC Apple Vision worker is deleted - the platform's vision-hybrid route (oap-vision-bridge) is now the primary OCR provider in 'auto', with keyed Google Vision as the fallback. The 'apple' provider pin is a deprecated alias for 'platform'. Structured observations arrive via the additive oap_ocr field (pixel vertices, boundingPoly order). Verified live: TOTAL extraction at [446,475] through run_ocr_core under default auto.
The platform exposes the vision-hybrid route (oap_ocr structured observations); whether ARTEMIS calls it is ARTEMIS's own choice. The original apple_vision_ocr.py worker and provider chain are restored.
Gemma's native box_2d is [ymin, xmin, ymax, xmax]/1000 — the detector's yx_1000 contract. The previous xy_1000 (correct for the qwen-vl fallback) axis-flipped every box into garbage; verified with a real model response normalized through _normalize_detected_box.
- Sniff image MIME from bytes (image_data_uri) instead of hardcoding image/jpeg; strict local gateways reject mislabeled PNG screenshots. - with_structured_output uses function_calling for the custom provider. - Flash: prune intermediate screenshots before invoke (platform caps images-per-request). - Planner/operator/checker focus reminders keep protocol rules inside the sliding attention window for small local models. - LLM_FLASH_HARD_TIMEOUT_SECONDS for slow local providers.
|
Thanks for your pull request! It looks like this may be your first contribution to a Google open source project. Before we can look at your pull request, you'll need to sign a Contributor License Agreement (CLA). View this failed invocation of the CLA check for more information. For the most up to date status, view the checks section at the bottom of the pull request. |
banggaoo
force-pushed
the
feat/gemma4-upstream
branch
from
October 10, 2026 03:50
fa22b22 to
37017be
Compare
- artemis model pull/serve/stop/status/list manages mlx_vlm.server as a background process (state under ~/.artemis/local_models/) - config/local_models.py: alias catalog (gemma4-e4b -> HF repo) - readiness probe now suggests artemis model pull/serve <alias> for catalog models, generic pull/serve hints otherwise - pip extra local (mlx-vlm, huggingface-hub; Apple Silicon gated) - docs/local_models.md leads with the self-operated path; BYO server remains for other platforms
banggaoo
force-pushed
the
feat/gemma4-upstream
branch
from
October 10, 2026 04:31
37017be to
2c39050
Compare
Gemma 4 image processor defaults to 280 soft tokens per image, which
downsamples a phone screenshot to ~528x1152 effective and produced
measured grounding failures (wrong thumbnail, echoed coordinates,
lookalike-anchored detections of absent elements). At 1120 tokens the
same suite goes 8/8: dead-on taps and correct empty-list refusals.
- LocalModelSpec.vision_soft_tokens + catalog default 1120 for e4b
- artemis model serve --vision-tokens {70,140,280,560,1120} patches the
pulled snapshot processor_config.json before spawning mlx_vlm.server
(idempotent; graceful noop when the snapshot or key is absent)
- docs/local_models.md: vision-detail tuning note
…/serve/stop endpoints
…s for model lifecycle
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Design & architecture: #177
Summary
Adds support for running ARTEMIS agents against a local OpenAI-compatible model server (tested with Gemma 4 E4B via MLX), enabling fully on-device multimodal agent execution.
artemis model pull/serve/status/stop/list— catalog aliases (gemma4-e4b) resolve to pinned HF repos;servelaunches a managedmlx_vlm.serverand applies the catalog's recommended vision soft-token budget (1120, ~1056×2352 effective detail vs the 280 default).customprovider path withfunction_callingstructured output; ready-madeconfig/artemis.gemma4.jsonc.local_model_probegates runs on reachability + served-alias checks forcustomendpoints; catalog aliases getartemis model pull/serveguidance, BYO servers get neutral tooling hints. Self-passes for cloud-only configs.mark_donetool gives the Pro/operator loop a structural completion claim (latched flag → convergence gate → settlement/final verification); false claims bounce back to execution. Plans whose top-level items are all resolved ([x]or[!]blocked) now route to settlement instead of looping.image_data_urisniffs PNG/JPEG/WebP magic bytes instead of hardcodingimage/jpeg; strict local gateways reject mislabeled payloads.LLM_FLASH_HARD_TIMEOUT_SECONDSfor slow local providers.tool_choice="required"oncustomproviders holds the one-tool-call-per-turn contract (eliminates narration-only loops on small local models);call:name{args}envelope repair recovers calls that mlx_vlm wraps into unparseableinvalid_tool_callsor leaks into text content; a consecutive no-tool-call cap (8) bounds the loop when a provider still refuses to call tools.yx_1000mode for Gemma-style outputs.Test plan
pytestfocused suites: object detector coords, local endpoint routing, optional OCR, operator focus reminder — passtests/unit/{agents,memory,tools}sweep: no new failures vs. base branchmark_doneclaim routing + blocked-plan termination — passgemma4-e4bon a directmlx_vlm.server: Flash tasks verified on-device — "Open Settings" x2 (click Settings icon → verified →report_task_status(completed)); complex multi-step "open Settings → Display & touch → Dark theme" completed in 7 turns withui_night_mode=2confirmed🤖 Generated with Devin