Skip to content

Let lcode look at images, and generate them through ComfyUI - #35

Merged
nasser1941 merged 2 commits into
mainfrom
feat/images
Oct 1, 2026
Merged

nasser1941 merged 2 commits into
mainfrom
feat/images

Conversation

@nasser1941

@nasser1941 nasser1941 commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

What

lcode can look at images:

❯ the layout breaks on mobile, see @screenshots/mobile.png
❯ make the settings page match @design/settings.png
  • Attaching: an image attached with @path is described in detail: all text transcribed, plus layout, colors and anything that looks wrong, with the user's question in mind. The description goes into the conversation.
  • view_image tool: lets the model open images itself, e.g. a screenshot a test wrote.
  • MCP screenshots: images returned by MCP tools (Playwright's browser_take_screenshot) are described too, instead of [image not shown].
  • Which model looks: with vision_model = "auto":
    • the session's model, if it can see; it keeps its options, so Ollama doesn't reload it
    • otherwise the original Ollama model behind lcode's text-only variant, which lcode setup already downloaded. Ollama confirms lcode-* lack vision and their base tags have it (qwen3.6:35b-a3b-coding, qwen3.8:27b, qwen3.5:*).
    • or any configured model (qwen3-vl:8b), or off
  • Doctor: lcode doctor shows which model looks at images.

Image generation and editing through MCP: lcode mcp add comfyui adds ComfyUI's official MCP server (Comfy-Org/comfy-mcp), run via uvx with comfy-cli.

  • Models: it works with local models such as FLUX.2 Klein 4B, FLUX.1 Dev/Kontext Dev, SDXL, SD 1.5 and Qwen-Image.
  • Local tools only: 17 of its 39 tools are enabled: generating, workflows and templates, models, uploading an input image for edits, jobs and memory. That's about 10K tokens instead of 21K. The Comfy Cloud partner tools are left out, so nothing leaves the machine.

How

  • src/lcode/vision.py: picks the model and handles the describe call (/api/chat with images, think: false, a 16K context unless it's the session's own model).
  • Ollama.chat(): a new non-streaming helper.
  • protocol.result_text(..., image_text=...): a hook that describes MCP image results.

Testing

  • 8 new tests in tests/test_vision.py:
    • model choice (own model, base tag, explicit, off)
    • image checks
    • view_image through a model turn, including the request payload
    • no reload for a model that can see
    • @image with the question passed along
    • the hint when no model can see
    • MCP image descriptions
    • the sandbox path guard
  • The catalog test covers the ComfyUI entry.
  • Real session with qwen3.6-35b (text-only) at 128K:
    • @docs/social-preview.png: the four feature tags were listed correctly. The image was described by qwen3.6:35b-a3b-coding, auto-picked; that took about 50 s, including loading it.
    • Playwright screenshot of example.com: described accurately. The page really has no <h1>, which the description got right.
    • ComfyUI: lcode mcp add comfyui connects (16 tools), and server_info correctly reported ComfyUI not running.
  • Not tested here: actual image generation. ComfyUI and its models (several GB) aren't installed on this machine.

Update: sharing one GPU, tested for real

A 12 GB GPU can't hold lcode's model and an image model at once. With qwen3.6-35b loaded, ComfyUI's FLUX.2 Klein run fails with VRAM grow failed.

  • New per-server setting free_gpu (true, or a list of tools). Before such a tool runs, lcode unloads its own Ollama models, meaning the session's model and the vision model (nobody else's). The next request reloads them.
  • The comfyui preset sets it for generate_image, run_workflow and run_template. Its setup starts ComfyUI with --disable-smart-memory, so ComfyUI gives the GPU back after each image.
  • fetch_template is now in the preset's allowlist; the model needs it to adjust a template.

ComfyUI 0.38 was installed on the 12 GB laptop with FLUX.2 Klein 4B and Base 4B (fp8):

Measured
Generating with lcode's model loaded fails (out of GPU memory)
Unloading lcode's model 0.2 s
One 512×512 image (Klein Base 4B, 20 steps) about 12 s
Reloading qwen3.6-35b (64K) about 20 s

Full lcode session (qwen3.6-35b, 128K, auto-edit), asked to make an app icon with ComfyUI and Klein, save it to assets/icon.png and judge it:

  • Workflow: the model found the template, worked around the fp8 file names, freed the GPU and ran the workflow.
  • Result: it copied the 512×512 PNG into the project, looked at it with qwen3.6:35b-a3b-coding, and gave an accurate verdict. The icon is a clean white paper plane on a rounded blue gradient square.
  • Time: about 13 minutes in total, mostly first-time trial and error with the template's settings. The docs now suggest saving the working workflow in the project for quick reuse.

Found along the way, not lcode's:

  • comfy-cli shifts subgraph parameters when noise_seed is set.
  • The machine's uv defaults to a Python 3.14.0 release candidate, which breaks pydantic; ComfyUI was rebuilt on 3.12.

Checklist

  • uv run pytest and uv run ruff check . pass
  • Tests cover the change
  • Docs: an "Images" section in usage.md; a ComfyUI section and catalog row (with honest notes for 12 GB GPUs) in mcp.md; models.md, configuration.md, how-it-works.md, README and CHANGELOG.md updated. mkdocs build --strict passes.
  • Scanned with gitleaks

🤖 Generated with Claude Code

nasser1941 and others added 2 commits October 1, 2026 13:05
Images in: attach a screenshot, mockup or diagram with @path, or the model
calls the new view_image tool. A model that can see describes the image
(all text transcribed, layout and colors, focused on the question) and the
description goes into the conversation, so any model can use it:

- vision_model = "auto": the session's model if it can see; otherwise the
  original Ollama model behind lcode's text-only variant (same weights plus
  the vision projector, already downloaded by lcode setup), such as
  qwen3.6:35b-a3b-coding; or a configured model (e.g. qwen3-vl:8b).
- A model that can see keeps the session's options, so Ollama doesn't
  reload it.
- Images returned by MCP tools (Playwright screenshots) are described too.
- read_file on an image points to view_image; the sandbox's path limits
  apply; lcode doctor shows the vision model.

Images out: `lcode mcp add comfyui` adds ComfyUI's official MCP server
(comfy-mcp via uvx, with comfy-cli), with only its 16 local tools enabled
(no Comfy Cloud partner tools), to generate and edit images with local
models such as FLUX.2 Klein, FLUX.1 Dev/Kontext, SDXL, SD 1.5 and
Qwen-Image.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A 12 GB GPU can't hold lcode's model and an image model at once: with
qwen3.6-35b loaded, ComfyUI's FLUX.2 Klein run fails with "VRAM grow failed".
New per-server setting free_gpu (true, or a list of tools): before such a tool
runs, lcode unloads its own Ollama models (the session's and the vision model,
nobody else's); the next request reloads them. The comfyui preset sets it for
generate_image, run_workflow and run_template, and its setup now starts
ComfyUI with --disable-smart-memory so ComfyUI hands the GPU back after each
image. fetch_template joins the preset's tools (needed to adjust a template).

Measured (RTX 4080 Laptop 12 GB, Klein Base 4B fp8, 512x512): unload 0.2 s,
about 12 s per image, about 20 s to reload qwen3.6-35b. A full lcode request
(find the template, adjust it, generate, save to assets/, look at the result)
worked end to end; documented with these numbers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@nasser1941
nasser1941 merged commit fb9daf5 into main Oct 1, 2026
12 checks passed
@nasser1941
nasser1941 deleted the feat/images branch October 1, 2026 13:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant