Let lcode look at images, and generate them through ComfyUI - #35
Merged
Merged
Conversation
Images in: attach a screenshot, mockup or diagram with @path, or the model calls the new view_image tool. A model that can see describes the image (all text transcribed, layout and colors, focused on the question) and the description goes into the conversation, so any model can use it: - vision_model = "auto": the session's model if it can see; otherwise the original Ollama model behind lcode's text-only variant (same weights plus the vision projector, already downloaded by lcode setup), such as qwen3.6:35b-a3b-coding; or a configured model (e.g. qwen3-vl:8b). - A model that can see keeps the session's options, so Ollama doesn't reload it. - Images returned by MCP tools (Playwright screenshots) are described too. - read_file on an image points to view_image; the sandbox's path limits apply; lcode doctor shows the vision model. Images out: `lcode mcp add comfyui` adds ComfyUI's official MCP server (comfy-mcp via uvx, with comfy-cli), with only its 16 local tools enabled (no Comfy Cloud partner tools), to generate and edit images with local models such as FLUX.2 Klein, FLUX.1 Dev/Kontext, SDXL, SD 1.5 and Qwen-Image. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A 12 GB GPU can't hold lcode's model and an image model at once: with qwen3.6-35b loaded, ComfyUI's FLUX.2 Klein run fails with "VRAM grow failed". New per-server setting free_gpu (true, or a list of tools): before such a tool runs, lcode unloads its own Ollama models (the session's and the vision model, nobody else's); the next request reloads them. The comfyui preset sets it for generate_image, run_workflow and run_template, and its setup now starts ComfyUI with --disable-smart-memory so ComfyUI hands the GPU back after each image. fetch_template joins the preset's tools (needed to adjust a template). Measured (RTX 4080 Laptop 12 GB, Klein Base 4B fp8, 512x512): unload 0.2 s, about 12 s per image, about 20 s to reload qwen3.6-35b. A full lcode request (find the template, adjust it, generate, save to assets/, look at the result) worked end to end; documented with these numbers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
lcode can look at images:
@pathis described in detail: all text transcribed, plus layout, colors and anything that looks wrong, with the user's question in mind. The description goes into the conversation.view_imagetool: lets the model open images itself, e.g. a screenshot a test wrote.browser_take_screenshot) are described too, instead of[image not shown].vision_model = "auto":lcode setupalready downloaded. Ollama confirmslcode-*lack vision and their base tags have it (qwen3.6:35b-a3b-coding,qwen3.8:27b,qwen3.5:*).qwen3-vl:8b), orofflcode doctorshows which model looks at images.Image generation and editing through MCP:
lcode mcp add comfyuiadds ComfyUI's official MCP server (Comfy-Org/comfy-mcp), run viauvxwithcomfy-cli.How
src/lcode/vision.py: picks the model and handles the describe call (/api/chatwithimages,think: false, a 16K context unless it's the session's own model).Ollama.chat(): a new non-streaming helper.protocol.result_text(..., image_text=...): a hook that describes MCP image results.Testing
tests/test_vision.py:view_imagethrough a model turn, including the request payload@imagewith the question passed along@docs/social-preview.png: the four feature tags were listed correctly. The image was described byqwen3.6:35b-a3b-coding, auto-picked; that took about 50 s, including loading it.<h1>, which the description got right.lcode mcp add comfyuiconnects (16 tools), andserver_infocorrectly reported ComfyUI not running.Update: sharing one GPU, tested for real
A 12 GB GPU can't hold lcode's model and an image model at once. With qwen3.6-35b loaded, ComfyUI's FLUX.2 Klein run fails with
VRAM grow failed.free_gpu(true, or a list of tools). Before such a tool runs, lcode unloads its own Ollama models, meaning the session's model and the vision model (nobody else's). The next request reloads them.comfyuipreset sets it forgenerate_image,run_workflowandrun_template. Its setup starts ComfyUI with--disable-smart-memory, so ComfyUI gives the GPU back after each image.fetch_templateis now in the preset's allowlist; the model needs it to adjust a template.ComfyUI 0.38 was installed on the 12 GB laptop with FLUX.2 Klein 4B and Base 4B (fp8):
Full lcode session (qwen3.6-35b, 128K, auto-edit), asked to make an app icon with ComfyUI and Klein, save it to
assets/icon.pngand judge it:qwen3.6:35b-a3b-coding, and gave an accurate verdict. The icon is a clean white paper plane on a rounded blue gradient square.Found along the way, not lcode's:
noise_seedis set.Checklist
uv run pytestanduv run ruff check .passusage.md; a ComfyUI section and catalog row (with honest notes for 12 GB GPUs) inmcp.md;models.md,configuration.md,how-it-works.md, README andCHANGELOG.mdupdated.mkdocs build --strictpasses.🤖 Generated with Claude Code