Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,18 @@ All notable changes to lcode are documented here. The format follows

## [Unreleased]

### Added

- lcode can look at images: attach a screenshot, mockup or diagram with `@path`, and the model can
open images itself with the new `view_image` tool. A model that can see describes the image in
detail (all text transcribed, with your question in mind): your model itself if it can see, or
the original model behind lcode's text-only variant, which `lcode setup` already downloaded
(`vision_model` setting; `lcode doctor` shows which). Screenshots returned by MCP tools, such as
Playwright's, are described too.
- `lcode mcp add comfyui`: generate and edit images with models you run locally in ComfyUI (FLUX,
SDXL, Stable Diffusion 1.5, Qwen-Image), through ComfyUI's official MCP server, with only its local
tools enabled.

## [0.5.0] - 2026-10-01

### Added
Expand Down
5 changes: 4 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,7 +52,10 @@ through [Ollama](https://ollama.com), and when it needs current information it c
with tool calling.
- **MCP servers, ready to go.** Connect Jira and Confluence, GitHub, AWS, Google Drive, Grafana,
Google Cloud, Sentry, Linear, Notion, Postgres, Kubernetes and more with one command
(`lcode mcp add atlassian`), or any other MCP server. Browser sign-in (OAuth) is built in.
(`lcode mcp add atlassian`), image generation with ComfyUI, or any other MCP server. Browser sign-in (OAuth) is built in.
- **Sees images.** Attach a screenshot or mockup with `@path` and lcode looks at it, using the
vision part of your model. With ComfyUI (`lcode mcp add comfyui`) it can also generate and edit
images with local models such as FLUX, SDXL and Qwen-Image.
- **Web search when needed.** Looks up the latest versions, docs and error messages with Ollama web
search, Brave, Tavily or your own SearXNG, and reads pages as clean text.
- **Private.** The model runs on your machine, lcode never uploads your files and has no telemetry.
Expand Down
1 change: 1 addition & 0 deletions docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,7 @@ lcode config path # print the file location
| `sandbox` | `off` | Run shell commands in a container: `docker` or `podman` ([Sandbox](sandbox.md)) |
| `sandbox_image` | lcode's image | Container image for the sandbox; any image with bash and setsid |
| `sandbox_network` | `false` | Let commands in the sandbox use the network |
| `vision_model` | `auto` | The model that [looks at images](usage.md#images): `auto`, `off` or an Ollama model that can see |
| `mcp_tools` | `auto` | How [MCP](mcp.md#context) tool definitions reach the model: `auto`, `direct` or `search` (on demand) |
| `checkpoints` | `true` | Save a checkpoint before the model changes files, so [`/undo`](usage.md#undo-and-checkpoints) can restore them |

Expand Down
1 change: 1 addition & 0 deletions docs/how-it-works.md
Original file line number Diff line number Diff line change
Expand Up @@ -77,6 +77,7 @@ the Ollama server you configure, it only contacts the web when the model searche
| `permissions.py` | Approval prompts and the read-only allowlist |
| `mcp/` | MCP client: stdio and HTTP transports, OAuth sign-in, server settings, the catalog, `lcode mcp` |
| `bench.py` | `lcode bench`: the benchmark tasks, their checks and the reports |
| `vision.py` | Looking at images with a model that can see |
| `sandbox.py` | The optional container for shell commands |
| `checkpoints.py` | Snapshots before the model changes files, for `/undo` and `/rewind` |
| `catalog.py`, `models.toml` | Model catalog and memory estimates |
Expand Down
53 changes: 53 additions & 0 deletions docs/mcp.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,7 @@ and find the code for the top one."*
| `gcp` | `gcloud` commands on your Google Cloud projects | `npx` and the `gcloud` CLI, signed in |
| `github` | Repositories, issues, pull requests, Actions | a GitHub token (or the GitHub CLI) |
| `playwright` | A real (headless) browser: open pages, click, fill forms | `npx` |
| `comfyui` | Generate and edit images with local models (FLUX, SDXL, SD 1.5, Qwen-Image) | `uvx` and [ComfyUI](#images-with-comfyui) running locally |
| `context7` | Up-to-date docs and examples for thousands of libraries | nothing (an API key is optional) |
| `sentry` | Errors, issues, traces and releases | a Sentry account (browser sign-in) |
| `postgres` | Schemas, read-only queries, query performance | `uvx`, a connection URL |
Expand Down Expand Up @@ -82,6 +83,57 @@ Google Cloud project:
See [Google's guide](https://developers.google.com/workspace/guides/configure-mcp-servers) for
details.

### Images with ComfyUI

[ComfyUI](https://github.com/comfyanonymous/ComfyUI) runs image generation and editing models
locally; its official MCP server lets lcode use it, for app icons, illustrations, mockups or
placeholder art.

```bash
uv tool install comfy-cli && comfy install # ComfyUI, once
comfy launch --background -- --disable-smart-memory # start it (http://127.0.0.1:8188)
lcode mcp add comfyui
```

Then add models in ComfyUI (its model manager, or ask lcode: *"download FLUX.2 Klein 4B in
ComfyUI"*), and ask for images: *"make a 512×512 app icon of a paper plane and save it in
assets/"*. To edit an image, ask lcode to upload it and run an editing workflow (Kontext or
Qwen-Image-Edit).

| Model | Good for | On a 12 GB GPU |
|---|---|---|
| FLUX.2 Klein 4B | Generation and editing, fast | Fits (about 8 GB) |
| SDXL and its community models | Generation, inpainting | Fits |
| Stable Diffusion 1.5 and its community models | Light generation, inpainting | Fits easily |
| FLUX.1 Dev and finetunes | High-quality generation | Needs an fp8 or GGUF version; slower |
| FLUX.1 Kontext Dev | Editing an image from instructions | Needs an fp8 or GGUF version; slower |
| Qwen-Image / Qwen-Image-Edit | Generation and editing, good text in images | Heavy: a GGUF version and RAM offloading |

Only ComfyUI's local tools are turned on; Comfy Cloud's partner tools are left out, so prompts and
images stay on your machine.

**Sharing one GPU.** A GPU rarely holds a coding model and an image model at the same time, so the
two take turns: before ComfyUI generates, lcode unloads its own model (the `free_gpu` setting), and
ComfyUI started with `--disable-smart-memory` gives the GPU back after each image. lcode's model
then reloads for the next step, which adds a few seconds to half a minute per image request. With
enough GPU memory for both (or on a Mac with plenty of memory), remove `free_gpu` from the server's
settings in `mcp.json`. lcode can also look at the results ([Images](usage.md#images)).

Measured on an RTX 4080 Laptop GPU (12 GB) with qwen3.6-35b at 64K context and FLUX.2 Klein Base 4B
(fp8, 512×512, 20 steps):

| | |
|---|---|
| Generating while lcode's model is loaded | Fails: ComfyUI runs out of GPU memory |
| Freeing the GPU (lcode unloads its model) | 0.2 s |
| Generating one image | about 12 s |
| Reloading lcode's model afterwards | about 20 s |

The first time, the model has to find the right template and fill in its settings (model file names,
size, prompt), which can take several minutes of trial and error. Once an image comes out right, ask
lcode to save that workflow in your project (for example `assets/icon.workflow.json`) and reuse it:
later images are a single `run_workflow` call.

## Adding your own servers

Any MCP server works. For a remote server, give its URL; for a local one, the command that starts it:
Expand Down Expand Up @@ -122,6 +174,7 @@ document, so you can also paste a server's example config into that file:
| `allow` | Tools that run without asking; `["*"]` for all of the server's tools |
| `timeout` | Seconds a tool call may take (default 300) |
| `disabled` | `true` to turn the server off (`lcode mcp disable <name>`) |
| `free_gpu` | Tools that need the GPU to themselves (`true` for all): lcode unloads its model first and reloads it afterwards |
| `oauth` | For servers that need a pre-registered OAuth client: `client_id`, `client_secret`, `scopes`, `authorize_params` |

Values can use environment variables: `${NAME}`, or `${NAME:-default}`. Keep tokens in environment
Expand Down
6 changes: 4 additions & 2 deletions docs/models.md
Original file line number Diff line number Diff line change
Expand Up @@ -225,7 +225,9 @@ GPU's share of unified memory. Results from other Macs are very welcome in the

## Text-only variants

Some models (the Qwen3.6 family, for example) ship with a vision encoder that a coding agent doesn't
use. `lcode setup` creates a text-only variant named `lcode-<key>` that reuses the downloaded
Some models (the Qwen3.6 family, for example) ship with a vision encoder that a coding agent rarely
needs. `lcode setup` creates a text-only variant named `lcode-<key>` that reuses the downloaded
weights, so it takes no extra disk space, and frees about 1 GB of GPU memory for the context cache
(and for a larger prompt batch, where it fits; see [tuning](configuration.md#tuning-for-speed-and-memory)).
When you show lcode an image, it borrows the original model, vision encoder included, to look at it
(see [Images](usage.md#images)).
28 changes: 28 additions & 0 deletions docs/usage.md
Original file line number Diff line number Diff line change
Expand Up @@ -86,6 +86,34 @@ reasoning is on.
lcode refuses to edit a file the model hasn't read in the session, or one that changed on disk since
it was read, so the model always edits the current version.

## Images

lcode can look at screenshots, mockups, diagrams and photos:

```text
❯ the layout breaks on mobile, see @screenshots/mobile.png
❯ make the settings page match @design/settings.png
```

- Attach an image with `@path`, like a file. lcode describes it in detail, with all visible text
transcribed and with your question in mind, and the model works from that description.
- The model can open images itself with the `view_image` tool, for example a screenshot a test wrote.
- Screenshots that [MCP](mcp.md) tools return, such as the Playwright browser's, are described too.

**Which model looks.** lcode's text-only model variants leave out the vision part to save GPU
memory, so lcode borrows the original model, which `lcode setup` already downloaded (for example
`qwen3.6:35b-a3b-coding` for `qwen3.6-35b`). If your model can see, it looks itself. `lcode doctor`
shows which one is used. With another model, install one that can see and point lcode to it:

```bash
ollama pull qwen3-vl:8b
lcode config set vision_model qwen3-vl:8b
```

Loading a second model takes a moment: on a 12 GB GPU, describing an image with the original
`qwen3.6:35b-a3b-coding` takes about a minute, and the next request reloads the main model. Turn
image support off with `lcode config set vision_model off`.

## Permissions

| Mode | File edits | Shell commands |
Expand Down
74 changes: 72 additions & 2 deletions src/lcode/agent.py
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,7 @@

from __future__ import annotations

import base64
import datetime as dt
import json
import platform
Expand All @@ -19,7 +20,7 @@
from rich.panel import Panel
from rich.text import Text

from lcode import catalog, limits, sessions, web
from lcode import catalog, limits, sessions, vision, web
from lcode.checkpoints import Checkpoints
from lcode.config import format_tokens
from lcode.mcp import McpManager
Expand All @@ -29,9 +30,11 @@
from lcode.sandbox import Sandbox, SandboxError, project_root
from lcode.tools import (
SCHEMAS,
VIEW_IMAGE_SCHEMA,
WEB_FETCH_SCHEMA,
WEB_SEARCH_SCHEMA,
Toolbox,
ToolError,
is_binary,
parse_text_tool_calls,
tree,
Expand Down Expand Up @@ -144,6 +147,7 @@ class Settings:
sandbox: str = "off" # off, docker or podman: where the model's shell commands run
sandbox_image: str | None = None
sandbox_network: bool = False
vision_model: str = "auto" # auto, off or an Ollama model that can see images


class Agent:
Expand All @@ -163,6 +167,7 @@ def __init__(self, ollama: Ollama, settings: Settings, cwd: Path, console: Conso
else None
)
self._mcp_prompt = "" # the MCP part at the end of the system prompt
self._vision: str | bool | None = False # the model that looks at images; False = not decided yet
self.session_name = ""
self.session_title = ""
self.ctx_used = 0
Expand Down Expand Up @@ -204,10 +209,65 @@ def tool_schemas(self) -> list[dict]:
schemas = list(SCHEMAS)
if self.settings.web != "off":
schemas += [*([WEB_SEARCH_SCHEMA] if self.search_backend() else []), WEB_FETCH_SCHEMA]
if self.vision_model():
schemas.append(VIEW_IMAGE_SCHEMA)
if self.mcp:
schemas += self.mcp.schemas(self.settings.context)
return schemas

def vision_model(self) -> str | None:
"""The model that looks at images for this session, if any (see lcode.vision)."""
if self._vision is False:
self._vision = vision.pick_model(self.ollama, self.settings.model, self.settings.vision_model)
return self._vision or None

def free_gpu(self) -> list[str]:
"""Unload lcode's models from Ollama so another program can use the GPU; the next request reloads.

Only lcode's own models: the session's and the one that looks at images.
"""
mine = {self.settings.model, *([self._vision] if isinstance(self._vision, str) else [])}
freed = []
for entry in self.ollama.running():
name = entry.get("name") or entry.get("model") or ""
if name in mine or name.removesuffix(":latest") in mine:
try:
self.ollama.unload(name)
freed.append(name.removesuffix(":latest"))
except OllamaError:
pass
return freed

def look(self, path: Path, question: str = "") -> str:
"""Describe an image with the vision model. Raises ToolError if it can't."""
model = self.vision_model()
if not model:
raise ToolError(f"can't look at {path.name}: {vision.INSTALL_HINT}")
try:
image = vision.read_image(path)
# The session's own model keeps its settings, so Ollama doesn't reload it.
options = self.options() if model == self.settings.model else None
loading = "" if model == self.settings.model else " (loads it; the next request reloads the main model)"
with self.console.status(f"Looking at {path.name} with {model}{loading}…", spinner="dots"):
return vision.describe(self.ollama, model, image, question, options, self.settings.keep_alive)
except vision.VisionError as e:
raise ToolError(str(e)) from e

def describe_image_data(self, data: str, mime: str) -> str:
"""For images that tools return (e.g. an MCP browser's screenshots): a description, or a note."""
model = self.vision_model()
if not model:
return f"[{mime} image not shown: {vision.INSTALL_HINT}]"
options = self.options() if model == self.settings.model else None
try:
with self.console.status(f"Looking at the {mime} image with {model}…", spinner="dots"):
text = vision.describe(
self.ollama, model, base64.b64decode(data), "", options, self.settings.keep_alive
)
except (vision.VisionError, ValueError) as e:
return f"[{mime} image not shown: {e}]"
return f"[{mime} image, as described by {model}]\n{text}"

def prepare_mcp(self) -> None:
"""Before a request: wait for MCP servers still starting and describe them in the system prompt."""
if not self.mcp:
Expand Down Expand Up @@ -559,7 +619,17 @@ def expand_mentions(self, text: str) -> str:
attached = []
for ref in re.findall(r"(?<!\S)@([\w./~\-]+)", text):
p = self.tools.resolve(ref)
if p.is_file() and not is_binary(p) and p.stat().st_size < 200_000:
if p.is_file() and vision.is_image(p):
try:
description = self.look(p, re.sub(r"(?<!\S)@[\w./~\-]+", "", text))
who = f' described_by="{self.vision_model()}"'
except ToolError as e:
description, who = f"(lcode couldn't look at this image: {e})", ""
self.console.print(Text(f" ⎿ {e}", style="yellow"))
else:
self.console.print(Text(f" ⎿ looked at {self.tools.rel(p)}", style="dim"))
attached.append(f'<image path="{self.tools.rel(p)}"{who}>\n{description}\n</image>')
elif p.is_file() and not is_binary(p) and p.stat().st_size < 200_000:
attached.append(f'<file path="{self.tools.rel(p)}">\n{p.read_text(errors="replace")}\n</file>')
self.tools.read_mtimes[str(p)] = p.stat().st_mtime
self.console.print(Text(f" ⎿ attached {self.tools.rel(p)}", style="dim"))
Expand Down
12 changes: 12 additions & 0 deletions src/lcode/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -273,6 +273,17 @@ def line(label: str, value: str, good: bool | None = True) -> None:
if spec:
detail += f" · ~{spec.memory_gib(ctx):.0f} GB needed, ~{hw.budget_gib:.0f} GB available"
line("Context", detail + (f" ({note})" if note else ""), None if note else True)
from lcode import vision

seer = vision.pick_model(ollama, model, cfg["vision_model"])
if seer:
line(
"Vision",
f"{seer} looks at screenshots and images" + ("" if seer == model else " (loaded when needed)"),
True,
)
else:
line("Vision", "off" if cfg["vision_model"] == "off" else vision.INSTALL_HINT, None)
if limits.get(model):
line(
"Limit",
Expand Down Expand Up @@ -497,6 +508,7 @@ def cmd_chat(args) -> None:
sandbox=(cfg["sandbox"] if cfg["sandbox"] != "off" else "docker") if args.sandbox else cfg["sandbox"],
sandbox_image=cfg["sandbox_image"],
sandbox_network=cfg["sandbox_network"],
vision_model=cfg["vision_model"],
)
agent = Agent(ollama, settings, cwd, console=console)
if agent.sandbox:
Expand Down
1 change: 1 addition & 0 deletions src/lcode/config.py
Original file line number Diff line number Diff line change
Expand Up @@ -43,6 +43,7 @@
"sandbox": ("off", str, "run the model's shell commands in a container: off | docker | podman"),
"sandbox_image": (None, str, "container image for the sandbox (default: lcode's, built on first use)"),
"sandbox_network": (False, bool, "let commands in the sandbox use the network"),
"vision_model": ("auto", str, "model that looks at images: auto | off | an Ollama model with vision"),
"mcp_tools": ("auto", str, "how MCP tools reach the model: auto | direct | search (on demand, saves context)"),
}
ENV_OVERRIDES = {
Expand Down
28 changes: 28 additions & 0 deletions src/lcode/mcp/catalog.toml
Original file line number Diff line number Diff line change
Expand Up @@ -130,6 +130,34 @@ requires = ["npx"]
setup = "Runs a headless browser; the first use may download it."
server = { command = "npx", args = ["-y", "@playwright/mcp@latest", "--headless"] }

[comfyui]
name = "ComfyUI"
description = "Generate and edit images with models you run locally in ComfyUI: FLUX, SDXL, SD 1.5, Qwen-Image"
homepage = "https://github.com/Comfy-Org/comfy-mcp"
requires = ["uvx"]
setup = """Needs ComfyUI on this machine (http://127.0.0.1:8188) with at least one image model:
uv tool install comfy-cli && comfy install
comfy launch --background -- --disable-smart-memory
--disable-smart-memory makes ComfyUI give the GPU back after each image. Before generating, lcode
unloads its own model so ComfyUI has the GPU; it reloads afterwards. Only the local tools are on;
Comfy Cloud's partner tools are left out, so nothing leaves this machine."""

[comfyui.server]
command = "uvx"
args = ["--python", "3.12", "--from", "comfy-mcp", "--with", "comfy-cli>=1.14.0", "comfy-mcp"]
env = { COMFYUI_URL = "${COMFYUI_URL}" }
tools = [
"server_info", "generate_image", "run_workflow", "run_template", "search_templates", "get_template", "fetch_template",
"list_workflow_slots", "set_workflow_slot", "job", "fetch_outputs", "upload_file", "search_models",
"download_model", "system_stats", "free_memory", "launch_comfyui",
]
free_gpu = ["generate_image", "run_workflow", "run_template"]

[[comfyui.inputs]]
var = "COMFYUI_URL"
prompt = "ComfyUI address (Enter for this machine, 127.0.0.1:8188)"
optional = true

[context7]
name = "Context7"
description = "Current documentation and code examples for thousands of libraries"
Expand Down
Loading
Loading