Skip to content

Repository files navigation

daapu

Name generated by https://uniq.site/

A PoC for an LLM chatbot with a memory system, split into a brain and a hand:

  • The brain — Kotlin/JVM (Gradle). The ktor HTTP API (Main.kt → src/main/kotlin/.../server/) plus the memory system on PostgreSQL (pgvector) via Exposed + Flyway. The brain decides when and how to run LLM requests: the turn loop, the system prompt, history and injection, compaction and memory extraction policies, tool advertisement, and persistence — every piece of content lives here.
  • The hand — hand-pi (hand-pi/, a stateless Node/TS service on @earendil-works/pi-ai). Like the name suggests, the hand is the part that executes: it performs the LLM calls the brain decides on — streaming, dialect handling, tool-call accumulation, retries, usage — and it is deliberately opinionless (no sessions, no prompts, no content decisions). It replaces the previous execution layers: the original hand-rolled koog implementation and the later langchain4j migration.

The frontend is a minimal Svelte app (frontend/, inspired by llama.cpp's own webui but deliberately tiny). The frontend dev server proxies /api to ktor; ktor serves the API only. Chat history is stored per chat id — pick an existing chat from the sidebar to resume it, or click "new chat". Text and image messages are supported; images only work with a vision-capable model (the turn loop rejects them with a clear error otherwise — see AGENTS.md).

Why pi, not deepseek-harness (dsh)

A phase-0 spike evaluated deepseek-harness (dsh) as an lc4j replacement; technically all five gates passed (reasoning dialect, MCP tools with auth, vision, auto-compaction, plugin dev), but it was rejected for this project:

  • Prompt sovereignty. daapu is a PoC for memory-system experiments: the prompt content is the independent variable. dsh composes its own system prompt sections, loads workspace AGENTS.md, ships its own tool schemas, makes session-title auxiliary calls, and injects its own compaction instruction — every shipped default is a confound that would have to be audited or stripped. @earendil-works/pi-ai — the LLM layer dsh itself wraps — is available standalone: no loop, no system prompt, no sessions, no prompt opinions. The gateway integration dsh proved (reasoning parse, tools, vision, cache accounting) comes for free without any of the sauce.
  • History ownership. dsh owns an append-only session event log: user messages commit durably before the turn runs, there is no message edit, and daapu's store-only-on-success semantics don't exist. Kotlin stays the history authority here.
  • Churn. rc releases (6 in its first 4 days at evaluation time), a vendored Cordis fork, single-user/no-auth gateway.

The decision: a stateless TS service (hand-pi/) on pi-ai owns LLM execution only — the hand, in the koog/lc4j lineage's place; Kotlin owns everything content — history, prompts, injection, compaction policy, extraction, tools, memory, persistence — and decides when and how the hand runs (every LLM call, chat loop or one-shot, goes through the same /v1/run round loop, see AGENTS.md).

Design Philosophy: One Topic Per Session

Daapu is designed around a "Micro-Session" (Task-Oriented) philosophy, rather than a single, infinitely long conversation.

  • One Topic Per Session: A chat session should ideally focus on a single task, feature, or discussion topic. This prevents the LLM from accumulating attention decay, hallucination snowballing, and excessive token costs associated with very long contexts.
  • Memory Isolation: There is no global Short-Term Memory (SSTM, which has been removed) that forces recent conversational context across all active chats. Mixing contexts (e.g., discussing "what's for dinner" in one session and "quantum physics" or "code style" in another) pollutes the LLM's attention. Instead, each session is isolated.
  • Task-Oriented & Extract: Once a task or topic is completed, the session should be deleted or truncated. This is not just UI cleanup; it acts as a system trigger. Dropped messages are fed into the MemoryExtractionService, where facts and events are distilled and persisted into the global ELTM (External Long-Term Memory).
  • On-Demand Context: Future sessions retrieve these persisted facts dynamically via vector search (QueryRewriteService), ensuring the LLM only sees relevant, unpolluted context when it needs it.

Development

Prerequisites

  • Docker + Docker Compose
  • JDK 25
  • Node.js ≥ 22.19 (for the hand and the frontend)

Run

Start PostgreSQL, the hand, then the API server:

docker compose up -d db
cd hand-pi && npm install && npm run build
HAND_PORT=3100 HAND_TOKEN=<same as hand.token in config.jsonc> npm --prefix hand-pi start   # the hand on http://127.0.0.1:3100
./gradlew run        # ktor API on http://localhost:8080

The hand's port and token — and this API server's address as the hand reaches it (hand.selfBaseUrl) — must match the hand section of config.jsonc (see below). In a second terminal, start the frontend dev server:

cd frontend
npm install
npm run dev          # Svelte app on http://localhost:5173, proxies /api to :8080

Open http://localhost:5173, pick a chat from the sidebar (or click "new chat"), pick a model, and chat.

Deployment (docker compose)

The whole stack ships via the checked-in compose.yaml (services db, hand, brain):

cp config.example.jsonc config.jsonc   # fill in the keys, see below
docker compose up -d --build

Local-only extras (personal bind mounts, per-host port changes) belong in a gitignored compose.override.yaml — Compose merges it over compose.yaml automatically (service volumes merge by container target path) — so the tracked file stays portable.

The UI and the API share one origin — http://localhost:8080 (the brain container); the host mapping is loopback-only (127.0.0.1:8080:8080) because the API carries no authentication — widen the binding deliberately if the host must serve a LAN. The frontend dist is baked into the brain image by the multi-stage Dockerfile (stage 1 builds the frontend, stage 2 copies it into the frontend classpath package before the Gradle build, stage 3 snapshots the build-context source — the worktree as filtered by .dockerignore, so uncommitted edits ship too while config.jsonc and .git never enter — baked to /root/daapu for the agent to read — and stage 4 runs the distribution on a toolbox base: zulu JDK 25 on Ubuntu, with node 24 and python3 + uv for the stdio MCP servers plus curl/wget/git — no separate frontend server exists. The container runs as root on purpose: the brain's bash tool installs CLI tools at runtime. Anything it installs is per-container ephemeral — dependencies the deployment relies on long-term belong in the Dockerfile's toolbox layer. The root agent can also read the mounted config.jsonc and use the container's network: the container is the isolation boundary. On kubernetes this needs securityContext: { runAsUser: 0 } (or root-equivalent) on the brain pod. Development differs by design: locally the resource package is empty (see .gitignore), the ktor server answers 404 for web-UI paths, and the UI comes from the vite dev server (server/WebServer.kt → staticWebUi).

The hand runs in its own container (hand-pi/Dockerfile, no host ports — the brain reaches it over the compose network, and its image carries a HEALTHCHECK on GET /v1/health — tokenless on purpose (probes carry no secrets; it answers {"ok": true, "version": ...}), which the compose brain waits on; the same URL is the kubernetes liveness/readiness probe, httpGet on port 3100). The mounted config.jsonc needs the deployment-specific values:

Field Value
database.url jdbc:postgresql://db:5432/postgres
hand.baseUrl http://hand:3100
hand.selfBaseUrl http://brain:8080
server.port must stay 8080 — the compose host port mapping and hand.selfBaseUrl both assume it
hand.token REQUIRED (no default): set the HAND_TOKEN env var for the hand service (compose refuses to start without it) and mirror the same value here

JVM tuning goes through JAVA_OPTS (the installDist start scripts honor it — add e.g. JAVA_OPTS: -Xmx4g to the brain service's environment). docker compose up -d db remains the development flow (the db keeps its loopback-only host port 5432 for ./gradlew run — reachable locally, invisible to the LAN).

Note the config's deployment sensitivity: mcp.proxy (a host-local proxy) is a dev-local concern — in a container deployment drop it and route egress at the host/network level instead. stdio MCP servers run fine in the container (node/uv are in the toolbox layer); stdio servers whose packages are NOT preinstalled resolve on first use via the bash tool's on-demand installs and vanish with the container.

Configuration

Configuration lives in config.jsonc (JSON with C-style // comments and trailing commas) at the repo root. It contains API keys, so it is gitignored — start from the checked-in example:

cp config.example.jsonc config.jsonc
# then fill in your keys
Section Fields Description
database url, user, password (required) PostgreSQL JDBC URL, user and password. url must be a DIRECT connection or a session-pooling proxy — a transaction-pooling proxy would drop the session-level chat locks mid-run (see db/AdvisoryChatLockManager.kt).
providers {<id>: {apiKey, baseUrl, llm[]?, embedding[]?}} (required) OpenAI-compatible gateways keyed by id (e.g. bifrost); each carries the catalog entries it serves — at least one llm[] entry overall is REQUIRED. baseUrl is used as-is and must carry the full /v1 root.
server port (default 8080) API port; the frontend dev server proxies /api to it.
mcp exa (required) + customs (default none) MCP tool servers, see below.
tool fs: enabled (false), allowedDirs (required when enabled), blacklists; bash: enabled (false), shellPath (/bin/bash), workdir (null), timeoutSeconds (120), maxCaptureBytes (1000000) The harness-owned tool providers: the read-only filesystem mock (fs__*, restricted to allowedDirs with blob/gitignore-style blacklists) and the bash tool (bash__*). SECURITY: bash runs arbitrary shell commands with the brain process's own privileges — no sandbox, no allowlist; enable only inside an isolated container/VM. Enabled providers join the investigator's pool only when agent.investigator.allowedNamespaces lists their namespace.
memory compactModel + eltm (required) Compaction + external long-term memory (ELTM) settings, see below.
agent investigator (required): model (required), allowedNamespaces (required), maxRounds (default 150) The investigate sub-agent settings, see below.
title model (required), lastNRound (default 0) Session-title generation (POST /api/chats/{id}/title).
hand baseUrl (required), selfBaseUrl (required), token (required), maxRounds (64), maxRetries (0), streamIdleTimeoutMs (300000) The hand-pi execution service endpoint (baseUrl), where the hand reaches this brain (selfBaseUrl — deployment-dependent: localhost in dev, a container/service name in docker), the shared static token (HAND_TOKEN on the hand's side — no default, both sides must carry the same value), and the run-policy knobs sent with every run (the hand holds no defaults).
observability oneShotTrace (default false) Console trace of every collect run (the one-shot pipelines: title generation, query rewrite, compaction summarization, memory extraction, the ELTM writer — plus the investigate sub-agent and its partial-history summarizer) under the dedicated OneShotTrace logger: system prompt, input messages, and the full per-round transcript (reasoning, text, tool calls with args, tool results, usage, finish reasons), including the partial history of failed runs. Verbatim and untruncated — chat content can be sensitive; attachments render as size placeholders.

config.schema.json is a JSON Schema (draft-07) mirroring the config models in config/Config.kt; the "$schema": "./config.schema.json" entry at the top of config.jsonc makes editors like VSCode/IntelliJ validate the file and autocomplete field names.

Memory pipeline (memory)

History compaction and memory extraction (see AGENTS.md and agent/pipeline/compaction/ + agent/pipeline/eltm/):

  • Compaction is triggered per model: the trigger fraction and the keep rounds live on each configured chat-model entry (config.jsonc → providers.<id>.llm[].compactionTriggerFraction / .compactionKeepRounds), because the headroom depends on the model's context size (agent/model/LLM.kt owns the [0, 1] / >= 1 contract). Before each round, the brain measures the prompt size (the last assistant message's provider-reported input-token snapshot) and compacts the history when it exceeds compactionTriggerFraction of the model's context window (0 disables the proactive path). A round that still exhausts the context compacts reactively and retries (every exhaustion triggers a compaction; a compaction that fails, returns a non-clean summary, or cannot enqueue the dropped messages fails the run). compactionKeepRounds complete rounds are kept verbatim at the tail of a compacted chat; everything older is replaced by one CONTEXT COMPACTION: -marked summary user message. When the chat has fewer rounds than this, the keep count shrinks (down to zero) so an overflowing chat always compacts.
  • compactModel + eltm.extractionModel — REQUIRED catalog model ids for the one-shot pipelines (summarizer, memory extractor); they are resolved once at startup and reused for every run — a chat run's own model is never used for these. The extractor/compactor see the raw history, images included, so the model must support the content — a model that cannot fails the run with a clear error (same as the chat model's own capability check), and unknown ids fail fast at startup.
  • The extracted facts are written into the ELTM diary directly by the ELTM writer agent (no intermediate short-term store): the dropped messages are snapshotted into the background extraction queue (memory/eltm/ExtractionQueue.kt — the same queue chat deletion feeds) and a background worker runs the extractor + writer off the request path. A failed extraction retries in the background (unlimited; the writer skips content that is already recorded) — a chat run never waits for, or fails on, the memory work; only the enqueue itself can fail the run (before the store, so the dropped messages stay).

ELTM (memory.eltm)

The external long-term memory (the diary model, see AGENTS.md and memory/eltm/). It is REQUIRED for every deployment: the extraction pipeline and the investigate agent's ELTM access depend on it, so the eltm section must be present and complete.

  • extractionModel, embeddingModel, writerModel, rewriteModel — REQUIRED catalog model ids for the memory extractor, the ELTM embedding, the ELTM writer, and the per-turn query rewrite one-shot; the writer model must support tool calls, and all are resolved once at startup (unknown ids fail fast). The writer's round cap is maxWriterRounds (default 150); rewriteRounds (REQUIRED) caps the chat tail fed to the query rewrite (the last N user rounds).
  • relatedEntitiesLimit / relatedNotesLimit — REQUIRED (the injected ELTM size is related to the main model's context, so they must be explicit): how many related entities and diary notes the ELTM context injection puts into the <injection>'s <memories> section (searched by the rewritten query); 0 skips that search (with both 0 the rewrite one-shot is skipped too).
  • entityMatchThreshold (0.5) / noteSearchThreshold (0.1) — the vector cosine floors for entity near-match candidates and note search.
  • The ELTM vector columns are FIXED at vector(2000) (pgvector's HNSW indexing limit for the vector type): the embedding model's output dimensions (≤ 2000, enforced at startup) only decide the nonzero prefix, and every vector/query is zero-padded to 2000 on write (cosine similarity is invariant under zero-padding), so switching embedding models never needs a schema change or DB reset.

Sub-agents (agent)

The investigate agent (agent/pipeline/investigate/InvestigatorService.kt): a runCollect tool loop that gathers information from the ELTM (read-only) and the web (the MCP tools) on behalf of the main agent. The main agent's tool set is the MCP servers, the harness-owned providers enabled under tool (fs, bash) plus the single gsg__investigate tool (agent/persist/GsgToolProvider.kt) — the granular ELTM read tools are not in the chat loop anymore; deep memory and web searches go through the sub-agent. The sub-agent runs on its OWN tool set (built separately from the chat loop's — its children are enumerated at the InvestigatorService registration in di/AppModule.kt), so gsg is not whitelistable for it (recursion is ruled out at boot).

GSG is the project's internal codename for this harness (the context-injection + compaction + ELTM machinery). It surfaces only as the reserved gsg tool namespace and the gsg__investigate tool name — personas whitelist gsg to grant access to the harness's sub-agent and the full memory injection (see the Personas section in AGENTS.md).

  • investigator.model — REQUIRED catalog model id (must support tool calls), resolved once at startup like the memory pipeline models (unknown ids and incapable models fail fast).
  • investigator.allowedNamespaces — REQUIRED, non-empty: the namespaces the investigate tool loop may execute, a whitelist over its own tool set (see above; entries validated like any tool namespace). Every entry must be a namespace that set serves — a typo fails fast at boot via the WhitelistedToolProvider construction.
  • investigator.maxRounds (default 150) — round cap for the investigate tool loop (0 = unlimited). A round_limit stop is recovered elastically: the whole partial history is summarized by a no-tools one-shot on the same model, and the summary (or the raw assistant texts when the summarization fails) is returned to the caller instead of failing the run. A context_exhausted stop is recovered the same way with a tool-call trace.

Session titles (title)

title.model is a REQUIRED catalog model id for the session-title generator (agent/pipeline/TitleGenerator.kt, wired to POST /api/chats/{id}/title): it is resolved once at startup like the memory pipeline models and never the chat run's own model. The generator summarizes the chat's stored history (empty chats short-circuit to a no-op — the title is left untouched and no LLM call happens; a text-only model with image history fails fast with a 400 as a config error), persists the new title, and the endpoint returns the updated {"id", "title"}. The endpoint takes no per-chat lock (the store upsert never touches the title), so a title generated mid-run reflects the last stored history — re-generate after the conversation moves on.

title.lastNRound (default 0) caps the history fed to the title model to the last N user rounds (a round is one user message plus the following assistant/tool messages), so long chats stay inside the title model's context window; 0 means the whole history.

MCP tool servers are configured under mcp. The dedicated exa server is REQUIRED at mcp.exa: an ordinary server entry whose namespace is hardcoded to exa (advertised tools become exa__web_search_exa, ...) — you fill the rest. Both Streamable HTTP and stdio transports are supported — config.example.jsonc shows a working example (an exa HTTP server with an Authorization header, plus a commented-out stdio entry to adapt; a self-hosted exa-mcp-server via npx with EXA_API_KEY in the environment works the same way).

Every other server lives under mcp.customs, keyed by namespace (the key IS the namespace prefix of the advertised tool names, e.g. fs__read — the separator is __, so it must not contain it; only [0-9a-z_-] is allowed, and the reserved namespaces system/inner/internal/gsg/eltm/harness — plus exa, owned by the dedicated server — are rejected). Each entry needs a type (http needs url + optional headers; stdio needs command + optional environment) and toolExecutionTimeoutSeconds (the execution budget of every advertised tool in seconds, 0 = no timeout — REQUIRED per server, enforced by the brain on the tool callback with withTimeout; the hand applies no deadline of its own and waits until the brain answers or the connection drops). It may also set initializationTimeoutSeconds, plus reconnectAttempts (total connect attempts including the first, default 3) and reconnectDelayMs (delay between attempts, default 1000). The provider connects eagerly at startup, so a server that cannot be reached aborts startup. A transport failure mid-execution drops the connection and answers an error tool-result the model can react to (no in-turn retry or reconnect); the next tool-list refresh before an LLM request reconnects, or fails the chat run with a clear error event when the server stays down. A tool that overruns its budget answers an error tool-result too (the execution is cancelled, the model can react in the next round).

An optional mcp.proxy (host + port) routes every http-type server's requests through an HTTP proxy (CONNECT tunneling for http and https endpoints; stdio servers are unaffected). Explicit only — no HTTP_PROXY-style env-var pickup, and the proxy config model has no authentication fields.

The main agent's system prompt is split into two parts: a user-managed persona (identity/personality text plus a tool-namespace whitelist, see the Personas tab) and the GSG harness introduction (harness mechanics, memory and injection documentation, rendered by agent/persist/MainAgentSystemPromptService.kt). The DEFAULT persona lives only in code (prompt updates need no data sync; its text carries the legacy policy/jailbreak sections); each chat records its persona (chats.persona_id, stamped by successful runs), and every message carries the selected persona id alongside the model — the server resolves it per run and restricts the chat loop's tools (MCP + gsg__investigate) to the persona's whitelist. An empty whitelist means ALL namespaces, and it is the intended shape: this harness targets small open-weight models that are prone to mistakes, and information-providing tools (web search, ELTM recall) help them produce correct output — a persona that disables every tool defeats the harness's purpose (plain chat UIs like Open WebUI already exist for that). The model catalog is configured in config.jsonc under providers.<id>.llm / providers.<id>.embedding (each entry carries its budgets, compaction knobs, and an explicit capability list; at least one LLM entry is required — see config.example.jsonc). The web UI can switch between the catalog models per message; the model is sent with every message (the UI defaults to the first catalog entry — the server has no default).

Verification

The authoritative build/test command list (JVM tests, hand-pi, frontend) lives in AGENTS.md ("Verification commands") — run it after any relevant source change; everything must exit clean.

Note: the schema is a fresh migration (V1__init.sql); if you had an older database, drop the volume (docker compose down -v) before starting again.

Utility scripts

One-off dev tools (e.g. the SillyTavern chat transformer) are not part of the server; they live under src/main/kotlin/info/skyblond/daapu/script/ and run through a generic Gradle runner — see that package's README.md for usage.

Divergences from the koog/langchain4j implementations

The hand-pi migration changed a few observable behaviors on purpose (the old turn loops were the hand-rolled koog implementation, later rewritten on langchain4j):

  • Transient upstream failures are retried by the hand. 5xx responses, mid-stream error chunks, and truncated streams retry with exponential backoff inside the hand (visible as retry SSE events); the old loop only retried a clean stream without a finish_reason. 4xx responses and content_filter still fail the run.
  • Gateway-side context rejections now compact. A 400/413 prompt-too-long response with a recognizable error body classifies as context_exhausted and triggers the reactive compaction path (the old loop failed the run with the raw HTTP error). Bodyless 400/413 rejections do not: pi-ai 0.86+ gates its bodyless-overflow pattern to the Cerebras provider and the hand's fixed provider id never matches, so they surface as plain upstream errors (see hand-pi/src/classification.ts).
  • Richer assistant metadata. Assistant messages now carry timestamps, and inputTokens is the full prompt size (input + cacheRead + cacheWrite).
  • A hand connection loss is terminal. The stateless hand cannot resume a dead run, so a dropped hand connection fails the run cleanly (nothing is stored) instead of being retried.

API

All endpoints are under /api (see server/WebServer.kt; the two internal hand routes live in server/endpoint/HandRoute.kt):

Method & path Purpose
GET /api/models Model catalog (vision, context, output limits).
GET /api/chats One page of chats as {"chats": [{"id", "title", "personaId"}], "nextCursor"} (newest first, 200 per page; keyset pagination — pass ?cursor=<previous nextCursor> to continue, absent nextCursor = no more pages, 400 on a malformed cursor).
POST /api/chats Create a chat (empty history, title New chat).
PUT /api/chats/{id} Rename a chat ({"title": "..."}).
POST /api/chats/{id}/title Generate a title from the chat history (no-op on an empty chat, 400 on a capability mismatch).
DELETE /api/chats/{id} Delete a chat (409 while a run is active, 503 during ELTM maintenance).
DELETE /api/chats/{id}/messages/{index} Truncate: drop the user message at index and everything after it (without memory extraction; 400 on a non-user/out-of-bounds index).
POST /api/chats/{id}/fork/{index} Fork: copy history up to and including the assistant message at index (finishReason "stop" required) into a new chat.
GET /api/chats/{id}/chat Full chat as neutral-format JSON (raw chat_json).
GET /api/chats/{id}/export Export: {"title", "messages"} (neutral-format history; no chat id) as an attachment named {id}.json.
POST /api/chats/import Import an exported payload: a NEW chat reusing the title (fresh fork-like ELTM/persona state), validated like any stored chat (400 on violation).
POST /api/chats/{id}/messages Run one agent turn; responds with an SSE stream (503 during ELTM maintenance).
GET /api/personas All personas: the code-only default persona first, then the personas rows.
POST /api/personas Create a persona ({"name", "systemPrompt", "allowedNamespaces"}; allowedNamespaces [] = all loop namespaces).
GET /api/personas/export Export all persona rows as an attachment named personas.json (payload shared with the import, details below).
POST /api/personas/import Import an exported payload (details below).
PUT/DELETE /api/personas/{id} Update/delete a persona row (400 on the reserved default persona id 0).
GET /api/eltm/entities Browse entities, paged (?limit, ?offset; max 500 per page).
GET /api/eltm/entities/{id} One entity's full view (attributes, counts, latest note; 404 unknown).
GET /api/eltm/entities/{id}/relationships The entity's relationships (?includeInvalid, default false).
GET /api/eltm/entities/{id}/notes The entity's notes, paged, optional ?from/?to date range.
GET /api/eltm/relationships Browse relationships, paged (?limit, ?offset).
GET /api/eltm/relationships/{id} One relationship's full view (counts, latest note; 404 unknown).
GET /api/eltm/relationships/{id}/notes The relationship's notes, paged, optional ?from/?to date range.
POST /api/eltm/digest Manual memory digest (503 during ELTM maintenance) — request/response details below.
GET /api/eltm/replay The ELTM replay job's status as {"state", "messagesTotal", "windowsCompacted", "jobsQueued", "error"} (never blocked).
POST /api/eltm/replay Start the background replay of an uploaded foreign chat — walk it through the compaction windows and enqueue every dropped region for memory extraction (503 during ELTM maintenance; 409 while a walk runs; details below).
GET /api/eltm/export Export the whole ELTM as an attachment named eltm.json (payload shared with the import, details below).
POST /api/eltm/import Import (merge) an exported payload (503 during ELTM maintenance); ?overwriteAttr controls attribute overwrites (details below).
GET /api/maintenance The ELTM maintenance-mode flag as {"enabled": bool}.
PUT /api/maintenance Set the flag ({"enabled": bool}, idempotent); answers the applied state.
GET /api/maintenance/reembed The re-embed job's status as {"state", "entities", "notes", "error"} (never blocked).
POST /api/maintenance/reembed Start the background re-embed job — re-embed every stored ELTM vector with the configured embedding model (409 unless maintenance mode is on or while a job runs; 202 + the running status).
GET /api/hand/tools Internal: the hand's per-round tool advertisement (?runId=...).
POST /api/hand/tool Internal: the hand's tool-execution callback (runId-scoped).

POST /api/chats/{id}/messages accepts {"text": "...", "images": [{"dataUrl": "data:image/png;base64,..."}], "model": "...", "personaId": 0} (text and/or images required, model and persona required — no server defaults) and streams SSE events: reasoning, text, tool_call, tool_result, retry (transient hiccup, the run will be retried), done, or error (terminal). Validation errors are plain 400/409 responses before the stream starts (and a plain 503 while ELTM maintenance mode is on). History compaction happens server-side with no dedicated event — the frontend's post-run resync (done/error) presents the compacted history.

GET /api/chats/{id}/export and POST /api/chats/import share one payload: {"title": "...", "messages": [...]}, where messages is exactly the stored neutral-format history (GET /api/chats/{id}/chat), so an exported file round-trips through the import unchanged. The payload carries no chat id (the import always mints a fresh one), ELTM fingerprint, or persona record — an imported chat starts fresh, like a fork. The export is a snapshot read (no lock); the import applies the same completeness validation as any stored chat (a non-empty history must end with a naturally finished assistant message, tool calls/results stay paired AND sit within one user round — the same round-local rule the ELTM replay enforces, because a straddling pair would be split by a later compaction) and answers 201 with the created chat's info. The file is named {chatId}.json — the id is filename-safe, the title is not.

GET /api/personas/export and POST /api/personas/import share one payload: a JSON array of {"name", "systemPrompt", "allowedNamespaces"} (the export answers it as an attachment named personas.json), one entry per persona row in creation order — the code-only default persona is not exported. Names are not unique, and the array preserves same-name rows losslessly. The import skips an entry only when an existing persona with the same name already carries the same (trimmed) system prompt and the same namespace set (order-insensitive); otherwise it creates a new persona with the full create validation. It is fail-fast and partial: the first invalid entry (e.g. a namespace this deployment's chat loop does not serve) fails the request with 400 and the personas created before it stick, so re-running the same file skips them and resumes. The import answers {"created": [...names], "skipped": [...names]}.

POST /api/eltm/digest accepts {"parts": [...], "date": "YYYY-MM-DD"} — parts in the user-message wire shape (text parts and/or image attachments; at least one non-blank text part or image is required, and only image attachments can be digested). date is the optional reference date for resolving relative dates (must not be in the future). The first part may be a dedicated text part carrying a <user-provided-context> block — the caller's free-form explanation of what the input is (an email, a PDF, a contract): the extractor uses it to interpret the rest and may draw facts from it. The request blocks for the memory-extraction one-shot plus the writer loop (minutes are normal) and responds 201 Created with an empty body on success — an empty extraction is an indistinguishable no-op — or 502 with the failure chain in the body when a stage fails terminally (whatever the writer already recorded sticks, and the writer deduplicates on retry). No chat lock: nothing here touches the chats table.

POST /api/eltm/replay (the web UI's ELTM Replay tab) replays a FOREIGN chat through this system's memory pipeline without importing it as a chat: the request body is the exported {"title": "...", "messages": [...]} payload — the ONE shape the system speaks (what GET /api/chats/{id}/export serves and POST /api/chats/import accepts, e.g. the SillyTavern transformer's single output file; the title must be present and any string value is accepted, the replay never uses it) — and the two window knobs ride the query params (?compactionRounds=8&contextRounds=3, the defaults). The body must arrive as application/json like every typed request body — a wrong Content-Type is ktor's body-less 415. The server walks the chat window by window with the production compaction stage (memory.compactModel) and enqueues every dropped region — the running summary included, exactly what a compaction of a stored chat would enqueue — into the background extraction queue: the worker's extractor + ELTM writer pipeline turns each region into memories off the request path, with the queue's retries, its maintenance-mode pause and its parallel drain. The residue (the last summary plus the trailing rounds, or the whole chat when it never exceeded one window) is enqueued as the final job; every raw round is extracted exactly once. Why not import-and-delete: a foreign chat can be far longer than the models' windows, and one giant extraction over the whole chat is exactly what the extractor cannot do reliably — the walk feeds it window-sized batches instead. The route validates synchronously and answers 400 for a body failing the {title, messages} decode (not JSON, a missing or non-string title, a message-init violation) or the stored-chat invariants, an empty chat, a chat without user messages, a chat whose tool_call/tool_result pairs straddle user rounds (the walk cuts regions at round boundaries; such a pair would be split across extraction jobs that can never decode — the check is deliberately conservative and also refuses pairs the walk itself would never split, e.g. one straddling only the leading prologue's boundary), bad knobs (compactionRounds >= 2, contextRounds >= 1) or a pipeline-model capability mismatch over the chat's content (e.g. images with a text-only model), 503 during ELTM maintenance, 409 while a walk is already active, and otherwise 202 with the walk's status. The walk runs in the background (its progress goes to the server log) and only tracks itself: GET /api/eltm/replay reports idle / running / finished (with the message, window and job counters) / failed (with the at-failure window and job counters — how much already went into the queue — and the reason) — "finished" means every region is QUEUED, not that the memories are recorded; they land asynchronously like a deleted chat's. A failed walk keeps its already-enqueued regions queued, and re-running the same file re-enqueues everything (safe: the writer deduplicates recorded content, at the cost of a full re-extraction). Single-flight per server instance (the status is in-memory; a restart resets it to idle, the queued jobs drain next boot). No chat lock: nothing here touches the chats table. Accepted PoC limits: no size cap on the body, and the knobs count user ROUNDS, not messages — a chat with huge single messages can still overflow the models' windows; lower the knobs in that case.

GET /api/eltm/export and POST /api/eltm/import share one payload: {"entities": {"<uuid>": {"name", "category", "attributes", "notes"}}, "relationships": [{"srcUuid", "verb", "dstUuid", "valid", "notes"}]}, where each note is {"date", "note"} (no createdAt — DB metadata). The entity keys are file-level uuids minted fresh at every export — pure join keys for the relationships' endpoint references, never used for matching. No db ids and no embeddings travel: the import re-derives both (fresh row ids, embeddings recomputed through the local hand), so a file transfers across instances with different embedding models. The import MERGES: entities match on (name, category) (missing ones are created); attributes set new keys always, existing keys only with ?overwriteAttr=true (default false; DB-only keys are never deleted); notes append when their (event date, trimmed text) pair is new for the subject — exact duplicates are skipped, whether already in the DB or repeated within one file; relationships match on the resolved (src, verb, dst) triple (missing ones are created), and the file's valid state applies only when the file's newest note date is strictly newer than the row's latest note date (a row created by this import always takes the file's state, and a row without notes loses to the file whenever the file carries notes; notes never replay validity, since a stored note carries no structural flag). The import is fail-fast and partial like the persona import: the whole file is validated BEFORE the first write (400 with the offending path on a broken file, nothing written — including a file that repeats an entity key, an attribute key, or a relationship triple after normalization, which would violate the DB's uniqueness constraints), then the first failure aborts the request — 400 for a validation reason, 502 for an upstream embedding failure — with everything already written sticking, so re-running the same file skips the existing content and resumes. It answers the per-kind split {"entitiesCreated", "entitiesMatched", "relationshipsCreated", "relationshipsMatched", "notesInserted", "notesSkipped", "attributesWritten", "attributesKept"}. No chat lock: nothing touches the chats table (the merge rides the same concurrent-writer-tolerant service methods as the extraction pipeline).

Maintenance mode

GET/PUT /api/maintenance read and set the ELTM maintenance mode (the gsg_meta_number.eltm_maintenance flag; the web UI's #/maintenance tab is its interface). While enabled, every operation that reads or writes the long-term memory answers 503 with the reason: POST /api/chats/{id}/messages (the memory injection and the investigator read the ELTM, compaction enqueues extraction), DELETE /api/chats/{id} (deletion enqueues the history for extraction), POST /api/eltm/digest, POST /api/eltm/import and POST /api/eltm/replay (direct ELTM writes; the replay enqueues extraction jobs — its status read stays open). The background extraction worker pauses too — no job is claimed while the flag is on, queued extractions keep waiting and drain once the mode is turned off. Everything else keeps working: listing/reading/renaming/truncating/forking/importing/exporting chats, all ELTM reads, personas, models — and the toggle itself is never blocked, so the mode can always be turned back off. Accepted limits: the check is advisory (a run that passed the guard just before the flag flips still proceeds), an extraction already in flight when the flag flips finishes normally, and a replay walk already running when the flag flips keeps enqueueing — its jobs drain once the mode is turned off.

Re-embedding all memory vectors

POST /api/maintenance/reembed re-embeds EVERY stored ELTM vector (entities and notes) with the embedding model the config currently points to (memory.eltm.embeddingModel) — run it after switching embedding models, since old vectors are useless to a new model and cosine similarities across models are not comparable. The typical flow: change memory.eltm.embeddingModel in config.jsonc, restart the server (the model resolves once at boot), enable maintenance mode, press the button in the #/maintenance tab, and turn the mode off once the job reports finished. The route answers 409 unless maintenance mode is on (the freeze is the run's safety — no ELTM writes happen while the job runs; the ELTM read endpoints stay open, see Maintenance mode above) or while a job is already running; otherwise 202 with the running status. The job is a background fire-and-forget: rows are processed in id order, page by page, each batch written in its own transaction, and the progress goes to the SERVER LOG (per-page "re-embedded N so far" lines), not the API. GET /api/maintenance/reembed reports the phase — idle, running, finished (with the entity/note counts), or failed (with the reason; already-written batches stay written, and the job is safely re-runnable). On full success the global ELTM version counter is bumped once, so every chat's next run flags eltm-updated. Turning maintenance mode off mid-run is allowed: the job keeps going, and a search in that window sees a mix of old and new vectors (mid-refresh writes already embed with the new model, so re-embedding them is a harmless no-op).

References

For Coding agents.

Exposed

pi-ai

MCP Kotlin SDK

About

Experimenting harness focused on long term memory

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages