Chat with your own documents, using a language model that runs on your own hardware. Upload PDF, DOCX, PPTX, XLSX, TXT, Markdown or images; LocalChat chunks and embeds them into PostgreSQL with pgvector, and answers questions from what it retrieves. Nothing leaves the machine unless you enable web search or a cloud fallback.
Built with FastAPI, Ollama, PostgreSQL + pgvector and Redis. Hybrid semantic and lexical retrieval, a cross-encoder reranker, tool calling, streaming answers, per-workspace document isolation, and RAG parameters tunable at runtime.
Hardened beta, for a specific thing. LocalChat is a single-node, self-hosted appliance for a small team of up to 25 users — see ADR-1. All eight exit criteria in the production plan are met and that hardening gate was lifted on 2026-08-31: fail-closed boot, authorisation enforced by default in CI, a concurrency budget, a mutation-tested security core, restore proven in CI, a reproducible tagged release, migrations executed rather than merely written, and documentation verified against the code.
It said "production-ready" until 2026-09-16. An external security audit that September found defects those eight criteria were never going to catch — the worst of them reachable by any authenticated user — and while every Critical and High finding is now fixed, remediation is not finished. The remaining work is tracked in ROADMAP.md; what is knowingly accepted is listed in SECURITY.md. Read those before putting it in front of people you do not already trust.
The scope is the important half of that sentence. Multi-tenant SaaS and horizontal scaling are out of scope — running more than one replica breaks cache coherence and rate limiting silently, and the debt register in ROADMAP says exactly where. Read the criteria before relying on the label; they are a floor, not a warranty.
git clone https://github.com/jwvanderstam/LocalChat
cd LocalChat
cp .env.example .env # then set the five values below — compose refuses to start without them
docker compose up -d # PostgreSQL, Redis, Ollama and the appdocker-compose.yml requires PG_PASSWORD, SECRET_KEY, JWT_SECRET_KEY,
ADMIN_PASSWORD and ENCRYPTION_KEY to be non-empty, and .env.example ships
ENCRYPTION_KEY empty on purpose — it is a key, not a placeholder. Generate it:
python -c "from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())"Open http://localhost:5000. You will be asked to sign in.
The app image is built on a hardened, distroless base: no shell, no package manager,
running as uid 65532. That changes how you debug it — docker exec ... sh will not work;
use docker compose run --rm --entrypoint python app. See
DEPLOYMENT.md.
Everything runs in Docker, on a private network where the services address each other
by name — the app reaches Ollama at http://ollama:11434 and Postgres at db. Compose
sets those itself, and compose's environment: beats .env, so OLLAMA_BASE_URL,
PG_HOST and REDIS_HOST in your .env have no effect on the containers. Set them in
docker-compose.yml (or an override file) if you need to point elsewhere. The .env values
that do matter here are the secrets and tuning: ADMIN_PASSWORD, SECRET_KEY,
JWT_SECRET_KEY, model names, limits.
Getting the first password. Sign in as admin with the ADMIN_PASSWORD you set;
first boot seeds the account from it. (Only the host-run path, python app.py, can leave
it empty — the app then generates one and logs it once. Under Docker, compose requires
the value.)
Then pull a model and select it under Models — without an active model, chat returns
400 No active model set:
docker compose exec ollama ollama pull llama3.2:latestUpload a document under Documents, and ask about it under Chat.
Running the app outside Docker
The backing services still run in Docker — only the app moves to the host, which is useful for a debugger or a fast edit loop.
pip install -r requirements.txt
cp .env.example .env
docker compose up -d db redis ollama # backing services only
python app.pyThis is the one case where .env's localhost URLs apply: the host process reaches the
containers over published loopback ports, 127.0.0.1:5432 for Postgres and
127.0.0.1:11434 for Ollama. Both are bound to loopback, not 0.0.0.0.
If something already owns one of those ports — a natively installed Ollama is the usual
culprit — move the published port instead of editing docker-compose.yml:
OLLAMA_BIND_PORT=21434 docker compose up -d ollama
# then in .env: OLLAMA_BASE_URL=http://localhost:21434BIND_HOST moves every published port at once, and BIND_PORT moves the app's own.
Retrieval. Hybrid search combines an independent semantic arm (pgvector) and a lexical arm (PostgreSQL tsvector), blended by weight, then reranked by a cross-encoder. GraphRAG expands queries one hop through entity co-occurrences. Chunking preserves table structure, and PDF tables are extracted rather than flattened.
Answers. Streaming responses over SSE, with tool calling, optional live web search, and long-term memory extracted from earlier conversations. Model routing picks a model class per query; a cloud model can be configured as fallback.
Isolation and access. Documents, conversations and memories are scoped to a workspace. Access is role-based, both globally (admin or user) and per workspace (viewer, editor, owner). Workspace API keys give a bot or workflow scoped, revocable access to exactly one workspace.
Integrity. Records other data refers to are never hard-deleted. A delete sets
deleted_at; purging is a separate, admin-only operation with preconditions. See the
Clark-Wilson section in CLAUDE.md.
Sources. Document connectors for local folders, SharePoint, OneDrive, Google Drive
and webhooks. Plugins — a .py file dropped into plugins/ that registers LLM-callable
tools, per plugins/README.md. The fuller
plugin contract (services, hooks, manifest) is designed and
not yet built; its one binding rule is that dependencies point inward.
flowchart TD
Browser["Browser / API client"]
subgraph FastAPI["FastAPI application"]
Routes["Routes (APIRouters)"]
Auth["Security — JWT · rate limit · CORS"]
Pydantic["Pydantic validation + sanitization"]
RAG["RAG pipeline — retrieval · reranking"]
Tools["Tool executor — function calling"]
SSE["SSE stream"]
end
subgraph Services["External services"]
PG["PostgreSQL + pgvector"]
Ollama["Ollama — LLM · embeddings"]
Redis["Redis — cache · rate limiting"]
end
Browser -->|HTTP request| Routes
Routes --> Auth
Auth --> Pydantic
Pydantic --> RAG
RAG -->|vector search| PG
RAG -->|embed query| Ollama
Pydantic --> Tools
Tools -->|tool-call loop| Ollama
Tools --> RAG
Ollama -->|stream tokens| SSE
SSE -->|text/event-stream| Browser
Routes -.->|cache r/w| Redis
A request is parsed and validated at the boundary, authorised against the workspace, then handed to a service. Routes hold no business logic and no SQL. Blocking work — retrieval, embedding, database writes — runs in a threadpool so one slow query cannot stall other requests.
| Layer | Technology |
|---|---|
| Web framework | FastAPI + Uvicorn |
| Database | PostgreSQL 16 + pgvector (psycopg3 pool) |
| LLM | Ollama, local; LiteLLM cloud fallback |
| Embeddings | nomic-embed-text |
| Cache | Redis, with an in-memory fallback |
| Auth | PyJWT (JWT), slowapi (rate limiting) |
| Validation | Pydantic 2 |
| Migrations | Alembic |
| Tests | pytest + pytest-asyncio |
Entry point is app.py → create_app() in src/app_fastapi.py. Every module is listed in
the module index.
Full index: docs/README.md — organised by Diátaxis, so what you need depends on what you are doing.
| I want to… | Go to |
|---|---|
| Deploy this properly | Deployment, Operations |
| Connect a bot or workflow | Workspace API keys, Discord via n8n |
| Look up a setting | Configuration, RAG settings |
| Understand a decision | ADRs, Lessons learned |
| Fix something broken | Troubleshooting |
| Contribute code | CLAUDE.md and .claude/rules/ |
The API is documented live at /api/docs/ (Swagger UI), and the same documentation is
browsable inside the application under Docs.
ruff check . # lint — the whole tree, as CI does
mypy src --ignore-missing-imports # types
bandit -r src/ -ll -q -c pyproject.toml # security
pytest -m "not (slow or ollama or db)" # fast suite, no external servicesAll four must be clean before a commit; CI enforces the same set. Merging to main
requires five checks to pass — unit-tests, integration-tests, repo-hygiene,
docker-smoke and perf-canary — and is a human decision; auto-merge is off
deliberately. (docker-smoke joined the required set on 2026-08-19 and perf-canary on
2026-08-24; this sentence still said three until 2026-08-27. Read the set back with
gh api repos/jwvanderstam/LocalChat/rulesets/14700924, not from the settings UI.)
Current state: 3,327 tests collected; the fast suite runs 3,221 of them (3,199 passed,
22 skipped) at 80.9% coverage — about 12 minutes on CI, twice that on a laptop. Integration
tests need PostgreSQL; some also need Ollama, and tests/e2e/ drives a real browser. Every
number here was measured on 2026-09-17 rather than remembered — see exit criterion 7 in
PRODUCTION_PLAN.
Notable changes per release are in CHANGELOG.md.
Coding standards live in .claude/rules/: architecture, Python, testing, plugins.
Report vulnerabilities per SECURITY.md. Route-by-route permissions are documented in PERMISSIONS.md.
Two properties worth knowing when reading logs or API errors: every user-controlled value
that reaches a log record passes through sanitize_log_value, which strips control
characters, the Unicode line separators and ESC so a crafted filename cannot forge a log
record or drive the terminal of whoever reads it; and API error bodies return fixed
messages, never exception text.
MIT — see LICENSE.