Repository navigation
Lazy ingest on first query + watch mode (incremental ingest and freshness already shipped) #175
Description
Activity
Part 1 shipped; this issue now tracks part 2 only. Checked against main at 52e0de1.
cgis ingest --incremental / -i exists: _process_file skips a file whose content hash still matches (files_state), _persist_incremental writes back only re-parsed files and deletes nodes of files that disappeared, and _is_noop_incremental exits early when nothing changed.
Part 2 has not shipped, and there is now direct evidence it should
stale_files is computed inside the pipeline during ingest (pipeline.py:201, 229) and never reaches a query. No tool returns a freshness field, so every caller still guesses.
The sharper evidence is that the same need has been solved twice, locally, in the last two merged PRs. OrphanReport now carries two hand-rolled staleness signals, each documented as "zero here means re-ingest":
test_sources(feat(audit): orphan-class query — nodes nothing builds or extends (needs an annotation edge to be precise) #415) — zero in a repo that has tests means the graph predates theis_testcolumn;generated_excluded(cgis orphans: repo-declared allowlist + nested-class scope (generated-stub filter shipped) #432/feat(orphans): generated classes are noise, not findings — hide them by default (#432) #441) — zero in a repo with generated code means it predatesis_generated.
Both exist because a query cannot ask "is this graph current?". A third column would mean a third such counter. That is the argument for a general graph_age / stale_files field rather than per-report proxies — and the proxies can then say what they actually mean instead of doubling as freshness probes.
Related friction from the same sessions: a graph built before a schema migration reads a silently-false column until re-ingest, and the is_generated migration had to blank files_state hashes precisely so that "re-ingest" was advice a user could follow.
Scope for part 2
- a freshness probe on the store — changed/stale file counts against the working tree, cheap enough to run per query;
- that count surfaced in query output (CLI and MCP), so a caller knows what it is reading;
- lazy ingest — missing db, or stale, triggers an incremental update with a one-line notice.
(1) and (2) are independently useful and are what the two ad-hoc counters are standing in for; (3) is the ergonomics win.
Part 2 shipped in #443 (0d6d197); this issue now tracks lazy ingest and watch mode only.
ingest_state records the root and ingest time; SQLiteStore.freshness() stats the tracked files and their directories — never walking — and returns FRESH / STALE / UNKNOWN. Every CLI read command and the graph-reading MCP tools carry it, and say nothing when the graph is fresh.
Two things this turned out to need that the issue did not anticipate
Three states, not two. UNKNOWN — no ingest record, or a root that has moved — is reported distinctly rather than folded into STALE or FRESH. A signal that reads the same when all is well and when it could not look is the failure #441 had just found in generated_excluded.
The signal's channel matters as much as its content. On stdout the CLI note broke thirteen tests at once (--format json is piped); it goes to stderr. In MCP the same hazard splits by return type — a freshness key inside JSON objects, a prose note for text tools, a YAML comment for cgis_init_ontology. cgis_find_symbol returns a JSON list with nowhere to put a key, so it deliberately carries no signal; giving it one means changing its documented shape.
Measured
| probe on owner-api (814 files, 102 dirs) | 6.8 ms, fresh and stale alike |
os.walk + stat over the same tree |
6.8 ms |
Not walking is a correctness choice, not a performance one — the two cost the same. A walk would have to reproduce IngestionPipeline's directory exclusions, so that set now lives once in core/paths.EXCLUDED_DIRS, shared by both.
Still open here
- Lazy / auto ingest. Making a read command write the graph is a behavioural change to every tool and deserves its own decision.
- Watch mode.
And a follow-up this enables
OrphanReport.test_sources and .generated_excluded were only ever freshness proxies — that is the argument in this issue's earlier comment. They can now say what they actually mean, but changing what two shipped fields report is its own reviewable diff.
Decision and what shipped. This will close with #541.
Measured on this repo (6.5k nodes / 28.6k edges): the freshness probe costs about 3 ms. An incremental ingest takes 0.7 s after a body edit and 5.6 s after a signature edit, which forces a full rebuild.
- fix(freshness): report a write made during ingest as stale #535: fixes an ingest race. A file written after the pipeline read it used to become the freshness mark itself, so the graph read FRESH while built from old text. mtimes are now observed before each read.
- feat(mcp): opt-in refresh of a stale graph before read tools answer #538: opt-in lazy refresh. With
CGIS_AUTO_REFRESH=1, every graph-reading MCP tool runs an incremental ingest first when the graph is STALE. One refresh at a time, with freshness re-read under the lock. The refresh reuses the recorded--source-root/--domains(otherwise it would rename every node). A missing database is never created. - Watch mode: not doing it. An agent's multi-file edit is a burst of saves, and each signature edit is a full rebuild; lazy refresh pays once, when something reads. docs(onboarding): refresh the graph from hooks instead of a watcher #541 documents hook recipes instead (Claude Code
Stophook, gitpost-checkout/post-merge/post-rewrite). - Split out: perf(pipeline): cache extracted nodes and raw edges per file hash so a cross-file rebuild skips re-extraction #539 caches raw edges per file so a cross-file rebuild skips re-extraction (59% of rebuild time). Decide: turn CGIS_AUTO_REFRESH on by default #540 decides whether to turn auto-refresh on by default after a release.
Generated by Claude Code
Summary
Make ingest incremental / automatic so users don't have to manually rebuild
graph.db.Motivation
Today the workflow requires the user (or agent) to remember to run
cgis_ingestbefore any query, and to re-run it after edits or the graph is stale. In practice this is the main day-to-day friction:trace_flowbeforeingestgets nothing.Proposal
Two parts, either independently useful:
Incremental ingest —
cgis_ingestdetects changed files (mtime/hash vs a stored manifest) and re-extracts only those, patching the graph. Full rebuild only when the schema/version changes.Auto/lazy ingest — query tools (
trace_flow,analyze_impact,drift, …) transparently ensure the graph exists and is fresh before running:db_pathmissing → ingest first (with a one-line notice)watchmode that keeps the graph live as files changeBonus: have query tools return a small
graph_age/stale_files: Nfield so callers know freshness.Impact
Removes the "did you ingest first?" footgun and the stale-graph trap — the biggest barriers to making cgis a default, always-on tool rather than a manual step. Critical for distribution (MCP server, CI action, IDE).