Skip to content

Lazy ingest on first query + watch mode (incremental ingest and freshness already shipped) #175

Description

@zaebee

Summary

Make ingest incremental / automatic so users don't have to manually rebuild graph.db.

Motivation

Today the workflow requires the user (or agent) to remember to run cgis_ingest before any query, and to re-run it after edits or the graph is stale. In practice this is the main day-to-day friction:

  • An agent that calls trace_flow before ingest gets nothing.
  • After editing files, every query silently runs against a stale graph until a full re-ingest.
  • Full re-ingest of a large repo on every change is wasteful (the FastAPI backend was 10k nodes / 38k edges).

Proposal

Two parts, either independently useful:

  1. Incremental ingest — cgis_ingest detects changed files (mtime/hash vs a stored manifest) and re-extracts only those, patching the graph. Full rebuild only when the schema/version changes.

  2. Auto/lazy ingest — query tools (trace_flow, analyze_impact, drift, …) transparently ensure the graph exists and is fresh before running:

    • if db_path missing → ingest first (with a one-line notice)
    • if stale (changed files since last ingest) → incremental update first
    • optional watch mode that keeps the graph live as files change

Bonus: have query tools return a small graph_age / stale_files: N field so callers know freshness.

Impact

Removes the "did you ingest first?" footgun and the stale-graph trap — the biggest barriers to making cgis a default, always-on tool rather than a manual step. Critical for distribution (MCP server, CI action, IDE).

Activity

zaebee commented on Sep 9, 2026

@zaebee
OwnerAuthor

Part 1 shipped; this issue now tracks part 2 only. Checked against main at 52e0de1.

cgis ingest --incremental / -i exists: _process_file skips a file whose content hash still matches (files_state), _persist_incremental writes back only re-parsed files and deletes nodes of files that disappeared, and _is_noop_incremental exits early when nothing changed.

Part 2 has not shipped, and there is now direct evidence it should

stale_files is computed inside the pipeline during ingest (pipeline.py:201, 229) and never reaches a query. No tool returns a freshness field, so every caller still guesses.

The sharper evidence is that the same need has been solved twice, locally, in the last two merged PRs. OrphanReport now carries two hand-rolled staleness signals, each documented as "zero here means re-ingest":

Both exist because a query cannot ask "is this graph current?". A third column would mean a third such counter. That is the argument for a general graph_age / stale_files field rather than per-report proxies — and the proxies can then say what they actually mean instead of doubling as freshness probes.

Related friction from the same sessions: a graph built before a schema migration reads a silently-false column until re-ingest, and the is_generated migration had to blank files_state hashes precisely so that "re-ingest" was advice a user could follow.

Scope for part 2

  1. a freshness probe on the store — changed/stale file counts against the working tree, cheap enough to run per query;
  2. that count surfaced in query output (CLI and MCP), so a caller knows what it is reading;
  3. lazy ingest — missing db, or stale, triggers an incremental update with a one-line notice.

(1) and (2) are independently useful and are what the two ad-hoc counters are standing in for; (3) is the ergonomics win.

zaebee commented on Sep 9, 2026

@zaebee
OwnerAuthor

Part 2 shipped in #443 (0d6d197); this issue now tracks lazy ingest and watch mode only.

ingest_state records the root and ingest time; SQLiteStore.freshness() stats the tracked files and their directories — never walking — and returns FRESH / STALE / UNKNOWN. Every CLI read command and the graph-reading MCP tools carry it, and say nothing when the graph is fresh.

Two things this turned out to need that the issue did not anticipate

Three states, not two. UNKNOWN — no ingest record, or a root that has moved — is reported distinctly rather than folded into STALE or FRESH. A signal that reads the same when all is well and when it could not look is the failure #441 had just found in generated_excluded.

The signal's channel matters as much as its content. On stdout the CLI note broke thirteen tests at once (--format json is piped); it goes to stderr. In MCP the same hazard splits by return type — a freshness key inside JSON objects, a prose note for text tools, a YAML comment for cgis_init_ontology. cgis_find_symbol returns a JSON list with nowhere to put a key, so it deliberately carries no signal; giving it one means changing its documented shape.

Measured

probe on owner-api (814 files, 102 dirs) 6.8 ms, fresh and stale alike
os.walk + stat over the same tree 6.8 ms

Not walking is a correctness choice, not a performance one — the two cost the same. A walk would have to reproduce IngestionPipeline's directory exclusions, so that set now lives once in core/paths.EXCLUDED_DIRS, shared by both.

Still open here

  • Lazy / auto ingest. Making a read command write the graph is a behavioural change to every tool and deserves its own decision.
  • Watch mode.

And a follow-up this enables

OrphanReport.test_sources and .generated_excluded were only ever freshness proxies — that is the argument in this issue's earlier comment. They can now say what they actually mean, but changing what two shipped fields report is its own reviewable diff.

changed the title [-]Incremental / auto ingest: keep graph.db fresh without manual re-ingest[/-] [+]Lazy ingest on first query + watch mode (incremental ingest and freshness already shipped)[/+] on Sep 16, 2026

zaebee commented on Sep 16, 2026

@zaebee
OwnerAuthor

Retitled to the remainder: --incremental and the freshness signal (#443) shipped; open are lazy/auto ingest on first query and watch mode. Note #38: incremental ingest currently leaves stale edges, which matters before making it automatic.

zaebee commented on Oct 2, 2026

@zaebee
OwnerAuthor

Decision and what shipped. This will close with #541.

Measured on this repo (6.5k nodes / 28.6k edges): the freshness probe costs about 3 ms. An incremental ingest takes 0.7 s after a body edit and 5.6 s after a signature edit, which forces a full rebuild.


Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions