cgis_metrics gained --exclude in #234 and --scope in #239. cgis_drift got neither, and on a repo with tests beside the code that turns the hygiene gate into a measurement of something other than architecture.
Measured on Ownima owner-api (2026-10-02)
Ingest root ownima-backend, 19 303 nodes / 96 805 edges, graph healthy (unresolved_ratio 0.1297 against a 0.3 threshold). Ontology: the repo's own docs/ontology/patterns.yaml, 17 domains.
Result: any_critical: true, 13 of 17 domains gate_failed. But only 4 exceed their own drift_tolerance:
| domain |
drift |
tolerance |
over? |
scripts |
0.229 |
0.11 |
yes, 2× |
common |
0.556 |
0.53 |
yes |
core |
0.424 |
0.39 |
yes |
mcp |
0.333 |
0.32 |
yes |
The other nine fail only on hygiene unresolved_ratio > 0.2. Here is what those unresolved edges actually are, from cgis_validate's top_unresolved on the same graph:
response.json 497
self.client.post 493
self.client.get 349
logger.info 279
logger.exception 212
logger.warning 181
resp.json 143
r.json 113
Two populations, neither architectural:
logger.* — logger = structlog.get_logger(__name__), a bound logger. Unresolvable by construction, and a domain's share of these is a function of how much it logs.
self.client.post / response.json — the test HTTP client. These exist only because tests are in the graph.
So the per-domain unresolved_ratio ranks domains by test volume and logging density. The node counts show the contamination directly: reservation 1130 nodes, finance 599, vehicle 485, notifications 476, chat 463 — these are mostly test files. app.vocabulary reports unresolved_ratio 0.50 on 39 nodes; app.worker 0.40.
Second symptom: the fit band goes to pure_utility for layered domains
Six domains declared layered_dag or dispatcher come back with pure_utility as fit.nearest_template:
| domain |
expected |
nearest |
residual |
band |
reservation |
layered_dag |
pure_utility |
0.351 |
weak |
vehicle |
layered_dag |
pure_utility |
0.470 |
none |
chat |
layered_dag |
pure_utility |
0.280 |
weak |
finance |
layered_dag |
pure_utility |
0.466 |
none |
notifications |
dispatcher |
pure_utility |
0.284 |
weak |
crud |
pure_utility |
layered_dag |
0.470 |
none |
Test modules are flat and leafy — each test function calls into the app and is called by nobody — so they pull the motif distribution toward pure_utility. When six clearly-layered service domains all land nearest to "pure utility", the basis is describing the test suite, not the design.
Ask
exclude: list[str] and scope: list[str] on cgis_drift, with the same dot-segment / dot-prefix semantics as cgis_metrics so one ontology can be scored the same way twice.
Two notes on semantics, since drift differs from metrics here:
- Exclusion must happen before the motif census and before
unresolved_ratio, not as a row filter on the output — both are aggregates over the domain's subgraph, so filtering afterwards changes nothing.
- Node counts should be reported post-exclusion in
actual, so a reader can see what was scored. Right now node_count: 1130 for reservation reads as a domain of 1130 production symbols.
Why it matters beyond cosmetics
A gate whose failures are dominated by an artifact teaches people to ignore it. We cannot put cgis_drift in CI in this state: 13 red domains, 9 of them red for reasons no code change would fix. With exclusion we would gate on the 4 that are real.
Related: #234 (exclude in metrics), #239 (scope in metrics), #178 (a mis-targeted ontology reading clean — the inverse failure), #170 (status semantics / hygiene gate).
cgis_metricsgained--excludein #234 and--scopein #239.cgis_driftgot neither, and on a repo with tests beside the code that turns the hygiene gate into a measurement of something other than architecture.Measured on Ownima owner-api (2026-10-02)
Ingest root
ownima-backend, 19 303 nodes / 96 805 edges, graph healthy (unresolved_ratio0.1297 against a 0.3 threshold). Ontology: the repo's owndocs/ontology/patterns.yaml, 17 domains.Result:
any_critical: true, 13 of 17 domainsgate_failed. But only 4 exceed their owndrift_tolerance:scriptscommoncoremcpThe other nine fail only on
hygiene unresolved_ratio > 0.2. Here is what those unresolved edges actually are, fromcgis_validate'stop_unresolvedon the same graph:Two populations, neither architectural:
logger.*—logger = structlog.get_logger(__name__), a bound logger. Unresolvable by construction, and a domain's share of these is a function of how much it logs.self.client.post/response.json— the test HTTP client. These exist only because tests are in the graph.So the per-domain
unresolved_ratioranks domains by test volume and logging density. The node counts show the contamination directly:reservation1130 nodes,finance599,vehicle485,notifications476,chat463 — these are mostly test files.app.vocabularyreportsunresolved_ratio0.50 on 39 nodes;app.worker0.40.Second symptom: the fit band goes to
pure_utilityfor layered domainsSix domains declared
layered_dagordispatchercome back withpure_utilityasfit.nearest_template:reservationvehiclechatfinancenotificationscrudTest modules are flat and leafy — each test function calls into the app and is called by nobody — so they pull the motif distribution toward
pure_utility. When six clearly-layered service domains all land nearest to "pure utility", the basis is describing the test suite, not the design.Ask
exclude: list[str]andscope: list[str]oncgis_drift, with the same dot-segment / dot-prefix semantics ascgis_metricsso one ontology can be scored the same way twice.Two notes on semantics, since drift differs from metrics here:
unresolved_ratio, not as a row filter on the output — both are aggregates over the domain's subgraph, so filtering afterwards changes nothing.actual, so a reader can see what was scored. Right nownode_count: 1130forreservationreads as a domain of 1130 production symbols.Why it matters beyond cosmetics
A gate whose failures are dominated by an artifact teaches people to ignore it. We cannot put
cgis_driftin CI in this state: 13 red domains, 9 of them red for reasons no code change would fix. With exclusion we would gate on the 4 that are real.Related: #234 (exclude in metrics), #239 (scope in metrics), #178 (a mis-targeted ontology reading clean — the inverse failure), #170 (status semantics / hygiene gate).