The subagent prompt in references/extraction-spec.md (unchanged through
v0.9.72) says:
Semantic similarity: if two concepts in this chunk solve the same problem ...
Chunks are dispatched in parallel and no subagent sees another's output, so a
similarity whose two ends land in different chunks cannot be emitted by
construction. Agents are not missing these edges, they are unable to produce
them, and each agent's output still reads as having followed the instruction.
Related: #65 (same diagnosis; closed after v0.3.17 added same-directory
grouping in Step B1, while "a full global linking pass remains on the table"),
#7 (local embeddings, open) and its PR #1126 (closed without merging), #537.
This report adds a concrete case from the current pipeline and what worked when
I built the controller-side pass.
Concrete case
A project's most consequential design rule is stated in two different data files
(two different directories, so same-directory grouping does not help). Both were
extracted correctly as rationale nodes, but in different chunks. Each has one
edge, to its own file; they are not connected to each other and ended up in
different communities. The graph holds the same fact twice and does not know it.
The more the corpus is split, the more of the cross-cutting edges fall between
chunks, so more parallelism means a sparser semantic layer.
Proposal
- Reconciliation pass after Step B3 collection, before Part C. With all
chunks in hand, compare rationale/concept node text across different
source files, and emit semantically_similar_to edges marked INFERRED with a
score derived from the measured similarity. Report the candidate pairs rather
than writing silently, since this is inference over content the controller
did not author.
- Amend the spec wording so agents know the cross-chunk case is handled
elsewhere.
What I learned implementing (1)
I checked the pass against the known pair above, chosen before picking a
metric:
- Word-shingle Jaccard found 5 pairs and missed the known pair. Shingles
measure near-verbatim copying; two texts stating one rule in different words
share almost no trigrams.
- TF-IDF cosine found 32 pairs and still missed it (0.371 against a 0.42
threshold). Both texts are long and surround the shared core with their own
material, which dilutes the cosine.
- What identified it was the number of rare terms in common (nine
low-frequency terms). Accepting a pair on either a strong cosine or a high
rare-term overlap with a moderate cosine caught it and added 12 more true
pairs, with no false positives on manual review.
Two things seem worth carrying into any implementation, embeddings included:
score paraphrase by shared low-frequency vocabulary, not string overlap, with a
second acceptance path for long texts whose similarity is concentrated; and
validate against a known pair fixed in advance, because 5 pairs and 32 pairs
both looked like success.
Environment: observed on graphifyy 0.9.55, graphify install --platform claude,
Claude Code on Windows 11; now on 0.9.65. Spec wording and absence of a linking
pass checked in v0.9.72.
The subagent prompt in
references/extraction-spec.md(unchanged throughv0.9.72) says:
Chunks are dispatched in parallel and no subagent sees another's output, so a
similarity whose two ends land in different chunks cannot be emitted by
construction. Agents are not missing these edges, they are unable to produce
them, and each agent's output still reads as having followed the instruction.
Related: #65 (same diagnosis; closed after v0.3.17 added same-directory
grouping in Step B1, while "a full global linking pass remains on the table"),
#7 (local embeddings, open) and its PR #1126 (closed without merging), #537.
This report adds a concrete case from the current pipeline and what worked when
I built the controller-side pass.
Concrete case
A project's most consequential design rule is stated in two different data files
(two different directories, so same-directory grouping does not help). Both were
extracted correctly as
rationalenodes, but in different chunks. Each has oneedge, to its own file; they are not connected to each other and ended up in
different communities. The graph holds the same fact twice and does not know it.
The more the corpus is split, the more of the cross-cutting edges fall between
chunks, so more parallelism means a sparser semantic layer.
Proposal
chunks in hand, compare
rationale/conceptnode text across differentsource files, and emit
semantically_similar_toedges markedINFERREDwith ascore derived from the measured similarity. Report the candidate pairs rather
than writing silently, since this is inference over content the controller
did not author.
elsewhere.
What I learned implementing (1)
I checked the pass against the known pair above, chosen before picking a
metric:
measure near-verbatim copying; two texts stating one rule in different words
share almost no trigrams.
threshold). Both texts are long and surround the shared core with their own
material, which dilutes the cosine.
low-frequency terms). Accepting a pair on either a strong cosine or a high
rare-term overlap with a moderate cosine caught it and added 12 more true
pairs, with no false positives on manual review.
Two things seem worth carrying into any implementation, embeddings included:
score paraphrase by shared low-frequency vocabulary, not string overlap, with a
second acceptance path for long texts whose similarity is concentrated; and
validate against a known pair fixed in advance, because 5 pairs and 32 pairs
both looked like success.
Environment: observed on graphifyy 0.9.55,
graphify install --platform claude,Claude Code on Windows 11; now on 0.9.65. Spec wording and absence of a linking
pass checked in v0.9.72.