You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
extraction-spec has no rule for referencing nodes outside the chunk: agents copy id fragments or mint stub nodes, and replace-on-re-extract deletes real content #3947
The subagent prompt in references/extraction-spec.md says how to build ids for
entities in the chunk's own files, and the merge in references/update.md relies
on source_file for replace-on-re-extract. The spec's source_file RULE fixes
the spelling of the value ("EXACTLY as it appears in FILE_LIST"), but nothing
tells a subagent what to do when a file in its chunk refers to an entity that
lives in a file outside its FILE_LIST, and nothing on the merge side
enforces membership. In one --update run that gap produced two separate
silent failures.
Related: #3253 and #537 (subagents can only relate files inside their own FILE_LIST), #1895 / #2217 / #2260 (out-of-scope source_file, guarded in the
LLM backend path but not in the skill's subagent path), #3004 and #1711
(replace-on-re-extract losing content by other routes).
Failure 1: an id prefix is used as a complete id
To let chunks link to nodes owned by other chunks, the dispatching agent listed
the id prefixes those nodes share (the stem part of {stem}_{entity}).
Several subagents used the prefix verbatim as a target id. 115 edges pointed
at ids that do not and cannot exist. They validate as well-formed strings and
only disappear when the graph is built, as dangling endpoints.
The spec's id rule invites this: it defines ids as stem plus entity, so a stem
reads like a usable id. Step 4.5's health check reports dangling endpoints after
the fact, but by then the edges are already gone.
Failure 2: a stub node for a foreign file triggers destructive replace
Another subagent, asked only to reference three documents handled elsewhere,
defined its own stub node for each so that its edges would not dangle, with source_file set to those documents. That is locally reasonable. But build_merge drops every existing node whose source_file appears in the
incoming extraction, so the merge would have deleted 23 rich nodes extracted
earlier for those three documents and replaced them with three stubs.
Nothing would have reported it: node count still rises, the health check passes.
It was caught by reading the subagent's summary, not by any check. The skill's
Part B3 merge and the build_merge call in update.md do not compare incoming source_file values against the files dispatched this run. The cache write does
(allowed_source_files=uncached), which is the check #1895 describes for the
backend path.
Expected
A subagent should never be able to change nodes for a file it was not given, and
a cross-chunk reference should either resolve to a real id or be reported, not
silently dropped.
Proposal
references/extraction-spec.md, prompt text:
Never create a node whose source_file is not in FILE_LIST. A reference to
something outside the chunk is expressed as an edge only.
A dangling cross-chunk reference is acceptable; the controller resolves it.
If the prompt ever gives example ids, give complete ids, and state that a stem
or prefix on its own is never a valid id.
SKILL.md Step B3 and references/update.md, before build_merge:
Optionally, resolve unresolved edge targets centrally against the full id
space (existing graph + all chunks), where it is known, instead of relying on
each agent's reconstruction of ids it cannot see.
references/update.md, prose: state that replace-on-re-extract makes source_file a destructive key, so every node's source_file must name the
file the node actually came from.
Environment: observed on graphifyy 0.9.55, graphify install --platform claude,
Claude Code on Windows 11; now on 0.9.65. Spec and merge flow checked in
v0.9.72 (both files unchanged since v0.9.61): no such rule or check. #2260,
still open, would add the dispatched-set check to the LLM backend path only.
The subagent prompt in
references/extraction-spec.mdsays how to build ids forentities in the chunk's own files, and the merge in
references/update.mdrelieson
source_filefor replace-on-re-extract. The spec'ssource_file RULEfixesthe spelling of the value ("EXACTLY as it appears in FILE_LIST"), but nothing
tells a subagent what to do when a file in its chunk refers to an entity that
lives in a file outside its
FILE_LIST, and nothing on the merge sideenforces membership. In one
--updaterun that gap produced two separatesilent failures.
Related: #3253 and #537 (subagents can only relate files inside their own
FILE_LIST), #1895 / #2217 / #2260 (out-of-scopesource_file, guarded in theLLM backend path but not in the skill's subagent path), #3004 and #1711
(replace-on-re-extract losing content by other routes).
Failure 1: an id prefix is used as a complete id
To let chunks link to nodes owned by other chunks, the dispatching agent listed
the id prefixes those nodes share (the stem part of
{stem}_{entity}).Several subagents used the prefix verbatim as a target id. 115 edges pointed
at ids that do not and cannot exist. They validate as well-formed strings and
only disappear when the graph is built, as dangling endpoints.
The spec's id rule invites this: it defines ids as stem plus entity, so a stem
reads like a usable id. Step 4.5's health check reports dangling endpoints after
the fact, but by then the edges are already gone.
Failure 2: a stub node for a foreign file triggers destructive replace
Another subagent, asked only to reference three documents handled elsewhere,
defined its own stub node for each so that its edges would not dangle, with
source_fileset to those documents. That is locally reasonable. Butbuild_mergedrops every existing node whosesource_fileappears in theincoming extraction, so the merge would have deleted 23 rich nodes extracted
earlier for those three documents and replaced them with three stubs.
Nothing would have reported it: node count still rises, the health check passes.
It was caught by reading the subagent's summary, not by any check. The skill's
Part B3 merge and the
build_mergecall inupdate.mddo not compare incomingsource_filevalues against the files dispatched this run. The cache write does(
allowed_source_files=uncached), which is the check #1895 describes for thebackend path.
Expected
A subagent should never be able to change nodes for a file it was not given, and
a cross-chunk reference should either resolve to a real id or be reported, not
silently dropped.
Proposal
references/extraction-spec.md, prompt text:source_fileis not inFILE_LIST. A reference tosomething outside the chunk is expressed as an edge only.
or prefix on its own is never a valid id.
SKILL.mdStep B3 andreferences/update.md, beforebuild_merge:hyperedge whose
source_fileis not in the set dispatched this run, and printhow many were dropped per file. The same invariant Semantic extraction: nonexistent model-invented source_file bypasses #1895 graph filter #2217 proposes for the
backend path.
space (existing graph + all chunks), where it is known, instead of relying on
each agent's reconstruction of ids it cannot see.
references/update.md, prose: state that replace-on-re-extract makessource_filea destructive key, so every node'ssource_filemust name thefile the node actually came from.
Environment: observed on graphifyy 0.9.55,
graphify install --platform claude,Claude Code on Windows 11; now on 0.9.65. Spec and merge flow checked in
v0.9.72 (both files unchanged since v0.9.61): no such rule or check. #2260,
still open, would add the dispatched-set check to the LLM backend path only.