From f157265cbabae910da1c59c4486611359761b748 Mon Sep 17 00:00:00 2001 From: Jonathan Borduas Date: Fri, 21 Aug 2026 22:37:39 +0100 Subject: [PATCH] =?UTF-8?q?gated-caller=20--stability:=20ask=20whether=20t?= =?UTF-8?q?he=20verdict=20depends=20on=20the=20STUB=20=E2=80=94=20it=20doe?= =?UTF-8?q?s,=20and=20it=20hid=20a=20fourth=20undercount?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit This figure has been published as a count three times today and was a property of the stub each time: 1 of 36 the stub died on AttributeError 2 of 36 it died on TypeError 4 of 36 it survives both Chasing fidelity is unbounded and gives no signal for when to stop. Running the SAME probe at TWO fidelities does: if the answer moves, the answer depends on the instrument rather than on the repository. Run once, it immediately found a fourth undercount in a field I had not revisited since this morning: SELFTEST MOVED rich=8 poor=5 only-rich: estate-provenance.py, pane-binding.py, pretooluse-guard.py SWEPT MOVED rich=4 poor=2 CALLED / REACHED stable at 25 So "4 of 34 have a DEDICATED suite running --self-test" — the first line this tool ever printed — was also low. It is 8. This is the error term the exit-code line could not supply. Truncation risk sat at 28 of 51 across a stub change that revealed two more real callers, because an exit code says a suite ended badly and never says WHERE. Two fidelities disagreeing is a fact about what went unseen. STUB_POOR is kept deliberately: it is not dead code, it is the second fidelity. STABLE means only that these two stubs agree, never that the count is right — a third fidelity may disagree with both, and the output says so. Costs 2x, so it is opt-in and never the default. Controls assert BOTH directions: the fixture hidden at one fidelity must report MOVED, and a tree where both agree must report STABLE — otherwise it is an alarm that always fires. Measured at b20df2d. DEV1 --- tools/README.md | 12 ++++- tools/gated-caller.py | 106 +++++++++++++++++++++++++++++++++++++++++- 2 files changed, 115 insertions(+), 3 deletions(-) diff --git a/tools/README.md b/tools/README.md index ce5520e..ffb20a4 100644 --- a/tools/README.md +++ b/tools/README.md @@ -345,7 +345,7 @@ of them, which is why it is stated here rather than in a docstring. | `gh-complete.py` | is this `gh api` list reading COMPLETE, or a silent prefix of its own population? | 0 complete · 1 **TRUNCATED — the reading is a prefix** | | `reference-check.py` | which recorded reference implementations have MOVED since we recorded them? | 0 every entry current · 1 MOVED or MISSING · **2 established nothing** | | `use-not-mention.py` | does this file CALL ``, or merely TALK ABOUT calling it? | 0 no call · 1 at least one CALL · **2 established nothing** | -| `gated-caller.py` | whose `--self-test` does CI actually invoke, and whose **sweep** does? | 0 all have a gated caller · 1 one does not · **2 established nothing** | +| `gated-caller.py` | whose `--self-test` does CI actually invoke, and whose **sweep** does? — `--stability` asks whether the answer depends on the stub | 0 all have a gated caller · 1 one does not · **2 established nothing** | | `hermetic-check.py` | is a suite the gate CALLS hermetic actually hermetic? | 0 all hermetic · 1 a suite moves when `gh` is shadowed · **2 established nothing** · 3 control failed | | `population-leg.py` | does each `--self-test` consult anything outside the repository — or the forge? | 0 all do · 1 a NO-REPO-INPUT control · **2 established nothing** · ⚠ NO-REPO-INPUT is a CANDIDATE for criterion 5, not a verdict | | `pointer-verified.py` | did this pane READ the artifact a pointer NAMED, before acting? | 0 all read · 1 at least one not · **2 established nothing** · **3 control failed** | @@ -753,6 +753,16 @@ continuing on the copy it loaded. That is the difference between a trigger and a ⚠ **2026-08-20: `role_of` promised the one thing it did not deliver.** *"The role a session was BOOTSTRAPPED as — a name can be changed; this cannot"* — and it scanned the **whole file** for `You are X.`, taking the first hit anywhere. Measured over nine live transcripts: **3 resolved, 2 of the 3 wrong.** One came from a **correction sent a day later** (*"your identity was wrong … You are DEV2"*, record 17155, against a bootstrap reading MAINTAINER); one from a **quotation** of someone else's prompt; and a session bootstrapped as `DX` was reported `DEV2` because it had spent the day discussing DEV2. ⇒ It returned **the mutable thing it promised immunity from**, and a **mention** rather than a use. ★ Now anchored to the bootstrap record, with three outcomes — `None` unreadable or no launch prompt, `""` read and names no role, a role otherwise. **6 of 9 after, all from bootstraps.** ⚠ The two accepted phrasings are a **measured snapshot**, not a closed set. +⚠ **`gated-caller.py --stability` exists because this figure was published as a count three times +and was a property of the STUB each time** — `1` (the stub died on `AttributeError`), `2` (on +`TypeError`), `4` (survives both); the dedicated-self-test figure moved `5 → 8` the same way. +⛔ **Chasing stub fidelity is unbounded and gives no signal for when to stop.** ⇒ Running the SAME +probe at TWO fidelities does: **if the answer moves, the answer depends on the instrument.** ★ That +is the error term an exit code cannot supply — *measured*: truncation risk sat at `28 of 51` across +a stub change that revealed two more real callers, because an exit code says a suite ended badly +and never says WHERE. ⚠ **`STABLE` means only that these two stubs agree**, never that the count is +right; a third fidelity may disagree with both. **Costs 2×, so it is opt-in and not the default.** + **`hermetic-check.py`** — the `hermetic suites (gating)` job decides membership by the **absence of a marker**: a suite is hermetic iff nobody wrote `# SUITE-DEPENDS` in it. ⛔ **Nothing verified the claim.** A suite that shells out to `gh` is declared hermetic by default, passes on every authenticated machine, diff --git a/tools/gated-caller.py b/tools/gated-caller.py index b353d10..a0ac79b 100644 --- a/tools/gated-caller.py +++ b/tools/gated-caller.py @@ -128,7 +128,16 @@ def gated_suites(tools_dir): return out -def probe(suites, tools_dir, instruments, timeout=120): +# ⛔ THE POOR STUB IS KEPT ON PURPOSE — it is not dead code, it is the second fidelity. +# ★ Three times today this figure was published as a count and was a property of the stub: +# 1 (died on AttributeError) -> 2 (died on TypeError) -> 4 (survives both). Chasing fidelity +# is unbounded and gives no signal for when to stop. ⇒ Running the SAME probe at TWO +# fidelities does: if the answer moves, the answer depends on the instrument, and THAT is the +# error term the exit-code line could not supply. +STUB_POOR = STUB.replace(" return _S(0)\n", " return 0\n") + + +def probe(suites, tools_dir, instruments, timeout=120, stub=None): """{instrument: {"selftest": [suites], "reached": [suites]}} — behavioural, never textual. Every instrument is stubbed at once; each suite runs ONCE. The stub records its own name, so @@ -142,7 +151,7 @@ def probe(suites, tools_dir, instruments, timeout=120): shutil.copytree(tools_dir, tmp, ignore=shutil.ignore_patterns("__pycache__", "*.json")) for n in instruments: - (tmp / n).write_text(STUB) + (tmp / n).write_text(stub if stub is not None else STUB) log = Path(d) / "calls.log" env = dict(os.environ, GC_LOG=str(log)) try: @@ -279,6 +288,62 @@ def g(*a): f" property of THIS CHECKOUT, not of the repository.") +def population(tools_dir): + """The instruments census() measures — factored out so stability() asks the same question.""" + me = Path(__file__).name + out = [] + for f in sorted(Path(tools_dir).glob("*.py")): + if f.name.startswith("test_") or f.name == me: + continue + try: + src = f.read_text(encoding="utf-8", errors="replace") + except OSError: + continue + if has_selftest(src): + out.append(f.name) + return out + + +def stability(tools_dir=None, timeout=120): + """Does the verdict depend on the STUB rather than on the repository? + + ⛔ THE ERROR TERM THE EXIT-CODE LINE COULD NOT SUPPLY. A suite's exit code says it ended + badly, never WHERE — measured unchanged at 28 of 51 across a stub change that revealed two + more real callers. ⇒ Running the SAME probe at two fidelities answers the question that + actually matters: *could this method have produced the other answer?* + ★ Had this existed this morning it would have caught all three undercounts automatically — + each was a case where the poorer stub disagreed with the richer one. + """ + tools_dir = Path(tools_dir) if tools_dir else (ROOT / "tools") + suites = gated_suites(tools_dir) + names = population(tools_dir) + if not suites or not names: + return 2, [" VOID no gated suite or no instrument — established nothing"] + rich = probe(suites, tools_dir, names, timeout, stub=STUB) + poor = probe(suites, tools_dir, names, timeout, stub=STUB_POOR) + lines, moved = [], [] + for key in ("selftest", "swept", "called", "reached"): + a = {n for n in names if rich.get(n, {}).get(key)} + b = {n for n in names if poor.get(n, {}).get(key)} + if a != b: + moved.append(key) + lines.append(f" ⛔ {key.upper():<9} MOVED rich={len(a)} poor={len(b)} " + f"only-rich={sorted(a - b) or '-'} only-poor={sorted(b - a) or '-'}") + else: + lines.append(f" ok {key.upper():<9} stable at {len(a)} across both fidelities") + lines.append("") + if moved: + lines.append(f" ⛔ THE VERDICT DEPENDS ON THE STUB for {', '.join(moved)}. Every count " + f"in a normal run is a LOWER BOUND, and the sets above are the part this " + f"tool can PROVE it was missing at the poorer fidelity — not the whole of " + f"it, because a third fidelity may move it again.") + else: + lines.append(" ⚠ STABLE ACROSS THESE TWO FIDELITIES ONLY. That is not 'correct' — it " + "is 'these two stubs agree'. A richer stub may still disagree with both.") + lines.append(tree_provenance()) + return (1 if moved else 0), lines + + def census(tools_dir=None, timeout=120): tools_dir = tools_dir or (ROOT / "tools") if not tools_dir.is_dir(): @@ -592,6 +657,30 @@ def self_test(): f"survives to reach its later subprocess run (swept={rl['late.py']['swept']}) — " f"with a bare 0 it died at [0] and the sweep was invisible") + # ⛔ --stability MUST BE ABLE TO SAY BOTH THINGS, or it is an alarm that always fires. + # ⇒ The `late.py` fixture above is hidden by the poor stub and seen by the rich one, so + # it MUST report MOVED. A tree without it must report stable. + rc_m, lines_m = stability(tools_dir=td, timeout=60) + moved_ok = rc_m == 1 and any("SWEPT" in l and "MOVED" in l for l in lines_m) + ok &= moved_ok + print(f" {'ok ' if moved_ok else 'FAIL'} --stability REPORTS MOVED when a caller is " + f"visible at one fidelity and not the other (rc={rc_m}) — the case it exists for") + + with tempfile.TemporaryDirectory() as d2: + st = Path(d2) / "tools" + st.mkdir() + (st / "plain.py").write_text('import sys\nif "--self-test" in sys.argv:\n' + ' raise SystemExit(0)\nraise SystemExit(0)\n') + (st / "test_plain.py").write_text( + "import subprocess, sys, os\n" + "here = os.path.dirname(os.path.abspath(__file__))\n" + "subprocess.run([sys.executable, os.path.join(here,'plain.py'), '--self-test'])\n" + "raise SystemExit(0)\n") + rc_s, _ = stability(tools_dir=st, timeout=60) + ok &= rc_s == 0 + print(f" {'ok ' if rc_s == 0 else 'FAIL'} --stability reports STABLE when the two " + f"fidelities agree (rc={rc_s}) — it is not an alarm that always fires") + # ⛔ TRUNCATION RISK MUST BE COUNTED, AND MUST NOT COUNT A CLEAN SUITE. Asserted as a # pair: "a dying suite is flagged" passes if every suite were flagged; "a clean suite is # not" passes if none were. @@ -730,6 +819,9 @@ def main(argv): ap = argparse.ArgumentParser(description=__doc__.split("\n")[0]) ap.add_argument("--self-test", action="store_true") ap.add_argument("--timeout", type=int, default=120) + ap.add_argument("--stability", action="store_true", + help="run the SAME probe at two stub fidelities and report whether the " + "verdict moves; 0 stable, 1 the answer depends on the stub. Costs 2x.") try: a = ap.parse_args(argv[1:]) except SystemExit as e: @@ -747,6 +839,16 @@ def main(argv): return 2 if a.self_test: return self_test() + if a.stability: + # ⇒ NOT the default. It doubles the runtime and answers a question ABOUT the instrument, + # not about the repository — but it is the only honest source of the error term. + rc, lines = stability(timeout=a.timeout) + print("\ngated caller — does the verdict depend on the STUB? (#551)") + for l in lines: + print(l) + print({0: " STABLE across the two fidelities tested", + 1: " FINDING — the answer moved", 2: " VOID"}[rc]) + return rc rc, lines = census(timeout=a.timeout) print("\ngated caller — whose --self-test does CI actually invoke? (#372, #381)") for l in lines: