fix(odd-status): recompute the verdict with the package's own script on the checkout instead of reading it from the model's answer - #52
Merged
Conversation
…on the checkout instead of reading it from the model's answer Closes #44 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
The gate is bound to the package's computation. The run's JSON block gains a
flagsarray (the flags it passed toodd_status.py --render); the verdict step runs the script the setup deployed (get-status/scripts/odd_status.pyunder the CLI's skills directory) again on the checkout and takesstatusandtodofrom that rendering, never from a line of the model's answer. What the run reports bounds the recomputation this far and no further:--service,--stack,--env) is kept only where each value is a whole word of thepromptinput; an empty prompt takes no scope; a scope that matches no stored report (the script's ownmatchedcount) is refused;--ruled,--runtime,--non-runtime) are its own judgment and are dropped; the log and the summary say how many;--repo,--repository,--today, ...) is refused;--render,--x vand--x=vinside one string are normalised.A refused flag or scope, a script not found or failing is an
errorand the step fails whateverfail-onsays. The log and the step summary printrecomputed with ...and say when the answer's own verdict line differs from the recomputed one; the summary carries the package's rendering.action.yml,odd-status/README.mdand the root catalog say so; the README states the binding holds against text, not against a run's shell.Why
Closes #44. The design and its three amendments (prompt-bounded scope, rulings out of the gate, fail closed on an empty match) are recorded on the issue.
How to test
221 tests off the runner (a fake package script under a temporary HOME; the regression case the issue asks for: the answer says
ok, the memory renderserror, exit 1). On the runner,tests/odd-status/test_runner.pygains a case assertingSTATUSequals the verdict line of a freshodd_status.py --renderon the checkout, on the real package, in every odd-status cell. Also verified by hand against the installed package:--service <name> --fullrenders,--service theis refused.Review
Review subagent: green after four rounds (the scope narrowing and the rulings, the substring bound,
--fullon the fact-sheet run). Security review: no findings after the same rounds.🤖 Generated with Claude Code