plexus pick: choose a tool for a request, or abstain - #24
Merged
Merged
Conversation
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scores every tool with BM25 over a document built from its manifest, turns scores into probabilities, and abstains under a threshold calibrated on a dev half of 100 independently written labels. On the 51-item test half P1, P2 and P3 pass and the shuffled-label control fails P3 as required. Numbers in docs/PICK.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds
plexus pick "REQUEST": a probability for every tool in the mesh, scored against what each manifest emits and consumes, and an abstain outcome when no tool fits or the top probability is under--threshold(default 0.23, the calibrated value).Bar, committed before the picker (commit cecd66c)
Labels: 100 requests from a separate labeller who saw only tool names, capability titles and a one-line purpose. Threshold calibrated on a seeded dev half; numbers from the 51-item test half. Baseline: dev majority tool.
Results
P1, P2 and P3 pass; the control fails P3 as required. 4 in 10 picks are still wrong, which
docs/PICK.mdstates.Tests
tests/test_plexus_pick.py, 11 tests in genuine-and-mutated pairs. Four source mutations (abstain rule, tool document, tokenizer, control shuffle) each fail a test. Full suite passes locally. CLI only; the MCP tool list is unchanged.🤖 Generated with Claude Code