Skip to content

Scope research skills to requested work and preserve authorized continuation - #62

Closed
nzy1997 wants to merge 2 commits into
mainfrom
improve/skill-scope-and-continuation
Closed

nzy1997 wants to merge 2 commits into
mainfrom
improve/skill-scope-and-continuation

Conversation

@nzy1997

@nzy1997 nzy1997 commented Sep 14, 2026 •

Copy link
Copy Markdown
Collaborator

Local edits, status questions, and already-scoped research tasks could trigger full workflows, repeated approval menus, and unrelated context loading. This change routes skills by the requested outcome while retaining shared interaction defaults: a grammar-only review stays local, an autoresearch status request stays read-only, and an authorized survey/report continues through its requested deliverable. New exploration and manuscript planning retain guidance; broad paper polish defaults to a marked proposal before application.

The design follows OpenAI’s Rethinking skills and prompts for GPT-6 Astra, especially precise decision boundaries, progressive disclosure, and explicit completion criteria. The README also records this reference.

Changes

  • Reuse supplied scope, sources, formats, narratives, and authorization across brainstorming, writing, review, indexing, and report handoffs. Complete full-text KB requests without adding an unrequested report, and limit targeted fact checks to the selected evidence.
  • Preserve model-independent interaction defaults: Socratic exploration with optional advisor discovery; marked proposals for broad revision/polish through either manuscript entry point; and initial venue, story, and outline checkpoints for new manuscripts. Prior choices, approved changes, scoped requests, and explicitly delegated drafting continue without duplicate approval. Story and outline can be approved together.
  • Move advisor/history instructions, download dependencies and optional acquisition/recovery details, figure/notation guidance, and legacy KB migration into task-specific references. Keep skill names and independently installed resource paths stable.
  • Remove blanket sentence-length and logical-connective replacements that conflicted with the shared writing guide. Preserve scientific meaning, verified citations, author decisions, sealed holdouts, acceptance gates, and attempt budgets. Goal relaxation requires an explicit user decision.
  • Fix the documented compile check so a nonzero compiler exit cannot appear clean merely because its diagnostic lacks the word “warning”.

Validation

  • python3 scripts/validate_skills.py: all 16 skills pass; skill-creator quick_validate.py also passes for all 16.
  • python3 -m pytest -q: 261 passed, 1 skipped.
  • New tests follow relative references from independently copied skill folders and execute the documented compile snippet with success, warning, error, and unavailable-compiler responses, including temporary-log cleanup.
  • npm pack --dry-run --json: all new skill references included.
  • git diff --check: clean.
  • Independent forward checks completed a grammar-only excerpt, a campaign-status answer with missing evidence, and a specified geometric-sum derivation without scenario-user questions. The campaign state and manuscript fixture were unchanged. The derivation’s ordinary conversation log was intentionally omitted by the fixture’s read-only constraint.
  • Independent implementation review identified two scope leaks in handoffs; both were corrected and rechecked. The follow-up review also caught broad polish bypassing the preview policy through write-paper; that route now delegates to review-paper and was rechecked.
  • Actual local Claude Code 2.1.209 calls reported claude-opus-4-7. Paired single-response checks against baseline 905ea3e exercised advisor discovery, a specified derivation, paper polish, accepted findings, and guided versus preapproved drafting. A toy training-paper fixture was not counted as a valid new-research checkpoint test; a completed computational-study fixture retained venue/story discussion in the revised version.
  • Isolated file checks compared broad English polish with already accepted edits. An intermediate revision changed the original during polish; the application-mode table was clarified. The retest preserved the original, and the accepted-edits control made exactly the two requested replacements without another approval. After adding an unavailable-tool fallback, the final polish run created main.proposed.md, left main.md byte-for-byte unchanged, and delivered numbered changes for acceptance; no shell diff or render was claimed. Content quality of the suggestions was not scored as a pass.

Limits

These are instruction and reference changes; executable research helpers and public skill names are unchanged. The forward checks are bounded examples, not broad model benchmarks or proof of Claude-wide non-regression. Claude ran with customizations disabled and synthetic data; the first-response checks had no tools, and the file checks had only Read/Write/Edit. They do not exercise normal plugin discovery, full literature acquisition, compilers, or a complete manuscript lifecycle. The samples also contain content errors: one revised derivation omitted a minus sign in an optional derivative, and both preapproved teaching drafts overgeneralized the noncontraction case by overlooking the zero initial condition. These are recorded as failures of generated content, not passed scientific validation or evidence that the skill change caused a regression. Live literature acquisition, publisher checks, and slide creation were not exercised. No external sharing or submission is authorized by the revised workflows alone.

@GiggleLiu

Copy link
Copy Markdown
Member

please split prs by features.

@nzy1997

nzy1997 commented Sep 21, 2026

Copy link
Copy Markdown
Collaborator Author

Thanks — split by feature as requested. This combined PR is superseded by:

Each PR cites the GPT-6 Astra skills-and-prompts article and carries its own validation. All PRs except #73 target main independently; #73 targets the shared-writing branch from #72. Closing this combined PR to keep review scoped.

@nzy1997 nzy1997 closed this Sep 21, 2026
@GiggleLiu
GiggleLiu deleted the improve/skill-scope-and-continuation branch September 23, 2026 07:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants