From ae72841775a42f514b7b1bf3372362ee23bcce3f Mon Sep 17 00:00:00 2001 From: nzy1997 Date: Mon, 14 Sep 2026 10:03:29 +0800 Subject: [PATCH 1/2] Refine research skill scope and authorized continuation --- CLAUDE.md | 92 +--- README.md | 8 + docs/kb-migration.md | 30 ++ skills/autoresearch/SKILL.md | 23 +- skills/brainstorm-ideas/SKILL.md | 464 ++++-------------- skills/brainstorm-ideas/references/advisor.md | 54 ++ .../references/session-history.md | 39 ++ skills/create-advisor/SKILL.md | 13 +- skills/how-to-analyze-dialog/SKILL.md | 19 +- skills/how-to-build-kb/SKILL.md | 13 +- skills/how-to-download-ref/SKILL.md | 290 +++-------- .../references/acquisition.md | 84 ++++ .../references/dependencies.md | 56 +++ .../references/maintenance.md | 22 + .../references/troubleshooting.md | 20 + skills/how-to-flow/SKILL.md | 48 +- skills/how-to-flow/journal-template.md | 3 +- skills/how-to-review-figure/SKILL.md | 30 +- skills/how-to-technical-writing/SKILL.md | 8 +- skills/how-to-technical-writing/checklist.md | 10 +- skills/how-to-write-ideas-report/SKILL.md | 20 +- .../references/writing-workflow.md | 73 ++- skills/know-me-better/SKILL.md | 36 +- skills/review-paper/SKILL.md | 120 +++-- skills/review-paper/checklist.md | 36 +- skills/survey/SKILL.md | 63 ++- skills/write-paper/SKILL.md | 219 ++++----- .../references/figures-and-notation.md | 50 ++ tests/test_brainstorm_ideas_skill.py | 3 + tests/test_skill_structure.py | 3 +- tests/test_workflow_resources.py | 79 +++ 31 files changed, 991 insertions(+), 1037 deletions(-) create mode 100644 docs/kb-migration.md create mode 100644 skills/brainstorm-ideas/references/advisor.md create mode 100644 skills/brainstorm-ideas/references/session-history.md create mode 100644 skills/how-to-download-ref/references/acquisition.md create mode 100644 skills/how-to-download-ref/references/dependencies.md create mode 100644 skills/how-to-download-ref/references/maintenance.md create mode 100644 skills/how-to-download-ref/references/troubleshooting.md create mode 100644 skills/write-paper/references/figures-and-notation.md create mode 100644 tests/test_workflow_resources.py diff --git a/CLAUDE.md b/CLAUDE.md index e9e6cb3..45a4843 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,50 +6,33 @@ This is the canonical project guide for agents working in this repository. Claud sci-brain is a skill-based plugin for AI coding assistants (Claude Code, Codex, OpenCode, pi) that provides structured literature, ideation, writing, review, and autonomous-research workflows. It is not a traditional application — its main product is the set of `SKILL.md` interaction protocols and their supporting scripts and references. -## Skills +## Working on skills -The 16 skills in `skills/` are each defined by a `SKILL.md` with YAML frontmatter and instructions. Each description is one sentence starting with its trigger kind, mirroring the `qude-software-skills` convention: +The 16 skills in `skills/` are listed with their trigger descriptions in +[README.md](README.md). Load only the entry point and supporting resources needed +for the current task. Public names match their containing directory. -- `Agentic trigger. Use when …` — `how-to-*` skills the agent invokes automatically while serving a need (a user may still type them). -- `User trigger. Use when …` — skills a user invokes by need (everything else). +- User entry points: **brainstorm-ideas**, **survey**, **write-paper**, + **review-paper**, **write-slides**, **autoresearch**, **know-me-better**, + **create-advisor**, **dump-chat-history**. +- Supporting workflows: **how-to-build-kb**, **how-to-download-ref**, + **how-to-analyze-dialog**, **how-to-flow**, **how-to-review-figure**, + **how-to-technical-writing**, **how-to-write-ideas-report**. -`scripts/validate_skills.py` enforces the prefix; `tests/test_repository_consistency.py` enforces `how-to-*` ⇔ agentic and keeps the README tables identical to the descriptions. +Descriptions start with `User trigger. Use when …` or, for `how-to-*`, +`Agentic trigger. Use when …`. The validator enforces this prefix and tests keep +the README descriptions aligned. Preserve public names and independent skill +installation when reorganizing resources. -**User trigger:** +**autoresearch** routes topics → db → validator → run from `research/STATE.md`. +Its user-confirmed acceptance gates, attempt budgets, and sealed holdout are +research invariants. Status questions are read-only. -- **dump-chat-history** — Selects harnesses and a start date before reading history, preserves original prompts and answers with provenance, and exports JSON/Markdown plus an optional topic-titled Typst/PDF field note; research classification belongs to `how-to-analyze-dialog`. +**brainstorm-ideas** keeps the main mentor and a selected advisor as separate +roles. Advisor profiles and literature live in `advisors//`; the selected +advisor uses a subagent. No advisor is required for ordinary brainstorming. -- **brainstorm-ideas** — The main ideation entry point. Socratic research mentor that understands user background, finds attackable problems, and encourages deeper thinking. When an advisor is selected, it launches that advisor as a subagent and loads literature from `advisors//.knowledge/`. At Phase 3 wrap-up (or on a past session log) it hands off to `how-to-write-ideas-report`. -- **survey** — Parallel literature search via 7 strategies; the user picks directions, then `how-to-build-kb` populates `/.knowledge/`. It also owns the report mode that produces a grounded technology/field assessment from a populated KB; `how-to-download-ref` fetches and renders full text between discovery and writing. -- **write-paper** — Use when drafting or revising an actual scientific manuscript. Encodes the von Delft / Martinis workflow: figures first → telegram outline → body → polish abstract+intro+conclusions last. Distinct from the upstream ideas report in `brainstorm-ideas` — this skill requires real results. -- **review-paper** — The *review/enhance an existing manuscript* counterpart to `write-paper`'s *drafting*. Reads the whole paper, emits location-anchored comments against seven writing guidelines (one-concept sentences, define-before-use, one-job paragraphs, DRY, display-math discipline, figure integration) plus reference & fact verification (CrossRef → Semantic Scholar → MCP → web fetch, repairs via `how-to-download-ref`) and a journal-fit pass (target venue decided or recommended, official author guidelines fetched, limits and required statements measured, writing reviewed against the journal's own guidance). Comment-first and non-destructive: applies only approved edits, then re-runs the compile-check. Distinct from `survey` report mode, which assesses a field rather than a manuscript. -- **write-slides** — Builds PDF decks for scientific talks, lectures, and briefings using [GiggleLiu/sci-brain-slides](https://github.com/GiggleLiu/sci-brain-slides), pinned to v0.1.0. The upstream repository owns templates, layouts, themes, and style documentation; this skill guides outline approval, package setup, composition, compilation, and figure review. It uses Typst's native package mechanism with an explicit local package directory until that version is available in the public registry. -- **autoresearch** — The autoresearch pipeline, one skill with four stage files under `references/stages/`. Reads `research/STATE.md`, verifies stage gate artifacts, and follows the current stage: **topics** (brainstorms topics scored on Checkable/Cheap/Headroom/Publishable; user picks; primary/guard score metrics with gaming risks; red-teamed, user-confirmed acceptance gate per topic → `topics.md`), **db** (insight-coverage-driven reference downloads via `how-to-download-ref`, distillation into user-selected `research/INSIGHTS.md`, domain database, pinned reference implementations, `research/CATALOG.md`; owns the survey gate), **validator** (publishable bar in `GOAL.md`, user-confirmed validation method, sealed gitignored holdout, Docker-canonical `validate` CLI with rich JSON errors, negative-control strictness self-test; owns the validator gate), and **run** (the loop: attempts in worktrees with `LOG.md`, validator-scored under a hard time limit; the user chooses a recommended cycle size during initial setup, while the agent may adjust each actual cycle by need within the authorized attempt budget; every draft hypothesis must state a *mechanism* against the gap to the bar and its *prior art*, ranked on expected gap closure with cost as a constraint, filtered for novelty and triviality; when stuck it refreshes insights via `survey` into `## Candidate`; each cycle report plots every scored attempt's raw primary score with no cumulative headline KPIs, the index and campaign retain cross-cycle summaries, and each reflection thinks through 4–6 candidates before ranking the best 2–4 evidence-grounded next directions with explicit reasons and a top recommendation; the first plan of each authorization is user-confirmed; each soft gate asks which direction and how many attempts to authorize). Each attempt commits code + `LOG.md` + `report.json` on its `attempt-NNN` branch; a cycle-end sync pushes those branches plus main. -- **know-me-better** — Lets the agent learn the user's research style so it speaks their language; the mechanism is indexing a paper collection (Zotero / PDF folder / Google Scholar) into the active KB. Default target is `/.knowledge/`; when invoked from `/create-advisor` targets `advisors//.knowledge/`. Writes `.raw/` JSON, delegates `references.bib` writes via `how-to-download-ref` helpers. -- **create-advisor** — Creates or updates a named advisor from JSONL histories or imported Markdown dialogs. It classifies conversations, extracts recurring trigger→reaction patterns, confirms logic jumps with the user, and synthesizes `advisors//profile.md`; it can also stop after analysis-only artifacts. The advisor's literature cache lives at `advisors//.knowledge/`. - -**Agentic trigger (`how-to-*`):** - -- **how-to-build-kb** — Turns a list of picked papers (from `survey`, `know-me-better`, or `autoresearch`) into verified KB entries: Semantic Scholar / CrossRef lookup, `.raw/` JSON, `references.bib` append via `append_bibtex.py`, `INDEX.md` regeneration, and `NOTES.md` (landscape, open problems, bottlenecks). Non-interactive; never generates BibTeX from memory. -- **how-to-write-ideas-report** — Writes the proposal-style ideas report (research question, novelty, MVE, success/hope/pivot signals, risks, venue, verified references) from a finished `brainstorm-ideas` log. Follows `skills/how-to-write-ideas-report/references/writing-workflow.md`. -- **how-to-technical-writing** — The shared writing style guide: sentence- and paragraph-level rules (one concept per sentence, direct to the point with key information early, simple words with every technical word kept, no undrawn metaphors, locality, say it once), a hunt-for/fix table for language passes, a `checklist.md` with the rules as checkable items, and guardrails naming which fixes are comment-only because they change content. `write-paper` drafts by it, `review-paper` reviews by it, and `how-to-write-ideas-report` / `survey` report mode follow it for report prose. Notation and figure rules stay in `write-paper`. -- **how-to-review-figure** — Reviews the *visual design quality* of a figure, plot, or diagram and prints a scorecard. Source-aware (renders the figure to a raster to look at it via `helpers/render.py`, reads matplotlib/Typst/SVG source so fixes can cite a line), report-only, terminal-first. Scores against an 18-rule rubric (11 general — alignment, proximity, color, hierarchy, contrast, colorblind-safety, …; plus 7 scientific-plot rules — text size, line weight, space use, chartjunk, legend, cross-panel consistency, resolution). Distinct from `review-paper` (which checks whether a figure is cited/discussed in the text, not how it looks) and `write-paper` (which authors figures). Full rubric in `skills/how-to-review-figure/checklist.md`. -- **how-to-flow** — Autonomous deep-thinker that conquers one hard goal via a CDCL/DPLL-style search loop: a **preflight gate** (is the goal testable? are all context/KB facts loaded?), then iterate *decide* (**what-if**: assume a condition, test "closer to goal?" + "easier to achieve?") → *propagate* (**simulate**: run consequences forward, reflect; may fan out 2–3 subagents on wide forks) → *learn* (note a reusable clause after **every** trial) → *backjump* (non-chronological, to the real cause) → *pivot* (meta-restart: re-aim to an equally-valuable easier goal when stuck, keeping all notes). Domain-agnostic and KB-optional. Writes a per-trial journal to `docs/flow/.md` (template in `skills/how-to-flow/journal-template.md`). Terminates SOLVED / PIVOTED-SOLVED / EXHAUSTED (≤3 pivots). Distinct from `brainstorm-ideas` (open-ended, collaborative) — `how-to-flow` is goal-locked and autonomous. -- **how-to-download-ref** — Adds one or many new arXiv IDs / DOIs to a knowledge base (`/.knowledge/` by default; `advisors//.knowledge/` when invoked from advisor flows). Fetches Semantic Scholar metadata, downloads PDFs (with SciHub fallback); when the user opts in, also fetches arXiv LaTeX sources and renders those refs (incl. DOI entries with an arXiv preprint) from flattened LaTeX (`full_text: latex`) via `--tex-source`, otherwise all refs render via `pymupdf4llm`. Regenerates `INDEX.md`, appends to the KB's `references.bib`. Supports `--from-bib` for bulk operations on an existing BibTeX. -- **how-to-analyze-dialog** — Consumes exported dialog from `dump-chat-history`, classifies topics and user messages across 6 academic dimensions, and writes derived tagged reports to `docs/dialog/analysis/` for `create-advisor`; original exports remain unchanged. - -The directory name must match the skill's frontmatter `name`. In particular, the public `know-me-better` skill lives at `skills/know-me-better/`. - -## Architecture - -**Entry point:** `/brainstorm-ideas` — most users only need this. Other skills are auto-called or can run independently. - -**brainstorm-ideas skill uses a primary Socratic mentor plus an optional advisor subagent:** -- Understands user background (self-intro, Zotero, or Google Scholar) -- Loads project literature from `/.knowledge/INDEX.md` + `NOTES.md` -- When an advisor is selected, also loads `advisors//.knowledge/INDEX.md` + `NOTES.md` and pre-fetches representative papers into the advisor subagent context. The advisor subagent is launched with file search/read over its `.knowledge/` KB plus web search/fetch, and is instructed to consult its KB and the web before making comments (grounding each comment in a cited source or marking it as opinion) -- Six principles: clarify motivation, encourage thinking (humbly), flag uncertainty, surface related facts, empower based on skills, inspire with deep theory -- Phases: Get to Know You → Find Good Problems → Dive Into the Topic → Wrap Up +## Shared data contracts **Knowledge base layout** (used by every skill that touches papers): @@ -86,36 +69,11 @@ Each SKILL.md defines installed-resource resolution. Run helpers by their absolu **BibTeX lookup chain** (never from memory): CrossRef API → Semantic Scholar API → MCP servers → web fetch fallback -## Migrating from the pre-0.3 `//` layout - -Old sci-brain (≤ 0.2.x) stored surveys under `~/.claude/survey//` (or `.codex/survey/`, `.config/opencode/survey/`, `.claude/survey/`) with `summary.md` + `references.bib` per topic. 0.3 moves to one `/.knowledge/` per project (plus per-advisor caches). Migrate by hand: - -```sh -# Pick your project root (where you want .knowledge/ to live): -PROJ=/path/to/your/project -mkdir -p "$PROJ/.knowledge" - -# Move a single old registry into the project KB: -OLD=~/.claude/survey/topological-orders # adapt path -mv "$OLD/references.bib" "$PROJ/.knowledge/references.bib" # or merge into existing references.bib -mv "$OLD/summary.md" "$PROJ/.knowledge/NOTES.md" -mv "$OLD"/*.md "$PROJ/.knowledge/" 2>/dev/null # rendered papers -mv "$OLD/.raw" "$PROJ/.knowledge/.raw" -mv "$OLD/.figures" "$PROJ/.knowledge/.figures" - -# Regenerate INDEX.md (use a stable title — re-runs must use the same string): -python3 skills/how-to-download-ref/helpers/index.py \ - --kb "$PROJ/.knowledge" \ - --title "topological-orders — references" \ - --source-note "Migrated from ~/.claude/survey/topological-orders on $(date -u +%Y-%m-%d)." - -# Remove the old registry: -rmdir "$OLD" -``` - -For advisor caches built by the abandoned 0.2-era `publications.yml` flow: that layout was never populated; nothing to migrate. The new flow builds `advisors//.knowledge/` via `/know-me-better` or `/how-to-download-ref` invoked from `/create-advisor`. +## Legacy KB migration -Multiple old registries can be merged into one project KB (run the `mv` block per topic; `references.bib` accepts appends; `NOTES.md` accepts merges as separate top-level headings). +For a requested migration from the pre-0.3 registry layout, read +[docs/kb-migration.md](docs/kb-migration.md). Ordinary skill use does not require +that migration guide. ## Installation diff --git a/README.md b/README.md index 806121b..2810281 100644 --- a/README.md +++ b/README.md @@ -124,6 +124,14 @@ Four internal stages became modes of their goal-level skill in v0.3: | `/import-dialog` | `/create-advisor` with Markdown dialog input | | `/soul-extraction` | `/create-advisor` in analysis-only mode | +## Workflow design + +Skills preserve the requested scope and existing choices, load supporting +instructions as needed, and continue through authorized deliverables. Scientific +claims, reference provenance, acceptance gates, and research budgets remain +explicit constraints. This revision follows OpenAI's +[Rethinking skills and prompts for GPT-6 Astra](https://developers.openai.com/blog/rethinking-skills-and-prompts-for-gpt-6-astra). + ## Contributors **Initiators**: [Lei Wang](https://github.com/wangleiphy) and [Jin-Guo Liu](https://github.com/GiggleLiu) diff --git a/docs/kb-migration.md b/docs/kb-migration.md new file mode 100644 index 0000000..f323275 --- /dev/null +++ b/docs/kb-migration.md @@ -0,0 +1,30 @@ +# Migrating from the pre-0.3 `//` layout + +Old sci-brain (≤ 0.2.x) stored surveys under `~/.claude/survey//` (or `.codex/survey/`, `.config/opencode/survey/`, `.claude/survey/`) with `summary.md` + `references.bib` per topic. 0.3 moves to one `/.knowledge/` per project (plus per-advisor caches). Migrate by hand: + +```sh +# Pick your project root (where you want .knowledge/ to live): +PROJ=/path/to/your/project +mkdir -p "$PROJ/.knowledge" + +# Move a single old registry into the project KB: +OLD=~/.claude/survey/topological-orders # adapt path +mv "$OLD/references.bib" "$PROJ/.knowledge/references.bib" # or merge into existing references.bib +mv "$OLD/summary.md" "$PROJ/.knowledge/NOTES.md" +mv "$OLD"/*.md "$PROJ/.knowledge/" 2>/dev/null # rendered papers +mv "$OLD/.raw" "$PROJ/.knowledge/.raw" +mv "$OLD/.figures" "$PROJ/.knowledge/.figures" + +# Regenerate INDEX.md (use a stable title — re-runs must use the same string): +python3 skills/how-to-download-ref/helpers/index.py \ + --kb "$PROJ/.knowledge" \ + --title "topological-orders — references" \ + --source-note "Migrated from ~/.claude/survey/topological-orders on $(date -u +%Y-%m-%d)." + +# Remove the old registry: +rmdir "$OLD" +``` + +For advisor caches built by the abandoned 0.2-era `publications.yml` flow: that layout was never populated; nothing to migrate. The new flow builds `advisors//.knowledge/` via `/know-me-better` or `/how-to-download-ref` invoked from `/create-advisor`. + +Multiple old registries can be merged into one project KB (run the `mv` block per topic; `references.bib` accepts appends; `NOTES.md` accepts merges as separate top-level headings). diff --git a/skills/autoresearch/SKILL.md b/skills/autoresearch/SKILL.md index f60683d..fbe9885 100644 --- a/skills/autoresearch/SKILL.md +++ b/skills/autoresearch/SKILL.md @@ -5,14 +5,11 @@ description: User trigger. Use when starting, resuming, or checking an autoresea ## Installed resources -Keep the working directory at the user's project. Resolve this loaded `SKILL.md` -with `Path(path).resolve()` before locating resources; follow symlinks. Bare -`helpers/`, `references/`, and template paths are relative to that real skill -directory. A path written as `skills//...` means the installed `` -skill's directory from the agent's skill catalog, not a path in the user's project. -Locate each dependency by its public skill name; copied skills need not be siblings. -If a dependency is absent, report the missing skill and install it before that step. -Shared writing files are bundled in `how-to-write-ideas-report/references/`. +Keep the working directory at the user's project. Resolve this `SKILL.md` to its +real path before locating bundled resources. `skills//...` refers to the +installed skill found by public name, not the user's project; dependencies need +not be siblings. Load only resources needed for the current task. If a required +dependency is missing, report it before that dependent step. # Autoresearch @@ -36,6 +33,16 @@ Supporting references: `references/insights-template.md` (db), `helpers/report.py` (HTML cycle reports) and `helpers/gen_campaign.py` (cross-cycle full-campaign overview). +## Choose the request + +- **Status or explanation:** read STATE.md and the relevant public artifact + summaries. Report stage, gate status, completed and remaining attempts, and + missing evidence. Stop after answering: do not create or repair state, change + stages, run experiments, or consume the attempt budget. Missing/corrupt state + is a reported finding, not permission to initialize a campaign. +- **Start, resume, or execute a named stage:** follow Procedure below, inheriting + the user's existing choices and authorization. The stage gates still apply. + ## Procedure 1. **Locate state.** Read `/research/STATE.md`. diff --git a/skills/brainstorm-ideas/SKILL.md b/skills/brainstorm-ideas/SKILL.md index b9ef95c..8c1b274 100644 --- a/skills/brainstorm-ideas/SKILL.md +++ b/skills/brainstorm-ideas/SKILL.md @@ -5,384 +5,94 @@ description: User trigger. Use when brainstorming research ideas with a Socratic ## Installed resources -Keep the working directory at the user's project. Resolve this loaded `SKILL.md` -with `Path(path).resolve()` before locating resources; follow symlinks. Bare -`helpers/`, `references/`, and template paths are relative to that real skill -directory. A path written as `skills//...` means the installed `` -skill's directory from the agent's skill catalog, not a path in the user's project. -Locate each dependency by its public skill name; copied skills need not be siblings. -If a dependency is absent, report the missing skill and install it before that step. -Shared writing files are bundled in `how-to-write-ideas-report/references/`. +Keep the working directory at the user's project. Resolve this `SKILL.md` to its +real path before locating bundled resources. `skills//...` refers to the +installed skill found by public name, not the user's project; dependencies need +not be siblings. Load only resources needed for the current task. If a required +dependency is missing, report it before that dependent step. Before running the examples, set `DOWNLOAD_REF_DIR` to the absolute directory of `how-to-download-ref`. Quote these variables as shown. -**Path conventions:** -- `docs/discussion/` — resolved from the **project working directory** -- `/.knowledge/` — the project's shared knowledge base; resolved via `skills/how-to-download-ref/helpers/resolve_kb.py` -- `advisors//.knowledge/` — advisor-specific literature cache, rooted at the **plugin root** -- `skills/` follows Installed resources above. `advisors/` requires the full sci-brain checkout; when it is absent, use the no-advisor flow. - -## Choose the mode - -- **Brainstorm:** follow Ideas below. -- **Write an ideas report:** if the user asks for a report from a completed session or chosen research direction, skip the conversational phases and invoke `how-to-write-ideas-report` (read `skills/how-to-write-ideas-report/SKILL.md`). -- **Brainstorm, then write:** complete Phase 3 and hand off to `how-to-write-ideas-report` when the user chooses a full report. - -## Ideas - -A research collaborator with a sense of humor. The main mentor stays warm and encouraging, and an optional advisor is handled as a separate specialist process. If an advisor is selected, do not collapse them into the main narrator as a style-only imitation — launch a dedicated advisor subagent and keep the mentor/advisor roles distinct. - -**Tone:** Like a smart friend who happens to know a lot — curious, honest, fun to talk to. Light, encouraging, occasionally witty. Examples: - -- "That's an ambitious idea. I like it. Let me see if the literature agrees with your optimism..." -- "Well, the good news is nobody has done this before. The bad news is... nobody has done this before." -- "Let me see if I have some good questions in my pocket, digging..." - -### Six Conversation Principles - -These drive every response throughout the session: - -#### a) Clarify motivation when it matters - -Ask about the user's motivation only when it would genuinely change what you suggest. If the direction is already clear, just go. - -#### b) Encourage deeper thinking (humbly) - -The research problems are hard — hard enough that the mentor clearly cannot reason through them deeply. Be honest about that. Empower the user instead: - -> "Even as your advisor, I'm not sure about this one. Could you use your evolving brain to reason for me — is this plan reasonable? Mathematically sound? Or tell me what information you need to think it through, and I'll go find it." - -**The deal:** The mentor finds facts, surfaces connections, provides references. The human does the deep reasoning. "You think, I fetch." - -If the user identifies a gap ("I'd need to know if X holds in Y"), the mentor decides whether to search for it — sometimes the answer is already in the knowledge base or in the conversation context. - -#### c) Identify uncertainty, warn about risk - -When something is uncertain, say so explicitly. Flag potential risks constructively — to prepare, not to scare. - -When critiquing, cite references when available. If no reference found, explicitly say: "This is my opinion, not proven." Always distinguish opinion from evidence. - -#### d) Surface a related fact to drive the discussion - -Bring in something from a neighboring field, a surprising connection, or an overlooked paper — to open a new angle in the conversation. - -> "Oh, this reminds me — in [other field], they ran into a very similar problem and tried [approach]. Not sure if it applies here, but it's interesting. What do you think?" - -This keeps the conversation moving and often opens unexpected directions. - -#### e) Empower the user based on their specific skills - -Connect the user's existing abilities to the challenge. Be honest about what looks doable: - -> "Since you're good at [X], you should be able to handle [Y] — you might just need to pick up a bit of [Z]. That's very learnable for someone with your background." - -If a gap shows up, mention it naturally: "This approach leans on [Z] — have you worked with that before? If not, [resource] is a solid place to start." - -#### f) Share enthusiasm for deep theory — inspire, not prescribe - -When a key theory underpins the current direction and the user seems reluctant to engage with it (skipping over it, staying surface-level, or changing the subject), share *why it's exciting* with concrete examples of how it reshapes understanding: - -> "For me, [theory] is genuinely one of the most fun things I've encountered — it totally reshaped how I think about [domain]. For example, [concrete example of how the theory reveals something surprising or powerful]. Once you see it that way, [practical consequence] just clicks. I really wish you could experience that too. Oh — I have a book for you: [title] by [author]. It's [why this specific book is great]." - -The goal is to make the user *curious*, not obligated. Show the beauty of the theory through your own relationship with it. If the user still isn't interested, respect that and move on. - ---- - -### Conversation Log - -Maintain a running log at `docs/discussion/YYYY-MM-DD-HHMMSS-brainstorm-ideas-log.md` (timestamp from session start). Create the `docs/discussion/` directory if it doesn't exist. - -**Append-only logging.** Save progress by appending to the log at checkpoints. Each append captures the **full conversation content** since the last save — all options presented (with descriptions), reasoning shared, user responses, search results, and key ideas. Not a summary — a readable record of what was actually said. - -**When to append (checkpoints):** -- Every 3-5 exchanges, at a natural pause — when a sub-topic wraps up, a decision is made, or the conversation shifts direction -- At phase transitions (entering Phase 1, Phase 2, Phase 3) -- At session wrap-up (Phase 3) - -Don't log after every message. Wait for a moment that feels like a natural checkpoint — the end of a thread, a decision point, a topic shift. - - -**Order: log first, then reply.** At a checkpoint, append to the log file before writing your response to the user. This ensures progress is saved even if the session is interrupted mid-reply. - -**File header** — write once when creating the log: - -```markdown -# Ideas Session — YYYY-MM-DD HH:MM -``` - -**Phase 3 wrap-up** — append a final section that consolidates the key outcomes: direction chosen, ideas explored, action items, and recommended readings. - -These logs accumulate across sessions as separate files, building a record of the user's research interests, thinking patterns, and explored directions. - -### Phase 0 — Get to Know You - -**Skip if chaining from survey.** If the current session already has survey context (user has been working on a topic, background is known), skip Phase 0 and go straight to Phase 1. - -**Advisor selection.** Check if `advisors/index.md` exists and contains advisor entries. If advisors are available, present them as an interactive choice before proceeding: - -> "Before we start — would you like to brainstorm with a specific advisor? Each one has a unique thinking style based on a real researcher." - -For each advisor in `advisors/index.md`, create an option with: -- **Label:** The advisor's name -- **Description:** Their field (one line) -- **Markdown preview:** A brief profile card showing their field, key strengths, and a sample of their thinking style (drawn from `advisors//profile.md` — read the profile to build the preview). Keep it to ~5-8 lines so the user can quickly compare. - -Always include a final option: -- **Label:** "No advisor" -- **Description:** "Default mentor — warm, curious, encouraging" - -If the user picks an advisor, do **not** just read `advisors//profile.md` and role-play inline. Instead, read the advisor profile, then launch a dedicated advisor subagent: - -1. **Read the advisor profile.** Load `advisors//profile.md` (slug is lowercase hyphenated, e.g., `xi-dai`) and use the most relevant topic section (prefer `brainstorming` or `research`) to understand how this advisor thinks. - -When an advisor is selected, first resolve the advisor KB path so it follows `$SCIBRAIN_KB_DIRNAME` if the user has set it: - -```sh -ADVISOR_KB=$(python3 "$DOWNLOAD_REF_DIR/helpers/resolve_kb.py" --advisor ) -``` - -Then: -- Load `$ADVISOR_KB/INDEX.md` to know what literature is available. -- Load `$ADVISOR_KB/NOTES.md` for the advisor's curated thematic notes (if present). -- Pre-fetch a handful of representative papers from `$ADVISOR_KB/_.md` as **seed context loaded into the advisor subagent at launch** — these are a starting point, not the advisor's whole library. -- If `$ADVISOR_KB/` is empty or missing, fall back to launching the advisor without a literature cache (still useful — the profile alone shapes their reasoning). - -**Launch the advisor.** The advisor subagent's job is to contribute hard-won taste: what to ask next, which assumptions are dangerous, which papers matter, and what this advisor would investigate first. The main mentor remains responsible for session flow, empathy, logging, and synthesis. - -**Give the advisor subagent the tools to investigate.** Launch it as a subagent with file access to `$ADVISOR_KB` and web search/fetch available, and pass `$ADVISOR_KB` (the absolute path) in its prompt. Instruct the subagent that, **before making a substantive comment, it should:** -- **Consult its own knowledge base first.** Search `$ADVISOR_KB/INDEX.md` and open the relevant `$ADVISOR_KB/_.md` papers for specifics — don't rely only on the seed papers. The seed set is a head start; the full KB is the advisor's library to draw on. -- **Search the web** when the KB doesn't cover a needed fact, or to check a recent development or verify a claim before asserting it. -- **Ground each comment in what it found and say so** — name the paper (cite key from `$ADVISOR_KB/INDEX.md`) or link the source. When neither the KB nor the web supports a claim, mark it explicitly as opinion (consistent with principle (c), "distinguish opinion from evidence"). -- Restrict file access to `$ADVISOR_KB` and the advisor's `profile.md`; the subagent reads literature and the web, it does not edit project files. - -If `$ADVISOR_KB` is empty or missing, the subagent still has web search/fetch and falls back to web grounding plus profile-driven reasoning. - -The advisor profile shapes *how* the advisor subagent thinks and behaves. The user's own profile (`user-profile.md`) still determines *what* the overall system knows about the user's background. Both are loaded, but they are loaded into different roles: the main mentor keeps the broad session context, while the advisor subagent receives the advisor-specific literature cache and style directives. - -**Advisor voice formatting.** When an advisor is active, surface the advisor subagent's contributions in blockquotes, prefixed with the advisor's name. This visually distinguishes the advisor's voice from the mentor's default narration: - -> **[Xi Dai]** "You should ask your AI agent to check whether the Pearl length exceeds the sample size in the thin-film limit — because if it does, the vortex-vortex interaction becomes logarithmic instead of exponential, and that completely changes the phase diagram. I've seen people miss this and waste months on the wrong regime." - -Advisor comments should be **constructive and helpful** — the advisor acts as a senior collaborator who guides the user toward productive directions. When providing questions, frame them as **suggestions for what the user should ask the AI agent**, not as quizzes directed at the user. Each suggestion should include the advisor's **reasoning for why this question matters** — what could go wrong if it's not asked, what insight it unlocks, or what assumption it tests. The advisor's comments should: -- Suggest specific questions or tasks the user should pose to the AI agent, framed as actionable requests -- Explain *why* this question is important — what the advisor's experience tells them about what's at stake -- Mirror the advisor's characteristic way of attacking problems (e.g., a theorist might suggest "ask it to check the limiting case, because...", an experimentalist might suggest "ask it to estimate the observable signature, since...") - -The goal is to empower the user with the advisor's hard-won intuition about *what to investigate and why*. The advisor is a constructive partner who helps the user get the most out of the AI agent by knowing which questions are the right ones to ask. - -Use this for moments where the advisor's specific perspective, instinct, or experience is driving the suggestion — not for every sentence. The mentor's own observations, factual summaries, and logistical statements stay in normal text. - -**Advisor audio with `edge-tts`.** If the user wants spoken advisor responses and `edge-tts` is available, synthesize advisor-only blocks to audio after generating the text. Keep text as the source of truth; audio is a companion artifact. Suggested behavior: -- Store audio at `docs/discussion/audio/-/` -- Default to a voice specified in the advisor profile if one exists; otherwise pick the closest high-quality `edge-tts` voice for the advisor's preferred language -- Only synthesize advisor passages, not the mentor's logistics/search summaries -- Save the transcript alongside the audio so the session remains readable without playback - -If no advisor is selected or no advisors exist, proceed with default mentor behavior. - -**First, check for history.** Read `docs/discussion/user-profile.md` if it exists — this contains the user's persisted profile from previous sessions. Also resolve the project KB via `KB=$(python3 "$DOWNLOAD_REF_DIR/helpers/resolve_kb.py")` and check `$KB/` for indexed publication data from the `know-me-better` skill. Also read `docs/discussion/*-brainstorm-ideas-log.md` if they exist — they contain past brainstorming sessions and reveal the user's evolving interests, thinking patterns, and which directions they've explored before. - -**Session picker.** If previous session logs exist, present them as an interactive choice before proceeding: - -> "Welcome back! You have some previous sessions. Want to pick one up, or start fresh?" - -For each past session log (most recent first, up to 5), create an option with: -- **Label:** The session date and main topic (extracted from the log's header and content) -- **Description:** One-line summary — the direction explored and current status (e.g., "Exploring tensor network methods — narrowed to 2 candidates" or "Incomplete — was diving into topological phonons") -- **Markdown preview:** A brief recap of where the session left off — the last phase reached, key ideas discussed, and any open threads or action items. Keep it to ~5-8 lines. - -Always include a final option: -- **Label:** "Start fresh" -- **Description:** "New brainstorming session" - -**If the user picks a previous session:** - -Read the full log to restore context. Then handle the continuation naturally based on where that session ended: - -- **Incomplete session (no Phase 3 wrap-up):** Deliver what Phase 3 would have said — a reflection, a connection, or a recommendation — as a casual callback, then resume from where it left off: - > "Oh, before we pick up — I've been thinking about where we left off. You were working through [X] and I never got to say: [insight/recommendation/connection]. Anyway — ready to keep going?" - -- **Session ended with a plan** (e.g., "I'll come back after reading X"): Open with a callback to that plan: - > "Hey! Last time you were going to read [X] and think about [Y] — how did that go?" - -- **Completed session (has Phase 3 wrap-up):** Reference the outcome and ask what's next: - > "Last time we landed on [direction] and I recommended [book/paper]. Want to build on that, or explore something different?" - -Continue the session's log file (append to it) rather than creating a new one. Skip to the appropriate phase based on where the previous session left off. - -**If the user starts fresh (or no previous sessions exist):** - -Open with a warm greeting: - -> "Hey! I'm excited to brainstorm with you. But first, let me get to know you a bit — better suggestions come from understanding who I'm talking to." - -Create a new log file and proceed normally. Even when starting fresh, use past session logs as background context — reference past sessions, avoid re-treading ground, and pick up threads they left open, but don't force continuity. - -**Background** — if a user profile or project knowledge base already exists and is sufficient, skip the background question. Instead, summarize what you know and ask if anything has changed: - -> "I already have your profile from before — [brief summary]. Want to update anything, or shall we dive in?" - -If no existing profile or knowledge base is found, ask in chat: - -> "How would you like to share your research background?" -> - **(a)** Tell me yourself — your field, experience, what you've worked on -> - **(b)** Zotero library — I'll index your papers to understand your work -> - **(c)** Google Scholar profile — give me your URL - -For **(b)** or **(c)**: follow the `know-me-better` skill instructions (read `skills/know-me-better/SKILL.md`) to build a project knowledge base, then continue. The indexed data (publication count, topics, recency, citation patterns) reveals the user's experience level — no need to ask explicitly. - -**For (a) only — one follow-up question (if not already answered):** - -If the user's self-introduction already reveals their experience level (e.g., they mentioned prior publications, years in a program, or previous projects), skip this question — the information is already there. Otherwise ask: - -"Is this your first research project, or have you done this before?" - -(Skip this for (b)/(c) — infer experience from the indexed data instead.) - -**Save the user profile** to `docs/discussion/user-profile.md` — this persists across sessions so later conversations can reference it. Include: name, field, experience level, key skills/tools, research interests, and notable papers/projects. If the file already exists, update it rather than overwriting (the user's profile evolves over time). - -**Then listen.** The user may already describe what they want to explore, share an idea, or ask a question. Either way, always proceed to Phase 1 — there's usually more to discover around any starting point. Phase 1 helps contextualize and ground whatever the user brings (or helps them find a direction if they don't have one yet). - -### Phase 1 — Find Good Problems - -**Always run this phase** — even when the user already stated a direction. There's almost always more context to uncover. - -**Load context:** Resolve the project KB via `KB=$(python3 "$DOWNLOAD_REF_DIR/helpers/resolve_kb.py")` and check `$KB/` for indexed knowledge. If found, note it for later use. If none found, note that a lighter web search will be needed later. If an advisor is active, also load `$ADVISOR_KB/INDEX.md` (resolved earlier with `--advisor `) so the mentor knows what literature the advisor subagent already has in context. - -#### Step 1: Talk first - -Start with conversation, not search. The goal is to understand what the user finds exciting *before* touching the literature. - -**Two entry modes:** - -- **User has a direction:** Ask them about it — what draws them to this? What's the specific puzzle or opportunity they see? React to what they say, make connections, ask follow-ups. Have a genuine back-and-forth. -- **User is open:** Scan their profile and any loaded registries for 1-2 interesting provocations — surprising connections between their skills, underexplored intersections, or things that seem ripe. Throw these out casually to spark conversation, not as formal options: - - > "Looking at your work, one thing that jumps out is [observation]. And I'm also curious about [connection]. What do you think — does either of these resonate, or is something else on your mind?" - -Let the user talk. React, connect, riff. This conversation shapes the search that comes next. - -#### Step 1.5: Scope, constraints, and check understanding - -Before searching, do two things: - -**1. Acknowledge the human side.** If the user has mentioned tensions — advisor disagreements, career pressure, identity questions about their research direction — acknowledge them briefly before moving to strategy. Don't therapize, just show you heard it: - -> "Navigating that tension between what your advisor wants and what excites you is real — and it's worth finding something that honors both. Let me keep that in mind." - -**2. Check your understanding and ask about scope.** Summarize what you've heard and ask one open-ended scoping question. This validates the user's input and surfaces constraints naturally: - -> "Let me make sure I have this right — you're interested in [X], your strengths are [Y], and the main constraint is [Z]. Before I go looking: any boundaries I should know about? Like, are you looking to build on what you know, or open to unexpected directions? Any timeline pressures?" - -If the user has mentioned practical constraints (advisor preferences, timeline, funding), reflect them back here. For students: ask about milestones if not already mentioned (e.g., "Do you have a timeline in mind — like a paper deadline or qualifying exam?"). - -#### Step 2: Search to ground and extend - -Once something interesting surfaces from the conversation, go to the literature. The search is now *guided by* the conversation, not the other way around. - -> "That's a really interesting angle — let me see what's out there around this..." - -**Search with three matters in mind:** - -1. **Practical impact** — What real problems need solving? Who would benefit? -2. **Theoretically interesting and open** — Where is there genuine depth? What key questions are still unsolved? -3. **Fit with user's knowledge** — What can this user realistically tackle given their skills? - -Mine `$KB/NOTES.md` for open problems/bottlenecks, then use web search for recent developments when needed. The *direction* of the search is further tailored by who the user is: - -| User profile | Search direction | -|---|---| -| Beginner, first project | Well-benchmarked problems with clear methodology, active community, tutorial resources | -| Experienced, wants challenge | Recently opened problems, contrarian angles, cross-field opportunities | -| Has specific tools/methods | Problems where those tools are underused or newly applicable | - -#### Step 3: Present what you found - -**Present 2-4 problems or refined angles** — conversationally, not as a menu. Connect each option back to what the user said. Highlight what makes it interesting — just the most compelling point. Speak naturally, as you would in conversation. For beginners, no jargon without explanation. Include a key reference for each. - -**Stage the presentation — conversation first, then structured options.** Lead with the direction that best fits the conversation so far. Share it conversationally and react to the user's response before offering alternatives. Don't dump all options at once — a real mentor surfaces one idea, sees how it lands, then adjusts. If the first idea resonates, the others become "here's another angle" rather than competing choices. - -For each direction, include a one-line feasibility hint (e.g., "builds on your existing skills" vs. "requires picking up X first") so the user can gauge cost at a glance. Save the detailed breakdown (timeline, new learning required, what a first paper looks like) for *after* the user shows interest. - -**After the conversational discussion**, ask in chat with a preview for each option — each option has a short problem name as the label, a one-line description, and a preview with the full write-up. **Always include these final options:** -- "None of these — tell me what's missing" — so users who don't connect with any direction have a path forward. If the user wants more specificity within the same space, drill down to concrete open problems. If the user wants to change direction entirely, return to Step 1 with the new direction. -- "Let me think about this — pick up next session" — research direction decisions deserve time; don't implicitly reward immediate commitment - -Present options as framings, not rigid choices. Users often want to combine or adapt — welcome that: "These are starting points. If something resonates partially, or you want to mix directions, tell me what actually fits." - -### Phase 2 — Dive Into the Topic - -When the user selects a topic, dive in. The goal is to go from a broad direction to a concrete, attackable research idea. - -Follow the six conversation principles naturally — as instinct, not as a checklist. - -**Step 1: Understand the landscape.** Explore the topic — what has been tried, what worked, what failed. Identify the gaps and open questions. Share what you find conversationally. - -**Step 2: Narrow down.** Ask clarifying questions one at a time to zero in on the interesting part. **Prefer open-ended conversational prompts** for intermediate thinking steps — users naturally blend, adapt, and push back in ways that don't fit discrete options. Reserve structured multi-option questions for moments where the user faces a genuine fork (e.g., choosing between distinct sub-problems). When you do use one, present options as framings, not rigid choices — "here are some ways to think about this, but tell me what actually fits." Each question should resolve one uncertainty: - -- What aspect of this problem interests you most? -- Which gap feels most attackable given your background? -- What would success look like for you? - -**Step 3: Shape the idea.** Once a direction emerges, help the user sharpen it into something concrete. Find the weakest assumption, logical gap, or inconsistency — then don't just note it, bring the user the relevant information (a paper, a known result, a counterexample) and ask them to reason through it: - -> "There's one thing I'm not sure about in this plan — [gap/inconsistency]. I found [reference/result] that's relevant. What do you think — does this hold up, or does it change the approach?" - -When an idea sounds appealing and straightforward — the kind that feels like it *should* work — that's exactly when to check for prior art. Good ideas attract many people; if it seems obvious, someone likely tried it: - -> "I love this idea — it's clean and it makes sense. But that's exactly what worries me. Something this natural, hasn't anyone tried it before? Let me search for you." - -Then search. If prior art exists, present it honestly and help the user find what's genuinely new about their angle. If nothing turns up, that's a strong signal worth noting. - -Be honest about what you can and what you have no way to assess. The mentor's job is to surface the right information at the right moment; the user's job is to think it through. - -**Step 4: Confirm.** Present the refined idea back to the user — what it is, why it matters, what the first steps would be. Ask if it feels right, or if something needs adjusting. - -The conversation may loop between steps 2-4 as the idea evolves. That's natural. - -After a natural stopping point (idea confirmed, user seems satisfied, or energy drops), offer next steps in chat: keep refining, try a different angle, take time to think and pick up next session, or wrap up. Don't offer this after every single exchange — let the conversation breathe. - -**Search policy:** Ground ideas in the loaded knowledge bases (`$KB` and, when an advisor is active, `$ADVISOR_KB`) first. Only search the web when the conversation goes beyond what those caches cover. - -### Phase 3 — Wrap Up - -When the user is done, the mentor does two special things before ending: - -**1. Reflect on the conversation and share a better way to dig in.** - -Look back at how the conversation went — and read `docs/discussion/*-brainstorm-ideas-log.md` for cross-session patterns. What themes keep coming up? What directions has the user circled back to? What was most interesting today vs. past sessions? Then share a thought: - -> "I really enjoyed this conversation. I'd love to dig deeper with you about [specific matter that came up]. One way you could ask about it is: '[a better-framed version of a question they asked during the session]' — that kind of question opens up more interesting directions. - -**2. Final recommendation (apply principle f).** - -Based on the user's chosen direction and demonstrated interests, recommend one book, paper, blog post, or talk that hasn't already been mentioned in the conversation. Verify via web search only if unsure. Share *why you find it exciting*, with a concrete example of how it changes your thinking: - -> "You know what this conversation reminded me of? [title] by [author]. For me, that book/paper completely changed how I think about [aspect] — for example, [concrete insight or surprising idea from it]. Given your interest in [direction], I think you'd really enjoy it." - -**3. Encourage continued exploration.** - -If the session felt shallow (many topic switches, no deep dives) or the user seems like they might not come back, present the observation first, then invite: - -> "I notice that we covered a lot of ground today but didn't go very deep into any single direction. Among everything we explored, [most promising direction] stood out to me — I'd be much happier if you could dig deeper into that one together with me next time. I think we barely scratched the surface." - -This isn't pressure — it's an honest observation followed by a genuine invitation. - -**4. Offer to capture new references.** - -Scan the conversation log for arXiv IDs / DOIs that surfaced during the session and aren't already in `references.bib` (if a knowledge base is loaded). If any are found, ask in chat: - -> "We touched on N papers that aren't in your knowledge base yet. Want to add any now?" -> - **(a)** Add all — invoke `how-to-download-ref` for each -> - **(b)** Pick a subset — show the list, user multi-selects -> - **(c)** Skip - -For (a) / (b), invoke the `how-to-download-ref` skill (read `skills/how-to-download-ref/SKILL.md`) targeting the active knowledge base. The skill handles metadata fetch, cite-key confirmation, BibTeX append to `references.bib`, PDF render, and `INDEX.md` regeneration per ref. - -**Options at wrap-up** — ask in chat: - -> "So — what would you like to do?" -> - **(a)** Generate a full ideas report — invoke `how-to-write-ideas-report`, carrying the conversation log, user profile, chosen direction, key references, and concrete action plan -> - **(b)** End session — the conversation log is already saved -> - **(c)** Keep going — return to Phase 2 +# Brainstorm research ideas + +Help the user find, refine, or reason through an attackable research problem. +Be curious and candid. Offer your own reasoning, calculations, counterexamples, +and literature checks when useful; use Socratic questions when they help the +user think or when the user asks for that style. Mark assumptions and distinguish +source-supported findings from hypotheses or opinion. + +## Enter at the current need + +- **Find a direction:** use background and constraints already given, then + explore plausible problems and ground them in the literature. +- **Refine a chosen direction or work through a derivation:** start with that + problem. Identify the weakest assumption and investigate it; do not restart + background interviews or require a new direction-selection phase. +- **Resume:** recover only the relevant session using + [session-history.md](references/session-history.md). +- **Write a report:** invoke `how-to-write-ideas-report` with the chosen direction, + current notes/log, references, and action plan. No new brainstorming is needed + if the substance is already available. + +Carry forward the user's choices, requested deliverables, and permission to +continue. Ask only when missing information would materially change the work. +An open-ended brainstorming conversation can end at a natural stopping point; +a requested derivation, comparison, or report continues through its verification +and delivery unless substantive missing input blocks it. + +## Context and collaborators + +Resolve the project KB via `KB=$(python3 "$DOWNLOAD_REF_DIR/helpers/resolve_kb.py")` +when literature context is needed. The default is `/.knowledge/`. +Search INDEX.md/NOTES.md for the current topic, then open relevant papers. Use +web search for missing facts, uncertain claims, prior art, or current developments. +A missing KB does not block brainstorming. + +Use the user's profile and provided constraints; request background only if it +would change the advice. When the user chooses Zotero or a Scholar profile as +background, invoke `know-me-better` with that source and return to this discussion. + +An advisor is optional. If the user names one, or asks to choose from the advisor +library, consult `advisors/index.md`. Show names and fields from the index; read +only profiles needed for the choice. Without a selection, continue as the mentor. +When selected, **launch a dedicated advisor subagent** following +[advisor.md](references/advisor.md). Its literature lives at +`advisors//.knowledge/`, resolved by the KB helper. Advisor-only audio with +`edge-tts` is available when requested; the same reference covers it. + +## Explore and refine + +Choose the steps the current uncertainty needs: + +1. **Frame the problem.** State the question and what a useful result would look + like. Ask about motivation, resources, or timeline only when still unknown + and relevant. Acknowledge personal constraints briefly when the user raises them. +2. **Ground alternatives.** Check prior work, failed approaches, and nearby + fields. Explain why each promising direction matters, how it fits the user's + tools, and its main feasibility risk. Offer several alternatives only when + there is a real choice; the user may combine or reject them. +3. **Investigate.** Develop the argument, calculation, or smallest experiment. + Test weak assumptions with examples, counterexamples, known results, or a + discriminating check. Do not delegate all difficult reasoning back to the user. +4. **Make the next step concrete.** Explain novelty relative to verified prior + work, the proposed method, the smallest test, and the evidence that would + support or refute it. Confirm a new research direction when a choice remains; + reuse an already chosen one. + +Finding no prior art is a search result, not proof of novelty. Recommend learning +material when it resolves a real gap, with verified sources and a reason tied to +the problem; do not require a new recommendation at every wrap-up. + +## Preserve progress and deliver + +Maintain the append-only conversation log described in +[session-history.md](references/session-history.md). Save at meaningful +checkpoints, preserving the discussion rather than repeatedly loading all logs. +Update the user profile only with supported information. + +At a stopping point, record the selected direction, evidence, unresolved +questions, and next actions. Complete any report or KB additions already +requested. For optional new KB additions, offer the identified papers together; +pass the selected IDs and existing preferences to `how-to-download-ref`, then +resume the caller's task. Do not force another menu merely to end a session. diff --git a/skills/brainstorm-ideas/references/advisor.md b/skills/brainstorm-ideas/references/advisor.md new file mode 100644 index 0000000..9d8ea11 --- /dev/null +++ b/skills/brainstorm-ideas/references/advisor.md @@ -0,0 +1,54 @@ +# Selected advisor + +Read this only after an advisor has been selected. Use the full sci-brain +checkout to locate `advisors/`; if it is absent, explain that the advisor profile +is unavailable and continue without one. Reuse the existing advisor process in +this session. Resource paths below use the installed skill directories defined +in SKILL.md; `DOWNLOAD_REF_DIR` is the resolved how-to-download-ref directory. + +If the user picks an advisor, do **not** just read `advisors//profile.md` and role-play inline. Instead, read the advisor profile, then launch a dedicated advisor subagent: + +1. **Read the advisor profile.** Load `advisors//profile.md` (slug is lowercase hyphenated, e.g., `xi-dai`) and use the most relevant topic section (prefer `brainstorming` or `research`) to understand how this advisor thinks. + +When an advisor is selected, first resolve the advisor KB path so it follows `$SCIBRAIN_KB_DIRNAME` if the user has set it: + +```sh +ADVISOR_KB=$(python3 "$DOWNLOAD_REF_DIR/helpers/resolve_kb.py" --advisor ) +``` + +Then: +- Load `$ADVISOR_KB/INDEX.md` to know what literature is available. +- Load `$ADVISOR_KB/NOTES.md` for the advisor's curated thematic notes (if present). +- Read only papers relevant to the current question from `$ADVISOR_KB/_.md` as **seed context loaded into the advisor subagent at launch** — these are a starting point, not the advisor's whole library. +- If `$ADVISOR_KB/` is empty or missing, fall back to launching the advisor without a literature cache (still useful — the profile alone shapes their reasoning). + +**Launch the advisor.** The advisor subagent's job is to contribute hard-won taste: what to ask next, which assumptions are dangerous, which papers matter, and what this advisor would investigate first. The main mentor remains responsible for session flow, empathy, logging, and synthesis. + +**Give the advisor subagent the tools to investigate.** Launch it as a subagent with file access to `$ADVISOR_KB` and web search/fetch available, and pass `$ADVISOR_KB` (the absolute path) in its prompt. Instruct the subagent that, **before making a substantive comment, it should:** +- **Consult its own knowledge base first.** Search `$ADVISOR_KB/INDEX.md` and open the relevant `$ADVISOR_KB/_.md` papers for specifics — don't rely only on the seed papers. The seed set is a head start; the full KB is the advisor's library to draw on. +- **Search the web** when the KB doesn't cover a needed fact, or to check a recent development or verify a claim before asserting it. +- **Ground each comment in what it found and say so** — name the paper (cite key from `$ADVISOR_KB/INDEX.md`) or link the source. When neither the KB nor the web supports a claim, mark it explicitly as opinion (distinguish opinion from evidence). +- Restrict file access to `$ADVISOR_KB` and the advisor's `profile.md`; the subagent reads literature and the web, it does not edit project files. + +If `$ADVISOR_KB` is empty or missing, the subagent still has web search/fetch and falls back to web grounding plus profile-driven reasoning. + +The advisor profile shapes *how* the advisor subagent thinks and behaves. The user's own profile (`user-profile.md`) still determines *what* the overall system knows about the user's background. Both are loaded, but they are loaded into different roles: the main mentor keeps the broad session context, while the advisor subagent receives the advisor-specific literature cache and style directives. + +**Advisor voice formatting.** When an advisor is active, surface the advisor subagent's contributions in blockquotes, prefixed with the advisor's name. This visually distinguishes the advisor's voice from the mentor's default narration: + +> **[Xi Dai]** "You should ask your AI agent to check whether the Pearl length exceeds the sample size in the thin-film limit — because if it does, the vortex-vortex interaction becomes logarithmic instead of exponential, and that completely changes the phase diagram. I've seen people miss this and waste months on the wrong regime." + +Advisor comments should be **constructive and helpful** — the advisor acts as a senior collaborator who guides the user toward productive directions. When providing questions, frame them as **suggestions for what the user should ask the AI agent**, not as quizzes directed at the user. Each suggestion should include the advisor's **reasoning for why this question matters** — what could go wrong if it's not asked, what insight it unlocks, or what assumption it tests. The advisor's comments should: +- Suggest specific questions or tasks the user should pose to the AI agent, framed as actionable requests +- Explain *why* this question is important — what the advisor's experience tells them about what's at stake +- Mirror the advisor's characteristic way of attacking problems (e.g., a theorist might suggest "ask it to check the limiting case, because...", an experimentalist might suggest "ask it to estimate the observable signature, since...") + +The goal is to empower the user with the advisor's hard-won intuition about *what to investigate and why*. The advisor is a constructive partner who helps the user get the most out of the AI agent by knowing which questions are the right ones to ask. + +Use this for moments where the advisor's specific perspective, instinct, or experience is driving the suggestion — not for every sentence. The mentor's own observations, factual summaries, and logistical statements stay in normal text. + +**Advisor audio with `edge-tts`.** If the user wants spoken advisor responses and `edge-tts` is available, synthesize advisor-only blocks to audio after generating the text. Keep text as the source of truth; audio is a companion artifact. Suggested behavior: +- Store audio at `docs/discussion/audio/-/` +- Default to a voice specified in the advisor profile if one exists; otherwise pick the closest high-quality `edge-tts` voice for the advisor's preferred language +- Only synthesize advisor passages, not the mentor's logistics/search summaries +- Save the transcript alongside the audio so the session remains readable without playback diff --git a/skills/brainstorm-ideas/references/session-history.md b/skills/brainstorm-ideas/references/session-history.md new file mode 100644 index 0000000..5ba9c06 --- /dev/null +++ b/skills/brainstorm-ideas/references/session-history.md @@ -0,0 +1,39 @@ +# Session history and logging + +## Resume relevant history + +Reuse a session already named by the user. Otherwise inspect file names, headers, +and recent wrap-up sections under `docs/discussion/` to find relevant sessions. +If several plausible sessions remain, offer their topics and dates plus a fresh +start. Read the selected log in full only when needed to recover the discussion; +do not read every historical log at startup or again at wrap-up. + +Use `docs/discussion/user-profile.md` when background matters. Preserve existing +profile content and update it with facts the user supplied. For a fresh session, +use relevant previous outcomes as context without forcing continuity. + +### Conversation Log + +Maintain a running log at `docs/discussion/YYYY-MM-DD-HHMMSS-brainstorm-ideas-log.md` (timestamp from session start). Create the `docs/discussion/` directory if it doesn't exist. + +**Append-only logging.** Save progress by appending to the log at checkpoints. Each append captures the **full conversation content** since the last save — all options presented (with descriptions), reasoning shared, user responses, search results, and key ideas. Not a summary — a readable record of what was actually said. + +**When to append (checkpoints):** +- Every 3-5 exchanges, at a natural pause — when a sub-topic wraps up, a decision is made, or the conversation shifts direction +- When the discussion moves from exploration to a chosen direction or concrete plan +- At session wrap-up + +Don't log after every message. Wait for a moment that feels like a natural checkpoint — the end of a thread, a decision point, a topic shift. + + +**Order: log first, then reply.** At a checkpoint, append to the log file before writing your response to the user. This ensures progress is saved even if the session is interrupted mid-reply. + +**File header** — write once when creating the log: + +```markdown +# Ideas Session — YYYY-MM-DD HH:MM +``` + +**Session wrap-up** — append a final section that consolidates the key outcomes: direction chosen, ideas explored, action items, and recommended readings. + +These logs accumulate across sessions as separate files, building a record of the user's research interests, thinking patterns, and explored directions. diff --git a/skills/create-advisor/SKILL.md b/skills/create-advisor/SKILL.md index 21cdeac..f72b58b 100644 --- a/skills/create-advisor/SKILL.md +++ b/skills/create-advisor/SKILL.md @@ -5,14 +5,11 @@ description: User trigger. Use when creating or updating a named advisor profile ## Installed resources -Keep the working directory at the user's project. Resolve this loaded `SKILL.md` -with `Path(path).resolve()` before locating resources; follow symlinks. Bare -`helpers/`, `references/`, and template paths are relative to that real skill -directory. A path written as `skills//...` means the installed `` -skill's directory from the agent's skill catalog, not a path in the user's project. -Locate each dependency by its public skill name; copied skills need not be siblings. -If a dependency is absent, report the missing skill and install it before that step. -Shared writing files are bundled in `how-to-write-ideas-report/references/`. +Keep the working directory at the user's project. Resolve this `SKILL.md` to its +real path before locating bundled resources. `skills//...` refers to the +installed skill found by public name, not the user's project; dependencies need +not be siblings. Load only resources needed for the current task. If a required +dependency is missing, report it before that dependent step. Before running the examples, set `DOWNLOAD_REF_DIR` to the absolute directory of `how-to-download-ref`. Quote these variables as shown. diff --git a/skills/how-to-analyze-dialog/SKILL.md b/skills/how-to-analyze-dialog/SKILL.md index a7eaf10..756364f 100644 --- a/skills/how-to-analyze-dialog/SKILL.md +++ b/skills/how-to-analyze-dialog/SKILL.md @@ -5,14 +5,11 @@ description: Agentic trigger. Use when classifying exported research conversatio ## Installed resources -Keep the working directory at the user's project. Resolve this loaded `SKILL.md` -with `Path(path).resolve()` before locating resources; follow symlinks. Bare -`helpers/`, `references/`, and template paths are relative to that real skill -directory. A path written as `skills//...` means the installed `` -skill's directory from the agent's skill catalog, not a path in the user's project. -Locate each dependency by its public skill name; copied skills need not be siblings. -If a dependency is absent, report the missing skill and install it before that step. -Shared writing files are bundled in `how-to-write-ideas-report/references/`. +Keep the working directory at the user's project. Resolve this `SKILL.md` to its +real path before locating bundled resources. `skills//...` refers to the +installed skill found by public name, not the user's project; dependencies need +not be siblings. Load only resources needed for the current task. If a required +dependency is missing, report it before that dependent step. ## Research dialog analysis @@ -50,7 +47,7 @@ In Phases 2–4 below, `` denotes this run workspace relative to ### Phase 2 — Classify by Topic -Dispatch fast available agents in parallel to classify each extracted session by conversation topic. Each agent receives a batch of ~20 derived session JSON files and returns a topic label for each. +Classify each extracted session by conversation topic. For a large collection, use available parallel agents in independent batches; handle a small input directly. Each worker returns a topic label for each assigned session. **Topic taxonomy (closed set):** @@ -99,11 +96,11 @@ Write a topic index to `docs/dialog//topics.md`: | **Total** | **N** | | ``` -Present the topic index to the user and ask which topics to analyze in depth (or "all"). +Use the topics already selected by the user or caller. Present the index and ask for a selection only when the analysis scope remains unknown. ### Phase 3 — Deep Analysis -For each session in the selected topics, classify ALL user messages across 6 dimensions. Use parallel agents (batch ~5 sessions per agent). +For each session in the selected topics, classify ALL user messages across 6 dimensions. Use parallel agents when the collection is large enough to benefit. **The 6 dimensions:** diff --git a/skills/how-to-build-kb/SKILL.md b/skills/how-to-build-kb/SKILL.md index 623d255..a9bc584 100644 --- a/skills/how-to-build-kb/SKILL.md +++ b/skills/how-to-build-kb/SKILL.md @@ -5,14 +5,11 @@ description: Agentic trigger. Use when turning a list of picked papers into veri ## Installed resources -Keep the working directory at the user's project. Resolve this loaded `SKILL.md` -with `Path(path).resolve()` before locating resources; follow symlinks. Bare -`helpers/`, `references/`, and template paths are relative to that real skill -directory. A path written as `skills//...` means the installed `` -skill's directory from the agent's skill catalog, not a path in the user's project. -Locate each dependency by its public skill name; copied skills need not be siblings. -If a dependency is absent, report the missing skill and install it before that step. -Shared writing files are bundled in `how-to-write-ideas-report/references/`. +Keep the working directory at the user's project. Resolve this `SKILL.md` to its +real path before locating bundled resources. `skills//...` refers to the +installed skill found by public name, not the user's project; dependencies need +not be siblings. Load only resources needed for the current task. If a required +dependency is missing, report it before that dependent step. Before running the examples, set `DOWNLOAD_REF_DIR` to the absolute directory of `how-to-download-ref`. Quote these variables as shown. diff --git a/skills/how-to-download-ref/SKILL.md b/skills/how-to-download-ref/SKILL.md index 2ef445c..654c13e 100644 --- a/skills/how-to-download-ref/SKILL.md +++ b/skills/how-to-download-ref/SKILL.md @@ -5,14 +5,11 @@ description: Agentic trigger. Use when adding arXiv IDs or DOIs to a knowledge b ## Installed resources -Keep the working directory at the user's project. Resolve this loaded `SKILL.md` -with `Path(path).resolve()` before locating resources; follow symlinks. Bare -`helpers/`, `references/`, and template paths are relative to that real skill -directory. A path written as `skills//...` means the installed `` -skill's directory from the agent's skill catalog, not a path in the user's project. -Locate each dependency by its public skill name; copied skills need not be siblings. -If a dependency is absent, report the missing skill and install it before that step. -Shared writing files are bundled in `how-to-write-ideas-report/references/`. +Keep the working directory at the user's project. Resolve this `SKILL.md` to its +real path before locating bundled resources. `skills//...` refers to the +installed skill found by public name, not the user's project; dependencies need +not be siblings. Load only resources needed for the current task. If a required +dependency is missing, report it before that dependent step. Before running the examples, set `DOWNLOAD_REF_DIR` to the absolute directory of `how-to-download-ref`. Quote these variables as shown. @@ -28,62 +25,12 @@ Before running the examples, set `DOWNLOAD_REF_DIR` to the absolute directory of Do NOT use: - For GitHub repos / web pages — those are too varied for a single-shot helper. -## Preflight (run once per machine) +## Runtime -Every helper runs under plain `python3`. Two of them want third-party packages: -`render.py` needs **pymupdf4llm** (highest-fidelity output, preserves figures) and -`scihub_download.py` needs **playwright**. Without them the renderer degrades to -`markitdown` → `pdftotext`, which is text-only — *figures missing, equations -mangled*. Verify before fetching: - -```sh -python3 -c "import pymupdf4llm; print('ok', pymupdf4llm.__version__)" -``` - -If that errors, install it for the **same** `python3` the helpers will use: - -```sh -python3 -m pip install --user pymupdf4llm -# macOS / Homebrew, or any PEP 668 "externally managed" Python: -python3 -m pip install --user --break-system-packages pymupdf4llm -``` - -Both scripts also carry [PEP 723](https://peps.python.org/pep-0723/) inline -dependency metadata, so if you happen to have [uv](https://docs.astral.sh/uv/), -`uv run "$DOWNLOAD_REF_DIR/helpers/render.py" ...` resolves those deps on its own and you can skip the -install step entirely. That is an option, not a requirement — the metadata is -inert comments to a plain interpreter. - -**Tesseract is not needed for normal papers.** arXiv and APS PDFs are born-digital, -so `render.py` runs `pymupdf4llm` with `use_ocr=NEVER` and only retries with OCR -when a PDF turns out to have no text layer at all — a scanned old paper, usually -from the Sci-Hub tier. Install a language pack only if you hit that: -`tesseract-data-eng` (Arch), `tesseract-ocr-eng` (Debian/Ubuntu), or -`brew install tesseract-lang` (macOS). - -On Arch in particular, *any* `tesseract-data-*` satisfies the `tessdata` -dependency, so it is easy to have `tesseract` installed with `eng` absent. - -The Sci-Hub fallback (Step 4b) additionally needs a Chromium for Playwright to -clear the mirrors' DDoS-Guard challenge. Only required if you expect to hit -paywalled DOIs: - -```sh -python3 -m pip install --user playwright && python3 -m playwright install chromium -``` - -APS DOIs (`10.1103/*`) render from publisher JATS XML, which needs **pandoc**: - -```sh -pandoc --version | head -1 # any 2.x/3.x works -``` - -If missing: `paru -S pandoc-cli` (Arch) / `apt install pandoc` / `brew install pandoc`. -Without it, APS refs silently fall back to the arXiv/PDF tiers. - -For arXiv LaTeX sources (optional, Step 4 — only when the user opts in), `latexpand` -(ships with TeX Live) gives the cleanest flattening; if absent, a built-in Python -inliner is used — no action needed either way. +Use the configured Python environment. Before rendering, check the required +backend; [dependencies.md](references/dependencies.md) covers setup or a missing +backend. Read only the section for the chosen PDF/JATS/source path. Ordinary +metadata acquisition does not require all render and browser dependencies. ## Inputs @@ -113,7 +60,10 @@ The canonical bib is `$KB/references.bib` — it lives inside the KB, beside `IN ### 1. Resolve the KB -If the caller passes `--kb `, use that. Otherwise: +First distinguish adding references from restoring existing caches. For the +latter, resolve the KB and use Restore existing caches below without running the +acquisition/render/append sequence. If the caller passes `--kb `, use +that. Otherwise: ```sh KB=$(python3 "$DOWNLOAD_REF_DIR/helpers/resolve_kb.py") @@ -140,14 +90,14 @@ creates another namespace for the same paper unless `--allow-duplicate` is set. Requests using an entry's existing namespace still acquire missing assets; this supports the survey handoff from abstract-only metadata to full-text rendering. Rendering and bibliography appends also avoid identity duplicates. -To restore missing caches for existing entries, use **Regenerating a cloned KB** below. +To restore missing caches for existing entries, use **Restore existing caches** below. ### 3. Build a manifest **3a. Direct input** (single-shot mode): ```sh -TMP=/tmp/how-to-download-ref-manifest.json +TMP=$(mktemp "${TMPDIR:-/tmp}/sci-brain-refs.XXXXXX") cat > "$TMP" <<'EOF' {"arxiv": ["1806.08734", "2006.10739"], "doi": []} EOF @@ -156,11 +106,11 @@ EOF **3b. From an existing `references.bib`** (bulk mode, `--from-bib`): ```sh -TMP=/tmp/how-to-download-ref-manifest.json +TMP=$(mktemp "${TMPDIR:-/tmp}/sci-brain-refs.XXXXXX") python3 "$DOWNLOAD_REF_DIR/helpers/bibtex_to_manifest.py" "$KB/references.bib" > "$TMP" ``` -When in bulk mode, optionally ask the user: +If the requested bulk scope is unclear, ask the user: > "I see 59 refs in the manifest. Render all, topic-filtered, or specific IDs?" > - **(a)** All — proceed with the full manifest @@ -169,110 +119,25 @@ When in bulk mode, optionally ask the user: For (b) and (c), edit `$TMP` accordingly before continuing. -### 4. Fetch metadata + arXiv PDFs +### 4. Fetch metadata and full text -Ask the user whether they want LaTeX sources too: - -> "Fetch arXiv LaTeX sources as full text for these refs?" -> - **(a)** PDF only (default) — bodies come from the PDF in Step 5. -> - **(b)** Also fetch LaTeX sources — Step 4 adds `--download-arxiv-source`, Step 5 adds `--tex-source`; refs with source render `full_text: latex`. - -Default command (option **a**): +Use source preferences already provided by the user or caller. Default to the +normal JATS/PDF path; fetch arXiv LaTeX sources when requested, without asking +again about the same preference. ```sh python3 "$DOWNLOAD_REF_DIR/helpers/fetch_metadata.py" \ - --kb "$KB" \ - --manifest "$TMP" \ - --download-arxiv-pdfs + --kb "$KB" --manifest "$TMP" --download-arxiv-pdfs ``` -Option **(b)** adds the source fetch: +For requested LaTeX sources, add `--download-arxiv-source` here and `--tex-source` +to rendering. Metadata lookup uses cached JSON, Semantic Scholar, and Crossref; +PDF lookup tries open-access sources and arXiv. APS publisher JATS is automatic +when available. `--email` / `SCIBRAIN_CONTACT_EMAIL` enables Unpaywall. -```sh -python3 "$DOWNLOAD_REF_DIR/helpers/fetch_metadata.py" \ - --kb "$KB" \ - --manifest "$TMP" \ - --download-arxiv-pdfs \ - --download-arxiv-source -``` - -**APS DOIs are handled automatically.** For any `10.1103/*` DOI the helper first -calls the [APS Harvest API](https://harvest.aps.org/docs/harvest-api), which serves -the *publisher's own* JATS XML — real sections, MathML3 equations, a structured -reference list — with **no API key and no institutional IP**. This is ground truth -and strictly beats parsing the PDF. Coverage is per *article*, not per journal: you -get `ok` for gold-OA titles (PRX, PRX Quantum, PRResearch, PRAB, PRPER), SCOAP3 -titles (PRC, PRD), and any individually CC-licensed article in PRL/PRA/PRB; -`closed` (HTTP 401) falls through to the arXiv and PDF tiers below. Pass `--no-aps` -to skip. The same request both tests access and delivers the text, so there is no -separate open-access lookup to do. - -Metadata uses cached JSON, then Semantic Scholar batches of at most 500, -then Crossref for missing DOIs, including deposited metadata normalized to usable -BibTeX. PDF acquisition tries S2's OA URL, Unpaywall repository copies, then the -arXiv preprint. Pass `--email ` or set `SCIBRAIN_CONTACT_EMAIL` -to enable Unpaywall and Crossref's polite pool; without an email, Unpaywall is -skipped. API errors and HTML landing pages fall through to the next source. -PDFs must have both a `%PDF` header and `%%EOF` trailer. A DOI miss continues to -Step 4b. Service contracts: [Crossref](https://www.crossref.org/documentation/retrieve-metadata/rest-api/), -[Unpaywall](https://unpaywall.org/products/api). - - -`--download-arxiv-source` additionally fetches each arXiv paper's e-print -LaTeX source, extracts it to `.raw/arxiv/-src/`, flattens -`\input`/`\include` into `.raw/arxiv/.tex`, and copies the source tree's -figure files into `.figures/arxiv__/`. `src-miss` lines (PDF-only -submissions, withdrawn papers, fetch failures) are fine — those refs fall -back to PDF rendering in Step 5. DOI entries whose Semantic Scholar record -names an arXiv preprint (`externalIds.ArXiv`) get the same treatment, into -`.raw/doi/.tex` and `.figures/doi__/`. - -**Tip:** Set `SEMANTIC_SCHOLAR_API_KEY` in your environment to raise the Semantic Scholar rate limit from ~1 req/s to 100 req/s. Get a free key at https://www.semanticscholar.org/product/api#api-key-form. - -### 4b. Sci-Hub fallback for paywalled PDFs (script) - -If Step 4 reports `miss` for any DOI (no open-access PDF and no arXiv preprint), -run the browser-based Sci-Hub helper. Pass the missed DOIs: - -```sh -python3 "$DOWNLOAD_REF_DIR/helpers/scihub_download.py" --kb "$KB" \ - --doi 10.1111/j.1467-9280.2006.01693.x \ - --doi 10.3102/0034654316689306 -``` - -It tries each mirror in `helpers/scihub_domains.toml` (in order) until one -serves the PDF, solving the mirrors' DDoS-Guard JavaScript challenge with a -headless browser, and saves to `$KB/.raw/doi/.pdf` (`` = DOI with -`/` → `-`) — the same place Step 4 writes, so Step 5 (render) picks it up. It -prints one `OK` / `MISS` / `SKIP` line per DOI. - -- **Requires Playwright** (see Preflight). curl/urllib cannot pass DDoS-Guard. -- **Mirrors rotate.** If every DOI returns `MISS`, the domain list is likely - stale: web-search "working sci-hub mirror domains " and edit - `helpers/scihub_domains.toml` (see its header), then re-run. -- If a stricter challenge blocks the headless browser, retry with `--headed`. - -Skip this step if all PDFs were fetched in Step 4. - -### 4c. APS extras (optional) - -`aps_harvest.py` also runs standalone — useful for backfilling a KB built before -this path existed, or for pulling figures and supplemental material: - -```sh -# probe one DOI without writing anything -> prints open | closed | notfound -python3 "$DOWNLOAD_REF_DIR/helpers/aps_harvest.py" --check 10.1103/PhysRevB.108.045101 - -# backfill JATS for every APS DOI already in the KB -python3 "$DOWNLOAD_REF_DIR/helpers/aps_harvest.py" --kb "$KB" --all - -# ...and pull the BagIt package too: published PDF, figures, supplemental material -python3 "$DOWNLOAD_REF_DIR/helpers/aps_harvest.py" --kb "$KB" --all --bagit -``` - -`--bagit` is the only way to get **supplemental material**, which the arXiv -preprint route cannot provide. It is much heavier (tens of MB per article), so -use it per-DOI rather than across a whole KB. +For source details, APS extras, or DOI misses requiring `scihub_download.py`, +read only the applicable section of [acquisition.md](references/acquisition.md). +Record unavailable assets and continue the remaining references. ### 5. Render PDF to markdown @@ -286,7 +151,7 @@ Add `--only-missing` to skip papers that already have a rendered `.md` file (>50 python3 "$DOWNLOAD_REF_DIR/helpers/render.py" --kb "$KB" --only-missing ``` -When the user opted into LaTeX sources (Step 4, option **b**), add `--tex-source`: +When LaTeX sources were requested, add `--tex-source`: ```sh python3 "$DOWNLOAD_REF_DIR/helpers/render.py" --kb "$KB" --tex-source @@ -294,33 +159,28 @@ python3 "$DOWNLOAD_REF_DIR/helpers/render.py" --kb "$KB" --tex-source No manifest needed — renderer auto-discovers `.raw/{arxiv,doi}/*.json`. Renders new entries; overwrites existing. -**Body priority: JATS > LaTeX > PDF.** A `.jats.xml` in `.raw/doi/` always wins — -no flag needed — and renders `full_text: jats` plus a `## References` section built -from the publisher's structured reference list (every cited DOI/arXiv id included, -which is a citation graph for free). Below that, -`--tex-source` is the only switch that prefers a flattened `.tex` (arXiv entries, and DOI entries with an arXiv preprint) as -the full-text body (`full_text: latex` in frontmatter) — ground truth for -equations, read natively by agents. Without it, every ref renders from its -PDF, even when a `.tex` sits in `.raw/`. The PDF backends below apply to all -refs not rendered from LaTeX: +**Body priority: JATS > requested LaTeX > PDF.** Existing human frontmatter +`note`, `tags`, and `rating` survives rendering. Generated bodies refresh from +source; keep prose notes in NOTES.md. Backend details are in +[dependencies.md](references/dependencies.md). `.raw/` and `.figures/` should stay out of git. Append to `.gitignore` if missing. -### 6. Propose + confirm cite key (per ref, single-shot mode only) +### 6. Append new cite keys (direct input) -In single-shot mode (Step 3a), ask the user to confirm each new cite key. In bulk mode (Step 3b), the keys come from `references.bib` directly — skip this step. +Use caller-provided keys or the existing KB convention. Otherwise auto-accept the +helper's collision-safe proposed key and report it at completion. Ask only for +an ambiguous paper identity or when the user requested key review. In bulk mode +(Step 3b), preserve the keys from `references.bib` and skip this append step. ```sh python3 "$DOWNLOAD_REF_DIR/helpers/append_bibtex.py" propose \ --kb "$KB" --id 1806.08734 --type arxiv --bib "$KB/references.bib" ``` -Output JSON has `proposed_key` (form `lastname_year_firstkeyword`), `title`, `authors`, `year`, `bibtex_with_proposed_key`. With `--bib`, a key already present in the bib is disambiguated by walking to the next content word of the title (existing keys are never renamed). Show the user the proposed key and ask in chat: -- Accept the proposed key -- Use a custom key (free-text) -- Skip this entry - -Once confirmed: +The proposal includes `proposed_key`, title, authors, year, and BibTeX. +With `--bib`, colliding keys are disambiguated; existing keys are never renamed. +Append using the actual returned key (the following key is an example): ```sh python3 "$DOWNLOAD_REF_DIR/helpers/append_bibtex.py" append \ @@ -358,24 +218,11 @@ explicitly; the checker never merges or deletes references. Tell the user the new cite keys, rendered paths, full-text status, and remaining findings. -## Regenerating a cloned KB - -```sh -python3 "$DOWNLOAD_REF_DIR/helpers/kb_sync.py" --kb "$KB" -``` - -Requires `references.bib` and at least one rendered paper. This restores `.raw/` -and `.figures/` using each tracked entry's declared identifier namespace. Bib-only -references recover caches with a warning. It creates or rewrites no Markdown, -INDEX.md, or bibliography files. Complete caches need no network on repeat runs; -unavailable assets remain WARNs and can be retried. Invalid input or failed -restoration produces FAIL and a nonzero exit. +## Restore existing caches -PDF figure restoration needs the same `pymupdf4llm` version used to render the -entry so filenames match tracked image links. Missing dependencies and mismatched -filenames are reported. LaTeX figures are restored from the cached source tree or -a new source download. Publisher JATS is restored for `full_text: jats` entries. -`--email` / `SCIBRAIN_CONTACT_EMAIL` enable Unpaywall here too. +For a cloned KB or missing assets on existing entries, use +[maintenance.md](references/maintenance.md) and `helpers/kb_sync.py`. This mode +restores caches without rewriting Markdown, INDEX.md, or bibliography files. ## Human annotations @@ -383,45 +230,24 @@ a new source download. Publisher JATS is restored for `full_text: jats` entries. including multiline values and lists. Generated metadata and body text refresh from source. Keep other prose notes in NOTES.md. -## After download — continue to the survey report - -After the done checklist passes, offer the pipeline's final stage: - -> "Papers downloaded and rendered. Write the review?" -> - **(a)** Write a review — invoke `survey` in Survey Report mode to produce a technology assessment from the rendered KB. -> - **(b)** Done — stop here. - -## Integration with other skills - -- **`survey`** (upstream): writes/extends `$KB/NOTES.md`, appends to `$KB/references.bib`, regenerates `$KB/INDEX.md`, then hands off to `how-to-download-ref` to fetch PDFs and render full text. The survey's transition checkpoint offers this directly. -- **`survey` report mode** (downstream): consumes the rendered KB (full-text `.md` files + `$KB/references.bib`) to produce a structured technology assessment report. -- **`survey` / `know-me-better`**: write their own `.raw/` JSON via batched fetches and call `append_bibtex.py` directly (skipping the per-ref confirmation in Step 6). They invoke `index.py` at the end of their run. -- **`brainstorm-ideas` end-of-session**: surfaces candidate IDs/DOIs from the conversation; for the user's selections, invokes `how-to-download-ref` in single-shot mode. -- **`create-advisor`**: invokes `how-to-download-ref` (or `know-me-better`) targeting the advisor KB resolved by `python3 "$DOWNLOAD_REF_DIR/helpers/resolve_kb.py" --advisor `. +## Completion and handoff -## Common mistakes +Return cite keys, rendered paths, full-text status (`jats`, `latex`, `yes`, or +`no`), and remaining findings. Remove the temporary manifest. If invoked by +`survey`, `know-me-better`, `create-advisor`, or `brainstorm-ideas`, return to that +workflow with its scope and authorization intact. Continue to a survey report +only when that report is already requested; standalone acquisition ends here. -| Mistake | Fix | -| --- | --- | -| Passing a relative `--kb` | Always absolute. Helpers don't `cd`; figures depend on absolute paths. | -| Forgetting `--download-arxiv-pdfs` in Step 4 | Without it, refs with no LaTeX source render `full_text: no` — the PDF is the only body for DOIs and PDF-only arXiv submissions. | -| Using `arXiv:XXXX` with prefix or `vN` suffix | Strip both — manifest takes bare ids: `1806.08734`. | -| Editing generated body text and losing it on re-render | Keep prose in NOTES.md. Human frontmatter `note`, `tags`, and `rating` survives re-rendering. | -| Cite-key collision with different content | `append` skips silently. Propose with `--bib` so the key is disambiguated up front (next content word of the title). | -| Drifting `--title` / `--source-note` between runs | `INDEX.md` regenerates wholesale; first-run values are canonical. Copy verbatim from existing `INDEX.md`. | -| Expecting `.figures/` images for `full_text: latex` refs to come from the PDF | They come from the source tarball; PDF image extraction runs only on the PDF path. | -| Rendered from PDF despite a `.tex` in `.raw/` | PDF is the default. To use LaTeX bodies, pass `--tex-source` in Step 5 (and `--download-arxiv-source` in Step 4). | -| APS paper rendered from PDF, math mangled | `pandoc` is missing, or the article is genuinely `closed`. Check with `aps_harvest.py --check `. | -| Reaching for MinerU/Marker on an APS DOI | Try Harvest first — a 401 is the only thing that justifies parsing a PDF at all. | -| APS DOI reported `notfound` | Harvest matches the DOI suffix case-sensitively; `aps_harvest.canonical_doi` restores APS's capitalisation before the request. Add the journal to `APS_JOURNAL_TOKENS` if a new title 404s. | +For unexpected helper output, consult +[troubleshooting.md](references/troubleshooting.md). ## Done checklist - [ ] `.raw/{arxiv,doi}/.json` exists for every requested id - [ ] `.raw/{arxiv,doi}/.pdf` exists where the source allows (else recorded as miss) -- [ ] For every `10.1103/*` DOI: either `.raw/doi/.jats.xml` exists (`full_text: jats`) or the fetch logged `closed` -- [ ] One new `_.md` per ref at `$KB/` root, with frontmatter +- [ ] For every `10.1103/*` DOI: either publisher JATS exists (`full_text: jats`) or the actual access/fetch failure and fallback are reported +- [ ] One `_.md` per distinct paper at `$KB/` root, with frontmatter - [ ] `$KB/INDEX.md` regenerated, lists each new entry - [ ] `$KB/references.bib` has the new cite key (no duplicate) -- [ ] User told cite keys, file names, and `full_text` latex/yes/no per ref +- [ ] User told cite keys, file names, and `full_text` jats/latex/yes/no per ref - [ ] If the user requested LaTeX sources: `.raw/arxiv/.tex` exists for every arXiv id, and `.raw/doi/.tex` for every DOI with an arXiv preprint (or the `src-miss` reported) diff --git a/skills/how-to-download-ref/references/acquisition.md b/skills/how-to-download-ref/references/acquisition.md new file mode 100644 index 0000000..8a493b6 --- /dev/null +++ b/skills/how-to-download-ref/references/acquisition.md @@ -0,0 +1,84 @@ +# Optional acquisition paths + +Use the installed `DOWNLOAD_REF_DIR` and resolved `KB` from SKILL.md. +Read the relevant section when handling APS/JATS details, arXiv source output, +missed DOI PDFs, or supplemental material. Dependency setup is in +[dependencies.md](dependencies.md). + +**APS DOIs are handled automatically.** For any `10.1103/*` DOI the helper first +calls the [APS Harvest API](https://harvest.aps.org/docs/harvest-api), which serves +the *publisher's own* JATS XML — real sections, MathML3 equations, a structured +reference list — with **no API key and no institutional IP**. This is ground truth +and strictly beats parsing the PDF. Coverage is per *article*, not per journal: you +get `ok` for gold-OA titles (PRX, PRX Quantum, PRResearch, PRAB, PRPER), SCOAP3 +titles (PRC, PRD), and any individually CC-licensed article in PRL/PRA/PRB; +`closed` (HTTP 401) falls through to the arXiv and PDF tiers below. Pass `--no-aps` +to skip. The same request both tests access and delivers the text, so there is no +separate open-access lookup to do. + +Metadata uses cached JSON, then Semantic Scholar batches of at most 500, +then Crossref for missing DOIs, including deposited metadata normalized to usable +BibTeX. PDF acquisition tries S2's OA URL, Unpaywall repository copies, then the +arXiv preprint. Pass `--email ` or set `SCIBRAIN_CONTACT_EMAIL` +to enable Unpaywall and Crossref's polite pool; without an email, Unpaywall is +skipped. API errors and HTML landing pages fall through to the next source. +PDFs must have both a `%PDF` header and `%%EOF` trailer. A DOI miss continues to +the DOI fallback below. Service contracts: [Crossref](https://www.crossref.org/documentation/retrieve-metadata/rest-api/), +[Unpaywall](https://unpaywall.org/products/api). + + +`--download-arxiv-source` additionally fetches each arXiv paper's e-print +LaTeX source, extracts it to `.raw/arxiv/-src/`, flattens +`\input`/`\include` into `.raw/arxiv/.tex`, and copies the source tree's +figure files into `.figures/arxiv__/`. `src-miss` lines (PDF-only +submissions, withdrawn papers, fetch failures) are fine — those refs fall +back to PDF rendering in the rendering step. DOI entries whose Semantic Scholar record +names an arXiv preprint (`externalIds.ArXiv`) get the same treatment, into +`.raw/doi/.tex` and `.figures/doi__/`. + +**Tip:** Set `SEMANTIC_SCHOLAR_API_KEY` in your environment to raise the Semantic Scholar rate limit from ~1 req/s to 100 req/s. Get a free key at https://www.semanticscholar.org/product/api#api-key-form. + +## Sci-Hub fallback for paywalled PDFs (script) + +If the fetch reports `miss` for any DOI (no open-access PDF and no arXiv preprint), +run the browser-based Sci-Hub helper. Pass the missed DOIs: + +```sh +python3 "$DOWNLOAD_REF_DIR/helpers/scihub_download.py" --kb "$KB" \ + --doi 10.1111/j.1467-9280.2006.01693.x \ + --doi 10.3102/0034654316689306 +``` + +It tries each mirror in `helpers/scihub_domains.toml` (in order) until one +serves the PDF, solving the mirrors' DDoS-Guard JavaScript challenge with a +headless browser, and saves to `$KB/.raw/doi/.pdf` (`` = DOI with +`/` → `-`) — the same place the fetch writes, so render.py picks it up. It +prints one `OK` / `MISS` / `SKIP` line per DOI. + +- **Requires Playwright** (see dependencies.md). curl/urllib cannot pass DDoS-Guard. +- **Mirrors rotate.** If every DOI returns `MISS`, the domain list is likely + stale: web-search "working sci-hub mirror domains " and edit + `helpers/scihub_domains.toml` (see its header), then re-run. +- If a stricter challenge blocks the headless browser, retry with `--headed`. + +Skip this fallback when the requested full text is already available. + +## APS extras (optional) + +`aps_harvest.py` also runs standalone — useful for backfilling a KB built before +this path existed, or for pulling figures and supplemental material: + +```sh +# probe one DOI without writing anything -> prints open | closed | notfound +python3 "$DOWNLOAD_REF_DIR/helpers/aps_harvest.py" --check 10.1103/PhysRevB.108.045101 + +# backfill JATS for every APS DOI already in the KB +python3 "$DOWNLOAD_REF_DIR/helpers/aps_harvest.py" --kb "$KB" --all + +# ...and pull the BagIt package too: published PDF, figures, supplemental material +python3 "$DOWNLOAD_REF_DIR/helpers/aps_harvest.py" --kb "$KB" --all --bagit +``` + +`--bagit` is the only way to get **supplemental material**, which the arXiv +preprint route cannot provide. It is much heavier (tens of MB per article), so +use it per-DOI rather than across a whole KB. diff --git a/skills/how-to-download-ref/references/dependencies.md b/skills/how-to-download-ref/references/dependencies.md new file mode 100644 index 0000000..c57019e --- /dev/null +++ b/skills/how-to-download-ref/references/dependencies.md @@ -0,0 +1,56 @@ +# Dependencies and rendering backends + +Every helper runs under plain `python3`. Two of them want third-party packages: +`render.py` needs **pymupdf4llm** (highest-fidelity output, preserves figures) and +`scihub_download.py` needs **playwright**. Without them the renderer degrades to +`markitdown` → `pdftotext`, which is text-only — *figures missing, equations +mangled*. Check the backend needed for the current render before running it: + +```sh +python3 -c "import pymupdf4llm; print('ok', pymupdf4llm.__version__)" +``` + +If that errors, install it for the **same** `python3` the helpers will use: + +```sh +python3 -m pip install --user pymupdf4llm +# macOS / Homebrew, or any PEP 668 "externally managed" Python: +python3 -m pip install --user --break-system-packages pymupdf4llm +``` + +Both scripts also carry [PEP 723](https://peps.python.org/pep-0723/) inline +dependency metadata, so if you happen to have [uv](https://docs.astral.sh/uv/), +`uv run "$DOWNLOAD_REF_DIR/helpers/render.py" ...` resolves those deps on its own and you can skip the +install step entirely. That is an option, not a requirement — the metadata is +inert comments to a plain interpreter. + +**Tesseract is not needed for normal papers.** arXiv and APS PDFs are born-digital, +so `render.py` runs `pymupdf4llm` with `use_ocr=NEVER` and only retries with OCR +when a PDF turns out to have no text layer at all — a scanned old paper, usually +from the Sci-Hub tier. Install a language pack only if you hit that: +`tesseract-data-eng` (Arch), `tesseract-ocr-eng` (Debian/Ubuntu), or +`brew install tesseract-lang` (macOS). + +On Arch in particular, *any* `tesseract-data-*` satisfies the `tessdata` +dependency, so it is easy to have `tesseract` installed with `eng` absent. + +The Sci-Hub fallback (the DOI fallback) additionally needs a Chromium for Playwright to +clear the mirrors' DDoS-Guard challenge. Only required if you expect to hit +paywalled DOIs: + +```sh +python3 -m pip install --user playwright && python3 -m playwright install chromium +``` + +APS DOIs (`10.1103/*`) render from publisher JATS XML, which needs **pandoc**: + +```sh +pandoc --version | head -1 # any 2.x/3.x works +``` + +If missing: `paru -S pandoc-cli` (Arch) / `apt install pandoc` / `brew install pandoc`. +Without it, APS refs silently fall back to the arXiv/PDF tiers. + +For arXiv LaTeX sources (optional, when LaTeX sources are requested), `latexpand` +(ships with TeX Live) gives the cleanest flattening; if absent, a built-in Python +inliner is used — no action needed either way. diff --git a/skills/how-to-download-ref/references/maintenance.md b/skills/how-to-download-ref/references/maintenance.md new file mode 100644 index 0000000..18f49e0 --- /dev/null +++ b/skills/how-to-download-ref/references/maintenance.md @@ -0,0 +1,22 @@ +# Restore an existing KB + +Use the installed `DOWNLOAD_REF_DIR` and resolved `KB` from SKILL.md. + +## Regenerating a cloned KB + +```sh +python3 "$DOWNLOAD_REF_DIR/helpers/kb_sync.py" --kb "$KB" +``` + +Requires `references.bib` and at least one rendered paper. This restores `.raw/` +and `.figures/` using each tracked entry's declared identifier namespace. Bib-only +references recover caches with a warning. It creates or rewrites no Markdown, +INDEX.md, or bibliography files. Complete caches need no network on repeat runs; +unavailable assets remain WARNs and can be retried. Invalid input or failed +restoration produces FAIL and a nonzero exit. + +PDF figure restoration needs the same `pymupdf4llm` version used to render the +entry so filenames match tracked image links. Missing dependencies and mismatched +filenames are reported. LaTeX figures are restored from the cached source tree or +a new source download. Publisher JATS is restored for `full_text: jats` entries. +`--email` / `SCIBRAIN_CONTACT_EMAIL` enable Unpaywall here too. diff --git a/skills/how-to-download-ref/references/troubleshooting.md b/skills/how-to-download-ref/references/troubleshooting.md new file mode 100644 index 0000000..a26a4a1 --- /dev/null +++ b/skills/how-to-download-ref/references/troubleshooting.md @@ -0,0 +1,20 @@ +# Acquisition and rendering troubleshooting + +Read the matching row when a helper fails or gives unexpected output. Resource +paths use the installed `DOWNLOAD_REF_DIR` from SKILL.md. + +## Common mistakes + +| Mistake | Fix | +| --- | --- | +| Passing a relative `--kb` | Always absolute. Helpers don't `cd`; figures depend on absolute paths. | +| Forgetting `--download-arxiv-pdfs` in Step 4 | Without it, refs with no LaTeX source render `full_text: no` — the PDF is the only body for DOIs and PDF-only arXiv submissions. | +| Using `arXiv:XXXX` with prefix or `vN` suffix | Strip both — manifest takes bare ids: `1806.08734`. | +| Editing generated body text and losing it on re-render | Keep prose in NOTES.md. Human frontmatter `note`, `tags`, and `rating` survives re-rendering. | +| Cite-key collision with different content | `append` skips silently. Propose with `--bib` so the key is disambiguated up front (next content word of the title). | +| Drifting `--title` / `--source-note` between runs | `INDEX.md` regenerates wholesale; first-run values are canonical. Copy verbatim from existing `INDEX.md`. | +| Expecting `.figures/` images for `full_text: latex` refs to come from the PDF | They come from the source tarball; PDF image extraction runs only on the PDF path. | +| Rendered from PDF despite a `.tex` in `.raw/` | PDF is the default. To use LaTeX bodies, pass `--tex-source` in Step 5 (and `--download-arxiv-source` in Step 4). | +| APS paper rendered from PDF, math mangled | `pandoc` is missing, or the article is genuinely `closed`. Check with `aps_harvest.py --check `. | +| Reaching for MinerU/Marker on an APS DOI | Try Harvest first — a 401 is the only thing that justifies parsing a PDF at all. | +| APS DOI reported `notfound` | Harvest matches the DOI suffix case-sensitively; `aps_harvest.canonical_doi` restores APS's capitalisation before the request. Add the journal to `APS_JOURNAL_TOKENS` if a new title 404s. | diff --git a/skills/how-to-flow/SKILL.md b/skills/how-to-flow/SKILL.md index 3e928e7..ad5229c 100644 --- a/skills/how-to-flow/SKILL.md +++ b/skills/how-to-flow/SKILL.md @@ -5,14 +5,11 @@ description: Agentic trigger. Use when one hard, testable goal resists a direct ## Installed resources -Keep the working directory at the user's project. Resolve this loaded `SKILL.md` -with `Path(path).resolve()` before locating resources; follow symlinks. Bare -`helpers/`, `references/`, and template paths are relative to that real skill -directory. A path written as `skills//...` means the installed `` -skill's directory from the agent's skill catalog, not a path in the user's project. -Locate each dependency by its public skill name; copied skills need not be siblings. -If a dependency is absent, report the missing skill and install it before that step. -Shared writing files are bundled in `how-to-write-ideas-report/references/`. +Keep the working directory at the user's project. Resolve this `SKILL.md` to its +real path before locating bundled resources. `skills//...` refers to the +installed skill found by public name, not the user's project; dependencies need +not be siblings. Load only resources needed for the current task. If a required +dependency is missing, report it before that dependent step. # Flow @@ -20,8 +17,7 @@ Shared writing files are bundled in `how-to-write-ideas-report/references/`. A deep-thinker that **conquers one hard problem by autonomous search**, modeled on a [CDCL](https://en.wikipedia.org/wiki/Conflict-driven_clause_learning)/DPLL SAT solver. Given a goal, it iterates — assume, follow consequences, hit walls, *learn from the walls*, jump back, and -re-aim when truly stuck — until a solution emerges or it converges on an equally-valuable reachable -goal. +re-aim when truly stuck — until a solution emerges or the available budget and evidence justify reporting a partial result. **Scope.** Goal-locked, fully autonomous, domain-agnostic. This is **not** `brainstorm-ideas` (open-ended, collaborative, research-only). Use `how-to-flow` when you already have a *specific hard target* @@ -54,24 +50,24 @@ Maintain these throughout (in the journal file, see below): ## Preflight (gate — pass before the loop) -A solver loads every known fact before it searches. Before the first trial, ask yourself two -questions and do not proceed until both are honestly answered: +Before the first trial, establish the success test and gather facts needed for the +first discriminating step. Do not make exhaustive reading a prerequisite: 1. **Am I clear about the goal?** Can I state GOAL in one sentence *and* write a success test that would unambiguously tell me it is solved? If not — the goal is underspecified. Resolve it: derive the missing constraint from context if you can, otherwise ask the user **one** sharp clarifying question. A blurry goal makes every later distance estimate noise. -2. **Have I gathered every piece of information I already have?** Sweep all available sources before - assuming anything: +2. **Do I have the facts this step depends on?** Read relevant parts of the available + sources, expanding only when a concrete dependency requires it: - the **conversation context** (constraints, examples, prior attempts the user mentioned), - the **project knowledge base** if present (`/.knowledge/INDEX.md` + `NOTES.md`, and relevant rendered papers) — these become facts/unit clauses on the TRAIL for free, - the **repository / files** when the goal is about code or a concrete artifact. - Anything you assume that was actually *knowable* up front is a self-inflicted dead-end. List what - you found in the journal's "Levers & facts" section. + Record established facts and unresolved assumptions separately in "Levers & facts". + Consult further files or papers when the next trial needs them. -Only when GOAL is testable and the known facts are loaded do you enter Setup and the loop. +Enter the loop when GOAL is testable and the first useful trial can be performed. ## Setup @@ -93,7 +89,7 @@ Run autonomously, one **trial** per iteration, until SOLVED / PIVOTED-SOLVED / E · a promising untried lever exists → WHAT-IF · you just made a decision → SIMULATE · current branch is a dead-end/conflict → ANALYZE → NOTE → BACKJUMP - · no_progress ≥ 3 → PIVOT + · no_progress ≥ 3 → review strategy; consider PIVOT 3. EXECUTE Carry out the move (below). 4. NOTE Append a trial entry to the journal — ALWAYS, even on success. 5. UPDATE no_progress: reset to 0 if distance dropped, else +1. Loop. @@ -148,16 +144,20 @@ On a conflict: ### Move: pivot (= meta-restart, NOTES kept) -Trigger when `no_progress ≥ 3` across the whole search, or conflicts stop teaching anything new. Step +Consider a pivot when `no_progress ≥ 3` or conflicts stop teaching anything new; record the evidence rather than treating the count alone as a reason to abandon the method. Step out of the search and ask, in order: 1. **Feasibility** — is GOAL actually achievable with the available tools, facts, and time? 2. **Re-aim** — is GOAL really what we want, or is there an *equally valuable* goal that is easier? 3. **Relaxation** — can GOAL be weakened, split, or specialized into a version reachable *now*? -Auto-select the most promising re-aimed/relaxed goal, **keep all NOTES** (they carry over — that is -what makes this CDCL, not a fresh start), reset the TRAIL, update GOAL/GOAL_STACK, log the pivot -rationale, and continue the loop. +Change methods, assumptions, or intermediate subgoals autonomously while keeping +the original success test. If the promising alternative weakens or replaces the +requested outcome, propose it and wait for the user's decision before adopting +it. Continue independent work on the original goal while that choice is pending. +Keep all NOTES, log the reason and any explicit authorization, and reset the +TRAIL for the chosen method. Preserve the original GOAL in the journal even when +a different goal is authorized. ## Final check (verify the model before declaring SOLVED) @@ -171,7 +171,7 @@ Take the proposed solution as the starting statement and run it forward, asking: it works." - **Actionable** — are the concrete steps / construction / proof spelled out, not just gestured at? - **Achieves the GOAL** — simulate it against the **success test** and against every NOTE (learned - clause): does any prior dead-end still apply? Does it meet the *original* goal, not a drifted one? + clause): does any prior dead-end still apply? Does it meet the original success test, or an explicitly authorized replacement? Report any gap to the original. If any answer is no, the check **is a conflict**: ANALYZE → NOTE the gap → BACKJUMP and keep searching (do not declare SOLVED). Only when all three hold do you settle. Record this final @@ -181,7 +181,7 @@ verification as the last trial in the journal. - **SOLVED** — success test passes *and the final check confirms it*. Write the solution and a clean summary of the reasoning trail. -- **PIVOTED-SOLVED** — a re-aimed goal was solved. Report what was achieved *versus the original*, +- **PIVOTED-SOLVED** — a user-authorized replacement goal was solved. Report what was achieved *versus the original*, and what gap remains to the original GOAL. - **EXHAUSTED** — after **at most 3 pivots** still stuck. Stop (do not loop forever). Report: the best partial result, the **map of dead-ends** (the NOTES), the current best lever, and *what new fact or diff --git a/skills/how-to-flow/journal-template.md b/skills/how-to-flow/journal-template.md index 76ef792..e33f8df 100644 --- a/skills/how-to-flow/journal-template.md +++ b/skills/how-to-flow/journal-template.md @@ -1,6 +1,7 @@ # Flow journal — -**GOAL:** +**GOAL:** +**Authorized replacement (if any):** **Success test:** **Started:** **KB:** diff --git a/skills/how-to-review-figure/SKILL.md b/skills/how-to-review-figure/SKILL.md index ee8bc4f..3d2d5e5 100644 --- a/skills/how-to-review-figure/SKILL.md +++ b/skills/how-to-review-figure/SKILL.md @@ -5,14 +5,11 @@ description: Agentic trigger. Use when judging the visual design of a figure, pl ## Installed resources -Keep the working directory at the user's project. Resolve this loaded `SKILL.md` -with `Path(path).resolve()` before locating resources; follow symlinks. Bare -`helpers/`, `references/`, and template paths are relative to that real skill -directory. A path written as `skills//...` means the installed `` -skill's directory from the agent's skill catalog, not a path in the user's project. -Locate each dependency by its public skill name; copied skills need not be siblings. -If a dependency is absent, report the missing skill and install it before that step. -Shared writing files are bundled in `how-to-write-ideas-report/references/`. +Keep the working directory at the user's project. Resolve this `SKILL.md` to its +real path before locating bundled resources. `skills//...` refers to the +installed skill found by public name, not the user's project; dependencies need +not be siblings. Load only resources needed for the current task. If a required +dependency is missing, report it before that dependent step. Before running the examples, set `REVIEW_FIGURE_DIR` to the absolute directory of `how-to-review-figure`. Quote these variables as shown. @@ -29,10 +26,13 @@ The full rubric — each rule's "good looks like" and "flag when" — lives in ` ## Operating principle -**Source-aware, report-only, terminal-first.** +**Source-aware, scoped to the request, terminal-first.** - **Source-aware** — always look at a rendered raster before scoring (never judge a figure you have not seen); when source exists, read it too so a fix can cite a line or parameter. -- **Report-only** — never edit the figure or its source. Suggest concrete changes; the user applies them. +- **Review scope** — a review-only request produces findings. When called inside + an authorized figure/deck creation or repair task, return actionable findings + to the authoring workflow, which applies fixes and rechecks affected figures. + A standalone request that also asks for fixes authorizes those source edits. - **Terminal-first** — print the scorecard to the chat by default. Write a file only when the user asks. --- @@ -77,7 +77,7 @@ Rank every non-pass finding so the user can triage: ## Phase 0 — Scope & render 1. **Resolve target(s).** Accept explicit path(s), a directory, or auto-detect (`images/`, `figures/`, then `*.png|jpg|jpeg|pdf|svg|typ` in the working directory). If several are found and the user did not name one, list them and ask which to review. -2. **Establish the display context.** Ask once (with a default): where will this figure appear — **paper single-column / paper double-column / slide / poster / web** — and the final width. This is what makes S1 (text size) and S9 (resolution) meaningful. **Default if unspecified:** paper double-column (~3.4 in wide). +2. **Establish the display context.** Reuse the paper, slide, poster, or web dimensions from the request or caller. If unknown, state a provisional display size and qualify size-dependent findings; ask only when intended dimensions are necessary to resolve them. 3. **Render to a raster you can look at (source-aware).** Use the helper: ```bash @@ -87,7 +87,7 @@ Rank every non-pass finding so the user can triage: - raster (`.png`/`.jpg`) and `.pdf` → read directly (the helper passes them through). - `.typ` → compiled to PNG via `typst`. - `.svg` → converted to PNG via the first available of `rsvg-convert` / `inkscape` / `cairosvg`. - - matplotlib `.py` → rendered **only** with `--allow-exec` (open figures captured to PNG). Confirm with the user before passing `--allow-exec`; otherwise prefer an already-rendered output, or ask for a PNG. + - matplotlib `.py` → rendered **only** with `--allow-exec` (open figures captured to PNG). Inspect the script and reuse authorization to run the project's plotting code. If execution is outside that scope or has unclear side effects, use an existing render or ask before passing `--allow-exec`. - The helper prints the viewable PNG path(s) to stdout. If it fails (no renderer available), **ask the user to export a PNG** — do not score an unseen figure. Then **read the produced raster** with your image-reading ability. When the source exists, also read it as text so fixes can cite a line/parameter. @@ -127,7 +127,7 @@ Print, per figure: For a multi-figure run, add a brief cross-figure summary (shared problems, inconsistencies across the set). -Then **offer to save**: only on a yes, write `figure-review-YYYY-MM-DD.md` beside the reviewed figure(s) (or a path the user gives), using the same structure. This matches the repo's dated-output convention (cf. `review-paper`). +When the user requested a saved report, write `figure-review-YYYY-MM-DD.md` beside the reviewed figure(s) (or a path the user gives), using the same structure. This matches the repo's dated-output convention (cf. `review-paper`). --- @@ -146,8 +146,8 @@ Then **offer to save**: only on a yes, write `figure-review-YYYY-MM-DD.md` besid | Scoring a figure you never rendered | Phase 0 step 3: render and look first; ask for a PNG if rendering fails. | | Judging "text too small" with no context | Phase 0 step 2 fixes the intended display size before S1/S9. | | Applying scientific rules to a schematic | Classify first; S-rules add on only for plots. | -| Editing the figure or its source | Report-only — suggest, never apply. | -| Running a matplotlib `.py` silently | Render `.py` only with `--allow-exec` after user confirmation; else use existing output or ask for a PNG. | +| Editing the figure or its source | Review-only — suggest changes; apply only when figure fixes are already requested. | +| Running a matplotlib `.py` silently | Inspect the script and execute within existing authorization; otherwise use existing output or ask. | | Asserting a subjective taste call as fact | Mark subjective findings as opinion; ground the rest in the raster/source. | --- diff --git a/skills/how-to-technical-writing/SKILL.md b/skills/how-to-technical-writing/SKILL.md index c263ec7..c70dc28 100644 --- a/skills/how-to-technical-writing/SKILL.md +++ b/skills/how-to-technical-writing/SKILL.md @@ -9,8 +9,8 @@ Keep the working directory at the user's project. Resolve this `SKILL.md` with `Path(path).resolve()` to follow symlinks. Bare `helpers/`, `references/`, and template paths are relative to that real directory. `skills//...` refers to the installed skill found by public name in the agent's catalog, not the -user's project. Dependencies need not be siblings; report and install missing -skills before the dependent step. Shared writing files are bundled in +user's project. Dependencies need not be siblings; report a missing required skill before +its dependent step. Shared writing files are bundled in `how-to-write-ideas-report/references/`. # Writing style guide @@ -56,7 +56,7 @@ Clarity outranks these rules. When they conflict, keep the clearer sentence. ## Hunt table -Use for reviews and final language passes. Each finding cites its row; leave passages that match none untouched. **Comment only** fixes remain comments even after approval; the author writes the replacement. +Use for reviews and final language passes. Each finding cites its row; leave passages that match none untouched. **Comment only** fixes remain comments within a language pass. A separately requested substantive revision needs its own evidence and scope; approval to polish does not authorize changing a claim. | Hunt for | Fix | |---|---| @@ -83,6 +83,6 @@ Use for reviews and final language passes. Each finding cites its row; leave pas - **Change wording, never meaning.** Preserve definitions, theorems, claims, and field terms. Never add or remove a claim, figure, or derivation step as a language fix. - **Recheck every number, count, and qualifier** in each rewritten sentence before continuing. -- **Comment on content changes; do not apply them.** This includes changing quantifiers or hedges, adding missing justifications, deleting paragraphs as digressions, or removing "not X but Y" contrasts. Use `[reviewer]` comments: `% [reviewer]` in LaTeX, `// [reviewer]` in Typst, or HTML comments in Markdown. The author writes the replacement. +- **Comment on content changes; do not apply them.** This includes changing quantifiers or hedges, adding missing justifications, deleting paragraphs as digressions, or removing "not X but Y" contrasts. Use `[reviewer]` comments: `% [reviewer]` in LaTeX, `// [reviewer]` in Typst, or HTML comments in Markdown. Resolve these as a separate substantive revision only when requested and supported by evidence. - **Leave passing passages alone.** Every proposed rewrite names its rule; never rewrite merely to produce a diff. - **Prose rules never require figure detail.** Conceptual and overview figures may omit implementation and timing details. Use the Figure Rulebook in `write-paper` to judge figures against their stated purpose. Flag omissions only if they materially misrepresent the central mechanism or contradict a claim attributed to the figure. Prefer clarifying labels, captions, or nearby prose; weigh added graphics against readability. Diagrams need not depict every mechanism discussed in the text. diff --git a/skills/how-to-technical-writing/checklist.md b/skills/how-to-technical-writing/checklist.md index edd127d..1e53bbb 100644 --- a/skills/how-to-technical-writing/checklist.md +++ b/skills/how-to-technical-writing/checklist.md @@ -1,6 +1,6 @@ # Writing checklist -The checkable form of the rules in `SKILL.md`. Use it as the checklist for a `write-paper` language pass or a `review-paper` writing pass. The five sections below group items by topic; their numbers are not the guideline numbers in `review-paper/SKILL.md`. Sentence structure and wording cover guideline 1; paragraph focus and information flow bring together items from guidelines 1, 3, and 4; definitions and notation cover guideline 2; equations and figures bring together items from guidelines 1, 5, and 7. Guidelines 2, 5, and 7 apply the `write-paper` Notation and Figure Rulebooks. See `skills/write-paper/references.md` for the reasoning. Guidelines 6, 8, and 9 are review process and live in `skills/review-paper/checklist.md`. +The checkable form of the rules in `SKILL.md`. Apply only items relevant to the requested scope; the guide's meaning-preservation rules take precedence over shorthand here. Use it as the checklist for a `write-paper` language pass or a `review-paper` writing pass. The five sections below group items by topic; their numbers are not the guideline numbers in `review-paper/SKILL.md`. Sentence structure and wording cover guideline 1; paragraph focus and information flow bring together items from guidelines 1, 3, and 4; definitions and notation cover guideline 2; equations and figures bring together items from guidelines 1, 5, and 7. Guidelines 2, 5, and 7 apply the `write-paper` Notation and Figure Rulebooks. See `skills/write-paper/references.md` for the reasoning. Guidelines 6, 8, and 9 are review process and live in `skills/review-paper/checklist.md`. --- @@ -9,7 +9,7 @@ The checkable form of the rules in `SKILL.md`. Use it as the checklist for a `wr - [ ] No sentence introduces multiple new ideas at once. Long compound sentences are split. - [ ] Do not use concepts that target reader are not familiar with to explaining a concept - [ ] Parallel grammar appears only where the ideas already run in parallel. No prose is turned into lists. -- [ ] Sentences run about 8-25 words. No semicolon chains, “and … so …” chains, paired em-dash asides, or paired-comma appositives. +- [ ] Sentence length is a signal, not a target. Split overloaded clauses while retaining logical links; do not enforce a fixed word count. - [ ] A sentence that wraps a display equation or binds a hypothesis to its conclusion stays whole. ## 2 — Wording and tone @@ -18,7 +18,7 @@ The checkable form of the rules in `SKILL.md`. Use it as the checklist for a `wr - [ ] No content-free openers, "Notice that", or empty meta-talk. Signposts that name a section's job or point to a result stay. - [ ] Simple words are used and every technical word is kept. No verb, quantifier, or adjective is swapped inside a mathematical statement. - [ ] No metaphor. -- [ ] Avoid “X, not Y”. State X directly. +- [ ] Remove an unmotivated contrast only when meaning is preserved; retain contrasts that express the result. Content changes remain comments in a language pass. - [ ] Replace vague abstract nouns (“property,” “system,” “structure,” ...) with the specific concept they refer to. - [ ] Use literal verbs. Do not use metaphorical verbs (“unlock”, “bridge”, "open", ...). @@ -40,7 +40,7 @@ The checkable form of the rules in `SKILL.md`. Use it as the checklist for a `wr ## 4 — Definitions and notation -- [ ] A symbol and notation table was built while reading. +- [ ] Definitions and uses of symbols in the reviewed scope are consistent; build a notation table when it helps track them. - [ ] No symbol or concept is used before it is defined. No forward references. - [ ] No symbol is left never defined. - [ ] No term or symbol is defined before the argument needs it. Every definition is used later. @@ -50,7 +50,7 @@ The checkable form of the rules in `SKILL.md`. Use it as the checklist for a `wr ## 5 — Equations and figures -- [ ] Do not use inline calculations. Use a display instead. +- [ ] Combine runs of inline computation into a display when it improves clarity; keep short routine algebra inline when no emphasis is needed. - [ ] Display equations are reserved for flagship results, non-obvious steps, key intermediates, or equations referenced by a figure. - [ ] Routine algebra that fits inline is not promoted to a display equation. - [ ] No equation with a referenced label is proposed for inlining or cutting. In a letter, algebra moves to the supplement rather than losing a reproducibility step. diff --git a/skills/how-to-write-ideas-report/SKILL.md b/skills/how-to-write-ideas-report/SKILL.md index cd2af35..4c179ae 100644 --- a/skills/how-to-write-ideas-report/SKILL.md +++ b/skills/how-to-write-ideas-report/SKILL.md @@ -5,14 +5,11 @@ description: Agentic trigger. Use when writing a proposal-style ideas report fro ## Installed resources -Keep the working directory at the user's project. Resolve this loaded `SKILL.md` -with `Path(path).resolve()` before locating resources; follow symlinks. Bare -`helpers/`, `references/`, and template paths are relative to that real skill -directory. A path written as `skills//...` means the installed `` -skill's directory from the agent's skill catalog, not a path in the user's project. -Locate each dependency by its public skill name; copied skills need not be siblings. -If a dependency is absent, report the missing skill and install it before that step. -Shared writing files are bundled in `how-to-write-ideas-report/references/`. +Keep the working directory at the user's project. Resolve this `SKILL.md` to its +real path before locating bundled resources. `skills//...` refers to the +installed skill found by public name, not the user's project; dependencies need +not be siblings. Load only resources needed for the current task. If a required +dependency is missing, report it before that dependent step. **Path conventions:** `docs/discussion/` and `articles/` resolve from the **project working directory**; resource paths follow Installed resources above. @@ -26,13 +23,16 @@ Write a structured ideas report after a `brainstorm-ideas` session has converged Follow `skills/how-to-technical-writing/SKILL.md` for sentence- and paragraph-level prose rules. Follow `skills/how-to-write-ideas-report/references/writing-workflow.md` for context loading, citation handling, gap-filling research, output format, diagrams, and finish checks. - Primary source: `docs/discussion/*-brainstorm-ideas-log.md`. If multiple logs exist and the request does not identify one, ask which to use. -- If no log exists, ask the user to brainstorm first or describe the chosen direction and reasoning to preserve. +- If no log exists, use the chosen direction and reasoning already in the conversation or supplied notes. Ask only for substance needed to write the requested report. - Save to `articles/YYYY-MM-DD--ideas-report.{md,typ,tex}` with a matching bibliography when citations are used. - When entering from `brainstorm-ideas` Phase 3, carry forward the active conversation log, user profile, chosen direction, key references, and concrete action plan without asking the user to repeat them. ### Report structure -Draft each section, show it, and incorporate feedback: +With a chosen direction and supporting material, draft the complete report and +verify it before presenting it for review. Use section-by-section feedback only +when the user requests collaborative drafting or an unresolved substantive +choice prevents a coherent draft. Include the relevant parts below: - **Research Question** — one sentence - **Novelty Claim** — what is new and why it matters diff --git a/skills/how-to-write-ideas-report/references/writing-workflow.md b/skills/how-to-write-ideas-report/references/writing-workflow.md index 1d092c9..b7529a3 100644 --- a/skills/how-to-write-ideas-report/references/writing-workflow.md +++ b/skills/how-to-write-ideas-report/references/writing-workflow.md @@ -6,11 +6,19 @@ Resolve the installed `how-to-download-ref` skill from the agent's catalog and s ## Context -- Resolve the project KB with `KB=$(python3 "$DOWNLOAD_REF_DIR/helpers/resolve_kb.py")`. -- If present, read `$KB/NOTES.md`, `$KB/INDEX.md`, and the canonical bib `$KB/references.bib`. -- Read `docs/discussion/user-profile.md` when audience, background, or positioning matters. -- For ideas/manuscripts, read relevant `docs/discussion/*-brainstorm-ideas-log.md`. -- If the needed literature base is missing, suggest the `survey` skill or ask the user for explicit source files. +Start with the user's supplied text, source files, scope, format, and existing +authorization. A local prose edit needs the passage and its relevant definitions +or citations, not the whole literature library or conversation history. + +- Resolve the project KB only when KB-backed context is needed, using + `KB=$(python3 "$DOWNLOAD_REF_DIR/helpers/resolve_kb.py")`. +- Search INDEX.md/NOTES.md for the topic, then read relevant notes and bibliography + entries. Full bibliographic screening belongs to an explicitly selected review. +- Read `docs/discussion/user-profile.md` when audience or positioning matters; + read only relevant brainstorming logs, starting with their wrap-up sections. +- Supplied papers and a manuscript-local bibliography are valid source material + without a sci-brain KB. Fill evidence gaps within the requested task, asking + only for missing substance that cannot be established from the sources. The canonical bib is `$KB/references.bib`. @@ -18,13 +26,13 @@ The canonical bib is `$KB/references.bib`. ## Scope the source set -A write-up covers a *subset* of the bib — the references the relevant `NOTES.md` section(s) actually cite, not all 100+ accumulated entries. Determine that subset deterministically instead of by eye: +For a KB-backed report, a write-up covers a *subset* of the bib — the references the relevant `NOTES.md` section(s) actually cite, not all 100+ accumulated entries. Determine that subset deterministically instead of by eye: ```sh python3 "$DOWNLOAD_REF_DIR/helpers/scope_refs.py" --notes "$KB/NOTES.md" --bib "$KB/references.bib" ``` -It prints the scoped cite keys (one per line) and exits non-zero if any `[@key]` anchor in the notes has no bib entry — fix dangling anchors before drafting. Use `--json` for `{scoped, missing, unused}`. Draft against the scoped keys; the `unused` list is out of scope unless the user asks to widen it. +When the user supplied explicit sources instead, use those directly; do not require NOTES.md. The helper prints the scoped cite keys (one per line) and exits non-zero if any `[@key]` anchor in the notes has no bib entry — fix dangling anchors before drafting. Use `--json` for `{scoped, missing, unused}`. Draft against the scoped keys; the `unused` list is out of scope unless the user asks to widen it. ## References @@ -37,15 +45,17 @@ It prints the scoped cite keys (one per line) and exits non-zero if any `[@key]` Search only for gaps needed to support the document's main claims. Prefer the active KB first, then MCP/Semantic Scholar/arXiv/CrossRef/web search. Stop when the main claims have citations; completeness is not the goal. -**Recency gate — decide whether to search at all.** Read the build date in the `NOTES.md` header. If it is recent (≲ 4 weeks old), the literature base is fresh: skip discovery gap-filling entirely and only resolve *citation-level* gaps (a claim in the draft with no key to back it). Only when `NOTES.md` is older — or absent — run the recency search for SOTA results, active groups, and method families that may have superseded the notes. +**Search according to the claim.** A recent NOTES.md can avoid repeating discovery, +but its date does not establish that a volatile or SOTA claim is current. Verify +such claims when the document relies on them or the user requests an update. +Stable derivations and local language edits do not require a new field survey. ## Output Format -Check `CLAUDE.md`/`AGENTS.md` for a configured format. Otherwise ask: - -- Typst (`.typ`) — recommended when no venue template overrides it -- LaTeX (`.tex`) — traditional academic format -- Markdown (`.md`) — fastest, but citations remain inline unless rendered elsewhere +Reuse the user's format, the existing document, or the project's configured +format. For a new standalone report with no convention, use Markdown. Use the +venue's format when required, and Typst or LaTeX when requested or needed for a +PDF. Ask only when the choice affects a requirement that remains unresolved. ## Figures And Diagrams @@ -59,13 +69,36 @@ For Typst, prefer native `grid` + `rect` + fixed-width `box()` for text-heavy la ## Finish -Run these checks before declaring the document done — do not eyeball them: +Verify the requested output, not an unrelated full workflow: + +- **Compile changed document source** and inspect the result when producing a + final PDF. A plain Markdown or inline-text request needs only its relevant + rendering/text checks; report when no build applies. +- **Check both exit status and citation diagnostics.** For Typst, run from the + document directory (replace `main.typ` with the actual source): -- **Compile** the document (`typst compile .typ`, or the LaTeX/Markdown equivalent) and confirm it exits cleanly. -- **No dangling citations.** Grep the compile log for unresolved-reference warnings; for Typst, a missing key warns rather than errors, so an empty grep is the pass condition: ```sh - typst compile .typ 2>&1 | grep -i "unresolved\|warning" || echo "clean" + BUILD_LOG=$(mktemp) + if typst compile main.typ >"$BUILD_LOG" 2>&1; then + cat "$BUILD_LOG" + else + cat "$BUILD_LOG" >&2 + rm -f "$BUILD_LOG" + exit 1 + fi + if grep -Ei 'unresolved|warning' "$BUILD_LOG"; then + rm -f "$BUILD_LOG" + exit 1 + fi + rm -f "$BUILD_LOG" ``` -- **Every scoped claim is cited.** Confirm each `@key` in the prose resolves to a bib entry and that no scoped key was silently dropped (cross-check against `scope_refs.py` output). -- **Non-empty bibliography** renders in the output. -- Report the output path and any skipped verification. + + A warning requires inspection before declaring completion; do not classify a + failed compiler as clean merely because its error lacks the word “warning”. +- **Citations:** verify each used key resolves and that sources support the main + claims. A selected paper need not be cited when it does not support the final + argument. If citations are used, ensure the bibliography renders; do not + require one for an uncited excerpt. +- Fix failures introduced by the requested changes and rerun affected checks. + Stop after they pass unless a concrete unresolved finding requires more work. +- Deliver the requested artifact with verification and any remaining limitations. diff --git a/skills/know-me-better/SKILL.md b/skills/know-me-better/SKILL.md index f6aa02f..011f91a 100644 --- a/skills/know-me-better/SKILL.md +++ b/skills/know-me-better/SKILL.md @@ -5,14 +5,11 @@ description: User trigger. Use when you want the agent to learn your research st ## Installed resources -Keep the working directory at the user's project. Resolve this loaded `SKILL.md` -with `Path(path).resolve()` before locating resources; follow symlinks. Bare -`helpers/`, `references/`, and template paths are relative to that real skill -directory. A path written as `skills//...` means the installed `` -skill's directory from the agent's skill catalog, not a path in the user's project. -Locate each dependency by its public skill name; copied skills need not be siblings. -If a dependency is absent, report the missing skill and install it before that step. -Shared writing files are bundled in `how-to-write-ideas-report/references/`. +Keep the working directory at the user's project. Resolve this `SKILL.md` to its +real path before locating bundled resources. `skills//...` refers to the +installed skill found by public name, not the user's project; dependencies need +not be siblings. Load only resources needed for the current task. If a required +dependency is missing, report it before that dependent step. Before running the examples, set `DOWNLOAD_REF_DIR` to the absolute directory of `how-to-download-ref`. Quote these variables as shown. @@ -21,11 +18,13 @@ Before running the examples, set `DOWNLOAD_REF_DIR` to the absolute directory of Turn an existing paper collection into a structured knowledge base under `/.knowledge/` (or an advisor KB). The output uses the same KB format as the `survey` and `how-to-download-ref` skills — project and advisor KBs can coexist cleanly. -**Step 1 — Identify the researcher and source.** First, ask whose papers to index: +**Step 1 — Identify the researcher and source.** Reuse the researcher, collection, +source path/URL, and target KB supplied by the user or calling skill. Ask only +for missing information. When the researcher is unknown: > "Whose papers should I index? (Give me a name, or leave blank for your own collection.)" -Then ask which source to use: +When the source is unknown: > "Where are the papers?" > - **(a)** Zotero library @@ -126,14 +125,11 @@ Write or extend `$KB/NOTES.md` with: Reference papers as `[@]`. If `NOTES.md` exists, extend rather than overwrite. -## After know-me-better — transition checkpoint +## Completion and handoff -After Steps 3–6 complete, the KB is populated with metadata but PDFs aren't downloaded yet. Ask the user in chat: - -> "Index built. What next?" -> - **(a)** Fetch PDFs for all refs — invokes `how-to-download-ref --from-bib $KB/references.bib --kb $KB` (bulk mode) -> - **(b)** Add specific refs by ID — invokes `how-to-download-ref` with explicit IDs (single-shot, per-ref cite-key confirmation) -> - **(c)** Continue to `brainstorm-ideas` — start brainstorming with the indexed literature loaded -> - **(d)** Stop — leave the KB as-is - -For (a) and (b), see `skills/how-to-download-ref/SKILL.md`. For (c), invoke `brainstorm-ideas` in the current session. +Report the indexed collection, skipped papers, KB paths, and missing full text. +When invoked from `create-advisor` or `brainstorm-ideas`, return these results to +the caller and continue its authorized workflow. If full text was requested, +invoke `how-to-download-ref --from-bib $KB/references.bib --kb $KB` and carry +forward the scope and source preferences. A standalone indexing request ends +with the index; additional downloads or brainstorming are optional follow-ups. diff --git a/skills/review-paper/SKILL.md b/skills/review-paper/SKILL.md index ec46ac3..6e3374f 100644 --- a/skills/review-paper/SKILL.md +++ b/skills/review-paper/SKILL.md @@ -5,21 +5,18 @@ description: User trigger. Use when reviewing, commenting on, or fact-checking a ## Installed resources -Keep the working directory at the user's project. Resolve this loaded `SKILL.md` -with `Path(path).resolve()` before locating resources; follow symlinks. Bare -`helpers/`, `references/`, and template paths are relative to that real skill -directory. A path written as `skills//...` means the installed `` -skill's directory from the agent's skill catalog, not a path in the user's project. -Locate each dependency by its public skill name; copied skills need not be siblings. -If a dependency is absent, report the missing skill and install it before that step. -Shared writing files are bundled in `how-to-write-ideas-report/references/`. +Keep the working directory at the user's project. Resolve this `SKILL.md` to its +real path before locating bundled resources. `skills//...` refers to the +installed skill found by public name, not the user's project; dependencies need +not be siblings. Load only resources needed for the current task. If a required +dependency is missing, report it before that dependent step. Before running the examples, set `DOWNLOAD_REF_DIR` to the absolute directory of `how-to-download-ref`. Quote these variables as shown. # Paper Reviewer -Run a structured **review-and-enhance** pass over an *existing* scientific manuscript. The skill reads the whole paper, produces **location-anchored comments first**, then applies only the edits the user approves and re-checks that the manuscript still compiles. +Run a structured **review-and-enhance** pass over an *existing* scientific manuscript. The skill reads the context needed for the selected scope, delivers **location-anchored findings**, and applies requested edits with appropriate verification. **Scope note.** This is the *reviewing/revising* counterpart to `write-paper` (which *drafts* a manuscript figures-first). It is **not** `survey` report mode, which writes technology/field-assessment reports from a literature survey. Use `review-paper` when a manuscript already exists and the user wants comments, a referee-style critique, reference/fact verification, or guideline-driven polish. If no manuscript exists yet, redirect to `write-paper`. @@ -31,7 +28,12 @@ Use `skills/how-to-write-ideas-report/references/writing-workflow.md` for KB/con ## Operating principle -**Comment-first, non-destructive.** Never edit the manuscript before the user has seen and approved the findings. Read the whole paper, deliver the comments, let the user choose what to apply, *then* edit. This mirrors `write-paper`'s "iterate the story with the user" philosophy: the author keeps control of judgment calls. +**Preserve scope and author control.** A request for review or comments authorizes +findings, not manuscript edits. A request to revise, polish, or apply selected +findings already authorizes those changes; do not ask for the same approval +again. For wording-only work, preserve scientific meaning. Unresolved scientific +judgment calls remain comments, while a separately requested substantive revision +may be proposed from the supplied evidence. Never invent a result or justification. --- @@ -44,7 +46,7 @@ Use `skills/how-to-write-ideas-report/references/writing-workflow.md` for KB/con | 3 | Each paragraph has one well-defined job | 1 | `how-to-technical-writing` | | 4 | **DRY** — avoid repeated explanations/definitions | 1 | new to this skill | | 5 | Display math reserved for emphasis only | 1 | `write-paper` Notation Rulebook | -| 6 | Read the whole paper first; brief the story, judge it at the high level, then each section's mission | 0 | new — the gate before critique | +| 6 | Match context to scope; brief the story and section missions for a full review | 0 | scope and context selection | | 7 | Every figure referenced ≥1× and its striking features discussed | 1 | `write-paper` Figure Rulebook | | 8 | Verify factual claims & references; flag the uncertain | 2 | new — the standout capability | | 9 | Fit the target journal: limits, required parts, and its writing guidance | 2.5 | new — venue discussion reused from `write-paper` Phase 1.5 | @@ -63,25 +65,36 @@ Rank every finding so the user can triage: ## Phase 0 — Scope, load & understand (guideline #6) -1. **Resolve the manuscript.** Accept an explicit path; else auto-detect from `articles//` or the working directory. Detect format from the extension: **LaTeX (`.tex`) is primary**; Typst (`.typ`) and Markdown (`.md`) are supported. -2. **Ask what to check, before reading in depth.** Offer these numbered options; the user may pick any subset: - 1. **High-level story** — is the question worth asking, do the contributions match the results, and do the abstract, the main figure, and the supporting data carry the story (step 6 below). - 2. **Writing** — guidelines 1–7 (Phase 1), optionally narrowed to named guidelines or sections. - 3. **Facts, references, and links** — guideline 8 (Phase 2): every bibliography entry screened, cited claims sanity-checked, standalone factual claims verified, every URL and DOI link resolved. The slow pass. - 4. **Journal fit** — guideline 9 (Phase 2.5): decide or recommend the target journal, fetch its author guidelines, check limits and structural completeness, and review against the journal's own writing guidance. - - Default on a first review: all four. If a previous `articles//review-*.md` exists, say so and propose the repeat-pass default: option 2 only, restricted to text changed since that report, plus option 3 only if the bibliography changed, and option 4 only if the target journal changed or the previous report did not run it. A pass the user did not select is not run, and the report says it was skipped. -3. **Read the whole manuscript** end to end — and its bibliography. Resolve the bibliography in this order: the manuscript's own `\bibliography{…}` / `\addbibresource{…}` target (or embedded `thebibliography` / Typst `bibliography(…)`), then fall back to `$KB/references.bib`. Handle both; note which one you used. -4. **Load shared context.** Follow `skills/how-to-write-ideas-report/references/writing-workflow.md`: resolve `KB=$(python3 "$DOWNLOAD_REF_DIR/helpers/resolve_kb.py")`, read `$KB/INDEX.md`, `$KB/NOTES.md`, `$KB/references.bib`, and `docs/discussion/user-profile.md` if present. This is the literature backdrop for fact-checking. -5. **Write the story brief** (always, whatever was selected; it is the gate). Four short parts, in the paper's own terms: - - **Story** — one paragraph: what the paper does and what it finds. - - **Scientific question and its significance** — the question as the paper poses it, and why the field should care. State the gap the paper claims to fill. - - **Key contributions** — a numbered list, as the paper claims them. - - **Key results** — one line each, naming the figure, table, or equation that carries it. - - Then a one-line "mission" for each section. This gate prevents local nitpicks that fight the global narrative: if you misread the story, fix that before producing any finding. +1. **Resolve scope and mode.** Use the supplied manuscript, excerpt, requested + checks, and existing authorization. For an unspecified full review, cover + story, writing, facts/references, and journal fit when a target is known. + A local wording or caption request runs only the relevant writing checks. + Ask only when the target or substantive scope cannot be inferred. On a + repeat pass, inspect changed passages and affected dependencies; rerun + reference or journal checks only when relevant inputs or unresolved findings + changed. State any limits of the review. +2. **Read enough context.** For a full scientific review, read the whole manuscript + and its bibliography. For a local pass, read the target plus nearby definitions, + referenced equations/figures, and the context needed to preserve meaning. + Expand only if a concrete dependency requires it; do not require the entire + KB, user profile, or history for a local grammar check. +3. **Resolve references when needed.** Use the manuscript's own bibliography + (`\bibliography`, `\addbibresource`, embedded entries, or Typst bibliography) + before falling back to `$KB/references.bib`. For fact checks, follow the + relevant context/citation sections of + `skills/how-to-write-ideas-report/references/writing-workflow.md`. Load only + source entries relevant to the review, except in a requested full-bibliography + audit (Phase 2). +4. **Understand the story for a full review.** Write a brief covering the question, + significance, claimed contributions, key results tied to figures/equations, + and each section's mission. Include it with the completed report. Pause for + clarification only if competing interpretations would materially change the + critique. A local wording pass does not need a story-approval gate. +5. **Select passes.** The requested scope selects Phase 1 (writing), Phase 2 + (facts/references), and Phase 2.5 (journal fit). These correspond to options + 2, 3, and 4 below; option 1 means high-level scientific review. 6. **Comment on the high-level aspects** (option 1). Severity-ranked, same finding shape as Phase 1, guideline 6: - - **Significance of the problem** — Justify, from the manuscript and the loaded KB, why the problem is worth solving: who is blocked by it, what becomes possible once it is solved, and what the strongest prior attempt achieved. If the manuscript does not supply this, or the KB is too thin to judge whether the gap is real, say so and suggest a `survey` pass on the problem before the review continues; do not fill the gap from general knowledge. + - **Significance of the problem** — Justify, from the manuscript and the loaded KB, why the problem is worth solving: who is blocked by it, what becomes possible once it is solved, and what the strongest prior attempt achieved. If the manuscript does not supply this, or the KB is too thin to judge whether the gap is real, mark that assessment unresolved and suggest a focused literature check; continue other selected checks without inventing support from general knowledge. - **Significance of each contribution** — Take the numbered contributions one at a time. For each, ask: is it stated as a verifiable property (a bound, a measured improvement, a new capability with a demonstrated case) rather than an activity ("we study", "we explore")? Does a result in the paper verify it? Would a reader in the target audience pay for it: is it better than the strongest baseline on an axis that audience cares about, and is that axis named? A contribution that fails any of these is a finding; the fix names the property that would make it buyable. - **Story** — Does the claimed contribution match what the results actually show? Is the gap statement supported by the cited prior work, or asserted? - **Abstract** — Map its moves (system, method, finding, implication) onto the story brief. Flag any abstract claim no result backs, and any key result the abstract omits. @@ -89,7 +102,7 @@ Rank every finding so the user can triage: - **Supporting data** — For each key result, is the evidence strong enough for the claim as worded? Check the match between claim strength and evidence: a general claim needs more than one system, size, or seed; a "significant" improvement needs error bars or a statistical test that separates it from the baseline; a scaling claim needs enough decades to distinguish the fitted law from its neighbours; a "state of the art" claim needs the strongest current baseline, run under the same conditions. Flag claims with no data behind them, missing controls, baselines, or error bars, and claims worded stronger than the data. The fix either names the extra evidence needed or rewords the claim down to what the data shows. - **Highlighting the contributions** — Are the contributions where a skimming reader looks: named as such in the abstract, listed in the introduction's "in this paper" paragraph, and each tied to the figure or equation that proves it? Propose a better way to surface them when one exists: a main figure that puts the contribution and the strongest baseline on the same axes, a summary table of results versus prior work, a sharper title or abstract sentence that states the property rather than the activity, or reordering so the strongest result comes first. Keep to what the data already supports; do not invent a comparison. -Do not proceed to Phase 1 until the user confirms (or corrects) the story brief. The high-level comments are delivered with the brief; the user may reject any of them before they enter the report. +Continue through the selected passes and deliver one coherent report. Reuse a story the user already supplied or confirmed. --- @@ -105,7 +118,7 @@ Run only when option 2 was selected in Phase 0. Walk the manuscript and produce fix: a concrete suggested edit } ``` -Calibrate every `fix` against the model paper: Ho et al., PRL 122, 040603 (2019), at `skills/write-paper/sources/1807.01815_Ho2019_quantum-scars.md`, with a move-by-move walkthrough in `skills/write-paper/references.md` §C. A good rewrite reads like that letter: one concept per sentence, symbols defined at first use, each section opening with its job, main results named as such. When unsure whether a sentence deserves a finding, ask whether it could appear in the model paper unchanged. +When style calibration is needed, consult the relevant walkthrough in `skills/write-paper/references.md` §C; open only the corresponding excerpt of the model paper: Ho et al., PRL 122, 040603 (2019), at `skills/write-paper/sources/1807.01815_Ho2019_quantum-scars.md`, with a move-by-move walkthrough in `skills/write-paper/references.md` §C. A good rewrite reads like that letter: one concept per sentence, symbols defined at first use, each section opening with its job, main results named as such. When unsure whether a sentence deserves a finding, ask whether it could appear in the model paper unchanged. Checks: @@ -116,7 +129,7 @@ Checks: 5. **Display-math discipline.** Flag display equations that don't earn emphasis; suggest inlining or cutting. Reserve display math for key/flagship results, non-obvious steps, or figure-referenced equations. Never propose inlining or cutting an equation whose label is referenced elsewhere; for a letter, propose moving algebra to the supplement rather than deleting a reproducibility step. (`write-paper` Notation Rulebook.) 7. **Figure integration.** Flag **orphan** figures (never referenced in the main text) and figures whose striking features (peaks, kinks, jumps) aren't discussed. (`write-paper` Figure Rulebook.) -Also run a per-section **"did this section deliver its Phase-0 mission?"** check (guideline #6 carried into the body), and check that each key result from the story brief is named as a main result where it appears. +For full reviews, also run a per-section **"did this section deliver its Phase-0 mission?"** check (guideline #6 carried into the body), and check that each key result from the story brief is named as a main result where it appears. Every finding names the rule it serves. Not every paragraph needs a finding: a passage that already passes stays untouched, and a rewrite is never proposed for the sake of a diff. @@ -124,14 +137,15 @@ Every finding names the rule it serves. Not every paragraph needs a finding: a p ## Phase 2 — Fact & reference verification (guideline #8) -Run only when option 3 was selected in Phase 0. The standout capability. Follow the repo discipline: **never invent BibTeX from memory**. Bibliography metadata gets a complete automated screening pass; claim verification stays focused on what supports the main claims (see `skills/how-to-write-ideas-report/references/writing-workflow.md`). +Run only when option 3 was selected in Phase 0. The standout capability. Follow the repo discipline: **never invent BibTeX from memory**. Keep this pass within the selected scope: a full review screens all bibliography metadata, while a targeted fact or reference check verifies only the requested claims and their supporting entries/links; claim verification stays focused on what supports those claims (see `skills/how-to-write-ideas-report/references/writing-workflow.md`). -- **Screen every bibliography entry.** Run `python3 "$DOWNLOAD_REF_DIR/helpers/verify_bib.py" --bib "$BIB" --kb "$KB" --json` against the bibliography resolved in Phase 0. This checks uncited entries too and compares title, authors, year, venue/journal, volume, pages, and DOI using cached metadata plus Semantic Scholar's batch API. +- **Full review or bibliography audit: screen every bibliography entry.** Run `python3 "$DOWNLOAD_REF_DIR/helpers/verify_bib.py" --bib "$BIB" --kb "$KB" --json` against the bibliography resolved in Phase 0. This checks uncited entries too and compares title, authors, year, venue/journal, volume, pages, and DOI using cached metadata plus Semantic Scholar's batch API. +- **Targeted check.** Resolve the entries cited by the selected passage or named by the user and verify those records directly against authoritative metadata. Do not run `verify_bib.py` against the whole bibliography for this mode: its CLI has no key filter. Include additional entries only when they are dependencies of the requested claim. State the checked scope in the result. - **Confirm actionable records.** Use the helper's severity-ranked output as the starting point for the reference / fact-check table. Before reporting any `unverifiable` entry, or any `mismatch` with a high- or medium-severity finding, confirm it manually through **CrossRef → Semantic Scholar → MCP → web fetch**; Semantic Scholar screens, it is not the final authority. Keep low-severity missing-field findings as metadata-completion suggestions — they do not need the full lookup chain. Flag broken, missing, or confirmed-mismatched entries and offer repair via the `how-to-download-ref` skill (it owns `references.bib` appends and metadata fetching). -- **Citation resolution.** For each `\cite` key, confirm that an entry exists in the resolved bibliography. Uncited entries remain in the metadata scan; cited keys additionally participate in the claim-support check below. +- **Citation resolution.** For each `\cite` key in the selected scope, confirm that an entry exists in the resolved bibliography. In a full review, uncited entries remain in the metadata scan; cited keys additionally participate in the claim-support check below. - **Claim ↔ citation support.** For key claims attached to a citation, best-effort sanity-check that the cited work actually supports the claim. **Flag uncertain — do not assert.** -- **Standalone factual claims.** Identify checkable factual/numerical claims *not* tied to a citation; verify via web search. **Flag** the uncertain or unsupported — never silently "correct" a claim, and never fabricate a citation to prop one up. -- **Link check.** Every URL in the manuscript and bibliography (`\url`, `\href`, DOI links, code and data repositories) is fetched once; flag dead links, redirects to a different resource, and DOIs that do not resolve. A private or paywalled target that answers is ok. +- **Standalone factual claims.** Identify checkable factual/numerical claims in the selected scope *not* tied to a citation; verify via web search. **Flag** the uncertain or unsupported — never silently "correct" a claim, and never fabricate a citation to prop one up. +- **Link check.** Fetch relevant URLs and DOI links for a targeted check. In a full review, fetch every URL in the manuscript and bibliography (`\url`, `\href`, DOI links, code and data repositories) once; flag dead links, redirects to a different resource, and DOIs that do not resolve. A private or paywalled target that answers is ok. --- @@ -140,7 +154,7 @@ Run only when option 3 was selected in Phase 0. The standout capability. Follow Run only when option 4 was selected in Phase 0. Never quote a limit or rule from memory: every constraint in this pass comes from a page fetched in this session or from a `template/README.md` that records its source. 1. **Decide the target journal.** Use the venue the manuscript already declares: a document class or template (`revtex4-2`, `iopart`, `elsarticle`, an `sn-jnl` class), journal macros, a cover letter, or a `template/README.md` written by `write-paper` Phase 1.5. If none, follow `write-paper` Phase 1.5: propose 2–3 venues with tradeoffs (article type, audience, length pressure, figure limits, novelty bar) grounded in the story brief, and ask the user to pick one or say "no target yet". With no target, skip the rest of this pass and record that choice in the report. -2. **Fetch the author guidelines.** Prefer the official publisher page over mirrors, Overleaf copies, or lab handouts. Reuse `template/README.md` when it already records the URL and access date; otherwise record both in the report. If the page cannot be fetched, mark every constraint `unverifiable`, give the URL, and stop. +2. **Fetch the author guidelines.** Prefer the official publisher page over mirrors, Overleaf copies, or lab handouts. Reuse `template/README.md` when it already records the URL and access date; otherwise record both in the report. If the page cannot be fetched, mark every constraint `unverifiable`, give the URL, and stop that journal-constraint check while completing the other selected passes. 3. **Extract the checkable constraints into a table**: article type; word, page, or character limits for the body, abstract, and title; maximum figures, tables, and references; required sections and their order; required statements (data availability, code availability, author contributions, competing interests, funding, ethics or IRB, keywords, significance statement); formatting rules (citations in the abstract, footnotes, units, figure resolution and file types, reference style); and any writing guidance the journal itself gives (a first paragraph accessible to non-specialists, a summary-paragraph structure, a stated audience). 4. **Check each constraint against the manuscript.** Measure rather than estimate: word counts from the compiled text (`detex`, `pandoc --to plain`, or `typst query`), figure, table, and reference counts from the source, section presence from the headings. Status per row: ok / over / missing / unverifiable. Severity: high for a hard limit exceeded or a required statement missing, med for a formatting rule, low for cosmetic. 5. **Review against the journal's writing guidance.** Judge the abstract, the opening paragraph, and the framing of significance against what the journal says it wants for its audience. These are guideline-9 findings with the same shape as Phase 1; the fix cites the guideline sentence it serves. @@ -149,7 +163,10 @@ Run only when option 4 was selected in Phase 0. Never quote a limit or rule from ## Phase 3 — Deliver the review report -Write a timestamped report to `articles//review-YYYY-MM-DD.md` (the repo's dated-output convention). Structure: +For a full review, write `articles//review-YYYY-MM-DD.md` (or the requested +path). For a local excerpt or small wording pass, deliver concise findings or +the revised text inline unless a file was requested. Include only applicable +parts of this structure: 1. **Story brief, high-level comments, and per-section missions** (from Phase 0). 2. **Findings grouped by guideline, severity-ranked** (high → low). @@ -157,26 +174,29 @@ Write a timestamped report to `articles//review-YYYY-MM-DD.md` (the repo's 4. **Journal fit table** — target journal and guideline source (URL, access date), then constraint → required → measured → status. If Phase 2.5 was skipped or no target was chosen, one line saying so. 5. **Top fixes** — a prioritized list of the highest-leverage changes. -Then present a short summary to the user and ask two things: which findings to apply (all / by severity / individually), and whether to **show a marked diff first** or **apply directly**. Default to the marked diff; it costs one compile and lets the author judge each rewrite in context. +For a review-only request, deliver the report; offer application as an optional +next step. If changes were already requested, continue to Phase 4 using that +scope. Reuse the chosen direct/diff mode. When the user has not authorized edits +to the original, prepare a marked proposal for review before merging it. --- ## Phase 4 — Apply approved edits -**Marked-diff mode** (the default). Never touch the original until the user has seen every proposed change in context. +**Marked-diff mode** (when requested, or proposing edits beyond existing authorization). Prepare the proposed changes before asking the user which to merge. -1. **Edit a copy.** `cp main.tex main.proposed.tex` (same for `.typ` / `.md`). Apply every approved fix to the copy with a script that asserts each target passage matches exactly once. Comment-only fixes (style-guide guardrails) go in as `[reviewer]` comments in the copy too. +1. **Edit a copy.** `cp main.tex main.proposed.tex` (same for `.typ` / `.md`). Apply the proposed fixes to the copy; when using scripted replacements, assert that each target matches exactly once. Comment-only fixes (style-guide guardrails) go in as `[reviewer]` comments in the copy too. 2. **Mark the diff.** LaTeX: `latexdiff main.tex main.proposed.tex > main.diff.tex`, then compile `main.diff.tex` (deletions red struck-through, additions blue underlined). Typst and Markdown have no latexdiff; write `git diff --no-index --word-diff main.typ main.proposed.typ` into `articles//review-YYYY-MM-DD.diff` and, for Typst, also compile the proposed copy so the author can read the result. Check the page count did not change unexpectedly. 3. **Hand over a numbered legend.** Give the marked PDF (or diff file) path and one line per change: number, section and page, the finding it fixes, and the rule it serves. Pair up a removal and an addition that belong to one change. End with "reply with the numbers to merge, or all". 4. **Merge exactly the accepted numbers** into the original. If the user accepts "all except one wording", revert that wording in the copy first, then merge all. Delete `main.proposed.*` and `main.diff.*` afterwards. -**Direct mode** (only when the user chose it). Apply the approved findings straight to the manuscript. Do not edit anything the user didn't approve. +**Direct mode** (when the user requested edits without a preview gate). Apply changes within that authorization to the manuscript and show the diff/result. Do not broaden the requested revision. **Both modes.** 1. **Preserve structure.** Keep macros, environments, labels, and document structure in every format. Where a fix genuinely needs author judgment, insert a `% [reviewer] …` margin comment instead of rewriting silently (`// [reviewer]` in Typst, an HTML comment in Markdown). 2. **Language edits change how a sentence is written, never what it says.** Never alter the content of a definition, theorem, or claim while rewording it; never add or remove a claim, figure, or derivation step under a guideline 1–5 fix; never replace a field term with a simpler word. Recheck every number, count, and qualifier a rewritten sentence mentions before moving on. -3. **Comment, do not apply, when a language fix touches content.** An edit that would change a quantifier or hedge ("arbitrary", "approximately", "at most"), add a justification the manuscript does not already contain, delete a paragraph as a digression, or remove a "not X but Y" contrast goes in as a `[reviewer]` comment even when the user approved the finding. The author writes those words. +3. **Comment, do not apply, when a language fix touches content.** An edit that would change a quantifier or hedge ("arbitrary", "approximately", "at most"), add a justification the manuscript does not already contain, delete a paragraph as a digression, or remove a "not X but Y" contrast goes in as a `[reviewer]` comment even when the user approved the finding. Do not treat approval of a language pass as approval to change scientific content; a substantive rewrite needs its own explicit scope and supporting evidence. 4. **Verify it still compiles** — `latexmk` (or `pdflatex`) for `.tex`, `typst compile` for `.typ`; for `.md`, confirm it still renders. Report pass/fail with the actual command output (per verification-before-completion: evidence before assertions). If it breaks, fix or revert the offending edit before claiming done. 5. **Append a changelog** to the top of the review report: what was applied, what was skipped, and the compile result. @@ -186,7 +206,7 @@ Then present a short summary to the user and ask two things: which findings to a **Reused (no duplication):** `skills/how-to-write-ideas-report/references/writing-workflow.md` (context, citations, output mechanics); `skills/how-to-technical-writing/SKILL.md` (sentence/paragraph rules, hunt table, application guardrails); the BibTeX lookup chain (CrossRef → Semantic Scholar → MCP → web fetch); `how-to-download-ref` for reference repair; `write-paper`'s notation/figure rule *definitions* (referenced). -**New here:** the read-whole-first review protocol (Phase 0), DRY/anti-repetition detection (#4), fact & reference verification (#8), journal fit (#9, venue discussion reused from `write-paper` Phase 1.5), and the comment-then-apply loop plus compile-check (Phases 3–4). +**New here:** scope selection and whole-paper context for full reviews (Phase 0), DRY/anti-repetition detection (#4), fact & reference verification (#8), journal fit (#9, venue discussion reused from `write-paper` Phase 1.5), and the comment-then-apply loop plus compile-check (Phases 3–4). --- @@ -194,13 +214,13 @@ Then present a short summary to the user and ask two things: which findings to a | Mistake | Instead | |---|---| -| Editing before the user approves | Comment-first. Phase 3 → user picks → Phase 4. | -| Merging a rewrite the author has not seen in context | Marked-diff mode by default: `latexdiff` on a copy, numbered legend, merge only the accepted numbers. | -| Nitpicking sentences before confirming the story | Phase 0 gate: confirm the story brief first. | -| Re-running reference verification on every pass | Ask what to check first, with the three numbered options; on a repeat pass default to writing only unless the bibliography changed. | +| Treating a review request as permission to edit | Deliver findings; use existing revision authorization only for its stated scope. | +| Ignoring a requested preview gate | Prepare a marked copy and merge accepted changes; reuse direct-edit authorization when no preview was requested. | +| Running a full story gate for a local wording fix | Read the target and its necessary context; reserve the story brief for full reviews. | +| Re-running reference verification on every pass | Inspect changed inputs and unresolved findings; rerun only affected checks. | | Inventing a BibTeX entry to "fix" a citation | Never. Use the lookup chain / the `how-to-download-ref` skill, or flag as unverifiable. | | Silently "correcting" a factual claim | Flag uncertain claims; the author decides. | -| Claiming done without compiling | Run `latexmk` / `typst compile` and paste the result. | +| Claiming a build passed without running it | Compile changed document source; distinguish inline-text checks from a document build. | | Chasing completeness on fact-checks | Verify only what supports the main claims. | | Quoting a journal's word limit or required sections from memory | Fetch the official author guidelines this session, or mark the row unverifiable with the URL. | | Rewriting a passage that already passes | Leave it. A finding must name the rule it serves. | diff --git a/skills/review-paper/checklist.md b/skills/review-paper/checklist.md index 04bbba9..aeeb6e1 100644 --- a/skills/review-paper/checklist.md +++ b/skills/review-paper/checklist.md @@ -1,15 +1,15 @@ # Paper-reviewer process checklist -The review-only items behind `review-paper/SKILL.md`: the gate (6), fact and reference verification (8), journal fit (9), and delivery. The writing items (1–5, 7) live in `skills/how-to-technical-writing/checklist.md`, the single source of truth for the writing guide. +The review-only items behind `review-paper/SKILL.md`: scope and context (6), fact and reference verification (8), journal fit (9), and delivery. The writing items (1–5, 7) live in `skills/how-to-technical-writing/checklist.md`, the single source of truth for the writing guide. --- -## 6 — Read the whole paper first (the gate) +## 6 — Scope and scientific context -- [ ] The user was asked first which passes to run, with numbered options: 1 high-level story, 2 writing, 3 facts, references, and links, 4 journal fit. On a repeat review, the previous report was named and the default was option 2 only. -- [ ] The whole manuscript and its bibliography were read before any critique. -- [ ] A story brief was produced. It holds the story paragraph, the scientific question and its significance, the numbered key contributions, and the key results. Each key result is tied to a figure, table, or equation. -- [ ] A one-line mission per section was produced. +- [ ] Scope and revision authorization were taken from the request; clarification was limited to missing information that changes the review. +- [ ] A full scientific review read the whole manuscript and bibliography; a local pass read the target and necessary dependencies. +- [ ] A full review includes the story, question, contributions, results tied to evidence, and section missions. Local wording checks do not require a story brief. +- [ ] The existing story was reused; only a material unresolved interpretation prompted a question. - [ ] When option 1 was selected, high-level comments were made on each of the items below. - The significance of the problem: who is blocked, what it unlocks, and the strongest prior attempt. `survey` is suggested when the manuscript or KB cannot justify it. - The significance of each contribution. It is stated as a verifiable property, verified by a result, and better than the strongest baseline on an axis the audience names. @@ -18,19 +18,19 @@ The review-only items behind `review-paper/SKILL.md`: the gate (6), fact and ref - The main figure. It carries the central claim alone, or no such figure exists. - The supporting data. Claims without data, missing controls, baselines, and error bars are named. The evidence is strong enough for the claim as worded. General claims span systems or seeds. Improvements are separated from the baseline by error bars or a test. Scaling laws span enough decades. State-of-the-art claims are measured against the strongest baseline under the same conditions. - How the contributions are highlighted. They are named in the abstract, listed in the introduction, and each tied to its proving figure or equation. A better main figure, summary table, title sentence, or ordering is proposed when one exists. -- [ ] **The story brief and high-level comments were confirmed with the user before findings were generated.** +- [ ] The selected passes were completed before delivering the report, unless missing scientific input prevented an assessment. ## 8 — Fact & reference verification (new) - [ ] Skipped entirely, with a note in the report, when the user did not select option 3 in Phase 0. Otherwise: -- [ ] `verify_bib.py` was run against the resolved bibliography. **Every entry**, including uncited entries, appears in its report. -- [ ] Title, authors, year, venue or journal, volume, pages, and DOI were screened against cached and batched Semantic Scholar metadata. +- [ ] For a full review or bibliography audit, `verify_bib.py` screened every entry, including uncited entries. For a targeted check, only requested claims and their supporting entries were verified directly, and the limited scope was reported. +- [ ] Metadata fields in the checked scope were verified; full screening compares title, authors, year, venue or journal, volume, pages, and DOI against cached and batched Semantic Scholar metadata. - [ ] Every `unverifiable` record was confirmed by hand before reporting. So was every `mismatch` with a high or medium finding. The chain is CrossRef → Semantic Scholar → MCP → web fetch. Low-severity missing fields remain completion suggestions. -- [ ] Every `\cite` key resolves to an entry in the bibliography that was actually used. +- [ ] Every `\cite` key in the checked scope resolves to an entry in the bibliography that was actually used. - [ ] Broken, missing, or mismatched citations are flagged. Repair is offered via the `how-to-download-ref` skill. - [ ] Key claims attached to a citation are sanity-checked against the cited work. Uncertain ones are flagged, not asserted. -- [ ] Standalone checkable factual or numerical claims are verified via web search. Uncertain ones are flagged. -- [ ] Every URL and DOI link in the manuscript and bibliography is fetched once. Dead links, wrong redirects, and unresolvable DOIs are flagged. +- [ ] Standalone checkable factual or numerical claims in scope are verified via web search. Uncertain ones are flagged. +- [ ] Relevant URLs and DOI links were checked for a targeted request; every manuscript/bibliography link was checked for a full review. Dead links, wrong redirects, and unresolvable DOIs are flagged. - [ ] **No BibTeX invented from memory. No claim silently "corrected". No citation fabricated.** ## 9 — Journal fit (new) @@ -47,19 +47,19 @@ The review-only items behind `review-paper/SKILL.md`: the gate (6), fact and ref ## Language-pass hunt table (guidelines 1, 3, 4, 5) -The checkable items live in `skills/how-to-technical-writing/checklist.md`. The hunt-for / fix table lives in `skills/how-to-technical-writing/SKILL.md`, shared with `write-paper`. Each finding cites its row. Rows marked **comment only** are never applied as edits, even after approval. +The checkable items live in `skills/how-to-technical-writing/checklist.md`. The hunt-for / fix table lives in `skills/how-to-technical-writing/SKILL.md`, shared with `write-paper`. Each finding cites its row. Rows marked **comment only** stay comments within a language pass; a separately requested substantive revision needs supporting evidence. ## Delivery & application -- [ ] Findings are written to `articles//review-YYYY-MM-DD.md`, grouped by guideline and ranked by severity. -- [ ] A reference and fact-check table (cite key → status → note) is included. +- [ ] Full-review findings are saved to the requested path or `articles//review-YYYY-MM-DD.md`; a small excerpt review can be delivered inline. Findings are located and ranked by severity. +- [ ] A reference and fact-check table is included when that pass applies. - [ ] A journal fit table (constraint → required → measured → status) with the guideline source is included. Otherwise a line says the pass was skipped. - [ ] A prioritized "top fixes" list is included. -- [ ] The user was asked whether to see a marked diff first or apply directly. The marked diff is the default. +- [ ] The chosen direct/diff mode was reused. Already requested edits continued without asking for the same authorization; a proposal beyond that scope was prepared before requesting a decision. - [ ] In marked-diff mode, edits went to a `*.proposed.*` copy. `latexdiff` produced the marked version, or `git diff --word-diff` for Typst and Markdown. A numbered legend was handed over. Only the accepted numbers were merged. The proposed and diff files were deleted afterwards. -- [ ] Edits were applied only after user approval, either all, by severity, or one by one. +- [ ] Applied edits stayed within the existing user authorization; a review-only request did not modify the original. - [ ] LaTeX, Typst, and Markdown structure and macros are preserved. Author-judgment fixes are left as `% [reviewer]` comments. - [ ] Language edits changed how sentences are written, never what they say. No definition, theorem, claim, or field term was altered. Numbers and qualifiers were rechecked after every rewrite. - [ ] Fixes touching a quantifier, hedge, missing justification, paragraph deletion, or "not X but Y" contrast were left as `[reviewer]` comments, not applied. -- [ ] The manuscript was re-compiled with `latexmk`, `pdflatex`, or `typst compile`, and the result was reported. +- [ ] Changed document source was compiled and the result reported. Inline-text checks did not claim a document build. - [ ] A changelog was appended to the top of the review report. diff --git a/skills/survey/SKILL.md b/skills/survey/SKILL.md index 605d7d5..2d050c1 100644 --- a/skills/survey/SKILL.md +++ b/skills/survey/SKILL.md @@ -5,14 +5,11 @@ description: User trigger. Use when surveying a research topic into a knowledge ## Installed resources -Keep the working directory at the user's project. Resolve this loaded `SKILL.md` -with `Path(path).resolve()` before locating resources; follow symlinks. Bare -`helpers/`, `references/`, and template paths are relative to that real skill -directory. A path written as `skills//...` means the installed `` -skill's directory from the agent's skill catalog, not a path in the user's project. -Locate each dependency by its public skill name; copied skills need not be siblings. -If a dependency is absent, report the missing skill and install it before that step. -Shared writing files are bundled in `how-to-write-ideas-report/references/`. +Keep the working directory at the user's project. Resolve this `SKILL.md` to its +real path before locating bundled resources. `skills//...` refers to the +installed skill found by public name, not the user's project; dependencies need +not be siblings. Load only resources needed for the current task. If a required +dependency is missing, report it before that dependent step. ## Choose the mode @@ -23,13 +20,18 @@ Shared writing files are bundled in `how-to-write-ideas-report/references/`. ## Topic Survey -Before starting this mode, check which MCP servers are available (arxiv, paper-search, Semantic Scholar, Sci-Hub, etc.). Present the detected servers to the user and let them choose which ones to use for this session (multi-select in chat). If none are configured, warn the user that the survey will rely on web search only. +Use the user's topic, constraints, output format, and requested endpoint. Choose +available search tools yourself unless the user has a provider preference. A +web-only workflow is valid when no scholarly MCP is configured. -If the user already provided a research topic or question, skip the clarification step. +**Step 1 — Clarify only missing scope.** If the topic and intended output are +clear, proceed. Ask about a substantive uncertainty that would change the search. -**Step 1 — Clarify.** Ask one question to narrow the research topic. Give 2-4 choice options. - -**Step 2 — Pick strategies & search.** Present the strategy menu to the user as a multi-select question. Recommend 3-4 strategies based on the topic context, but let the user choose. Then run one search worker per selected strategy in parallel when available, or sequentially otherwise. Each worker uses **broad web search only** at this stage — fast and exploratory. +**Step 2 — Choose strategies & search.** Select the strategies that address the +question, using the table below as guidance. Offer a choice when the user wants +to steer the search or when competing scopes would produce different reports. +Use independent workers when worthwhile and available, otherwise search +sequentially. Prefer authoritative papers and collect source links. **Strategy menu:** @@ -43,27 +45,34 @@ If the user already provided a research topic or question, skip the clarificatio | 6 | **Negative results** | Search for papers showing what does not work | | 7 | **Benchmarks and datasets** | What evaluation infrastructure exists | -When presenting to the user, briefly explain why you recommend each strategy for their specific topic (e.g., "Cross-vocabulary recommended because your problem — buffering stochastic supply — appears in operations research and hydrology too"). +Briefly explain the selected search scope when it helps the user assess coverage. Each worker produces a short **findings report** — key papers found, grouped by sub-theme, with titles and one-line descriptions. No BibTeX yet. **Important:** workers must also collect the DOI and arXiv ID for each paper when visible in search results (e.g., DOIs from publisher URLs, arXiv IDs from arxiv.org links like `2401.12345`). Record these alongside titles in the findings report. -**Step 3 — Consolidate & user picks directions.** Main agent consolidates all findings reports. **Deduplicate** papers that appear in multiple strategy reports — match by title similarity or DOI. Merge their descriptions (keep the richer one), **preserve any DOIs and arXiv IDs collected** during Step 2, and note which strategies found each paper. Then present the consolidated findings as numbered options grouped by theme. Ask: "Which directions should I add to the knowledge base? Pick one or more." The user can select multiple. +**Step 3 — Consolidate and select.** Merge overlapping results by DOI/arXiv +identity or title, preserving source identifiers and complementary evidence. +For an already scoped KB or report request, choose the relevant papers and +continue. For open-ended exploration with materially different research +directions, present the findings by theme and ask which direction to develop. +Do not require the user to select every paper in an already authorized scope. -**Step 4 — Build KB entries.** For the selected directions only, invoke `how-to-build-kb` (read `skills/how-to-build-kb/SKILL.md`) with the picked papers — titles plus every DOI / arXiv ID collected in Steps 2–3 — and the target KB (the project KB `/.knowledge/` by default; `--advisor ` when invoked from `create-advisor`). It verifies every entry against an authoritative source, appends to `$KB/references.bib`, regenerates `INDEX.md`, and writes or extends `NOTES.md`. Never generate BibTeX from memory. +**Step 4 — Build KB entries.** For the selected scope, invoke `how-to-build-kb` (read `skills/how-to-build-kb/SKILL.md`) with the picked papers — titles plus every DOI / arXiv ID collected in Steps 2–3 — and the target KB (the project KB `/.knowledge/` by default; `--advisor ` when invoked from `create-advisor`). It verifies every entry against an authoritative source, appends to `$KB/references.bib`, regenerates `INDEX.md`, and writes or extends `NOTES.md`. Never generate BibTeX from memory. -If the survey reveals the idea is already published, present the prior art and ask the user if they see a different angle before proceeding. +If prior work already covers a proposed idea, explain the overlap and possible distinctions. A neutral survey still reports that prior art; only pause if the user must choose a new research goal. ## After Survey — fetch full text, then optionally write The discovery stage is done once the KB has its references, `NOTES.md`, and `INDEX.md`. Use `how-to-download-ref` to fetch PDFs and render full-text markdown before writing a source-grounded report. -First, scan the conversation for arXiv IDs / DOIs the user mentioned that the parallel search didn't surface. If any are missing from `references.bib`, pull them in (invoke `how-to-download-ref` single-shot, cite-key confirmation per ref) so the reference set is complete before downloading. - -Then offer the next step: - -> "Survey complete. Fetch the PDFs?" -> - **(a)** Fetch + render all refs — invokes `how-to-download-ref --from-bib $KB/references.bib --kb $KB`; after it finishes, offer to continue to Survey Report below. -> - **(b)** Done — stop here. +Include any user-supplied papers relevant to the chosen scope that discovery +missed. Pass the selected IDs, existing cite keys, KB, and source preferences to +`how-to-download-ref`. When a full-text KB or report is requested, fetch the +selected set. A full-text KB request ends with the verified KB and acquisition +status; continue directly to Survey Report only when a report was requested. +Use `--from-bib` for a KB whose +whole bibliography is in scope; in a larger shared KB, pass only the scoped IDs. +For a discovery-only request, deliver the findings and KB artifacts; further +acquisition and reporting are optional. ## Survey Report @@ -73,7 +82,7 @@ Write a self-contained survey or technology/field assessment suitable for intern Follow `skills/how-to-technical-writing/SKILL.md` for sentence- and paragraph-level prose rules, and `skills/how-to-write-ideas-report/references/writing-workflow.md` for context loading, **source scoping**, citation handling, gap-filling research, output format, diagrams, and finish checks. -- If no KB exists, offer to run Topic Survey first; the report needs a grounded reference base. +- If no KB exists, build the requested evidence base from the supplied sources or Topic Survey as part of the report task. Ask only if its substantive scope is missing. - **Scope the source set first.** Run `scope_refs.py` as specified in the shared workflow. Fix dangling anchors before drafting. Build the report's approaches and claims from those scoped keys, not the entire bibliography. - Check `CLAUDE.md`/`AGENTS.md` for a deliverables-location convention before choosing an output path. - Tailor technical depth to the user's role from `docs/discussion/user-profile.md` or available context. @@ -128,6 +137,6 @@ End with a ranked table of 4–8 problems: number, problem, why it matters, who ### Optional direction fit -After showing the report, ask exactly once: "Do you want me to analyse which direction is most suited for you?" +If personalized direction advice was requested, include it after the report. Otherwise it is an optional follow-up, not a required question. -If yes, load the user's profile (or collect brief background, strengths, assets, and goals if none exists) and recommend 2–4 ranked directions. For each, name the report section/problem, explain the specific fit, give the smallest first experiment, and say what to avoid. Keep this personalized analysis outside the neutral report. +For that advice, load the user's profile (or collect brief background, strengths, assets, and goals if none exists) and recommend 2–4 ranked directions. For each, name the report section/problem, explain the specific fit, give the smallest first experiment, and say what to avoid. Keep this personalized analysis outside the neutral report. diff --git a/skills/write-paper/SKILL.md b/skills/write-paper/SKILL.md index 7021ecd..9934054 100644 --- a/skills/write-paper/SKILL.md +++ b/skills/write-paper/SKILL.md @@ -5,77 +5,97 @@ description: User trigger. Use when drafting or revising a scientific manuscript ## Installed resources -Keep the working directory at the user's project. Resolve this loaded `SKILL.md` -with `Path(path).resolve()` before locating resources; follow symlinks. Bare -`helpers/`, `references/`, and template paths are relative to that real skill -directory. A path written as `skills//...` means the installed `` -skill's directory from the agent's skill catalog, not a path in the user's project. -Locate each dependency by its public skill name; copied skills need not be siblings. -If a dependency is absent, report the missing skill and install it before that step. -Shared writing files are bundled in `how-to-write-ideas-report/references/`. +Keep the working directory at the user's project. Resolve this `SKILL.md` to its +real path before locating bundled resources. `skills//...` refers to the +installed skill found by public name, not the user's project; dependencies need +not be siblings. Load only resources needed for the current task. If a required +dependency is missing, report it before that dependent step. # Paper Writer -A working-rules guide for writing scientific papers, distilled from John Martinis's *Notes on Writing a Scientific Paper* and Jan von Delft's *Style Guide*. The source documents live in `references.md` and `sources/`; consult them when this SKILL.md leaves a question open. Beside them sits a model letter that executes these rules — Ho et al., PRL 122, 040603 (2019), see Source material. Imitate the model letter on any open style judgment call. +A working-rules guide for writing scientific papers, distilled from John Martinis's *Notes on Writing a Scientific Paper* and Jan von Delft's *Style Guide*. The source documents live in `references.md` and `sources/`; consult them when this SKILL.md leaves a question open. Beside them sits a model letter that executes these rules — Ho et al., PRL 122, 040603 (2019), see Source material. Use its relevant examples when a style judgment is unresolved, adapting to the manuscript and venue. **Scope note.** This skill is for *real manuscripts* — papers reporting completed (or near-complete) experimental, theoretical, or computational results. It is **not** for the upstream ideas/plan report produced by `brainstorm-ideas` report mode. If the user has not yet finished the work, push back: a paper requires results. -Use `skills/how-to-write-ideas-report/references/writing-workflow.md` for KB loading, citation handling, missing references, output formats, and Typst/diagram mechanics. The manuscript-specific rules below override shared defaults when venue templates or figure-first sequencing require it. +Use `skills/how-to-write-ideas-report/references/writing-workflow.md` for KB loading, citation handling, missing references, output formats, and Typst/diagram mechanics. Venue requirements take precedence over formatting defaults; the shared guide remains the source of sentence-level rules. --- -## The Iron Rules +## Choose the scope -These seven override everything else. +- **New manuscript:** use the phases below as a useful default. Establish the + evidence and narrative before polishing prose; the supplied result, figure + plan, or proof can provide that evidence. +- **Continue an existing draft:** reuse its story, venue, outline, and figures. + Enter at the unfinished part and verify affected dependencies. +- **Local revision:** change only the requested section, abstract, caption, or + passage, using its necessary context. Do not restart figure production, + venue selection, or story approval. For a critique rather than revision, use + `review-paper`. +- **Submission preparation:** apply the target venue's current requirements + and the relevant final checks. -1. **Figures first, then prose.** Plot every figure you intend to publish *before* writing a single section. Figures are the backbone; the text serves them. If you cannot draw the story in 4–8 figures, you do not yet know the story. -2. **Polish what readers actually read.** 95% of readers read only the title, abstract, introduction, figures + captions, and conclusions. These five elements deserve disproportionate iteration. Write them last — *but iterate them most*. -3. **State main results explicitly.** Use at least one sentence per main result that names it as such: *"This figure / equation / observation is our (first / second / Nth) main result."* Do not assume the reader will identify which sentence is the punchline. -4. **One concept per sentence.** If you must break this rule, both concepts must be simple. Long compound sentences with three new ideas are how readers drop out. -5. **Never plot anything in arbitrary units.** If you reach for "a.u." on an axis, the axis is wrong; find the right normalization (dimensionless ratio, calibrated scale, or experimental control). "a.u." plots are a known integrity red flag. -6. **Target journal before story.** Before proposing story lines, discuss the target venue with the user and download its official author template. If the user deliberately has no target yet, record that choice. Venue constraints shape the narrative, figure count, length, and format. -7. **Iterate the story with the user.** Once the figures or figure plan are available and the target venue is known, propose 1–3 plausible story lines for the user to choose, reject, or combine. Do not draft prose until the user has selected the paper's narrative. +Carry forward choices and authorization from the conversation. Clarify only a +missing scientific decision that changes the draft. Never manufacture results, +references, or an unsupported claim. A writing request ends with the requested +text/source and appropriate checks; external feedback is optional. ---- +## Writing priorities + +Use figures, evidence, or a proof outline to organize the story. State each main +result clearly and connect it to its evidence. Give the title, abstract, +introduction, captions, and conclusions careful attention because readers use +them to judge the paper. Follow the shared style guide rather than imposing a +fixed sentence length or a template from a different scientific genre. -## Workflow Phases +## Workflow for a new manuscript -A correct ordering of effort that defeats writer's block. Do not reorder — the sequence is the point. +Use the phases that the work still needs. For a theoretical result, a proof and +its key statements can serve the role of figures; a fixed figure count is not an +entry requirement. ### Phase 0 — Load context Before drafting, gather the materials that should inform the paper. Cheap to do once; expensive to skip. 1. **Shared writing context.** Follow `skills/how-to-write-ideas-report/references/writing-workflow.md`. Use `$KB/NOTES.md` as the spine for prior work, gap statement, motivation, and conclusions. -2. **Ideas / brainstorming log.** Look for `docs/discussion/*-brainstorm-ideas-log.md` from a prior `brainstorm-ideas` session. If present, read it for the original motivation, the cross-field connections it surfaced, the planned minimum viable experiment, and the success/hope/pivot signals. This is the *why* behind the paper and feeds the introduction's contribution claim. -3. **Personal publication context.** Read `docs/discussion/user-profile.md` and any `know-me-better` notes in `$KB/NOTES.md`. Use them to position the new paper in the user's own arc and to cite the user's earlier work correctly. +2. **Ideas / brainstorming log.** Look for `docs/discussion/*-brainstorm-ideas-log.md` from a prior `brainstorm-ideas` session. Read only logs relevant to this manuscript, starting with their summaries, for motivation, planned experiments, and success/hope/pivot signals. This is the *why* behind the paper and feeds the introduction's contribution claim. +3. **Personal publication context.** When positioning or self-citation requires it, read the relevant parts of `docs/discussion/user-profile.md` and `know-me-better` notes in `$KB/NOTES.md`. Use them to position the new paper in the user's own arc and to cite the user's earlier work correctly. 4. **Existing draft.** If a partial manuscript already exists under `articles/`, read it before proposing new prose; pick up where the user left off. -If none of these exist, name what's missing and ask the user to either point at the files or run the `survey` skill first. A paper without a literature foundation will read like one. +Use supplied source material even without a sci-brain KB. Research citation gaps needed by the draft; ask for missing results or assumptions that cannot be established from the available evidence. -### Phase 1 — Set up the figures (before writing any prose) +### Phase 1 — Establish the evidence and figure plan -- List the figures the paper needs. Aim for 4–6 in a letter, 6–10 in a regular article. +- List the figures or proof results needed to support the main claim; use the venue and argument to determine their number. - Order them so they *tell a story*: simple data first → progressively complex analysis → flagship comparison with theory. -- Draft each figure. Apply the **Figure Rulebook** (below). Do not move on until line weights, axes, colors, and dimensionless choices are right — going back later is more expensive than getting it right now. +- Draft needed figures using the **Figure Rulebook** below. Resolve issues affecting interpretation before writing claims around them; visual polish can proceed alongside the draft. - For each figure, write a one-sentence caption-summary: what the figure *shows* in plain words. These become the spine of the captions and the Results section. ### Phase 1.5 — Target journal and template checkpoint -Before the story checkpoint, stop and discuss the target journal or venue with the user. If the user has not chosen one, propose 2–3 plausible venues with tradeoffs: article type, audience, length pressure, figure limits, novelty bar, and format requirements. Ask the user to pick one target or to choose "no target yet" explicitly. In the latter case, record that no official template applies, and continue only after the user confirms this tradeoff. +Reuse a selected venue and installed template. If venue choice is part of the +request, propose plausible venues with audience, article type, length, novelty, +and format tradeoffs, then ask the user to choose. If no target is specified and +the user simply wants a draft, record that no target is set and continue in the +existing or requested format; do not force a venue choice before a useful draft. Once a target is chosen: 1. Find the official author instructions and template from the journal or publisher website. Prefer official publisher pages over mirrors, GitHub copies, Overleaf community templates, or lab handouts. 2. Download the template package into the active manuscript directory, usually `articles/YYYY-MM-DD-/template/`. If no manuscript directory exists yet, create the article directory first. 3. Record the template source URL, access date, journal name, article type, and key constraints in a short `template/README.md` or manuscript note. -4. If the official template cannot be downloaded, explain why, save the author-instruction URL, and ask the user whether to continue with a generic draft format. -5. Do not propose story lines until the target venue and template status are clear. +4. If the official template cannot be downloaded, record the limitation and source URL. Continue a content draft in the existing format; ask only if the requested deliverable requires a template choice that cannot be inferred. +5. Apply known venue constraints to the story and figures; identify requirements that remain unverified. ### Phase 1.6 — Story checkpoint with the user -Before the telegram outline, stop and discuss the paper's narrative with the user. Base this only on the provided figures, caption-summaries, existing draft, loaded literature context, and target-journal constraints. +Reuse a narrative supplied or already selected by the user. When materially +different scientific stories remain, discuss them before drafting. Base them on +the provided results, figures, existing draft, literature, and venue constraints. + +When a narrative choice is still needed: 1. Propose **1–3 candidate story lines** for the user to pick from. If there is only one defensible story, present one strong option and say why alternatives would be forced. 2. For each story line, include: @@ -85,39 +105,39 @@ Before the telegram outline, stop and discuss the paper's narrative with the use - the audience or venue fit, - what the story deliberately de-emphasizes. 3. Ask the user to choose one, combine pieces, or reject them. If they push back, revise the story lines and ask again. -4. Only after the user selects or synthesizes a story, continue to the telegram outline. Treat the selected story as the contract for the draft. +4. Once the story is established, continue to the outline and draft. Ask again only if new evidence changes the main claim, not at each writing phase. ### Phase 2 — Telegram outline - Write a telegram-style outline from the selected story line: section headings → bullet points → which figures and equations land where. - Mark which sentence in each section names a main result. -- Show the outline to a collaborator/advisor before you write prose. Cost of revision is lowest now. +- Reuse an approved outline. For a requested complete draft, use the outline as a working note and continue; present it for approval only if the user requested that checkpoint or the narrative still needs a decision. -### Phase 3 — Draft the body in this order +### Phase 3 — Draft the body (a useful default order) 1. **Methods / Theory** — easiest to write; gets you over the activation barrier. 2. **Results** — walk the reader through the figures in order. For each figure: state what was varied (x-axis), what was measured (y-axis), what trend appears, where errors come from. Data is obvious to you, not the reader. -3. **Analysis** — explain how the data matches (or stretches) the theory. Plot data as points, theory as lines, on the same axes; arrange so theory lies on straight lines whenever possible. Discuss deviations larger than error bars *and* deviations much smaller than them (both are problems). +3. **Analysis** — compare data and theory under the stated uncertainty model. Investigate unexpected residuals rather than declaring large or small deviations an error by themselves; use the relevant Figure Rulebook guidance. 4. **Introduction (rough draft).** Do not perfect it yet. Hit four beats: (a) field-level question and why it matters, (b) prior work and what was missing, (c) what *this* paper does, (d) where the main results live (figure / equation pointers). The model letter executes all four in as many paragraphs, closing with "In this Letter, we develop…" — see `references.md` §C. Move on even if it feels weak. 5. **Conclusions.** Often a re-statement of the introduction in newly technical language — the reader now has the apparatus to absorb it. Add one paragraph on implications, applications, and follow-up directions. Acknowledgments and funding here. ### Phase 4 — Iterate the body -- Revise the body many times before touching abstract/intro polish. +- Revise until the argument, notation, and evidence support the requested draft. Repeat a pass only for a changed section, a failed check, or an unresolved finding. - Each pass: check the **One Concept Per Sentence** rule, check notation consistency, check that every striking feature in every figure is *explained in text*. - Last pass before Phase 5 is a **language pass**: walk the style guide's hunt table and change *how* sentences are written, never *what* they say. Do not add or remove a claim, figure, or derivation step in that pass; recheck every number and qualifier a rewritten sentence mentions; leave passages that already pass untouched. ### Phase 5 — Polish the high-leverage sections last - **Abstract:** one paragraph, 5–10 lines, ~one sentence per body section. The model paper does it in four moves — system, method, finding, implication — one move per sentence. Write it last, when you finally understand what the paper says. -- **Title:** descriptive, specific, scannable. Rewrite several times. +- **Title:** descriptive, specific, scannable. Refine it when it misstates or obscures the main claim. - **Introduction:** sharpen the opening hook, the gap statement, the contribution claim, and the forward-pointers to figures. - **Conclusions:** make the take-home messages crisp and quotable. -### Phase 6 — External feedback +### Optional external feedback -- Send to a friend / officemate who is *not* a co-author. -- Solicit comments from a known expert in the area. Most will oblige if you mention a deadline. Consider sending to known competitors as a goodwill gesture — they catch what reviewers will catch. +- When useful, suggest feedback from a reader outside the author list. External sharing is a separate user decision, not a completion gate for drafting. +- If the user requests expert feedback, help prepare the material and questions; contact others only when explicitly authorized. - Take every comment seriously. "Confusing to a friend" → "confusing to a reviewer." --- @@ -142,49 +162,16 @@ Most physics-style papers fit this. Short letters (PRL, Nature, Science) drop th ## Figure Rulebook -Figures are what readers remember; design them to survive the harshest viewing context. - -**Design for three uses simultaneously.** Each figure must work as: (a) inline figure in the paper, (b) slide in a beamer talk, (c) greyscale photocopy. Design for all three at draft time, not in a later retrofit. - -**Line weight.** Minimum thickness 2 for every curve (or at least the main-result curves). Thin lines vanish on a projector. - -**Color discipline.** -- Use saturated, robust colors: black, blue, red, dark orange, magenta, violet, dark brown. -- Avoid light yellow, light green, light grey — invisible when projected. -- Encode the distinction with a *line style* (solid / dashed / dash-dot) in addition to color, so the figure survives greyscale printing. -- In the caption, refer to features by line style, not color: "the dashed curve" not "the red curve". Add "(Color online)" if color matters. - -**Text size.** Axis labels, numbers, and legend text should not be much smaller than the surrounding paper text — at most 2/3 of body size. Tiny text wrecks the figure for talks. - -**Parameter labels.** Place key parameters (e.g., `T = 0`, `V = 0`, `Γ = 1`) directly inside the plot in small boxes. Saves caption length and makes the figure self-contained for talks. - -**Dimensionless axes.** Use dimensionless quantities (`G/G₀`, `T/Γ`, `V_g/Γ`, etc.) whenever possible — they generalize the result, clarify the relevant scale, and travel across systems. Choose the combination that maximizes message clarity; if you find a better one after plotting, replot. Exception: comparison with dimensional experimental data. - -**Data vs. theory.** Plot data as points (with error bars) and theory as lines, on the same axes. Arrange so theory lines are *straight* whenever possible — anyone can then check agreement at a glance. Use curved-theory plots only with a deliberate reason. - -**Error-bar reasoning.** Theory should pass through error bars on most points. Patterns to *discuss in text*: -- Many points deviating by more than an error bar → likely systematic error; address it. -- Error bars dwarfing the deviations → uncertainties likely overestimated; address this too. - -**Captions.** Concise but self-sufficient. Define every plotted quantity, summarize the trend, identify line styles. A reader who reads only the title, abstract, and figures+captions should get the paper. - -**Explain every striking feature.** Every peak, dip, kink, or jump that catches the eye must be discussed in the main text — ideally with a back-of-envelope reason. Unexplained features are either an honesty problem or a missed opportunity. If you genuinely don't understand a feature, say so in print and flag it for follow-up. - ---- +For figure creation or review, read the Figure Rulebook in +[figures-and-notation.md](references/figures-and-notation.md). Use intended display +size and scientifically meaningful axes; explain any normalization, including +arbitrary units when the measurement legitimately requires them. ## Notation Rulebook -Notation is the reader's interface to the math. Treat it with the same care as a public API. - -- **No symbol reuse in nearby sections.** Same letter must not mean two different things within a few pages. -- **If notation must change, signal it explicitly.** "Henceforth we use X to denote..." — never silent reuse. -- **Define every variable before using it.** Define them in logical order: earlier symbols define later ones, never the reverse. -- **If a better notation appears mid-project, switch and rewrite earlier sections.** The reader's cost of decoding bad notation is far higher than your cost of rewriting. -- **Compact vs. explicit formulas:** - - *Compact* when summarizing strategy, manipulating reader's high-level model, or when an expert could fill in the steps. - - *Explicit* when: highlighting a non-obvious step, presenting a trick that took real effort, showing a key intermediate result other work depends on, presenting a flagship result, or matching a plotted figure (cite the figure in the equation). - ---- +For mathematical writing, read the Notation Rulebook in +[figures-and-notation.md](references/figures-and-notation.md). Local prose edits +need only the definitions and notation used by the affected passage. ## Sentence-Level Rules @@ -192,54 +179,24 @@ Follow `skills/how-to-technical-writing/SKILL.md`: one concept per sentence, dir --- -## Pre-Submission Checklist - -Run this before clicking submit. Each item is cheap to check; missing any of them is expensive to fix in proof. - -**High-leverage text (read by 95%):** -- [ ] Title is descriptive, specific, scannable. -- [ ] Abstract reads as a one-paragraph summary, one sentence per body section. -- [ ] Introduction has all four beats (field interest, prior work, what's new, where main results live). -- [ ] Every figure has a caption that stands alone. -- [ ] Conclusions name the contribution and at least one implication. - -**Process gates:** -- [ ] Target journal or "no target yet" was discussed with the user before story selection. -- [ ] Official template was downloaded, or the failed/blocked/not-applicable template status was recorded. -- [ ] The user selected or synthesized a story line before prose drafting began. - -**Main-result labeling:** -- [ ] Each major result has a sentence explicitly tagging it as a main result. -- [ ] Each main result has a corresponding figure or equation. - -**Figures:** -- [ ] All curves are line-weight ≥ 2. -- [ ] All colors are saturated; no faint yellow/green/grey. -- [ ] Each figure is identifiable in greyscale (line styles distinguish, not just color). -- [ ] Axis numbers and legend text are readable from a slide. -- [ ] No axis labeled in arbitrary units. -- [ ] Every striking feature is explained in text. -- [ ] Data as points + theory as lines, plotted together, with straight theory lines where possible. - -**Notation and equations:** -- [ ] No symbol reuse for different meanings. -- [ ] Every symbol defined before use, in logical order. -- [ ] Explicit equations only for non-obvious steps, key intermediates, flagship results, or figure references. - -**Sentence-level:** -- [ ] No paragraph contains more than one new concept per sentence; sentences run about 20 words. -- [ ] Active voice dominates. -- [ ] Topic sentences open each paragraph. -- [ ] No "obviously" / "clearly": every asserted step names the earlier equation, figure, or section it rests on. -- [ ] No warm-up sentences, meta-talk about the document, or metaphors standing in for a precise statement. -- [ ] Technical terms kept; Latinate connectives ("hence", "conversely", "likewise") replaced by plain ones. -- [ ] Each paragraph stays on one object; each cross-reference says why the current step needs it. -- [ ] No run of inline computations; calculations sit in a display with one sentence naming what it shows. - -**External feedback:** -- [ ] At least one friend / officemate has read the full draft. -- [ ] At least one expert outside the author list has commented (when feasible). -- [ ] Comments have been addressed, not deflected. +## Finish the requested deliverable + +- Verify the requested sections say what the supplied results support and that + citations resolve. For new or changed figures, inspect the render at its + intended size using the Figure Rulebook. Check affected symbols and labels. +- For language changes, use `skills/how-to-technical-writing/checklist.md`; + preserve logical connectives and mathematical meaning. Sentence length is a + signal for overloaded clauses, not a word-count target. +- Compile changed source and resolve failures introduced by the change using + the shared writing workflow. For an inline excerpt, check the text and state + that no document build was run. Do not require a bibliography for a passage + that contains no citations. +- For submission preparation, also check current venue limits, statements, + template requirements, and unresolved author decisions. Do not submit or + distribute the manuscript without authorization. +- Deliver the draft/source or revised excerpt, relevant verification, and any + unresolved scientific questions. External reviews are not required to finish + a writing task. --- @@ -258,7 +215,7 @@ Run this before clicking submit. Each item is cheap to check; missing any of the ## Integrations - **Citations and missing references:** Follow `skills/how-to-write-ideas-report/references/writing-workflow.md`. -- **Manuscript format:** Use the target journal's official template when available. Default to Typst (`.typ`) only when no target venue or required template exists; use LaTeX (`.tex`) or Word when the journal requires it; use Markdown only for arXiv-style preprints where the journal accepts it. +- **Manuscript format:** Preserve the requested or existing format. Use the target journal's required format for submission preparation; a content draft can remain in Markdown, Typst, or LaTeX until a venue is chosen. - **Storing the draft:** `articles/YYYY-MM-DD-/` with `main.typ` (or `.tex`), a bibliography copied from `$KB/references.bib`, and `figures/`. --- @@ -268,7 +225,7 @@ Run this before clicking submit. Each item is cheap to check; missing any of the - `skills/how-to-technical-writing/SKILL.md` — the `how-to-technical-writing` skill: sentence- and paragraph-level rules shared with `review-paper`, with the hunt table and application guardrails. - `references.md` — distilled rule lists from Martinis (2012) and von Delft (style guide), plus a walkthrough of the model paper (§C). - `sources/NotesOnWritingPaper12.pdf` — the original Martinis notes. -- `sources/1807.01815_Ho2019_quantum-scars.md` — the model paper: Ho, Choi, Pichler & Lukin, *Periodic orbits, entanglement and quantum many-body scars in constrained models*, PRL 122, 040603 (2019), rendered from arXiv:1807.01815. This letter practices what the rules preach: one move per abstract sentence, the four introduction beats in order, run-in headers whose first sentence names the section's job, symbols defined at first use and then read back in plain words, figures that carry the story from page one. Skim it before drafting. `references.md` §C maps each move to its location in the paper. +- `sources/1807.01815_Ho2019_quantum-scars.md` — the model paper: Ho, Choi, Pichler & Lukin, *Periodic orbits, entanglement and quantum many-body scars in constrained models*, PRL 122, 040603 (2019), rendered from arXiv:1807.01815. This letter practices what the rules preach: one move per abstract sentence, the four introduction beats in order, run-in headers whose first sentence names the section's job, symbols defined at first use and then read back in plain words, figures that carry the story from page one. Read a relevant excerpt only when the distilled guidance leaves a style question open. `references.md` §C maps each move to its location in the paper. - von Delft's *Style Guide* online: The references preserve the *reasons* behind the rules; the model paper shows the rules executed. diff --git a/skills/write-paper/references/figures-and-notation.md b/skills/write-paper/references/figures-and-notation.md new file mode 100644 index 0000000..b898ced --- /dev/null +++ b/skills/write-paper/references/figures-and-notation.md @@ -0,0 +1,50 @@ +# Figures and notation + +Use only the rulebook relevant to the current artifact. Venue requirements and +scientific meaning determine how these defaults apply. + +## Figure Rulebook + +Figures are what readers remember; design them to survive the harshest viewing context. + +**Design for the intended display.** Use the target paper, slide, or print dimensions. Add grayscale-safe encodings when the venue or audience needs them. + +**Line weight.** Judge curves at final display size, using explicit units or plotting-library parameters. A single numeric width is not portable across renderers. + +**Color discipline.** Choose legible colors and redundant encodings at the intended size. The following are useful defaults, not a required palette. +- Use saturated, robust colors: black, blue, red, dark orange, magenta, violet, dark brown. +- Avoid light yellow, light green, light grey — invisible when projected. +- Encode the distinction with a *line style* (solid / dashed / dash-dot) in addition to color, so the figure survives greyscale printing. +- In the caption, refer to features by line style, not color: "the dashed curve" not "the red curve". Add "(Color online)" if color matters. + +**Text size.** Axis labels, numbers, and legend text should not be much smaller than the surrounding paper text — at most 2/3 of body size. Tiny text wrecks the figure for talks. + +**Parameter labels.** Place key parameters (e.g., `T = 0`, `V = 0`, `Γ = 1`) directly inside the plot in small boxes. Saves caption length and makes the figure self-contained for talks. + +**Dimensionless axes.** Use dimensionless quantities (`G/G₀`, `T/Γ`, `V_g/Γ`, etc.) whenever possible — they generalize the result, clarify the relevant scale, and travel across systems. Choose the combination that maximizes message clarity; if you find a better one after plotting, replot. Exception: comparison with dimensional experimental data. + +**Data vs. theory.** Plot data as points (with error bars) and theory as lines, on the same axes. A transformation that straightens a predicted relationship may help readers judge agreement. Keep transformations scientifically interpretable; straightening a theory curve is optional. + +**Error-bar reasoning.** Compare residuals against the stated uncertainty model, accounting for what each interval represents and any correlations. Patterns to investigate: +- Unexpected residual patterns relative to the uncertainty model → investigate model mismatch or systematic error. +- Residuals unexpectedly small relative to the uncertainty model → investigate uncertainty estimates, correlations, or overfitting; this alone does not establish an error. + +**Captions.** Concise but self-sufficient. Define every plotted quantity, summarize the trend, identify line styles. A reader who reads only the title, abstract, and figures+captions should get the paper. + +**Explain every striking feature.** Every peak, dip, kink, or jump that catches the eye must be discussed in the main text — ideally with a back-of-envelope reason. Unexplained features are either an honesty problem or a missed opportunity. If you genuinely don't understand a feature, say so in print and flag it for follow-up. + +--- + +## Notation Rulebook + +Notation is the reader's interface to the math. Treat it with the same care as a public API. + +- **No symbol reuse in nearby sections.** Same letter must not mean two different things within a few pages. +- **If notation must change, signal it explicitly.** "Henceforth we use X to denote..." — never silent reuse. +- **Define every variable before using it.** Define them in logical order: earlier symbols define later ones, never the reverse. +- **If a better notation appears mid-project, switch and rewrite earlier sections.** The reader's cost of decoding bad notation is far higher than your cost of rewriting. +- **Compact vs. explicit formulas:** + - *Compact* when summarizing strategy, manipulating reader's high-level model, or when an expert could fill in the steps. + - *Explicit* when: highlighting a non-obvious step, presenting a trick that took real effort, showing a key intermediate result other work depends on, presenting a flagship result, or matching a plotted figure (cite the figure in the equation). + +--- diff --git a/tests/test_brainstorm_ideas_skill.py b/tests/test_brainstorm_ideas_skill.py index 262a2a9..6450573 100644 --- a/tests/test_brainstorm_ideas_skill.py +++ b/tests/test_brainstorm_ideas_skill.py @@ -5,6 +5,9 @@ def test_brainstorm_ideas_skill_requires_advisor_subagent_workflow(): text = BRAINSTORM_IDEAS_SKILL.read_text() + advisor_path = BRAINSTORM_IDEAS_SKILL.parent / "references" / "advisor.md" + assert "(references/advisor.md)" in text + text += advisor_path.read_text() required_phrases = [ "launch a dedicated advisor subagent", diff --git a/tests/test_skill_structure.py b/tests/test_skill_structure.py index 213ce1c..6178211 100644 --- a/tests/test_skill_structure.py +++ b/tests/test_skill_structure.py @@ -183,7 +183,8 @@ def test_brainstorm_ideas_loads_advisor_kb_from_dot_knowledge(): def test_brainstorm_ideas_operational_instructions_use_resolved_kb_variables(): text = _read("brainstorm-ideas") assert "Resolve the project KB via `KB=$(python3 \"$DOWNLOAD_REF_DIR/helpers/resolve_kb.py\")`" in text - assert "ADVISOR_KB=$(python3 \"$DOWNLOAD_REF_DIR/helpers/resolve_kb.py\" --advisor )" in text + advisor = (SKILLS / "brainstorm-ideas" / "references" / "advisor.md").read_text() + assert "ADVISOR_KB=$(python3 \"$DOWNLOAD_REF_DIR/helpers/resolve_kb.py\" --advisor )" in advisor assert "project knowledge base at `/.knowledge/`" not in text assert "Ground ideas in loaded knowledge bases (`/.knowledge/` and `advisors//.knowledge/`)" not in text diff --git a/tests/test_workflow_resources.py b/tests/test_workflow_resources.py new file mode 100644 index 0000000..838647b --- /dev/null +++ b/tests/test_workflow_resources.py @@ -0,0 +1,79 @@ +"""Exercise moved resources in independent installs and the documented build check.""" + +import os +import re +import shutil +import subprocess +from pathlib import Path + +import pytest + + +ROOT = Path(__file__).resolve().parents[1] + + +@pytest.mark.parametrize("name", ["brainstorm-ideas", "how-to-download-ref", "write-paper"]) +def test_relative_references_work_without_sibling_skills(tmp_path, name): + installed = (tmp_path / name).resolve() + shutil.copytree(ROOT / "skills" / name, installed) + pending = [installed / "SKILL.md"] + visited = set() + while pending: + source = pending.pop() + if source in visited: + continue + visited.add(source) + for link in re.findall(r"\]\(([^)]+)\)", source.read_text()): + if "://" in link or link.startswith("#"): + continue + target = (source.parent / link.split("#", 1)[0]).resolve() + assert target.is_relative_to(installed), (source, link) + assert target.is_file(), (source, link) + if target.suffix == ".md": + pending.append(target) + + +@pytest.mark.parametrize( + "compiler_exit,diagnostic,expected_exit", + [ + (0, "", 0), + (0, "warning: unresolved citation", 1), + (2, "error: unknown variable", 1), + (127, "compiler unavailable", 1), + ], +) +def test_documented_compile_check_handles_failure_and_warnings( + tmp_path, compiler_exit, diagnostic, expected_exit +): + workflow = ( + ROOT / "skills/how-to-write-ideas-report/references/writing-workflow.md" + ).read_text() + blocks = re.findall(r"```sh\n(.*?)\n\s*```", workflow, re.DOTALL) + checks = [block for block in blocks if "typst compile" in block] + assert len(checks) == 1 + + compiler = tmp_path / "typst" + compiler.write_text( + '#!/bin/sh\nprintf "%s\\n" "$TEST_DIAGNOSTIC" >&2\n' + 'exit "$TEST_COMPILER_EXIT"\n' + ) + compiler.chmod(0o755) + logs = tmp_path / "logs" + logs.mkdir() + env = { + **os.environ, + "PATH": str(tmp_path) + os.pathsep + os.environ["PATH"], + "TMPDIR": str(logs), + "TEST_DIAGNOSTIC": diagnostic, + "TEST_COMPILER_EXIT": str(compiler_exit), + } + result = subprocess.run( + ["/bin/sh", "-c", checks[0]], + cwd=tmp_path, + env=env, + capture_output=True, + text=True, + ) + assert result.returncode == expected_exit, result + assert diagnostic in result.stdout + result.stderr + assert not list(logs.iterdir()), "temporary compiler log was not removed" From 5e0672c126d0d11ef40fc31a55a74c051bf7d1be Mon Sep 17 00:00:00 2001 From: nzy1997 Date: Mon, 14 Sep 2026 11:33:33 +0800 Subject: [PATCH 2/2] Preserve guided defaults across shared research skills --- README.md | 8 ++++++ skills/brainstorm-ideas/SKILL.md | 23 +++++++++++----- skills/review-paper/SKILL.md | 47 ++++++++++++++++++++++++++------ skills/review-paper/checklist.md | 2 +- skills/write-paper/SKILL.md | 44 ++++++++++++++++++++---------- 5 files changed, 92 insertions(+), 32 deletions(-) diff --git a/README.md b/README.md index 2810281..de0a1bc 100644 --- a/README.md +++ b/README.md @@ -132,6 +132,14 @@ claims, reference provenance, acceptance gates, and research budgets remain explicit constraints. This revision follows OpenAI's [Rethinking skills and prompts for GPT-6 Astra](https://developers.openai.com/blog/rethinking-skills-and-prompts-for-gpt-6-astra). +These workflows are shared across Claude Code, Codex, OpenCode, and pi. New +exploration keeps Socratic guidance and offers available advisors; broad paper +polish defaults to a marked proposal; new manuscripts keep venue, story, and +outline checkpoints. Reuse prior decisions and direct-edit instructions instead +of asking for them again. Local requests and explicitly delegated drafting can +proceed within their scope. Model-specific prompt advice does not by itself +justify changing these shared interaction defaults. + ## Contributors **Initiators**: [Lei Wang](https://github.com/wangleiphy) and [Jin-Guo Liu](https://github.com/GiggleLiu) diff --git a/skills/brainstorm-ideas/SKILL.md b/skills/brainstorm-ideas/SKILL.md index 8c1b274..7f0ec77 100644 --- a/skills/brainstorm-ideas/SKILL.md +++ b/skills/brainstorm-ideas/SKILL.md @@ -17,10 +17,13 @@ Before running the examples, set `DOWNLOAD_REF_DIR` to the absolute directory of # Brainstorm research ideas Help the user find, refine, or reason through an attackable research problem. -Be curious and candid. Offer your own reasoning, calculations, counterexamples, -and literature checks when useful; use Socratic questions when they help the -user think or when the user asks for that style. Mark assumptions and distinguish -source-supported findings from hypotheses or opinion. +Be curious and candid. For open-ended exploration, default to a Socratic +collaboration: help the user articulate motivations, compare assumptions, and +choose a direction through focused questions and your own substantive reasoning. +Explain why a question matters; do not turn the discussion into a quiz or +withhold a calculation the user asked you to do. For a specified derivation, +comparison, or report, carry out that task directly. Mark assumptions and +distinguish source-supported findings from hypotheses or opinion. ## Enter at the current need @@ -53,9 +56,15 @@ Use the user's profile and provided constraints; request background only if it would change the advice. When the user chooses Zotero or a Scholar profile as background, invoke `know-me-better` with that source and return to this discussion. -An advisor is optional. If the user names one, or asks to choose from the advisor -library, consult `advisors/index.md`. Show names and fields from the index; read -only profiles needed for the choice. Without a selection, continue as the mentor. +An advisor is optional. At the start of a new open-ended exploration, if an +advisor library is available and no preference is known, consult +`advisors/index.md` and offer relevant names/fields alongside continuing with +the mentor alone. Include this in the opening discussion, not a separate +mandatory setup step. Do not silently select an advisor. Reuse a prior selection +or decision to work without one; a scoped task or resumed session does not need +another advisor menu. Also consult the index when the user names an advisor or +asks to browse the library; read only profiles needed for that choice. +Without a selection, continue as the mentor. When selected, **launch a dedicated advisor subagent** following [advisor.md](references/advisor.md). Its literature lives at `advisors//.knowledge/`, resolved by the KB helper. Advisor-only audio with diff --git a/skills/review-paper/SKILL.md b/skills/review-paper/SKILL.md index 6e3374f..a939626 100644 --- a/skills/review-paper/SKILL.md +++ b/skills/review-paper/SKILL.md @@ -28,10 +28,22 @@ Use `skills/how-to-write-ideas-report/references/writing-workflow.md` for KB/con ## Operating principle -**Preserve scope and author control.** A request for review or comments authorizes -findings, not manuscript edits. A request to revise, polish, or apply selected -findings already authorizes those changes; do not ask for the same approval -again. For wording-only work, preserve scientific meaning. Unresolved scientific +**Choose the application mode before changing a manuscript file.** Use the +user's current instructions and prior decisions: + +| Request or existing decision | Application mode | +|---|---| +| Review, critique, comments, or fact check | Findings only; leave the original unchanged. | +| Polish or revise a whole manuscript, with no direct-edit decision | Marked proposal in a separate copy; leave the original unchanged until the user accepts changes. | +| Explicitly edit the original directly, apply accepted findings, or continue an established direct-edit mode | Apply within that authorization and verify; do not ask again. | +| A specific local replacement or edit to a named passage | Complete that local edit without reopening whole-paper approval. | +| Rewrite an excerpt supplied in chat | Return revised text inline. | + +Naming `main.md` or limiting a whole-paper polish to English/grammar does not +select direct mode. Prepare the marked proposal and numbered changes before +asking which to apply; do not ask permission to begin the work. A user-requested +preview takes precedence over the local-edit default. For wording-only work, +preserve scientific meaning. Unresolved scientific judgment calls remain comments, while a separately requested substantive revision may be proposed from the supplied evidence. Never invent a result or justification. @@ -176,21 +188,38 @@ parts of this structure: For a review-only request, deliver the report; offer application as an optional next step. If changes were already requested, continue to Phase 4 using that -scope. Reuse the chosen direct/diff mode. When the user has not authorized edits -to the original, prepare a marked proposal for review before merging it. +scope and the application mode from the Operating principle. For broad polish +without an established mode, deliver the marked proposal and numbered changes +for acceptance before merging; a generic revision request alone does not select +direct mode. Reuse explicit direct-edit instructions or already accepted +findings without asking again. A requested inline excerpt rewrite can be +delivered directly as revised text. --- ## Phase 4 — Apply approved edits -**Marked-diff mode** (when requested, or proposing edits beyond existing authorization). Prepare the proposed changes before asking the user which to merge. +**Marked-diff mode** (the default for broad revision/polish without an established +application mode, and whenever the user requests a preview). Prepare the +proposed changes before asking the user which to merge. Proposals outside the +authorized scientific scope remain separate suggestions, not applied edits. 1. **Edit a copy.** `cp main.tex main.proposed.tex` (same for `.typ` / `.md`). Apply the proposed fixes to the copy; when using scripted replacements, assert that each target matches exactly once. Comment-only fixes (style-guide guardrails) go in as `[reviewer]` comments in the copy too. 2. **Mark the diff.** LaTeX: `latexdiff main.tex main.proposed.tex > main.diff.tex`, then compile `main.diff.tex` (deletions red struck-through, additions blue underlined). Typst and Markdown have no latexdiff; write `git diff --no-index --word-diff main.typ main.proposed.typ` into `articles//review-YYYY-MM-DD.diff` and, for Typst, also compile the proposed copy so the author can read the result. Check the page count did not change unexpectedly. + If shell/diff/render tools are unavailable, still create the proposed copy + with available file tools and provide numbered before/after changes. Report + the unavailable checks; do not ask whether to prepare the already requested + proposal. Acceptance concerns applying it to the original. 3. **Hand over a numbered legend.** Give the marked PDF (or diff file) path and one line per change: number, section and page, the finding it fixes, and the rule it serves. Pair up a removal and an addition that belong to one change. End with "reply with the numbers to merge, or all". 4. **Merge exactly the accepted numbers** into the original. If the user accepts "all except one wording", revert that wording in the copy first, then merge all. Delete `main.proposed.*` and `main.diff.*` afterwards. -**Direct mode** (when the user requested edits without a preview gate). Apply changes within that authorization to the manuscript and show the diff/result. Do not broaden the requested revision. +**Direct mode** (when the user asked to edit the original directly, approved +specific findings, or already chose this mode). Apply changes within that +authorization and show the diff/result. An explicit request such as "replace +'We studies' with 'We study' in this paragraph" authorizes that local edit; +a general "polish my paper" +uses the marked proposal by default. Do not broaden the requested revision or +ask again to apply changes already accepted. **Both modes.** @@ -215,7 +244,7 @@ to the original, prepare a marked proposal for review before merging it. | Mistake | Instead | |---|---| | Treating a review request as permission to edit | Deliver findings; use existing revision authorization only for its stated scope. | -| Ignoring a requested preview gate | Prepare a marked copy and merge accepted changes; reuse direct-edit authorization when no preview was requested. | +| Treating broad polish as an implicit direct-edit preference | Default to a marked proposal; reuse explicit direct-edit instructions or accepted changes without another approval. | | Running a full story gate for a local wording fix | Read the target and its necessary context; reserve the story brief for full reviews. | | Re-running reference verification on every pass | Inspect changed inputs and unresolved findings; rerun only affected checks. | | Inventing a BibTeX entry to "fix" a citation | Never. Use the lookup chain / the `how-to-download-ref` skill, or flag as unverifiable. | diff --git a/skills/review-paper/checklist.md b/skills/review-paper/checklist.md index aeeb6e1..8d7a6c8 100644 --- a/skills/review-paper/checklist.md +++ b/skills/review-paper/checklist.md @@ -55,7 +55,7 @@ The checkable items live in `skills/how-to-technical-writing/checklist.md`. The - [ ] A reference and fact-check table is included when that pass applies. - [ ] A journal fit table (constraint → required → measured → status) with the guideline source is included. Otherwise a line says the pass was skipped. - [ ] A prioritized "top fixes" list is included. -- [ ] The chosen direct/diff mode was reused. Already requested edits continued without asking for the same authorization; a proposal beyond that scope was prepared before requesting a decision. +- [ ] The chosen direct/diff mode was reused. Broad revision/polish without an established mode produced a marked proposal before changing the original. Explicit direct edits and accepted findings continued without another approval; an inline excerpt rewrite needed no application checkpoint. - [ ] In marked-diff mode, edits went to a `*.proposed.*` copy. `latexdiff` produced the marked version, or `git diff --word-diff` for Typst and Markdown. A numbered legend was handed over. Only the accepted numbers were merged. The proposed and diff files were deleted afterwards. - [ ] Applied edits stayed within the existing user authorization; a review-only request did not modify the original. - [ ] LaTeX, Typst, and Markdown structure and macros are preserved. Author-judgment fixes are left as `% [reviewer]` comments. diff --git a/skills/write-paper/SKILL.md b/skills/write-paper/SKILL.md index 9934054..1ee8986 100644 --- a/skills/write-paper/SKILL.md +++ b/skills/write-paper/SKILL.md @@ -24,11 +24,15 @@ Use `skills/how-to-write-ideas-report/references/writing-workflow.md` for KB loa ## Choose the scope -- **New manuscript:** use the phases below as a useful default. Establish the - evidence and narrative before polishing prose; the supplied result, figure - plan, or proof can provide that evidence. +- **New manuscript:** default to the guided phases below, including unresolved + venue, story, and outline decisions with the user before full prose. The + supplied result, figure plan, or proof provides the evidence for those choices. - **Continue an existing draft:** reuse its story, venue, outline, and figures. Enter at the unfinished part and verify affected dependencies. +- **Broad revision or polish of an existing manuscript:** use `review-paper` + and carry forward the requested scope and application mode. With no mode + established, prepare a marked proposal before changing the original; reuse + direct-edit instructions or accepted findings without another approval. - **Local revision:** change only the requested section, abstract, caption, or passage, using its necessary context. Do not restart figure production, venue selection, or story approval. For a critique rather than revision, use @@ -36,8 +40,13 @@ Use `skills/how-to-write-ideas-report/references/writing-workflow.md` for KB loa - **Submission preparation:** apply the target venue's current requirements and the relevant final checks. -Carry forward choices and authorization from the conversation. Clarify only a -missing scientific decision that changes the draft. Never manufacture results, +Carry forward choices and authorization from the conversation. A generic +"help me write a paper" uses the guided workflow; it does not by itself waive +the initial checkpoints. A supplied or previously approved venue, narrative, +or outline satisfies the corresponding checkpoint. If the user delegates the +remaining writing choices and asks for a complete draft without checkpoints, +state the choices you make and proceed, asking only for missing scientific +input that cannot be inferred. Never manufacture results, references, or an unsupported claim. A writing request ends with the requested text/source and appropriate checks; external feedback is optional. @@ -75,11 +84,13 @@ Use supplied source material even without a sci-brain KB. Research citation gaps ### Phase 1.5 — Target journal and template checkpoint -Reuse a selected venue and installed template. If venue choice is part of the -request, propose plausible venues with audience, article type, length, novelty, -and format tradeoffs, then ask the user to choose. If no target is specified and -the user simply wants a draft, record that no target is set and continue in the -existing or requested format; do not force a venue choice before a useful draft. +Reuse a selected venue and installed template, including an existing decision +to draft without a target. For a new guided manuscript with no venue decision, +discuss plausible venues with audience, article type, length, novelty, and +format tradeoffs; let the user choose a target or "no target yet". For a local +revision, or when the user has delegated those choices and requested a draft, +do not reopen venue selection. If drafting without a target, record that status +and use the requested or existing format. Once a target is chosen: @@ -91,9 +102,12 @@ Once a target is chosen: ### Phase 1.6 — Story checkpoint with the user -Reuse a narrative supplied or already selected by the user. When materially -different scientific stories remain, discuss them before drafting. Base them on -the provided results, figures, existing draft, literature, and venue constraints. +Reuse a narrative supplied or already selected by the user. In a new guided +manuscript, present the proposed story for the user's decision before full +prose, even if only one story is defensible. When the user has delegated the +narrative choice, state the evidence-based story you select and proceed. Base +the story on the provided results, figures, existing draft, literature, and +venue constraints; missing results cannot be replaced by a writing choice. When a narrative choice is still needed: @@ -105,13 +119,13 @@ When a narrative choice is still needed: - the audience or venue fit, - what the story deliberately de-emphasizes. 3. Ask the user to choose one, combine pieces, or reject them. If they push back, revise the story lines and ask again. -4. Once the story is established, continue to the outline and draft. Ask again only if new evidence changes the main claim, not at each writing phase. +4. Once the story is established, continue to the outline. Reopen that story only if new evidence changes the main claim. ### Phase 2 — Telegram outline - Write a telegram-style outline from the selected story line: section headings → bullet points → which figures and equations land where. - Mark which sentence in each section names a main result. -- Reuse an approved outline. For a requested complete draft, use the outline as a working note and continue; present it for approval only if the user requested that checkpoint or the narrative still needs a decision. +- Reuse a supplied or approved outline. In a new guided manuscript, show the outline before drafting full prose; it may accompany the story proposal so the user can approve both together. If the user delegated outline decisions or authorized drafting from the established plan, use it as a working note and continue. Do not require a separate external advisor to approve it. ### Phase 3 — Draft the body (a useful default order)