Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 21 additions & 0 deletions components.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
{
"$schema": "https://ui.shadcn.com/schema.json",
"style": "new-york",
"rsc": false,
"tsx": true,
"tailwind": {
"config": "",
"css": "src/ui/index.css",
"baseColor": "neutral",
"cssVariables": true,
"prefix": ""
},
"iconLibrary": "lucide",
"aliases": {
"components": "@/components",
"utils": "@/lib/utils",
"ui": "@/components/shadcn",
"lib": "@/lib",
"hooks": "@/hooks"
}
}
145 changes: 145 additions & 0 deletions docs/architecture/ai-context-providers.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,145 @@
# AI hint context providers

What the hint assistant is told about the student, where each piece comes from, and why the
boundaries are drawn where they are. Placement, gating and the tutoring policy are in `ai-hints.md`.

## 1. Problem

A model that sees only the exercise text can only paraphrase it back, and, lacking any view of the
student's work, falls back on asking what they already tried. The student came here because they
are stuck; being questioned about their own history is the wrong response when that history is on
disk.

So each turn carries a fresh snapshot of the exercise, and the system prompt tells the model to
infer progress from it rather than ask.

## 2. Shape

`src/electron/ai/context.ts` holds an array of providers:

```ts
type ContextProvider = {
id: string;
label: string;
collect: (ctx: {
exerciseId: string;
location: ExerciseLocation | null;
}) => Promise<string | null>;
};
```

`collectContext` resolves the exercise location once (`locateExercise`), runs every provider in
parallel, drops anything that fails, times out or is empty, and truncates each to `MAX_BLOCK_CHARS`.

| Provider | Block label | Source |
| ----------------- | ---------------- | ------------------------------------------------------------ |
| `instructions` | Instructions | live scrape of the lesson page, else the copy cached on open |
| `exercise-folder` | Exercise folder | exercise manifest, progress store, names-only directory tree |
| `git-state` | Repository state | read-only `git` subprocesses in each repository found |

Adding a signal is one file under `ai/context/` plus one array entry. The system prompt, IPC and the
panel's context disclosure need no changes.

### Why separate blocks rather than one

Truncation at `MAX_BLOCK_CHARS` is per-block and silent. Concatenating sources means whichever lands
last is cut mid-line, in an order-dependent way. Separate blocks make the cap a per-source guarantee,
let a dead WebContentsView cost only the instructions, and keep the disclosure's labels granular.
Those labels are the student-facing record of what was sent, so they must name the actual sources.

Block text is fenced in `buildSystemPrompt` with a backtick run longer than any inside it. Porcelain
status lines begin with `##`, and the scraped instructions carry their own headings and code fences.
Without the fence, both would be read as structure of the surrounding prompt.

### Deadlines are not optional

Collection gates the turn, and every provider reaches something that can stall: a WebContentsView
mid-navigation, or a git subprocess. Each provider races `PROVIDER_TIMEOUT_MS` (2500 ms), and
`collectContext` takes the turn's abort signal. A partial context always beats a turn that will not
start.

### Instructions survive navigation

The instructions are scraped when the student clicks AI Hints and cached per exercise
(`session.ts`). Later turns try a live scrape first, and fall back to the cache when the student has
navigated to another page. Otherwise, reading ahead in the lesson would silently strip the model of
the task.

## 3. Directory resolution

The location comes from `resolveExerciseCwd`, mirroring `resolveReadyExerciseCwd` in
`ipc/gitmastery.ts`. The pty's tracked cwd (`getCwd()`) is deliberately **not** used. It is
regex-guessed from typed `cd` commands and drifts on `pushd`, `cd -`, subshells and chained commands,
so by the time a hint is requested it can point anywhere. See `exercise-directory-resolution.md` §9
for the same problem in `verify`.

`ExerciseLocation` carries the exercises root, the exercise root and the working folder (where Start
`cd`s). For hands-on practicals the working folder is the exercise root, but the repository usually
sits one level down (`hp-init-repo/things`). The git provider therefore searches rather than
assuming (§5).

## 4. What is never sent

- **File contents.** The folder provider lists names only; `.git` is summarised, not walked.
- **Anything outside the exercise folder.** The tree walk and repository search do not follow
symlinks, and every repository is containment-checked (§5).
- **The student's home path or username.** Paths are rendered relative to the folder above
`gitmastery-exercises`, e.g. `gitmastery-exercises/under-control`.
- **Terminal output and scrollback.** Output can echo anything the student ran, including
credentials, and its size is unbounded.

Adding any of these would cross the boundary stated in `../llm-integration.md` and needs a new
decision, not an extension of this one.

## 5. The git provider

### Repository discovery

Candidates are the working folder plus every directory within `SEARCH_DEPTH` (2) of the exercise
root that contains `.git`, nearest first. Hidden directories are skipped and symlinks are not
followed. Each candidate must pass the containment check. At most `MAX_REPOS` (3) are described,
each under its own `### Repository at …` heading, so exercises that set up more than one
repository are all visible. If none pass, the block says so explicitly. For `under-control`,
where `git init` _is_ the task, "no repository yet" is the most useful fact the model can hold.

### Containment is a privacy control

A directory existing proves nothing about whether it is a repository, or **whose**. Git searches
upwards. In `under-control`, and in every `repo_type: "ignore"` exercise, a bare `git status`
succeeds against whatever repository sits above the exercises folder. A student whose home or
coursework directory is a repository would have unrelated branch names and filenames sent to a
third-party free tier.

So each candidate runs `git rev-parse --show-toplevel` first, and `fs.realpathSync` of the result
must equal the candidate exactly. Anything else means "not a repository here", and nothing more
runs. `GIT_CEILING_DIRECTORIES` is set to the candidate's parent as a second line of defence.

### Commands

`status --porcelain=v1 -b`, `log --oneline --decorate --all -n 15`, `branch -vv`, `remote -v`,
`tag --list` and `stash list` run in parallel via `execFile` with an args array (no shell, matching
the rest of the codebase).

- **Porcelain, not plain `git status`.** Plain status advertises commands in its hint lines
(`use "git restore --staged <file>..."`), which works against a model meant to stay inside the
course. Porcelain suggests nothing. The cost is a short legend, because the XY column semantics are
what a small model misreads.
- **`--no-optional-locks`.** Status runs on every send, exactly when the student may be part-way
through their own `git add`. Without the flag, status rewrites the index and can race `index.lock`.
- **No `--graph`.** The ASCII art is noise to a small model; `--decorate --all` plus `branch -vv`
carries the same topology more legibly.
- **Paused operations.** A paused merge, rebase, cherry-pick, revert or bisect is reported from
marker files in `rev-parse --absolute-git-dir`. Porcelain status alone does not say which
operation is in progress, and "you are mid-rebase" is often the whole answer.
- `GIT_TERMINAL_PROMPT=0`, a 1500 ms timeout and a 1 MB buffer.

## 6. Cost

Context is collected per send, so the system prompt is rebuilt every turn. `data-*` parts are dropped
by `convertToModelMessages`, so context does not accumulate across turns. Cost is linear in turns,
not quadratic, and the model never sees earlier turns' repository state. History sent to the model is
capped at the last 20 messages.

Added latency per send is a directory walk, one `rev-parse` per candidate, six parallel git commands
per repository and one `executeJavaScript`. That is typically well under 200 ms, bounded at 2500 ms,
against free-tier first-token latency.
166 changes: 166 additions & 0 deletions docs/architecture/ai-hints.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,166 @@
# AI Hints

How the AI Hints feature is placed, gated, wired to a model, and kept from handing out answers.
What the model is told about the student is in `ai-context-providers.md`; how the app obtains
inference at all is assessed in `../llm-integration.md`.

## 1. Entry point

Every exercise card and every hands-on practical on the embedded lesson site gets an **AI Hints**
button, injected by `ipc/webContentsView.ts` next to the existing download controls. A click sends
the exercise id over `wcv-ai-hints`. Main re-validates it (format, exercise on disk, AI configured),
scrapes the title and instructions from the page, and announces an `AiHintsSession` to the renderer
with `ai-hints-open`.

Hands-on practicals have no identifier on the page beyond the `hp-…` id in their download command,
so the injected script tags the wrapper with `data-gm-hands-on-id` when it adds its buttons. The
instruction scrape reads `innerText` with the injected controls (`data-gm-actions`) hidden, so the
model never sees "AI Hints" or "Start Exercise" as part of the task.

## 2. Enablement

The button is only enabled when there is something to ground a hint in:

1. **AI is configured**: the active provider has everything it needs (key, base URL, model).
2. **The exercise is on disk**: `resolveExerciseCwd` returns `ready` for that id.

A disabled button carries a tooltip naming the missing step ("Click Start Exercise first…", "Set up
AI hints in Settings first"). A hint without the student's repository could only paraphrase the
instructions back at them, which is the failure `ai-context-providers.md` §1 describes.

Main pushes `{ configured, ready: string[] }` into the page (`__gmSetAiState`) on dom-ready, after a
download or Start, after AI settings are saved, when the exercises root changes, and when the window
regains focus. Focus covers the one change the app cannot observe: the student deleting a folder in
Finder. The click handler re-checks rather than trusting page state, since the page is remote content.

## 3. Placement

The chat is a pane **stacked above the terminal** in the right-hand work column, with a draggable
divider between them. Closing it returns the full height to the terminal; the conversation survives
close/reopen and is only cleared by **Clear history** or opening hints for a different exercise. The split
defaults to 60% of the column and becomes a fixed pixel height once dragged.

The layout never exceeds three columns: lessons nav, lesson page, work column. The lesson page's
width does not change when hints open, so instructions don't reflow, and the hint sits directly
above the prompt it will be acted on in.

### Rejected

- **A docked fourth column** between the lesson page and the terminal. Built first, and too
crowded: with the lessons nav open the window held four panes, and making room for the chat meant
squeezing the lesson page or auto-collapsing the nav.
- **A floating modal or popover over the lesson page.** The lesson site is a native
`WebContentsView`, which paints above all DOM. A DOM overlay would need to suppress the view (via
`useEmbeddedSuppressed`), hiding the very instructions the student is asking about. It would also
cover the page while the student works through a hint.
- **A floating second `WebContentsView` for the chat.** This was the previous implementation: its own
renderer, route, drag handling and window-level IPC. It solves the stacking problem, but it is a
second web contents with its own lifecycle, and it still covers the lesson.
- **Tabs in the work column ("Terminal | AI Hints").** Cheaper on space, but students would have to
switch away from the chat to run the command it just explained, then switch back to read the next
step.

## 4. Providers

`ai/llmProviders.ts` is a registry of `ProviderSpec`s. Each spec has display metadata, `createModel`,
a key `validate` probe and optional `providerOptions`. Shipped: OpenRouter (default), OpenAI,
Anthropic, Google, and **Custom** for any OpenAI-compatible base URL (Groq, Ollama, a course proxy).
Adding a provider is one spec; settings, IPC and the settings UI read the registry.

- **The model is not a setting, except for Custom.** Each named provider always uses its spec's
default model, and any model saved earlier is ignored. Students can't judge which model tutors
well, and a free-text model field let a saved `openrouter/free` bypass the routing below. Custom
has no sensible default, so it keeps a required Model field.
- **Per-provider settings.** `config.ai.providers[id]` stores the key (plus base URL and model for
Custom) for each provider separately, so switching providers doesn't discard a key already
entered.
- **Keys are encrypted with `safeStorage`.** Where that is unavailable, keys fall back to plaintext
and the settings panel says so.
- **Keys are validated on save.** Each spec probes a cheap authenticated endpoint. 401/403 means the
key was rejected, and a network failure is reported as such, so a typo surfaces in Settings rather
than as a failed chat. Custom endpoints treat a 404 on `/models` as acceptable because not every
compatible server implements it.
- **All traffic goes through `net.fetch`.** Requests use Chromium's network stack, which trusts the
macOS Keychain. Corporate TLS inspection certificates therefore work, where Node's bundled CA list
would reject them.

The earlier OpenRouter-only key field (`openRouterApiKeyEnc`) was dropped without migration. The
feature had not shipped.

### OpenRouter model routing

The default provider has to work with a free key, and free models are the least reliable part of
the whole feature. Each request sends a primary `model` plus OpenRouter's `models` fallback array,
and OpenRouter tries them in order, moving on when one is down, rate-limited, or has left the free
tier:

1. `google/gemma-4-31b-it:free` (primary)
2. `qwen/qwen3.8-27b:free`
3. `nvidia/nemotron-3-super-120b-a12b:free`
4. `openrouter/free` (last resort)

The three fallbacks are the most OpenRouter is documented to accept (its equivalent `fallbacks`
parameter caps at three).

**Why named models come first.** On its own, `openrouter/free` picks from every zero-cost model, and
in testing it produced unusable replies:

- a ~2B agent-tuned model (`liquid/lfm-2.5-2.6b`) answered in its raw tool-call syntax
(`<|tool_call_start|>[read(filePath=…)]<|tool_call_end|>`), because it wanted to read files and no
tools were offered;
- a content-safety classifier (`nvidia/nemotron-3.5-content-safety`) answered "User Safety: safe".

The named models are mid-sized, instruction-tuned, and follow the hint policy reasonably well.

**Why `openrouter/free` is still the last resort.** The tradeoff is availability against quality.
Free models are rate-limited per model, and at busy times all three named models can be throttled
together. Without the router, the student gets a rate-limit error and has to wait. With it, they
usually still get an answer, but occasionally from a model that can't tutor. Availability wins: the
feature is only useful if it answers, and a bad reply is visible and easy to retry.

What limits the damage when the router picks badly:

- The system prompt says the model has no tools and must reply only in text. Weaker agent-tuned
models then answer in prose more often, but not reliably.
- **Regenerate** retries the whole chain. That lands on a named model once its rate limit has reset,
or on a different random model if it hasn't.
- The classifier and code-only models can't be steered by the prompt. When they answer, the reply is
obviously wrong rather than subtly misleading, which is the lesser failure for a tutor.

**Alternatives rejected:**

- **Stripping tool-call tokens from the output.** This treats one model family's symptom, and the
answer underneath still comes from a model too small to tutor.
- **Choosing a free model at runtime from `/models`.** The listing has price and context length but
no quality signal. It can't tell a 2B model or a classifier from a 30B instruct model.
- **A paid default model.** It would be reliable, but every student would need a funded account,
which is what the free default exists to avoid.

**Upkeep.** Free-tier membership changes without notice (see `../llm-integration.md`, Option 9).
When replies start coming from the router regularly, check the named models against
`https://openrouter.ai/api/v1/models` and replace any that have left the free tier. The response's
`model` field shows which model actually answered.

## 5. Tutoring policy

`ai/prompt.ts` builds the system prompt. The policy depends on the session kind:

- **Exercises are graded.** The model gives the smallest useful nudge, never the full command
sequence or exercise-specific values, and never answers for `answers.txt`. It escalates one step at
a time, and repeated pressure does not change this.
- **Hands-on practicals are guided walkthroughs.** The model may point to the command the
instructions give for the current step, but must not run ahead of where the repository shows the
student to be.

Both kinds share the same rules:

- **Scope.** Answers are limited to this exercise, Git and GitHub, the Git-Mastery app and CLI, and
basic terminal use. Anything else is politely declined.
- **Injection resistance.** Context blocks are fenced data, never instructions, and the prompt is
not revealed.
- **Grounding.** The model infers progress from the repository snapshot instead of asking what the
student tried. When the work looks complete, it suggests running Verify.

The prompt is guidance rather than enforcement: a determined student can extract more than the
policy intends. Grading is what actually protects exercises, and the policy keeps the default
interaction pedagogical.
Loading
Loading