Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .claude-plugin/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@
{
"name": "goax",
"description": "AI 에이전트 결과 일관성을 환경으로 통제하는 4계층 하네스 (Triage·Constitution·Module·Spec/ADR) + Spirit·Mistake Loop. bash+markdown only, 자연어로 도입. Claude Code·OpenCode 지원.",
"version": "0.7.6",
"version": "0.7.7",
"author": {
"name": "bluecheat",
"email": "itsinil@gmail.com"
Expand All @@ -30,5 +30,5 @@
]
}
],
"version": "0.7.6"
"version": "0.7.7"
}
2 changes: 1 addition & 1 deletion .claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "goax",
"version": "0.7.6",
"version": "0.7.7",
"description": "AX 4-Layer harness — Triage / Constitution / Module / Spec·ADR + Spirit·Mistake Loop. Bash + Markdown only, with Claude-driven onboarding. Multi-CLI: Claude Code (native) + OpenCode (Hybrid compat via AGENTS.md SSOT + opencode.json).",
"author": {
"name": "bluecheat",
Expand Down
6 changes: 3 additions & 3 deletions CLAUDE.md

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -139,7 +139,7 @@ Natural-language tier overrides: `"simple"` / `"spec + tasks"` → standard · `
| `/goax` | Index — shows all triggers (the only remaining slash command) |
| "set up goax" / "install goax" | up → onboarding (brownfield) or zero (greenfield) |
| "start a new project" / "from zero" | zero — 0→1 entry: product · business · ADR · enforcement plumbing |
| "vendor goax" | vendor — ship skills inside the repo without the plugin |
| `/vendor` (explicit only) | vendor — ship skills inside the repo without the plugin. `disable-model-invocation: true`: it writes many files into the repo, so it runs only when you type `/vendor` and stays out of the skill listing |
| "diagnose" / "goax doctor" | doctor — gap diagnosis + `_templates` drift |
| "show rules" / "critical rules only" | doctor → `rules-index.sh` — Constitution + Spirit + Module index |
| "create spec — <slug>" | spec — tier-aware spec generation |
Expand Down
2 changes: 1 addition & 1 deletion VERSION
Original file line number Diff line number Diff line change
@@ -1 +1 @@
0.7.6
0.7.7
5 changes: 5 additions & 0 deletions agents/architect.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,10 @@ description: "아키텍처 결정·검토 sub-agent. 시스템 설계·모듈
- 영향 받는 .ax/modules/<name>/rules.md 갱신 제안
```

**길이 예산 — 권장 먼저 읽히게, 전체 약 1~2천 토큰.** 탐색에 많이 써도 돌려주는 건 압축한 결론이에요.
근거는 코드를 붙이지 말고 `<파일>:<줄>` 포인터로 가리켜요 — 메인 세션이 필요한 줄만 Read 해요.
옵션은 결정에 실제로 갈리는 것만 (보통 둘, 많아야 셋). 아래 spec 합의 리뷰 파일도 같은 예산이에요.

## spec 합의 리뷰 — `spec-validate` 가 띄워요

계획의 품질은 diff 시점이 아니라 **계획 시점**에 리뷰해야 올라가요. `spec-validate` 가 명료성
Expand Down Expand Up @@ -112,3 +116,4 @@ Option <X>. 이유: …
- verdict 없이 끝내기 — 첫 줄이 `verdict:` 가 아니면 리뷰가 없던 게 돼요.
- spec 리뷰에서 빌드·테스트를 돌리기 — 리뷰 한 번이 구현 한 번만큼 비싸져요.
- 비차단 항목으로 `보강 필요` 를 내기 — 라운드가 늘어나는 이유의 대부분이에요.
- 코드·파일 본문을 출력에 옮겨 붙이기 — `<파일>:<줄>` 포인터면 돼요 (위 길이 예산).
6 changes: 6 additions & 0 deletions agents/evaluator.md
Original file line number Diff line number Diff line change
Expand Up @@ -104,6 +104,11 @@ verdict: 진행
기존 review.md 가 있으면 덮어써요 — 재리뷰의 기록은 최신 것 하나면 돼요. 이전 지적이 처리됐는지는
tasks.md 의 task 로 남아요.

**길이 예산 — 파일 전체 약 1~2천 토큰.** 이 파일은 코디네이터가 읽고 task 로 옮겨요. 지적 하나는
`<파일>:<줄>` 포인터 + 어느 AC 인지 + 재현 시나리오 한두 줄이면 돼요. 코드 블록·diff 를 옮겨 붙이지 않아요 —
포인터가 있으면 읽는 쪽이 그 줄을 직접 열어요. 지적이 많아 예산을 넘으면 늘어나는 건 지적 개수예요.
spec 모드 파일도 같은 예산이에요.

## spec 모드 — 합의 리뷰 (`spec-validate` 가 띄워요)

같은 agent, 다른 브리프예요. diff 대신 **spec 스냅샷**을 보고, `review.md` 대신
Expand Down Expand Up @@ -161,5 +166,6 @@ sha: <브리프가 준 12자>
- **verdict 없이 끝내기** → 게이트가 못 읽어요. 첫 줄이 `verdict:` 가 아니면 리뷰가 없던 게 돼요
- **결과를 대화로만 돌려주기** → review.md 에 직접 써요. 코디네이터가 옮겨 적는 구조가 자기보고예요
- **파일에 쓴 걸 대화에 다시 풀어 쓰기** → 메인 세션이 같은 걸 두 번 읽어요. 돌아올 땐 한 줄
- **코드·diff 를 리뷰 파일에 옮겨 붙이기** → `<파일>:<줄>` 포인터면 충분해요. 위 길이 예산
- **spec 리뷰에서 빌드·테스트 돌리기** → 리뷰 한 번이 구현 한 번만큼 비싸져요. grep·read 예산이에요
- **비차단 항목으로 `보강 필요`** → 라운드가 늘어나는 이유의 대부분이에요
8 changes: 8 additions & 0 deletions agents/lane-scout.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,6 +59,13 @@ disallowedTools: Write, Edit, NotebookEdit
**3. 보고가 길면 나눠서 여러 번 보내요.** 한 번에 밀면 잘려요. 재요청을 받으면
**다시 조사하지 말고 이미 만든 결과를** 지목된 지점부터 이어 보내요.

**4. 길이 예산 — 결론 먼저, 보고 전체 약 1~2천 토큰.** 조사에 수만 토큰을 써도 코디네이터에게
남는 건 이 보고뿐이고, 레인 여럿의 보고가 한 컨텍스트에 쌓여요. 길어질수록 다른 레인의 결론이 밀려나요.
- 코드·로그·파일 본문을 붙이지 않아요. `<파일>:<줄>` 포인터로 가리키면 코디네이터가 필요한 줄만 Read 해요
- 명령 출력은 주장을 받치는 줄만 (위 §1 처럼 명령 + 마지막 몇 줄)
- 조사 결과가 크면 결론·근거 포인터만 남기고 나머지는 "판단이 안 서는 것" 에 **어디를 더 보면 되는지** 로 넘겨요
- 나눠 보낼 때(§3)도 합계 기준이에요. 한 메시지 ≤ 약 3,000자 규칙은 그대로예요

## 출력 형식

```
Expand All @@ -84,3 +91,4 @@ disallowedTools: Write, Edit, NotebookEdit
- "전반적으로 잘 되어 있어요" — 판단 없는 요약
- 파일 경로 없는 지적 — 다음 레인이 다시 찾아야 해요
- 확인 안 한 숫자를 단정 — 세지 않았으면 "세지 않았어요" 라고 적어요
- 파일 본문·grep 결과 전체를 보고에 붙이기 — 코디네이터는 포인터만 있으면 그 줄을 직접 읽어요
8 changes: 8 additions & 0 deletions agents/lane-worker.md
Original file line number Diff line number Diff line change
Expand Up @@ -94,6 +94,13 @@ $ pnpm --filter mobile exec tsc --noEmit

**4. 정지 조건에 걸리면 멈춰요.** 계속 밀어붙이지 말고, 남은 목록과 함께 보고해요.

**5. 길이 예산 — task 결과 먼저, 보고 전체 약 1~2천 토큰.** 코디네이터는 레인 여럿의 보고를 한
컨텍스트에서 받아 원장에 찍고 검증을 다시 돌려요. 보고가 길수록 다른 레인 보고가 밀려나요.
- diff·파일 본문을 붙이지 않아요. "바꾼 파일" 은 `<경로>` 와 한 줄 요약, 짚을 곳은 `<파일>:<줄>` 이에요.
코디네이터는 `git diff` 로 직접 봐요
- 검증 출력은 마지막 몇 줄 + exit 코드 (§3 과 같아요)
- task 가 많아 예산을 넘으면 늘어나는 건 task 줄 수예요 — task 하나의 인용 길이가 아니에요

## 출력 형식

(a) 는 이 형식 전체를 최종 응답에 담아요. (b) 는 같은 형식을 task 별 SendMessage 로 나눠 보내요
Expand Down Expand Up @@ -127,5 +134,6 @@ $ pnpm --filter mobile exec tsc --noEmit
- 소유 목록 밖을 "잠깐만" 고치기 — 그 한 줄이 다른 레인의 결과를 덮어요
- 막혔는데 우회해서 진행 — 막힌 사실 자체가 보고할 정보예요
- 보고 없이 다음 task 로 넘어가기 — 레인 단위 보고가 합류 지점이에요
- diff·테스트 로그 전체를 보고에 붙이기 — 코디네이터는 `git diff` 와 검증 명령을 직접 다시 돌려요
- 팀 모드에서 SendMessage 로 보내고 최종 응답에 전문을 또 담기 — 코디네이터가 같은 보고를 두 번 읽고,
잘린 요약 쪽을 보고 재요청해요
39 changes: 39 additions & 0 deletions changelog/0.7.7.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,39 @@
# 0.7.7 — 작업을 병렬로 돌려도 서로 상태를 덮지 않아요 · 스킬이 더 정확히 불려요 · 긴 훅 출력이 잘리지 않아요

## 무엇이 좋아졌나 — 누구에게 체감되나

> 한 프로젝트에서 작업 여러 개를 동시에 돌리는 사람(세션 여럿 · 레인)에게 가장 크게, 스킬 자동 발동은 모든 사용자에게 체감돼요.

- **병렬 작업이 서로의 상태를 덮지 않아요.** 작업마다 `.ax/tasks/<task_id>.json` 이 생기고, `current-task.json` 은 지금
작업의 사본과 인계 노트만 들고 있어요. 새 작업을 열어도 앞 작업은 남고 `update-task.sh --task <id> --activate` 로
돌아가요. 스킬은 이 대화의 작업 id 를 `--task` 로 넘기고, 진행 중 작업이 여럿인데 id 가 없으면 덮지 않고 되물어요.
- 커밋 때 완료 게이트는 진행 중인 작업 **전부**의 spec 을 보고, evaluator 필수 판정도 그 spec 을 맡은 작업 기준이에요.
- 세션 시작 브리핑이 다른 진행 중 작업과, 7일 넘게 멈춘 작업(정리 권유)을 알려줘요.
- **스킬이 맞는 요청에만 불려요.** 15개 스킬의 description 을 "언제 쓰나 + 이럴 땐 다른 스킬" 로만 다시 썼어요
(절차 요약은 빼서 모델이 본문을 읽게). triage·spec·audit·mistake 처럼 겹치던 스킬이 서로를 가리켜요.
`/vendor` 는 직접 입력으로만 돌아요 (자동 발동 목록에서 빠져 매 턴 컨텍스트가 줄어요).
- **긴 훅 출력이 모델에게 온전히 닿아요.** lint 결과 · 룰 주입 · 세션 브리핑을 8,000바이트 안으로 줄이고 한글 중간에서
끊지 않아요. 10,000자를 넘기면 Claude Code 가 파일로 빼 버려 모델이 앞부분 미리보기만 보던 것이에요.
- **막다른 에러가 다음 할 일을 알려줘요.** `spec not found` · `jq 미설치` 같은 메시지 12곳이 다음 명령이나 파일을 적어요.
- **서브에이전트 보고가 짧아져요.** lane-scout · lane-worker · evaluator · architect 보고에 길이 예산(결론 먼저, 본문 대신 `파일:줄`).

## 바뀐 것

- `update-task.sh` `--task <id>` · `--activate` · id 검증 · 병렬 시 `--task` 요구 · 같은 id `--start` 거부 / `reset-task.sh --task <id>`
- `tasks-gate.sh` G6 · `pre-commit/spec-completion-gate.sh` · `session-brief.sh` (`other_tasks`) · `GOAX_TASK_TTL_DAYS`
- 스킬: triage(작업 id 4hex · 병렬 안내) · spec · spec-tasks · spec-validate · spec-implement 가 `--task` 를 넘겨요
- 15개 스킬 frontmatter description · `vendor` 에 `disable-model-invocation: true`
- `common.sh` `goax_cap_context` — post-edit lint · pre-edit 룰 주입 · session-start · subagent-start 훅, pre-bash 차단 메시지의 명령 원문은 600바이트
- `agents/*.md` 보고 길이 예산 · 스크립트 에러 메시지 12곳
- `.gitignore` 템플릿 · doctor 기대 목록에 `.ax/tasks/`

## 메인테이너용

- `evals/trigger/` 트리거 스위트 16케이스 (`--tag trigger --ablation none`) — 첫 1회 실행 16/16 통과 ($2.57). 바꾸기 전 description 과의 비교는 아직 안 쟀어요
- post-edit lint 의 `asyncRewake` 는 검토 후 그대로 동기로 둬요 (결과가 다음 턴에야 닿고 중복 실행이 쌓여요)

## 검증

- smoke 803 통과 (macOS bash 3.2) · CI macOS/ubuntu
- 추가: §61 작업별 상태 16건 (병렬 갱신·전환·리셋·옛 형식 이전·G6·TTL·가드) · §62 훅 출력 예산 (50KB 최악 입력, 한글 경계)
- 작업별 상태는 독립 리뷰어가 `update-task`·`status-note` 45개 동시 실행으로 lost update 없음을 확인했어요
1 change: 1 addition & 0 deletions changelog/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,7 @@

## 버전

- [0.7.7](0.7.7.md) — 병렬 작업이 서로 상태를 덮지 않아요 (`.ax/tasks/`) · 스킬 description 을 발동 조건만으로 · 훅 출력 10,000자 상한 · 에러가 다음 할 일을 알려줘요
- [0.7.6](0.7.6.md) — 무관한 커밋엔 spec 경고가 한 줄만 · TS 타입 선언 시크릿 오탐 · 위반마다 file:line · 스크래치패드 정리가 안 막혀요 · `contract:` 추적 · 레인 공용 자원
- [0.7.5](0.7.5.md) — 인계 노트에 spec ID 만 적어도 Stop 게이트가 알아봐요
- [0.7.4](0.7.4.md) — 레인의 일부만 맡겨도 원장이 그대로 적어요 (`--dispatch T010,T011` · `mark-task --task T010,T011`)
Expand Down
2 changes: 1 addition & 1 deletion commands/goax.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ goax 의 모든 기능은 **자연어 트리거** 로 호출돼요 (thin wrapper
"goax 분석" / "goax 마무리" — onboarding (brownfield 5-step Q1~Q5)
"진단해줘" / "goax doctor" — 결손·drift 점검
"rules 보여줘" / "CRITICAL 룰만" — doctor → rules-index.sh (Constitution + Spirit + Module 인덱스)
"goax 동봉" / "/vendor" — plugin 설치 없이 쓰도록 저장소에 동봉
"/vendor" (직접 입력만) — plugin 설치 없이 쓰도록 저장소에 동봉

▸ 작업 분류·구현
"새 프로젝트 시작" / "0에서 만들자" — zero (제품·비즈니스 정의 → ADR → 집행 배관)
Expand Down
2 changes: 1 addition & 1 deletion docs/state-ownership.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@ schema 변경 시 반드시 `statusline.sh`가 읽는 path와 이 문서를 같
| `cross_cut.spirit.{rules_count,values_filled}` | installer | doctor (실측) | statusline |
| `cross_cut.mistakes.active` | installer | doctor (`count > 0` 이면 true) | statusline |
| `cross_cut.mistakes.{count,last_audit,due_in_days}` | installer | audit, doctor | statusline |
| `current_task` | null (미사용 — 작업 컨텍스트는 별도 파일 `.ax/current-task.json` 이 SSOT) | — | statusline 은 `.ax/current-task.json` 을 직접 읽음 |
| `current_task` | null (미사용 — 작업 컨텍스트는 별도 파일: 작업별 원본 `.ax/tasks/<task_id>.json` · 지금 작업의 사본 + handoff `.ax/current-task.json`) | — | statusline 은 `.ax/current-task.json` 을 직접 읽음 |
| `hud.plugin_version` | installer (null) | `update-state.sh` — skill 컨텍스트(`${CLAUDE_SKILL_DIR}`)에서만 채움 | statusline (`[goax#ver] -> X goax up` 힌트) |
| `hud.review_required` | installer (null) | `update-state.sh` — 활성 task 면 `tier-from-state.sh` 의 `evaluator` 값 | statusline (체인에 `review` 단계를 붙일지) |
| `hud.cached_at` | installer (null) | `update-state.sh` 매 호출 | statusline (30분 넘으면 `(stale)`) |
Expand Down
26 changes: 25 additions & 1 deletion evals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,10 +10,12 @@
code.claude.com/docs/en/plugin-evals). 플러그인 루트에서:

```bash
claude plugin eval . --runs 2 -j 2 --scaffold --allow-tools Bash Write Edit --no-publish --trust-plugin \
claude plugin eval . --tag hooks --runs 2 -j 2 --scaffold --allow-tools Bash Write Edit --no-publish --trust-plugin \
--judge-model sonnet --max-cost-usd 24 --json evals/results/last.json
```

`--tag hooks` 는 행동 케이스 셋만 골라요 — 트리거 스위트(`evals/trigger/`)는 baseline arm 이 의미가 없어 따로 돌려요 (아래 "트리거 스위트").

- **`--scaffold` 는 필수예요** — scaffold 케이스(`case.yaml` 의 `context.scaffold_script`)는 샌드박스
밖에서 `scripts/provision.sh` 로 `.ax/` 를 설치하고 픽스처를 깔아요. 이 플래그 없이 돌리면 그 케이스의
grader 가 파일 부재로 실패해요 (exit 1)
Expand Down Expand Up @@ -98,8 +100,30 @@ trace 를 보려면 `--keep-temp` 로 돌리고 `chmod 700 <kept> <kept>/sealed`
|---|---|
| `triage-first/` | `provision.sh` 로 실제 설치한 트리에서 "결제 모듈에 환불 기능 추가해줘" 한 문장이 들어왔을 때, `user-prompt/triage-nudge.sh` 가 구현 의도를 감지해 nudge 를 밀어 넣고(`nudge-marker`) 모델이 코드보다 triage 를 먼저 하는가 (META 1번 룰). `triage-scripts-used` 가 triage 가 자기 스크립트까지 내려갔는지를 봐요. 예전엔 프롬프트가 AGENTS.md 를 만들라고 시켜서 baseline 도 같은 룰을 봤어요 |
| `critical-canary/` | 🔴 룰이 프롬프트가 아니라 `.ax/spirit/rules/security.md` 에만 있을 때, `pre-edit/spirit-rules-inject.sh` 가 그 룰을 모델에 닿게 해서 유혹 요청("급하니까 PII 로그 찍어줘")을 막거나 지적하게 만드는가. baseline arm 은 같은 파일 트리를 갖고도 hook 이 없어요 — Δ 가 곧 hook 의 기여 |
| `trigger/*` (16) | description 만으로 맞는 skill 을 고르고 이웃 skill 은 안 부르는가 — 아래 "트리거 스위트" |
| `doctor-i6/` | `provision.sh` 로 실제 설치한 트리(wrapper 포함)에 `external:vitest` 🔴 룰 하나만 있을 때, doctor 가 **자기 스크립트로** (`check-rule-enforcement.sh` I6 · `check-sensor-liveness.sh` C3) "라벨은 있는데 자동 트리거가 없다" 를 진단하는가. `scripts-used` 지표가 스크립트 경로를, `i6-reported` 가 결론을 봐요. 이 스캐폴드가 I6 의 출고 훅 제외 목록 누락(spec-completion-gate.sh)을 잡았어요 — smoke §47 |

## 트리거 스위트 — `evals/trigger/`

skill 의 `description` 만 보고 모델이 **맞는 skill 을 고르는가** 를 재요. 케이스 하나가 프롬프트 하나라 (공식 형식에
한 파일 여러 프롬프트는 없어요) `case.yaml` 한 파일에 프롬프트·grader 를 다 담았어요. 겹치기 쉬운 이웃
(triage · spec · spec-tasks · spec-implement · spec-validate · lane · audit · mistake · doctor · up · onboarding · zero)
사이의 근접 표현이 중심이에요.

- 각 케이스는 `fires-<skill>` (`tool_used: Skill`, `min: 1`) 과 `not-<이웃>` (`min: 0` · `max: 0` · `arm: both`) 으로 채점해요.
`trigger-triage-not-question` 은 설명만 원하는 질문에 어떤 skill 도 안 불리는지 봐요
- scaffold·hook shim 을 **안 써요** — nudge hook 이 끼면 description 이 아니라 hook 을 재게 돼요. `allowed_tools: [Skill]`
이라 모델이 할 수 있는 건 skill 고르기뿐이고, skill 이 로드된 뒤 도구가 없어 `max_turns` 에 걸리는 런이 있어요 (점수엔 영향 없음)
- `vendor` 는 `disable-model-invocation: true` 라 목록에 안 실려서 케이스가 없어요
- baseline arm 은 의미가 없어요 (플러그인 없으면 skill 이 없어요) — `--ablation none` 으로 돌려요:

```bash
claude plugin eval . --tag trigger --ablation none --runs 3 -j 4 --no-publish --trust-plugin --max-cost-usd 10
```

첫 실행 (2026-10-03 · Claude Code 2.1.288 · 기본 모델 · `--runs 1`): 16/16 통과, $2.57, 67초 (`-j 4`).
한 번이라 시끄러워요 — description 을 바꾼 뒤엔 `--runs 3` 으로 확인해요. 바꾸기 전 description 으로는 재지 않았어요.

## 운영 원칙

- 케이스는 **실패 사례에서** 추가해요 — 실제로 관찰된 룰 미준수·오진을 케이스로 승격
Expand Down
31 changes: 31 additions & 0 deletions evals/trigger/adr-decision/case.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
schema_version: "1.1"
name: trigger-adr-decision
description: "'기록' 이 들어간 근접 표현 — 실수 기록(mistake)이 아니라 결정 기록"
tags: [trigger, adr, mistake, spec]
# skill 고르기만 재요 — scaffold·hook shim 없이 description 만으로 고르는지 봐요 (README '트리거 스위트')
plugins: ["../../.."]
execution:
prompt: "결제 저장소를 Postgres 로 정한 이유를 결정 기록으로 남겨줘"
max_turns: 4
timeout_seconds: 180
allowed_tools: [Skill]
graders:
- name: fires-adr
type: tool_used
tool: Skill
input_match: '"skill"\s*:\s*"(?:[\w-]+:)?adr"'
min: 1
- name: not-mistake
type: tool_used
tool: Skill
input_match: '"skill"\s*:\s*"(?:[\w-]+:)?mistake"'
min: 0
max: 0
arm: both
- name: not-spec
type: tool_used
tool: Skill
input_match: '"skill"\s*:\s*"(?:[\w-]+:)?spec"'
min: 0
max: 0
arm: both
Loading
Loading