Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 31 additions & 0 deletions .github/workflows/scheduled-quality.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
name: Scheduled quality

# Runs the editorial quality gates weekly on the default branch, with no code
# change needed. `make check-waiver-expiry` fails while a quality waiver has 30
# days or fewer left, so an expiry is noticed here (GitHub emails the failure)
# a month before `make verify` would start failing every pull request.
on:
schedule:
- cron: '23 6 * * 1'
workflow_dispatch:

permissions:
contents: read

jobs:
quality:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: astral-sh/setup-uv@c18668ad3cf93ea998bef934396af7bb5c839dc7 # v10.2.0
with:
enable-cache: false
- uses: actions/setup-python@v7
with:
python-version: '3.13'
- name: Install locked dependencies
run: uv sync --locked --all-groups
- name: Quality gates
run: make quality-checks
- name: Fail 30 days before a quality waiver expires
run: make check-waiver-expiry
2 changes: 1 addition & 1 deletion CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -97,7 +97,7 @@ The single source of truth for the registries is `docs/quality-registries.toml`.

## Secrets and deploy configuration

Production deployment is manual through an authenticated Wrangler session with `make deploy`; public `workers.dev` and version preview URLs are disabled. Runtime Worker secrets (`TURNSTILE_SECRET_KEY`, `TURNSTILE_CLEARANCE_SECRET`, `PBE_SMOKE_BYPASS_SECRET`) are managed with `wrangler secret put`; see `docs/turnstile-runner-protection-spec.md`.
Production deployment is manual through an authenticated Wrangler session with `make deploy`, which smoke-tests the deployed origin afterwards and fails if the smoke fails; public `workers.dev` and version preview URLs are disabled. Runtime Worker secrets (`TURNSTILE_SECRET_KEY`, `TURNSTILE_CLEARANCE_SECRET`, `PBE_SMOKE_BYPASS_SECRET`) are managed with `wrangler secret put`; see `docs/turnstile-runner-protection-spec.md`.

Generated output is prevented from drifting before merge: install the local hooks with `scripts/install-git-hooks.sh`, and keep `main` protected so pull requests require the `verify` status check to pass against the current base. CI enforces the same `make check-generated` contract for contributors without local hooks.

Expand Down
21 changes: 19 additions & 2 deletions Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
PY := uv run --python 3.13
NODE_DEPS_STAMP := node_modules/.package-lock.json

.PHONY: check-node-version node-deps test embed-examples embed-editorial-registry build-search-index build check-generated fingerprint prototypes browser-layout-test search-ranking-test social-cards check-social-cards seo-cache-lint verify-examples check-registry-integrity check-confusable-pairs check-broad-surface-tours check-footgun-coverage check-notes-supported check-program-covers-cells check-prose-duplication check-inline-links score-example-criteria check-quality-scores check-no-figure-rationales check-journey-outcomes audit-example-graph quality-checks rubric-audit format-examples verify-python-version verify smoke-deployment dev deploy upgrade-runtime-deps lint
.PHONY: check-node-version node-deps test embed-examples embed-editorial-registry build-search-index build check-generated fingerprint prototypes browser-layout-test search-ranking-test social-cards check-social-cards seo-cache-lint verify-examples check-registry-integrity check-confusable-pairs check-broad-surface-tours check-footgun-coverage check-notes-supported check-program-covers-cells check-prose-duplication check-inline-links score-example-criteria check-quality-scores check-waiver-expiry check-no-figure-rationales check-journey-outcomes audit-example-graph quality-checks rubric-audit format-examples verify-python-version verify smoke-deployment post-deploy-smoke dev deploy upgrade-runtime-deps lint

check-node-version:
@major="$$(node -p 'process.versions.node.split(".")[0]')"; \
Expand Down Expand Up @@ -92,6 +92,11 @@ score-example-criteria:
check-quality-scores:
$(PY) scripts/check_quality_scores.py

# Scheduled (weekly) gate: fail while a quality waiver has 30 days or fewer
# left, so the expiry surfaces before it turns every pull request red.
check-waiver-expiry:
$(PY) scripts/check_quality_scores.py --fail-within-days 30

check-no-figure-rationales:
$(PY) scripts/check_no_figure_rationales.py

Expand Down Expand Up @@ -121,12 +126,24 @@ dev: node-deps
uv run --group workers pywrangler dev --port 9696

smoke-deployment:
$(PY) scripts/smoke_deployment.py $(URL)
$(PY) scripts/smoke_deployment.py $(URL) $(SMOKE_ARGS)

# Origin that `make deploy` smoke-tests once Wrangler has deployed. With
# Turnstile challenges enabled, export PBE_SMOKE_BYPASS_SECRET so the POST
# checks can run; SMOKE_ARGS passes extra flags (e.g. --skip-post) explicitly.
DEPLOY_URL ?= https://www.pythonbyexample.dev

post-deploy-smoke:
@$(MAKE) --no-print-directory smoke-deployment URL=$(DEPLOY_URL) || { \
echo "Deployment smoke FAILED for $(DEPLOY_URL). The new version is already live: investigate now, or roll back with 'uv run --group workers pywrangler rollback'." >&2; \
exit 1; \
}

deploy: node-deps check-generated
uv run --group workers pywrangler sync --force
git diff --exit-code -- pylock.toml
uv run --group workers pywrangler deploy
@$(MAKE) --no-print-directory post-deploy-smoke

# Production vendors pylock.toml; tests run against uv.lock. Refresh both together.
upgrade-runtime-deps: node-deps
Expand Down
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -125,7 +125,7 @@ git diff --check
- Worker Cache API keys include the HTML version
- prototype layout pages are not cached

`make browser-layout-test` launches headless Chrome and checks the rendered Shiki code-block layout so generated line markup does not create visual blank rows.
`make browser-layout-test` launches headless Chrome against the local Worker. It checks the rendered Shiki code-block layout so generated line markup does not create visual blank rows, drives the runner, sharing, copy, search and keyboard navigation through the page, and measures computed styles (pressed states, touch targets, contrast in both themes, the reader's font size, reduced-transparency and more-contrast fallbacks) in place of unit tests that matched CSS text.

## Asset fingerprinting and cache busting

Expand Down Expand Up @@ -174,7 +174,7 @@ scripts/format_examples.py --check
make deploy
```

`make deploy` first runs `make check-generated`, which rebuilds and rejects any generated output not committed to the branch. It then syncs the ignored Python Workers dependency bundle from the committed `pylock.toml`, refusing to deploy if the sync would change that lock, before Wrangler deploys.
`make deploy` first runs `make check-generated`, which rebuilds and rejects any generated output not committed to the branch. It then syncs the ignored Python Workers dependency bundle from the committed `pylock.toml`, refusing to deploy if the sync would change that lock, before Wrangler deploys. Finally it runs `scripts/smoke_deployment.py` against `DEPLOY_URL` (default `https://www.pythonbyexample.dev`) and exits non-zero if the deployed Worker fails any GET or POST check. Export `PBE_SMOKE_BYPASS_SECRET` when Turnstile challenges are enabled so the POST checks can run; `SMOKE_ARGS` passes extra flags such as `--skip-post` explicitly.

## Updating dependencies

Expand Down
5 changes: 3 additions & 2 deletions docs/lessons-learned.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,7 +40,7 @@ This document records project lessons that should guide future changes to Python
- Keep the complete editable program visible because it is the thing that actually runs.
- When storing examples as Markdown, keep the full editable program in `:::program` and teaching fragments in separate `:::cell` blocks. Do not concatenate cells to recreate the editor source.
- Fine-grained cells can restate definitions to stay executable. This is better than collapsing class, property, recursion, match, or type-hint examples into one large cell.
- Preserve a frozen golden catalog while migrating source formats. Full stdout parity is not enough; rendered teaching-cell structure must also match. (After the migration, the fixture's job changed: it is now a reviewed structural snapshot, refreshed explicitly by `scripts/refresh_golden_fixture.py`, that catches loader/parser regressions rather than guarding a migration.)
- Preserve a frozen golden catalog while migrating source formats. Full stdout parity is not enough; rendered teaching-cell structure must also match. (After the migration the frozen fixture and its refresh script were retired in 0929c25, 2026-06-13; loader and parser regressions are now caught by the frontmatter property tests in `tests/test_parser_properties.py` and the byte-exact output check in `scripts/verify_examples.py`.)
- Beware broad-surface titles that quietly teach only one narrow slice. Pages named `Testing`, `Packages`, `Regular Expressions`, `Type Hints`, `Async Await`, or `Special Methods` must either cover the forms a reader reasonably expects or explicitly frame themselves as a first pass and link to focused neighbors.
- Do not let an important syntax form live only in a separate page if another page title strongly implies it. An umbrella `Operators` page should at least point to and lightly show `:=`, even though `Assignment Expressions` remains the focused lesson.
- Journey order should follow prerequisite thinking, not catalog order. Put booleans before truthiness and conditionals, scope before closures, bytes before networking, and environment boundaries before subprocess/thread boundaries.
Expand All @@ -64,6 +64,7 @@ This document records project lessons that should guide future changes to Python
- Avoid layout shifts after execution. Reserve space for metadata such as execution time before a run occurs.
- When two columns use the orange rail, source and output need identical rail spacing. Differences in border/padding make the page feel broken even if the content is correct.
- Use browser screenshot tests for visual bugs. Static HTML/CSS assertions are useful but can miss actual rendered layout behavior.
- Assert what the browser computes, not what `site.css` says. Unit tests that matched literal CSS or JS text broke on harmless refactors and passed on real regressions, so `scripts/check_browser_layout.mjs` now measures the same intentions: computed transforms under a forced `:active`, touch-target heights, contrast in both themes, the reader's font size, emulated `prefers-reduced-transparency`/`prefers-contrast`, and runner behaviour driven through the page. To add a visual rule, add a measurement there that fails when the rule is removed.

## Testing and verification

Expand Down Expand Up @@ -108,7 +109,7 @@ git diff --check
- **Score what's shipping, not what was designed.** A scoring dict on the gestalt is design-time review. Production figures live in `src/marginalia.py` `FIGURES` and may have been redesigned during promotion. Scoring should track the production version with the gestalt as separate history.
- **Semantic diagram review has four objects: page, cell, paint function, caption.** Geometry contracts can prove that a figure renders, but not that it still teaches the adjacent cell. On every diagram pass, read the current example cell, the attached paint function, the caption, and the score rationale together. The same audit caught `container-protocols` still showing `iter()/next()` after the page had shifted to `__setitem__`/`__contains__`/`__getitem__`, `structured-data-shapes` over-focusing on `TypedDict` while anchored to a dataclass cell, and `object-lifecycle` mentioning `__del__` after the lesson had been reframed around references.
- **Some examples should never have figures.** Constraint-shaped, infrastructure-shaped, and aggregator-shaped slugs lack a single mechanism to depict. Force-fitting figures on them scores below the gate. Leave them figure-less and document why rather than ship weak figures.
- **Audits without contracts rot; bug classes need automated gates.** When you find a bug class — clipping, collision, off-palette colour, drifting twin coordinates, duplicate caption — write a unit test that asserts the invariant across every figure. The geometry contracts in `tests/test_marginalia_geometry.py` started as ad-hoc scripts and were promoted to CI gates after the same bug class recurred. 54 tests today cover 9 contract families; new figures pass them automatically because each test iterates `FIGURES`.
- **Audits without contracts rot; bug classes need automated gates.** When you find a bug class — clipping, collision, off-palette colour, drifting twin coordinates, duplicate caption — write a unit test that asserts the invariant across every figure. The geometry contracts in `tests/test_marginalia_geometry.py` started as ad-hoc scripts and were promoted to CI gates after the same bug class recurred. Each contract family is a test class there; new figures pass them automatically because each test iterates `FIGURES`. (Count the classes rather than quoting a number here: an earlier "54 tests, 9 families" claim drifted from the file.)
- **Clipping ≠ collision; padding the viewBox fixes one but not the other.** The `value-types` bug had two components: the first `INT` tag was clipped above the viewBox (geometry escapes its frame), and the `STR`/`LIST`/`DICT` tags overlapped the boxes above them (geometry collides internally). Padding the viewBox solved (1) and disguised (2). Element-element collision needs its own audit that walks every text-rect, text-text, and (for label-on-edge cases) text-line bounding-box pair.
- **Heuristic audits over-flag; trust the design, not the regex.** Probes for "prose duplication" (SVG text matching caption substring), "text crossing a line" (label bbox bisected by a hairline), and "text overlapping a circle" each surfaced ~3-10 hits — all false positives. Diagrammatic labels naturally appear in captions (`__getattr__`, `yield from inner`); dashed strikes through `.append` are deliberate; text inside node circles is the design. Don't promote heuristic findings to contracts without confirming each is a real bug.
- **Structural twins must share coordinates exactly.** When two figures depict parallel concepts — `kw-only-separator` and `positional-only-separator`, `class-triangle` and `metaclass-triangle` — they read as a pair. A single-pixel drift in one breaks the visual rhyme. Treat a coordinate change in one as a forced change in both, in the same commit. The audit caught `kw-only-separator` at the old `x=82` after `positional-only-separator` had moved to the corrected `x=75`.
Expand Down
Loading
Loading