Skip to content

Merge flaskfrontback: adds Eyal's 4-subject corpus + prior session's feature/corpus work - #30

Merged
femtechtk merged 30 commits into
mainfrom
flaskfrontback
Sep 4, 2026
Merged

femtechtk merged 30 commits into
mainfrom
flaskfrontback

Conversation

@femtechtk

Copy link
Copy Markdown
Collaborator

Summary

This branch is 30 commits ahead of main and has diverged significantly in both game features and corpus content. Highlights relevant to review:

  • New this PR: adds geography/history/literature/computer_science (2,400 questions, 200 topics x 4 subjects) to the live corpus (cf-pages/functions/_lib/data.json). This content went through multi-round independent Gemma re-verification after an initial 47% defect rate was found in what Gemini had scored as clean (dominant issue: hint-tier inversion). Final state: 0 flagged topics, 91.3% of individual hint scores high-quality.
  • Pre-existing drift this PR also surfaces: main's corpus for the original 4 subjects (biology/math/chemistry/physics) is only 200 questions each (800 total) — much smaller than what's already on this branch (766-798 each, 3,152 total) before the new subjects are even added. That expansion happened across this branch's history and never made it back to main. Reviewers should be aware this PR brings substantially more than just the 4 new subjects — merging it means adopting this branch's corpus and feature state as the new baseline.
  • Also includes prior feature work already on this branch (Google Sign-In fix, per-tier question timers, upgrade shop, answer-position-bias shuffle fix, holographic cards, etc. -- see individual commits).

Content verification detail (new subjects)

  • backend/validate_corpus.py passes: 5,552 questions across 8 realms
  • backend/data.json mirror regenerated via backend/sync_corpus.py to keep CI's corpus-drift check green
  • New subjects ship with plain educational feedback text (matching what was verified all session) rather than the other subjects' combat-flavor voice, to avoid layering new unaudited content on top of what was just verified — a follow-up pass could add flavor text later if desired

Test plan

  • Confirm CI passes (corpus-validation, corpus-drift, js-syntax, readme-links)
  • Manually verify the 4 new subjects appear in the syllabus selector and are playable end-to-end
  • Confirm hint lookup resolves correctly for the new subjects (via get-hint.js's exact-text match against data.json)
  • Review whether main's outdated corpus/feature baseline is intentional or should be reconciled before merge

🤖 Generated with Claude Code

Talia Kohen and others added 30 commits August 25, 2026 09:41
…y/math/physics)

Hand-written tiered question+hint content: 200 topics per subject, each
with easy/medium/hard question variants and matching 3-tier hints.
Includes the Gemini-as-judge script (judge_tiered_content.py) and its
in-progress results (512/800 judged so far, avg 9.98/10).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Issue #18: the single-select answer path (submitQuizAnswer, the most
common one) never wrote anything to the battle log; the multi-select
path did. Both now populate it identically.

Issue #19 (partial): replaced the 6 in-combat alert() calls (Attack/
Ability/Recharge failures and network errors) with a new shared
addBattleLogEntry() helper that renders each turn as a comic-caption-
style panel instead of a bare text line or a jarring native dialog.
Tone/burst badges are inferred from the backend's existing narrated
message text -- no combat numbers or balance changed. Battle log now
has role="log" aria-live="polite" for screen readers, and the entry
animation respects prefers-reduced-motion.

The remaining 6 alert() calls (session-expired, initial deploy-into-
combat failure, reset-game failure) fire when no battle screen is
showing, so they're left as native alerts pending a separate fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Final results: 800/800 topics judged, global average 9.98/10 (biology
9.95, chemistry 10.00, math 9.99, physics 10.00). Only 2 topics scored
below 9: a real hint/question mismatch bug in claude_tiered_batch4_
biology.json (7.8) and a cosmetic wordiness issue in
claude_tiered_batch100_biology.json (8.7). The version of this file
committed earlier was mid-run (~512/800); this replaces it with the
finished evaluation.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds 2,352 Gemini-judged questions (9.98/10 avg) across biology,
chemistry, math, and physics -- each of the 3 difficulty-tiered
questions per topic becomes its own entry in the syllabus's existing
question list, tagged with `difficulty`, carrying its own 3-tier
hints object in the exact pre-existing hints:{hard,medium,easy}
schema. No changes needed to combat-action.js, get-hint.js, or the
frontend -- this reuses the existing random-question-draw mechanic
and hint system exactly, just with a much larger, higher-quality pool.
48 of the theoretical 2,400 were skipped as exact-text duplicates of
already-present questions (backend/merge_tiered_into_live_data.py is
idempotent by question text, safe to re-run).

Verified end-to-end against a local wrangler pages dev instance:
start-combat, attack (drew both pre-existing and newly-merged
questions), answer submission, and get-hint all work correctly on
the merged pool.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Difficulty is picked once when entering a realm (replacing the single
"Initialize Sync" button with three difficulty buttons on each syllabus
card) and locked in for the whole battle -- not a per-turn toggle.

Backend: start-combat.js now filters the question pool by the chosen
difficulty before building question_order (questions with no difficulty
tag -- the original, pre-tiered content -- are treated as "medium" so
they stay reachable). combat-action.js surfaces the active difficulty
in every combat_state response so the frontend can display it.
Falls back to the full unfiltered pool if a subject/difficulty
combination ever has zero matches, so a player is never left with no
questions to answer.

Frontend: the selected difficulty now shows in the persistent nav bar
("Biology · Hard") for the rest of the battle.

Verified end-to-end against a local wrangler pages dev instance:
selecting "hard" only ever drew questions tagged difficulty:"hard" in
the live data, and "easy" only drew difficulty:"easy" ones -- confirmed
by cross-referencing each drawn question's text against the merged
data.json rather than just eyeballing question style.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
… buttons

The previous UI had two buttons (Simple Hint, Deep Hint), where Deep
Hint cascaded through all three tiers one at a time via a "Need More
Help?" button. Replaced with three independently-clickable buttons,
one per tier, matching the difficulty-picker's color convention
(green/yellow/red).

Backend budget semantics unchanged: tier="simple" (easy only, cheap
budget) backs the Easy button; tier="hard" (all three tiers in one
call, pricier budget) backs both Medium and Hard, which now share a
single cached fetch promise so clicking both only spends the "deep
hint" budget once, not twice.

Verified locally with Playwright: clicking tiers out of order (Medium
first, then Hard, then Easy) correctly reveals each one independently
and displays all three in a consistent Easy/Medium/Hard order
regardless of click order.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…he UI

These two words had been used loosely/interchangeably in conversation
and in the UI itself, which risked confusing players about what each
control actually does. Now explicit:

- Syllabus card caption: "Level = how advanced this realm's questions
  are (Easy: intro-level, Medium: standard coursework, Hard: advanced/
  exam-level)" -- calibrated to the actual criteria used when writing
  the tiered corpus (easy = recall/definitions, medium = explaining a
  mechanism/relationship, hard = multi-step reasoning/synthesis).
- Hint box caption: "Hint tiers for THIS question (not the realm's
  difficulty level) -- pick how much help you want" -- directly
  distinguishes hint tiers from the syllabus-level difficulty choice.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Easy = Gifted Leo Baeck (high school), Medium = Technion (undergrad),
Hard = Technion (graduate) -- per explicit direction. Note: the actual
question content at "Hard" tops out around AP-high-school/early-college
rigor, well below genuine Technion coursework difficulty at either
level; this labeling is intentional flavor/branding, not a rigor claim
verified against the content.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…hool

Replaces the previous Leo Baeck / Technion labeling with the user's
own personal educational background (Ramaz for high school, Cornell
for college) plus a generic "graduate school" for Hard.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…ption

Ramaz/Cornell/graduate-school labeling is kept as an internal code
comment for reference, but the visible caption reverts to generic
"Choose a difficulty" copy -- players shouldn't see the specific
institution names.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Parallel to judge_tiered_content.py (Gemini), reusing the identical
prompt/schema/rubric so the two judges' outputs are directly
comparable. Separate output file, fully resumable. Includes
progressive backoff (up to 90s) and a higher consecutive-failure
threshold (10) to ride through Gemma's intermittent transient
500/503 errors without needing a manual resume every time.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
#27)

- Deleted the dead showFeedbackModal() stub (unreachable -- the real
  definition later in the same file wins at parse time, and the stub's
  only body was alert(), one of the flagged alert() calls).
- Removed pushRecentMessages() and its 3 call sites -- it wrote to
  #recent-results-list, which doesn't exist anywhere in cf-pages/public/.
  Deleted per the issue's stated default (the battle log already covers
  this need), along with its orphaned CSS (.neural-recent,
  .neural-recent-title, .recent-list-entry).
- Removed all [DEBUG]-tagged console.log/warn calls. Verified first
  that the "full question object" logged by openQuizModal is NOT
  answer-revealing -- combat-action.js already strips isCorrect/
  feedback before sending a pending question to the client -- but
  removed the noise anyway per the issue's tidiness ask.
- Removed two stale markers: the "VERSION label removed" HTML comment
  and an already-gone "...existing code..." placeholder (removed in an
  earlier edit this session).

Verified end-to-end via Playwright against a local wrangler dev
instance: menu -> realm select -> combat -> single-select answer ->
multi-select answer -> recharge -> reset via nav Home, zero JS page
errors throughout, battle log correctly populated.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…uping (#18)

The core bug (single-select never wrote to the log) was fixed earlier
this session; this closes the rest of the issue's acceptance criteria
that a fuller read of the issue text revealed weren't yet met:

- Added a visible "Battle Log" heading above the panel.
- Added a CSS-only empty-state placeholder ("Your actions will appear
  here.") via :empty::before -- shows automatically when the log has
  no entries, disappears on the first real one, can't drift out of
  sync with actual content since there's no JS state to maintain.
- Turn-grouping: addBattleLogEntry() is now addBattleLogTurn(), taking
  a whole action's message array and rendering it as ONE grouped block
  (your result, then the enemy's response) instead of separate floating
  lines. All 3 call sites updated; addBattleLogEntry() now delegates to
  it for single free-standing messages (errors etc).
- Non-color differentiation: each line in a turn gets a text "You:" /
  "Enemy:" speaker label (position-based, matching combat-action.js's
  consistent message ordering) in addition to the existing border-color
  coding, so message kind no longer depends on color alone.

Verified via Playwright: heading and placeholder render correctly;
25 synthetic turns produce a 2678px-tall scrollable log inside its
142px container with zero page-height growth and zero horizontal
overflow; a full combat run (single-select + multi-select) renders
grouped turns with correct speaker labels and zero JS errors.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The dead showFeedbackModal() stub and 6 in-combat alerts were already
handled (#27, earlier this session). This closes the rest:

- Discovered and fixed a real latent bug while working on this:
  setStatus() has been targeting a #test-output element that doesn't
  exist anywhere in index.html, so every setStatus() call in the
  codebase was silently a no-op. Added the missing element (role=
  "status" aria-live="polite") so it actually works. This mattered
  here specifically because handleInvalidSession()'s existing "UI
  path" the issue assumed was already showing something turned out
  to be silently broken -- removing its alert() without this fix
  would have made session-expiry produce zero visible feedback.
- Removed the redundant session-expired alert() -- the function ends
  with setStatus() showing the same message, now that setStatus
  actually works. Also fixed it to use the caller's specific message
  instead of a hardcoded string that was silently discarding it.
- Fixed a pre-existing setStatus('Error: ' + e.message) call (start-
  game's catch block) that raised the same "never show e.message to a
  player" problem this issue is about, even though it wasn't one of
  the originally-counted 12 alerts.
- Converted the remaining 4 alerts (start-combat failure/exception,
  reset-game failure/exception) to the existing feedback modal
  infrastructure. Added an optional `title` field to showFeedbackModal
  so these can show "Error" instead of the quiz-specific "Correct!/
  Incorrect" title that would otherwise be misleading for a system
  error. Raw exception text stays console-only; players see plain-
  language messages.
- "Not enough CAP" reclassified as guidance, not just an error string:
  both backend message sites now explicitly name Recharge as the next
  move ("Not enough CAP -- Recharge to continue.").

Verified via Playwright: draining CAP naturally triggers the new
guidance text in the battle log (not a native dialog), zero alert()
calls remain in cf-pages/public/static/js/ (grep-confirmed), zero JS
page errors.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Moves session-status/reconstruction docs and benchmark/diagnostic
reports out of the repo root into docs/history/ and docs/reports/ so
they read as historical record instead of live top-level state.
MVP_TESTING_PLAN.md is included -- it describes a 2026-07-22 session's
outstanding test gaps against an older build and is stale the same way.

Removes confirmed-dead files: next-env.d.ts (no Next.js/root
package.json in this repo), and three unused frontend/templates/
variants (only 'index.html' is ever passed to render_template() in
app.py). Moves the unrelated standalone rag_quiz.py prototype into
backend/archive/.

Adds backend/README.md indexing the ~75-script pile by purpose
(content generation, judging/QA, bake-offs, corpus maintenance,
extraction, throwaway Playwright verification scripts) rather than
physically relocating them into subfolders -- many assume backend/ as
their working directory for relative paths to data.json and their own
intermediate output files, and moving them without auditing each one
risked breaking that. Updates the root README's repo layout section to
match (docs/history, docs/reports, backend/archive, backend/README.md
pointer).

Verified backend/app.py still imports cleanly after the root-level
moves.
Found while building the corpus validator for issue #26: two questions
were typed multiple_choice_single but had 2 options genuinely correct
(both 6 and 8 are factors of 24; both "elements in both sets" and "the
intersection of the sets" describe a Venn overlap) -- retyped both to
multiple_choice_multiple rather than arbitrarily picking one answer to
mark wrong. The third (line through (2,3) and (4,5)) had 3 garbled,
apparently glued-together distractor options, with the actual correct
answer "y = x + 1" marked false and two nonsense options marked true --
removed the corrupted options and fixed the isCorrect flag on the
correct one, keeping the 3 legitimate distractors.

This is the same bug class flagged earlier during the 800-topic corpus
work (a wrong isCorrect flag found in the corpus); a full audit is
still outstanding, but these 3 needed fixing now so the new corpus
validator (backend/validate_corpus.py) has a clean baseline to check
future changes against instead of failing on day one.
Issue templates (.github/ISSUE_TEMPLATE/): bug.yml, feature.yml, and
content.yml (for corpus-specific reports -- wrong answer, hint leaks
the answer, broken formula), all following the Problem/Proposal/
Acceptance-criteria structure the existing issues already use. Blank
issues stay enabled via config.yml.

PR template (.github/pull_request_template.md): what/why/how-verified
plus an explicit "redeployed to Cloudflare Pages: yes/no" checkbox --
deploys are manual (no Git integration on the Cloudflare side), so a
merged PR is not a shipped PR and that distinction was previously
invisible.

CI (.github/workflows/ci.yml), runs on PRs and pushes to main:
- Corpus validation (backend/validate_corpus.py): schema check plus the
  invariants the game logic assumes -- exactly one correct option per
  multiple_choice_single question, at least one for _multiple, all 3
  hint tiers present and non-empty. This is a new script, not the
  drift-check mentioned in the issue text -- see below.
- JS syntax check (cf-pages/check-js-syntax.sh): node --check across
  every classic script and ES module in cf-pages/, since the frontend
  is unbundled with no build step and a syntax error would otherwise
  only surface live in the browser.
- Python lint (ruff over backend/), non-blocking (--exit-zero) given
  ~75 inherited scripts never linted before.
- README link check (check-readme-links.py): relative links and
  in-page anchors only, no external URLs (there are none in the README
  currently, and checking those in CI is flaky).

Deliberately NOT included: a corpus-drift check comparing
backend/data.json against cf-pages/functions/_lib/data.json. The two
have already diverged and which one is authoritative is an open
question tracked in #24 -- adding the check now would just fail every
PR. Wired up once #24 is resolved.

Verified corpus-validation and JS-syntax-check both fail correctly
against deliberately broken input (a duplicated isCorrect flag; a
truncated JS file), per the issue's acceptance criteria.

README documents the new CI table and reiterates that CI does not
replace the manual `wrangler pages deploy` step.
After 10s of no pointer/keyboard/focus input inside #combat-screen (or
right after any completed action, via updateCombatHUD), pulse the
button the player is expected to press next -- Attack when CAP >= 3
(covers both the "cheapest progress" and "only affordable option"
cases from the issue), Recharge otherwise, since Attack/Ability are
disabled below 3 CAP and nudging a disabled button would read as a
broken screen. No nudge while a modal is open -- the player's
attention belongs to the question there.

CSS owns the animation (nudge-pulse keyframes on .neural-action-btn.nudge,
animating box-shadow/transform off currentColor so each button keeps its
own accent color) gated under prefers-reduced-motion: no-preference; the
reduce branch substitutes a static outline instead of removing the cue.
A visually-hidden #idle-nudge-announcer (role="status", aria-live="polite")
announces "Suggested next action: <label>" once per idle period, not once
per pulse cycle, and the nudge never moves focus or triggers an action.

Verified with two Playwright scripts against local wrangler pages dev:
timer fires at 11s idle and targets Attack; any keydown clears it and
restarts the timer; no nudge appears while the quiz modal is active even
past 11s; animation-name resolves to 'none' under emulated
prefers-reduced-motion: reduce; draining CAP below 3 moves the nudge to
Recharge and never lands on a disabled button; no horizontal overflow at
390x844 while the pulse animates. (Playwright's click-stability check
needed force:true against the pulsing button, since Playwright treats a
continuously-transformed element as "not stable" for actionability
purposes -- expected given the animation, not a functional issue for
real pointer input.)
Each combat button now shows a smaller second line stating its cost
and effect ("3 CAP · 15 dmg", "free · +5 CAP", "5 CAP · 25 dmg")
instead of communicating affordability through opacity alone.
Recharge (CAP) is renamed to plain Recharge now that its cost line
explains the mechanic.

Cost/damage/gain values are no longer duplicated: cf-pages/functions/
_lib/game.js now exports ACTIONS and actionCosts() as the single
source, combat-action.js imports ACTIONS instead of redefining it
locally, and every endpoint that returns combat state (start-combat,
combat-action, auth-resume) includes action_costs in it. The client
caches this onto window.actionCosts inside updateCombatHUD() and reads
it for both the affordability checks (previously hardcoded `cap >= 3`
/ `cap >= 5`) and the idle nudge's target selection -- a hardcoded
DEFAULT_ACTION_COSTS constant is kept only as a pre-first-response
fallback, mirrored to match ACTIONS server-side.

Accessibility: the cost line lives in an aria-hidden span so it never
runs into the accessible name; a separate visually-hidden span per
button (referenced via aria-describedby) gives screen readers "Attack"
as the name and "Costs 3 CAP, deals 15 damage" as the description.
Verified via Playwright's aria_snapshot() that button names resolve to
just "Attack" / "Recharge" / "Use Ability".

Disabled state now states a reason instead of just dimming: a status
line under the action row reads "Not enough CAP -- Recharge to
continue." whenever CAP drops below Attack's cost, live-announced via
aria-live="polite". A permanent caption below it discloses that a
wrong answer still costs CAP, since combat-action.js deducts the cost
before checking correctness and players would otherwise find out by
feeling cheated.

Buttons switched to a flex-column layout to fit the second line;
verified via Playwright at 390px width that the row still fits with no
horizontal overflow and the Attack button's tap target stays above
44px (measured ~114x50).

"CAP" itself is still unexpanded in plain language on this screen --
that's issue #17 (How to play panel), landing next.
Expanded the two invented-abbreviation labels in place: "CAP" becomes
"CAP (actions)", and "HP" gains a "Your"/"Enemy" prefix on the player
and enemy cards respectively so the two bars are distinguishable
without relying on colour or card position (they were previously both
just "HP:").

Added a "How to Play" button to the persistent combat nav (one click,
always reachable) that opens a modal explaining HP, CAP, and Resolve
in plain language. Content is rendered from a single array,
COMBAT_STAT_EXPLANATIONS in game-simple.js, so this and the issue #16
tutorial (landing next) read from one source instead of maintaining
two copies of the same rules text that could drift.

Resolve is now legible while it moves: a delta annotation ("-18" /
"+14") appears next to the bar after each turn, plus a text state
label at the extremes matching combat-action.js's own thresholds --
"Rattled" at <=15 (the hesitate-chance threshold) and "Emboldened" at
>=85 (the heavy-hit-chance threshold). The delta's fade-in is gated
under prefers-reduced-motion: no-preference; under reduce it just
appears/clears with no transition. Found and fixed a real bug while
wiring this up: submitQuizAnswer's existing (pre-existing, unrelated
to this issue) defensive re-render of updateCombatHUD() in its
`finally` block was immediately overwriting the delta with a blank
"no change" result, since it ran a second time against the
already-applied state -- fixed by making updateResolveAnnotation a
no-op when called again with a resolve value matching the last one
actually rendered.

All 4 HUD bars now expose role="progressbar" with accurate
aria-valuenow/valuemax/aria-label (factored through one new
setProgressBar() helper instead of repeating the width-and-nothing-else
update inline per bar), so a screen reader reports e.g. "Enemy
Resolve, 50" instead of an unlabelled div.

Verified via Playwright at 390px: no horizontal overflow, all 4
progressbar roles/values correct, How-to-Play panel opens on the first
click and closes correctly, and the resolve delta/state-label
computation produces "-18"/"+14" after a real turn and "Rattled"/
"Emboldened" at the extremes.
A five-step tutorial runs once on a player's first-ever combat screen,
gated on a localStorage flag (studysaga_tutorial_seen, read/write
wrapped in try/catch like holo-card.js's existing localStorage use --
private browsing means it may reappear next session, which the issue
explicitly allows rather than throwing). Each step dims the screen via
a spotlight cutout (an oversized box-shadow on the highlight box, so
the target itself stays undimmed without a separate masked overlay),
highlights one real element -- player card, CAP bar, Attack, Recharge,
enemy Resolve bar -- and states the mechanic in plain language, closing
with "Your turn -- press Attack" pointing at Attack again rather than
highlighting nothing.

No sandbox battle, no fake state: the overlay sits visually on top of
the real buttons and captures every click itself, so nothing
underneath is ever actually pressed while a step is showing --
verified CAP is unchanged after clicking through all six steps.

Reuses the existing trapFocus()/deactivateModal() modal infrastructure
(focus trap, Tab wrapping, Escape handling, the transitionend-based
focus-first workaround) instead of a second implementation. Found and
fixed a real bug while wiring this up: trapFocus() auto-focuses the
first focusable element in DOM order, which was the Skip button --
meaning Enter would skip the tutorial instead of advancing it. Fixed
by reordering the buttons so Next is first in the DOM (Skip keeps its
original visual position via CSS `order`).

Step content lives in one array, TUTORIAL_STEPS, and the highlight
position is derived from each target's live getBoundingClientRect() on
render and on resize -- no hardcoded coordinates or tooltip side, so
mobile's stacked card layout doesn't leave a tooltip pointing at empty
space. The delta-fade on the highlight box is skipped under
prefers-reduced-motion: reduce.

Reachable on demand via a new "Replay Tutorial" button inside the
issue #17 How-to-Play panel (closes that modal and restarts the
tutorial) -- a tutorial a player can't reopen becomes a support
question. Closing the tutorial (finish, skip, or replay-then-skip)
re-arms the idle nudge from that moment rather than leaving it primed
against whatever time had already elapsed while the tutorial was
showing; pickIdleNudgeTarget() also now suppresses the nudge entirely
while the tutorial is active, matching how it already suppresses
during quiz/feedback modals.

Verified via Playwright: auto-appears on a fresh context's first
combat entry, does not reappear on a second run in the same context,
Skip closes it from step 1, Escape closes it from any step (once the
focus trap has actually engaged -- trapFocus's own 450ms
focus-establishment fallback, not a tutorial-specific issue), Replay
Tutorial from the How-to-Play panel reopens it, CAP is unchanged
throughout, and it still runs and is still skippable with
localStorage's getter made to throw (no uncaught JS errors). Tooltip
positioning verified in-bounds at 390x844, 768x1024, and 1920x1080 for
all six steps.
…#20)

player.score has existed in the session model since the port to Pages
Functions but nothing ever incremented it. Adds scoreForAnswer(),
VICTORY_BONUS, and hpRemainingBonus() to cf-pages/functions/_lib/game.js
as the single scoring implementation, called from the one place
combat-action.js already shares for grading both single- and
multi-select answers (they'd already drifted once, per the pre-#11
battle-log bug, so this had to not be a second copy):

- Correct answer: +100 base
- Streak bonus: +25 per consecutive correct beyond the first, capped
  at +100 (a 5+ streak is the ceiling)
- Wrong answer: 0, never negative, streak resets to 0
- A hint used on that question halves its award -- get-hint.js now
  sets session.pending_q_hint_used on any non-blocked hint grant,
  which combat-action.js reads and clears when that question is graded
- Enemy defeated: +250; surviving to victory adds +1 per HP remaining
- difficultyMultiplier defaults to 1x, a hook for issue #9's Easy/
  Medium/Hard multipliers rather than combat-action.js reimplementing
  scoring whenever that lands

Score and streak are computed and stored server-side only (session.streak,
alongside player.score) -- no client code writes either, so the score
can't be edited from the client the way a client-computed one could be.
Streak is deliberately reset on every new encounter in start-combat.js/
reset-game.js (a fresh battle's run of correct answers, not one
inherited from a fight that already ended), while score is deliberately
carried forward across encounters within a session via the existing
freshPlayer(existingScore) -- both now called out explicitly in comments
rather than left incidental, per the issue's ask. Works for signed-out
guests since it all lives in the KV session guests already have.

Client: score and streak ride inside combat_state (same pattern as
issue #15's action_costs) so no new fetch or call-site plumbing was
needed. Added a Score display to the combat HUD with a live "+100"
delta annotation and a "Streak x2"+ indicator once it's worth showing.
_holoStreak (the existing holo-card shine intensity driver) is now a
mirror of the server's streak instead of an independent client-side
counter that could drift from it.

Found and fixed a real bug while wiring the delta annotation: several
call sites (submitQuizAnswer's `finally` block, its multi-select
counterpart) re-invoke updateCombatHUD a second time with the exact
same already-applied combat_state as a defensive re-render. The #17
Resolve delta got away with a simple value-diff guard against that
because Resolve always changes on a graded turn -- score doesn't: a
wrong answer is a real, fresh turn that legitimately scores a delta of
0, which a value-diff can't distinguish from "redundant call, nothing
changed" and would leave a stale "+100" showing after a miss. Fixed by
computing isFreshState once per updateCombatHUD call via object
identity against the last-rendered combat_state (a real new server
response is always a distinct object; the redundant call reuses
window.combatState's existing reference) and gating both the Resolve
and score delta annotations on that instead of on value equality.

Verified via Playwright: streak-bonus math against real API responses
(+100 at streak 1, +125 at streak 2, matching the formula), never a
negative delta on a wrong answer, hint-halving via direct API calls
(score_delta 50 after a hint on an otherwise-100-point question,
resolved against the real corpus for a deterministic correct answer),
and that the delta annotation no longer sticks at a stale value across
the redundant-re-render turns that exposed the bug above.
Both end screens previously said "Victory!"/"Defeat" and listed
questions with a tick or cross -- a receipt, not a result: no score,
accuracy, streak, XP, or path to understanding what was missed.

combat-action.js computes a run_summary object (score, accuracy,
correct/total, best_streak, xp_earned, hints_used, hp_remaining) once
outcome is set, extending the existing end-of-run payload rather than
requiring a second request. best_streak is new session state, tracked
alongside streak and reset per-encounter same as it (start-combat.js,
reset-game.js). xp_earned is floor(score / 10) per the scoring issue's
own suggested starting divisor. level_results entries now also carry
correctAnswer and the question's easy-tier hint (every one of the 800
questions has one) so a missed question's review renders from data
already being returned, not a second lookup.

Client: one shared renderRunSummary()/handleRunOutcome() replaces three
separate copies of the victory/defeat transition logic (in performAction,
submitQuizAnswer, and submitQuizAnswerMulti) that had already drifted
once -- the single-select path needed its own fix for the 'active'-class
toggle that the multi-select path already had. Score is the largest
element on the screen (clamp(2.75rem, 9vw, 5.5rem)), above the existing
flavour-text heading. Missed questions expand via native <details>/
<summary> (keyboard/screen-reader support from the browser, no custom
expand-collapse JS) to reveal the correct answer and hint.

renderLevelResults() now runs question/answer/hint text through
normalizeMathText (escapes HTML, then adds KaTeX $...$ delimiters, same
treatment the quiz modal already gets) instead of interpolating raw
corpus text -- unescaped "<" or "&" used to break this markup outright.

Added "Play Again (Same Realm)" and "Change Realm" next to the existing
"Return to Main Menu". Neither needs a confirmation dialog -- starting a
new run from an end screen discards nothing the player hasn't already
finished, unlike the nav Home button's mid-run abandonment -- noted
explicitly in playAgainSameRealm()'s comment so it doesn't get added
reflexively later. Play Again reuses the just-finished run's syllabus_id/
difficulty from window.combatState and calls the existing selectSyllabus()
directly; Change Realm shows the syllabus-select screen without a full
reset. Both correctly hide the end-screen overlay first (it doesn't
clear on its own the way combat-screen's display toggle does).

Focus moves onto the summary container (tabindex="-1") when the screen
appears, using the same transitionend-or-450ms-fallback pattern as
trapFocus's own focus-establishment workaround, since the screen's
opacity/visibility change is behind an 0.8s CSS transition.

Verified via Playwright: run_summary math against real API responses
(a 6-correct-in-a-row victory scored exactly 1280 -- streak-bonus sum
950 + victory bonus 250 + 80 HP remaining, matching the formula by
hand); score/accuracy/streak/XP/hints render correctly and match the
API response; missed-question review expands to real corpus answer +
hint text; focus lands on the summary container after the screen
transition; no horizontal overflow at 390px; Play Again returns to
combat with score carried forward; Change Realm shows the syllabus grid.
#22)

profile.js previously stored one field ({uid, active_game_id}) with its
own header comment naming the points economy as deliberately deferred.
Extends the schema to lifetime_xp, xp_balance, totals (runs/questions
answered/correct), per-realm records (best_score, runs, accuracy inputs,
best_streak, keyed by realm name so #9's difficulty tiers can nest under
each one later without a migration), and recent_runs capped at 10
(a Firestore document has a 1MB limit; an uncapped array is a slow leak).

toFirestoreFields()/fromFirestoreFields() were hand-mapping 2 flat
fields -- rewritten as a generic, recursive encoder/decoder
(toFirestoreValue/fromFirestoreValue) that handles strings, numbers,
booleans, null, arrays (arrayValue), and nested objects (mapValue) for
any shape, not just today's. This is what makes the realm-tier-nesting
and recent_runs-array requirements possible without hand-writing a
converter per field. Verified round-trip correctness (including nested
maps/arrays) by mocking fetch and decoding what putProfile() actually
sends back through getProfile()'s path -- see
backend/test_issue22_firestore_roundtrip.mjs.

applyRunToProfile() applies one finished run to a profile in place
(pure data manipulation, no I/O) and mergeProfiles() does the guest-to-
account merge on first sign-in (sums totals, takes the max of bests,
merges recent_runs by recency) rather than an account overwriting a
week of guest play.

Write timing: combat-action.js writes the profile exactly once, at run
end, from the same place run_summary (#21) is computed -- not per
answer. XP/score come only from that server-computed run_summary, never
the request body, so a client can't post an arbitrary XP total; the
Firestore write is wrapped in try/catch so a failure there can't break
the run's own response (the player still sees their summary). The
client now sends its Firebase ID token on every combat-action call so
the server can verify auth.uid matches session.uid before writing.

Guests get the identical shape in localStorage (studysaga_guest_profile)
via a deliberate line-for-line mirror of applyRunToProfile() in
game-simple.js -- there's no shared module system between Pages
Functions and the unbundled static frontend, so the duplication is
by-hand and called out in a comment on both sides. A new
/api/sync-profile endpoint both fetches a signed-in player's profile
(so a second device shows the same lifetime totals) and, when passed
local_profile, merges it in exactly once right after sign-in --
mergeGuestProfileOnSignIn() only clears the local copy after the server
confirms the merge saved, so a network failure leaves it intact to
retry rather than silently losing a guest's history.

A "My Profile" panel on the main menu (reusing the existing modal/
trapFocus infrastructure) shows lifetime XP, XP balance, total runs,
overall and per-realm accuracy, per-realm best score, and the last 10
runs -- reads localStorage directly for guests, calls /api/sync-profile
for signed-in players.

Security: kept the existing model (client's own ID token, Functions
verifies it, no service account). The Firestore rules still give a
user full write access to their own document, which profile.js's
comment now states explicitly as an accepted tradeoff for a
single-player study game with no leaderboards -- the game itself can't
be cheated through the normal client since XP is always server-derived
from session state, only a player's own private stat display could be
edited directly through the Firebase SDK, which only harms their own
record-keeping. Revisit before any social comparison feature ships.

Verified via Playwright: a full guest run correctly populates
localStorage with the right shape and values; the profile panel renders
real accumulated data (lifetime XP, per-realm table, recent runs) for
guests; recording 12 runs caps recent_runs at exactly 10;
/api/sync-profile returns 401 on a missing or invalid token.
XP only counted upward until now. Depends on scoring (#20) and the
persistent profile (#22), both shipped earlier this session, and is
explicitly gated by its own issue text against trivializing the game
before real per-tier enemy scaling (#9) exists -- addressed with a
documented, conservative balance target rather than shipping unlimited
power creep (see MAX_TOTAL_UPGRADE_LEVELS's comment in game.js): the
catalogue's 6 upgrades sum to 23 possible levels, capped at 12 total
purchased, so a player can reach roughly half of any one upgrade's
ceiling and never max every category simultaneously. Placeholder tuned
by inspection, not playtest data -- revisit once #9 ships.

Catalogue (game.js UPGRADE_CATALOG, deterministic -- no randomness, no
loot boxes, no real money): Neural Capacity (+1 max CAP/level, 5
levels), Resilience (+10 max HP/level, 5), Efficient Recall (+1
Recharge CAP/level, 3), Extra Insight (+1 Simple hint/run/level, 3),
Deep Insight (+1 Deep hint/run/level, 2), Focused Strike (+2 Attack
damage/level, 5 -- Ability's damage is never upgraded, per the
catalogue). Doubling cost curves per upgrade.

The real architectural work this issue asked for: action costs/damage,
max HP/CAP, and hint budgets were module-level constants
(ACTIONS, SIMPLE_HINT_MAX, HARD_HINT_MAX, CONFIG.players.default_kk)
read at each use site -- upgrades require them to be per-session.
effectiveStats() in game.js resolves a player's purchased levels into
actual numbers once, at start-combat time, stored as
session.effective_stats and threaded through: freshPlayer() takes it
for max_hp/max_cap, actionCosts() for attack damage/recharge gain
(both already returned to the client via issue #15's mechanism, so the
self-describing buttons automatically show upgraded numbers with no
separate client change), and hintsSummary()/get-hint.js's budget check
for simple/hard hint maximums. Verified via direct API calls: a player
with Resilience x2 + Neural Capacity x1 gets max_hp 130/max_cap 11
(base 110/10); Focused Strike x1 deals 17 real damage in an actual
graded turn (base 15); Extra/Deep Insight x1 each raise the hint
budget to 4 simple / 2 hard.

Where upgrades come from: signed-in players' levels are read
server-side from their own Firestore profile via id_token (same
accepted tradeoff already documented in profile.js -- this can only
affect a player's own run). Guests have no server-side account, so
their levels ride in the start-combat request body, same trust model
as the rest of the guest profile.

New /api/buy-upgrade endpoint validates a purchase entirely against the
player's own stored profile -- current level, next cost, the total-
level cap, and XP balance are all re-read there, never trusted from the
request, so a client can't grant itself an upgrade or spend XP it
doesn't have. Both halves of a purchase (XP deduction, level increment)
are applied to one in-memory object before a single PATCH write, so a
failed write leaves the untouched old profile in place -- atomic by
construction, not by a transaction API. Guests buy against their own
localStorage balance client-side (buyUpgradeGuest(), mirroring the
server logic the same by-hand-sync way applyRunToProfile() already
does for issue #22) since there's nothing server-side to validate a
guest purchase against.

profile.js's schema gains upgrades: { [key]: level }; mergeProfiles()
merges it by max-per-key (not summed -- these are permanent levels
already paid for once, so a guest-then-signed-in player must not get
both sets of levels for the price of one).

UI: a card grid matching .syllabus-card's look (name, plain-language
effect, level/max, next cost, affordability) reachable from the main
menu and both run-summary screens. Unaffordable and maxed cards stay
visible with their state shown, never hidden. Buying requires a
confirm() dialog -- spending XP is irreversible, unlike issue #21's
Play Again/Change Realm which discard nothing and deliberately don't
prompt.

Verified via Playwright: shop renders all 6 catalogue entries with
live cost/affordability; a purchase deducts the correct XP and
increments the correct level; the 12-level cap blocks further
purchases with the capped-message shown; insufficient XP disables the
buy button; /api/buy-upgrade rejects a missing or invalid token.
The corpus has a severe generation-time skew: 88.7% of single-select
questions (2641/2976) store the correct answer at options[0]. The
server was never shuffling option order before sending a question to
the client -- sanitizedOpts was a straight .map() over the stored
order -- so every player saw that same skew, live, in production.
Reported by the user as "almost all correct answers were choice A."

Fixed by shuffling each question's option order once per serve (both
places a question is dispatched: the "next question after grading" and
"initial question fetch" branches), using the existing shuffle()
helper on an index array so the mapping back to the original,
unshuffled question.options can be kept. That mapping is stashed as
session.pending_option_order and used at grading time to translate the
client's submitted answer_index/answer_indices (positions in the
shuffled order the player actually saw) back to the corpus's original
indices before comparing against the answer key, which is keyed to
that original order.

Verified via direct API calls across 45 sampled questions (15 games):
served option[0] matched the corpus's stored option[0] only 31.1% of
the time post-fix (vs. the 88.7% baseline, and in line with chance for
mostly-4-to-5-option questions), and all 45 submissions of the
corpus-correct answer at its new shuffled position were graded correct
-- the translation logic doesn't break grading.

This is a data/product-integrity bug, not a regression from this
session's other work -- the bias has been live in production since
before today. Not yet deployed; needs a manual `wrangler pages deploy`
alongside the rest of this session's undeployed work.
…gating

#9 (Easy/Medium/Hard difficulty tiers) was only partially done -- the
difficulty picker UI and question-tagging existed, but the rest of the
issue's scope did not, contrary to an earlier assumption in this
session that it was fully shipped. This closes the gap except for the
per-question timer, which the issue's own comment thread says is
worth splitting into its own issue (no timer of any kind exists
anywhere in cf-pages/ today) -- filed separately.

Score multiplier: DIFFICULTY_MULTIPLIERS (Easy 1x / Medium 1.5x / Hard
2x) in game.js, wired into scoreForAnswer()'s call site in combat-
action.js via session.difficulty (locked in for the encounter at
start-combat time). Stated on the difficulty-picker buttons themselves
("2x score") so it's visible before a player commits to a tier, not
just reflected in the score afterward. Verified: a correct answer at
Hard scored exactly 200 (100 base x 2x).

Per-realm-per-tier persistence: profile.js's realm records are now
nested by tier (realms.biology.hard.best_score etc.) -- the exact shape
issue #22 was deliberately left able to hold without a migration. Each
tier record also tracks victories, driving a "Cleared"/"Not cleared"
badge per realm x tier in the profile panel. Profiles saved before this
nesting existed (a flat record directly on profile.realms[realm]) are
migrated to 'medium' the first time they're touched
(migrateLegacyRealmRecord(), mirrored client-side for guests) rather
than requiring a one-time migration script or silently losing the old
data -- verified the migration merges rather than overwrites.

Tier-selection gating: /api/syllabi now returns tier_counts and the
15-question floor (MIN_TIER_QUESTIONS in game.js); a difficulty button
below the floor is disabled outright instead of silently falling back
to the full question pool after a player has already committed to it
(that fallback stays as a server-side defensive safety net). All four
realms currently sit well above the floor in every tier (lowest is
biology/hard at 172), so nothing is disabled in practice today, but the
mechanism is real and tested.

Docs: added a difficulty tier guidelines section to the README (tier
definitions, score multipliers, the question floor) for future question
authors, per the issue's acceptance criteria.

Verified via Playwright and direct API calls: tier_counts/min_tier_
questions returned correctly; Hard-tier scoring multiplies correctly;
per-tier victories/best_score/accuracy track and render correctly in
the profile panel with correct Cleared/Not cleared/Not played states;
legacy flat realm records migrate into the medium tier and merge
(not overwrite) on the next run recorded against them.
Split out of #9's difficulty-modifier scope since no timer, countdown,
or deadline logic existed anywhere in cf-pages/ before this. Starting
values (tunable in one place, game.js's QUESTION_TIME_LIMIT_MS, not
playtested yet): Easy 30s, Medium 20s, Hard 12s.

Server: computed once per question serve (both the "next question" and
"initial fetch" branches in combat-action.js) from the question's
difficulty tag, stored as session.pending_q_deadline, and returned to
the client as question.time_limit_ms. At grading time this is enforced
as a grace-padded (3s) backstop -- forces isCorrect = false if a real
answer somehow arrives after deadline+grace -- consistent with scoring
being server-authoritative everywhere else in this codebase (#20) and
never trusting client timing for the actual grade.

Client: a countdown bar + numeric text in the quiz modal
(startQuizTimer()/stopQuizTimer(), cleared on every modal close path so
a stray timer can't fire against a question the player already
answered). On expiry, auto-submits an empty answer through the exact
same submitQuizAnswer()/submitQuizAnswerMulti() path a real click uses
-- null/[] never matches the answer key, so no separate "timed_out"
signal needs to travel to the server, and CAP is deducted exactly like
any other wrong answer (existing "wrong answer still costs CAP" rule,
issue #15, applies unchanged). Countdown bar sweep is a CSS transition
disabled under prefers-reduced-motion; the numeric seconds-remaining
text is the primary signal either way.

Found and fixed a message-framing bug while verifying end-to-end: the
"Time's up!" message (vs. plain "Incorrect") was keyed only to the
grace-padded timedOut check, which the client's own auto-submit
almost always lands inside (it fires right at the nominal deadline,
which the grace window is specifically there to treat leniently) -- so
a genuine timeout would silently read as an ordinary wrong answer.
Fixed by keying the message on whether an answer was actually
submitted at all (only ever empty via the auto-submit-on-expiry path),
keeping timedOut purely as the scoring backstop the two checks were
never meant to share.

Verified via Playwright: time_limit_ms served correctly per tier (12s
for Hard); the countdown UI renders and updates; a real, un-answered
12s Hard-tier question auto-closes the quiz modal, deducts CAP exactly
like a wrong answer, and shows "Time's up!" with the correct answer
text.
cf-pages/functions/_lib/data.json becomes the single source of truth
for the question corpus -- it's what's actually served to players
(3,152 questions vs backend's ~800) and already carries bug fixes
that never made it back into backend's copy. backend/data.json is now
a generated mirror for the local Flask dev server, never edited
directly.

- merge_corpus_reconciliation.py (one-time): pulled forward 10
  questions' better-written hints from backend/data.json before
  retiring it to a mirror; everything else backend/data.json had that
  differed (difficulty tags from a classification pass that found
  zero Hard-tier questions anywhere) was deliberately not imported
- sync_corpus.py (ongoing): regenerates backend/data.json from the
  master; CI's new corpus-drift job runs it and fails the build if
  the committed mirror doesn't match
- validate_corpus.py: added the 3 checks issue #24's acceptance
  criteria call for that were missing -- duplicate question text
  within a realm, duplicate option text within a question, and
  answer_index/answer_indices range + consistency with isCorrect
  flags
- README.md / backend/README.md: documented the master/mirror
  relationship and indexed the two new scripts

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Merges Eyal's 800-topic tiered corpus (200 topics x 4 subjects) into the
production data.json served by cf-pages Functions, adding 2,400 questions
across 4 new subjects. Content went through a multi-round independent
Gemma re-verification after an initial 47% defect rate was found in what
Gemini had scored as clean (dominant issue: hint-tier inversion). Final
state: 0 flagged topics, 91.3% of individual hint scores high-quality.

Existing biology/math/chemistry/physics content is untouched. Feedback
text ships in its verified plain-educational tone rather than the other
subjects' combat-flavor voice, to avoid layering new unaudited content on
top of what was just verified.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@femtechtk
femtechtk merged commit 5fd368f into main Sep 4, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant