Merge flaskfrontback: adds Eyal's 4-subject corpus + prior session's feature/corpus work - #30
Merged
Merged
Conversation
…y/math/physics) Hand-written tiered question+hint content: 200 topics per subject, each with easy/medium/hard question variants and matching 3-tier hints. Includes the Gemini-as-judge script (judge_tiered_content.py) and its in-progress results (512/800 judged so far, avg 9.98/10). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Issue #18: the single-select answer path (submitQuizAnswer, the most common one) never wrote anything to the battle log; the multi-select path did. Both now populate it identically. Issue #19 (partial): replaced the 6 in-combat alert() calls (Attack/ Ability/Recharge failures and network errors) with a new shared addBattleLogEntry() helper that renders each turn as a comic-caption- style panel instead of a bare text line or a jarring native dialog. Tone/burst badges are inferred from the backend's existing narrated message text -- no combat numbers or balance changed. Battle log now has role="log" aria-live="polite" for screen readers, and the entry animation respects prefers-reduced-motion. The remaining 6 alert() calls (session-expired, initial deploy-into- combat failure, reset-game failure) fire when no battle screen is showing, so they're left as native alerts pending a separate fix. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Final results: 800/800 topics judged, global average 9.98/10 (biology 9.95, chemistry 10.00, math 9.99, physics 10.00). Only 2 topics scored below 9: a real hint/question mismatch bug in claude_tiered_batch4_ biology.json (7.8) and a cosmetic wordiness issue in claude_tiered_batch100_biology.json (8.7). The version of this file committed earlier was mid-run (~512/800); this replaces it with the finished evaluation. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds 2,352 Gemini-judged questions (9.98/10 avg) across biology,
chemistry, math, and physics -- each of the 3 difficulty-tiered
questions per topic becomes its own entry in the syllabus's existing
question list, tagged with `difficulty`, carrying its own 3-tier
hints object in the exact pre-existing hints:{hard,medium,easy}
schema. No changes needed to combat-action.js, get-hint.js, or the
frontend -- this reuses the existing random-question-draw mechanic
and hint system exactly, just with a much larger, higher-quality pool.
48 of the theoretical 2,400 were skipped as exact-text duplicates of
already-present questions (backend/merge_tiered_into_live_data.py is
idempotent by question text, safe to re-run).
Verified end-to-end against a local wrangler pages dev instance:
start-combat, attack (drew both pre-existing and newly-merged
questions), answer submission, and get-hint all work correctly on
the merged pool.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Difficulty is picked once when entering a realm (replacing the single
"Initialize Sync" button with three difficulty buttons on each syllabus
card) and locked in for the whole battle -- not a per-turn toggle.
Backend: start-combat.js now filters the question pool by the chosen
difficulty before building question_order (questions with no difficulty
tag -- the original, pre-tiered content -- are treated as "medium" so
they stay reachable). combat-action.js surfaces the active difficulty
in every combat_state response so the frontend can display it.
Falls back to the full unfiltered pool if a subject/difficulty
combination ever has zero matches, so a player is never left with no
questions to answer.
Frontend: the selected difficulty now shows in the persistent nav bar
("Biology · Hard") for the rest of the battle.
Verified end-to-end against a local wrangler pages dev instance:
selecting "hard" only ever drew questions tagged difficulty:"hard" in
the live data, and "easy" only drew difficulty:"easy" ones -- confirmed
by cross-referencing each drawn question's text against the merged
data.json rather than just eyeballing question style.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
… buttons The previous UI had two buttons (Simple Hint, Deep Hint), where Deep Hint cascaded through all three tiers one at a time via a "Need More Help?" button. Replaced with three independently-clickable buttons, one per tier, matching the difficulty-picker's color convention (green/yellow/red). Backend budget semantics unchanged: tier="simple" (easy only, cheap budget) backs the Easy button; tier="hard" (all three tiers in one call, pricier budget) backs both Medium and Hard, which now share a single cached fetch promise so clicking both only spends the "deep hint" budget once, not twice. Verified locally with Playwright: clicking tiers out of order (Medium first, then Hard, then Easy) correctly reveals each one independently and displays all three in a consistent Easy/Medium/Hard order regardless of click order. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…he UI These two words had been used loosely/interchangeably in conversation and in the UI itself, which risked confusing players about what each control actually does. Now explicit: - Syllabus card caption: "Level = how advanced this realm's questions are (Easy: intro-level, Medium: standard coursework, Hard: advanced/ exam-level)" -- calibrated to the actual criteria used when writing the tiered corpus (easy = recall/definitions, medium = explaining a mechanism/relationship, hard = multi-step reasoning/synthesis). - Hint box caption: "Hint tiers for THIS question (not the realm's difficulty level) -- pick how much help you want" -- directly distinguishes hint tiers from the syllabus-level difficulty choice. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Easy = Gifted Leo Baeck (high school), Medium = Technion (undergrad), Hard = Technion (graduate) -- per explicit direction. Note: the actual question content at "Hard" tops out around AP-high-school/early-college rigor, well below genuine Technion coursework difficulty at either level; this labeling is intentional flavor/branding, not a rigor claim verified against the content. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…hool Replaces the previous Leo Baeck / Technion labeling with the user's own personal educational background (Ramaz for high school, Cornell for college) plus a generic "graduate school" for Hard. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…ption Ramaz/Cornell/graduate-school labeling is kept as an internal code comment for reference, but the visible caption reverts to generic "Choose a difficulty" copy -- players shouldn't see the specific institution names. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Parallel to judge_tiered_content.py (Gemini), reusing the identical prompt/schema/rubric so the two judges' outputs are directly comparable. Separate output file, fully resumable. Includes progressive backoff (up to 90s) and a higher consecutive-failure threshold (10) to ride through Gemma's intermittent transient 500/503 errors without needing a manual resume every time. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
#27) - Deleted the dead showFeedbackModal() stub (unreachable -- the real definition later in the same file wins at parse time, and the stub's only body was alert(), one of the flagged alert() calls). - Removed pushRecentMessages() and its 3 call sites -- it wrote to #recent-results-list, which doesn't exist anywhere in cf-pages/public/. Deleted per the issue's stated default (the battle log already covers this need), along with its orphaned CSS (.neural-recent, .neural-recent-title, .recent-list-entry). - Removed all [DEBUG]-tagged console.log/warn calls. Verified first that the "full question object" logged by openQuizModal is NOT answer-revealing -- combat-action.js already strips isCorrect/ feedback before sending a pending question to the client -- but removed the noise anyway per the issue's tidiness ask. - Removed two stale markers: the "VERSION label removed" HTML comment and an already-gone "...existing code..." placeholder (removed in an earlier edit this session). Verified end-to-end via Playwright against a local wrangler dev instance: menu -> realm select -> combat -> single-select answer -> multi-select answer -> recharge -> reset via nav Home, zero JS page errors throughout, battle log correctly populated. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…uping (#18) The core bug (single-select never wrote to the log) was fixed earlier this session; this closes the rest of the issue's acceptance criteria that a fuller read of the issue text revealed weren't yet met: - Added a visible "Battle Log" heading above the panel. - Added a CSS-only empty-state placeholder ("Your actions will appear here.") via :empty::before -- shows automatically when the log has no entries, disappears on the first real one, can't drift out of sync with actual content since there's no JS state to maintain. - Turn-grouping: addBattleLogEntry() is now addBattleLogTurn(), taking a whole action's message array and rendering it as ONE grouped block (your result, then the enemy's response) instead of separate floating lines. All 3 call sites updated; addBattleLogEntry() now delegates to it for single free-standing messages (errors etc). - Non-color differentiation: each line in a turn gets a text "You:" / "Enemy:" speaker label (position-based, matching combat-action.js's consistent message ordering) in addition to the existing border-color coding, so message kind no longer depends on color alone. Verified via Playwright: heading and placeholder render correctly; 25 synthetic turns produce a 2678px-tall scrollable log inside its 142px container with zero page-height growth and zero horizontal overflow; a full combat run (single-select + multi-select) renders grouped turns with correct speaker labels and zero JS errors. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The dead showFeedbackModal() stub and 6 in-combat alerts were already handled (#27, earlier this session). This closes the rest: - Discovered and fixed a real latent bug while working on this: setStatus() has been targeting a #test-output element that doesn't exist anywhere in index.html, so every setStatus() call in the codebase was silently a no-op. Added the missing element (role= "status" aria-live="polite") so it actually works. This mattered here specifically because handleInvalidSession()'s existing "UI path" the issue assumed was already showing something turned out to be silently broken -- removing its alert() without this fix would have made session-expiry produce zero visible feedback. - Removed the redundant session-expired alert() -- the function ends with setStatus() showing the same message, now that setStatus actually works. Also fixed it to use the caller's specific message instead of a hardcoded string that was silently discarding it. - Fixed a pre-existing setStatus('Error: ' + e.message) call (start- game's catch block) that raised the same "never show e.message to a player" problem this issue is about, even though it wasn't one of the originally-counted 12 alerts. - Converted the remaining 4 alerts (start-combat failure/exception, reset-game failure/exception) to the existing feedback modal infrastructure. Added an optional `title` field to showFeedbackModal so these can show "Error" instead of the quiz-specific "Correct!/ Incorrect" title that would otherwise be misleading for a system error. Raw exception text stays console-only; players see plain- language messages. - "Not enough CAP" reclassified as guidance, not just an error string: both backend message sites now explicitly name Recharge as the next move ("Not enough CAP -- Recharge to continue."). Verified via Playwright: draining CAP naturally triggers the new guidance text in the battle log (not a native dialog), zero alert() calls remain in cf-pages/public/static/js/ (grep-confirmed), zero JS page errors. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Moves session-status/reconstruction docs and benchmark/diagnostic reports out of the repo root into docs/history/ and docs/reports/ so they read as historical record instead of live top-level state. MVP_TESTING_PLAN.md is included -- it describes a 2026-07-22 session's outstanding test gaps against an older build and is stale the same way. Removes confirmed-dead files: next-env.d.ts (no Next.js/root package.json in this repo), and three unused frontend/templates/ variants (only 'index.html' is ever passed to render_template() in app.py). Moves the unrelated standalone rag_quiz.py prototype into backend/archive/. Adds backend/README.md indexing the ~75-script pile by purpose (content generation, judging/QA, bake-offs, corpus maintenance, extraction, throwaway Playwright verification scripts) rather than physically relocating them into subfolders -- many assume backend/ as their working directory for relative paths to data.json and their own intermediate output files, and moving them without auditing each one risked breaking that. Updates the root README's repo layout section to match (docs/history, docs/reports, backend/archive, backend/README.md pointer). Verified backend/app.py still imports cleanly after the root-level moves.
Found while building the corpus validator for issue #26: two questions were typed multiple_choice_single but had 2 options genuinely correct (both 6 and 8 are factors of 24; both "elements in both sets" and "the intersection of the sets" describe a Venn overlap) -- retyped both to multiple_choice_multiple rather than arbitrarily picking one answer to mark wrong. The third (line through (2,3) and (4,5)) had 3 garbled, apparently glued-together distractor options, with the actual correct answer "y = x + 1" marked false and two nonsense options marked true -- removed the corrupted options and fixed the isCorrect flag on the correct one, keeping the 3 legitimate distractors. This is the same bug class flagged earlier during the 800-topic corpus work (a wrong isCorrect flag found in the corpus); a full audit is still outstanding, but these 3 needed fixing now so the new corpus validator (backend/validate_corpus.py) has a clean baseline to check future changes against instead of failing on day one.
Issue templates (.github/ISSUE_TEMPLATE/): bug.yml, feature.yml, and content.yml (for corpus-specific reports -- wrong answer, hint leaks the answer, broken formula), all following the Problem/Proposal/ Acceptance-criteria structure the existing issues already use. Blank issues stay enabled via config.yml. PR template (.github/pull_request_template.md): what/why/how-verified plus an explicit "redeployed to Cloudflare Pages: yes/no" checkbox -- deploys are manual (no Git integration on the Cloudflare side), so a merged PR is not a shipped PR and that distinction was previously invisible. CI (.github/workflows/ci.yml), runs on PRs and pushes to main: - Corpus validation (backend/validate_corpus.py): schema check plus the invariants the game logic assumes -- exactly one correct option per multiple_choice_single question, at least one for _multiple, all 3 hint tiers present and non-empty. This is a new script, not the drift-check mentioned in the issue text -- see below. - JS syntax check (cf-pages/check-js-syntax.sh): node --check across every classic script and ES module in cf-pages/, since the frontend is unbundled with no build step and a syntax error would otherwise only surface live in the browser. - Python lint (ruff over backend/), non-blocking (--exit-zero) given ~75 inherited scripts never linted before. - README link check (check-readme-links.py): relative links and in-page anchors only, no external URLs (there are none in the README currently, and checking those in CI is flaky). Deliberately NOT included: a corpus-drift check comparing backend/data.json against cf-pages/functions/_lib/data.json. The two have already diverged and which one is authoritative is an open question tracked in #24 -- adding the check now would just fail every PR. Wired up once #24 is resolved. Verified corpus-validation and JS-syntax-check both fail correctly against deliberately broken input (a duplicated isCorrect flag; a truncated JS file), per the issue's acceptance criteria. README documents the new CI table and reiterates that CI does not replace the manual `wrangler pages deploy` step.
After 10s of no pointer/keyboard/focus input inside #combat-screen (or right after any completed action, via updateCombatHUD), pulse the button the player is expected to press next -- Attack when CAP >= 3 (covers both the "cheapest progress" and "only affordable option" cases from the issue), Recharge otherwise, since Attack/Ability are disabled below 3 CAP and nudging a disabled button would read as a broken screen. No nudge while a modal is open -- the player's attention belongs to the question there. CSS owns the animation (nudge-pulse keyframes on .neural-action-btn.nudge, animating box-shadow/transform off currentColor so each button keeps its own accent color) gated under prefers-reduced-motion: no-preference; the reduce branch substitutes a static outline instead of removing the cue. A visually-hidden #idle-nudge-announcer (role="status", aria-live="polite") announces "Suggested next action: <label>" once per idle period, not once per pulse cycle, and the nudge never moves focus or triggers an action. Verified with two Playwright scripts against local wrangler pages dev: timer fires at 11s idle and targets Attack; any keydown clears it and restarts the timer; no nudge appears while the quiz modal is active even past 11s; animation-name resolves to 'none' under emulated prefers-reduced-motion: reduce; draining CAP below 3 moves the nudge to Recharge and never lands on a disabled button; no horizontal overflow at 390x844 while the pulse animates. (Playwright's click-stability check needed force:true against the pulsing button, since Playwright treats a continuously-transformed element as "not stable" for actionability purposes -- expected given the animation, not a functional issue for real pointer input.)
Each combat button now shows a smaller second line stating its cost
and effect ("3 CAP · 15 dmg", "free · +5 CAP", "5 CAP · 25 dmg")
instead of communicating affordability through opacity alone.
Recharge (CAP) is renamed to plain Recharge now that its cost line
explains the mechanic.
Cost/damage/gain values are no longer duplicated: cf-pages/functions/
_lib/game.js now exports ACTIONS and actionCosts() as the single
source, combat-action.js imports ACTIONS instead of redefining it
locally, and every endpoint that returns combat state (start-combat,
combat-action, auth-resume) includes action_costs in it. The client
caches this onto window.actionCosts inside updateCombatHUD() and reads
it for both the affordability checks (previously hardcoded `cap >= 3`
/ `cap >= 5`) and the idle nudge's target selection -- a hardcoded
DEFAULT_ACTION_COSTS constant is kept only as a pre-first-response
fallback, mirrored to match ACTIONS server-side.
Accessibility: the cost line lives in an aria-hidden span so it never
runs into the accessible name; a separate visually-hidden span per
button (referenced via aria-describedby) gives screen readers "Attack"
as the name and "Costs 3 CAP, deals 15 damage" as the description.
Verified via Playwright's aria_snapshot() that button names resolve to
just "Attack" / "Recharge" / "Use Ability".
Disabled state now states a reason instead of just dimming: a status
line under the action row reads "Not enough CAP -- Recharge to
continue." whenever CAP drops below Attack's cost, live-announced via
aria-live="polite". A permanent caption below it discloses that a
wrong answer still costs CAP, since combat-action.js deducts the cost
before checking correctness and players would otherwise find out by
feeling cheated.
Buttons switched to a flex-column layout to fit the second line;
verified via Playwright at 390px width that the row still fits with no
horizontal overflow and the Attack button's tap target stays above
44px (measured ~114x50).
"CAP" itself is still unexpanded in plain language on this screen --
that's issue #17 (How to play panel), landing next.
Expanded the two invented-abbreviation labels in place: "CAP" becomes "CAP (actions)", and "HP" gains a "Your"/"Enemy" prefix on the player and enemy cards respectively so the two bars are distinguishable without relying on colour or card position (they were previously both just "HP:"). Added a "How to Play" button to the persistent combat nav (one click, always reachable) that opens a modal explaining HP, CAP, and Resolve in plain language. Content is rendered from a single array, COMBAT_STAT_EXPLANATIONS in game-simple.js, so this and the issue #16 tutorial (landing next) read from one source instead of maintaining two copies of the same rules text that could drift. Resolve is now legible while it moves: a delta annotation ("-18" / "+14") appears next to the bar after each turn, plus a text state label at the extremes matching combat-action.js's own thresholds -- "Rattled" at <=15 (the hesitate-chance threshold) and "Emboldened" at >=85 (the heavy-hit-chance threshold). The delta's fade-in is gated under prefers-reduced-motion: no-preference; under reduce it just appears/clears with no transition. Found and fixed a real bug while wiring this up: submitQuizAnswer's existing (pre-existing, unrelated to this issue) defensive re-render of updateCombatHUD() in its `finally` block was immediately overwriting the delta with a blank "no change" result, since it ran a second time against the already-applied state -- fixed by making updateResolveAnnotation a no-op when called again with a resolve value matching the last one actually rendered. All 4 HUD bars now expose role="progressbar" with accurate aria-valuenow/valuemax/aria-label (factored through one new setProgressBar() helper instead of repeating the width-and-nothing-else update inline per bar), so a screen reader reports e.g. "Enemy Resolve, 50" instead of an unlabelled div. Verified via Playwright at 390px: no horizontal overflow, all 4 progressbar roles/values correct, How-to-Play panel opens on the first click and closes correctly, and the resolve delta/state-label computation produces "-18"/"+14" after a real turn and "Rattled"/ "Emboldened" at the extremes.
A five-step tutorial runs once on a player's first-ever combat screen, gated on a localStorage flag (studysaga_tutorial_seen, read/write wrapped in try/catch like holo-card.js's existing localStorage use -- private browsing means it may reappear next session, which the issue explicitly allows rather than throwing). Each step dims the screen via a spotlight cutout (an oversized box-shadow on the highlight box, so the target itself stays undimmed without a separate masked overlay), highlights one real element -- player card, CAP bar, Attack, Recharge, enemy Resolve bar -- and states the mechanic in plain language, closing with "Your turn -- press Attack" pointing at Attack again rather than highlighting nothing. No sandbox battle, no fake state: the overlay sits visually on top of the real buttons and captures every click itself, so nothing underneath is ever actually pressed while a step is showing -- verified CAP is unchanged after clicking through all six steps. Reuses the existing trapFocus()/deactivateModal() modal infrastructure (focus trap, Tab wrapping, Escape handling, the transitionend-based focus-first workaround) instead of a second implementation. Found and fixed a real bug while wiring this up: trapFocus() auto-focuses the first focusable element in DOM order, which was the Skip button -- meaning Enter would skip the tutorial instead of advancing it. Fixed by reordering the buttons so Next is first in the DOM (Skip keeps its original visual position via CSS `order`). Step content lives in one array, TUTORIAL_STEPS, and the highlight position is derived from each target's live getBoundingClientRect() on render and on resize -- no hardcoded coordinates or tooltip side, so mobile's stacked card layout doesn't leave a tooltip pointing at empty space. The delta-fade on the highlight box is skipped under prefers-reduced-motion: reduce. Reachable on demand via a new "Replay Tutorial" button inside the issue #17 How-to-Play panel (closes that modal and restarts the tutorial) -- a tutorial a player can't reopen becomes a support question. Closing the tutorial (finish, skip, or replay-then-skip) re-arms the idle nudge from that moment rather than leaving it primed against whatever time had already elapsed while the tutorial was showing; pickIdleNudgeTarget() also now suppresses the nudge entirely while the tutorial is active, matching how it already suppresses during quiz/feedback modals. Verified via Playwright: auto-appears on a fresh context's first combat entry, does not reappear on a second run in the same context, Skip closes it from step 1, Escape closes it from any step (once the focus trap has actually engaged -- trapFocus's own 450ms focus-establishment fallback, not a tutorial-specific issue), Replay Tutorial from the How-to-Play panel reopens it, CAP is unchanged throughout, and it still runs and is still skippable with localStorage's getter made to throw (no uncaught JS errors). Tooltip positioning verified in-bounds at 390x844, 768x1024, and 1920x1080 for all six steps.
…#20) player.score has existed in the session model since the port to Pages Functions but nothing ever incremented it. Adds scoreForAnswer(), VICTORY_BONUS, and hpRemainingBonus() to cf-pages/functions/_lib/game.js as the single scoring implementation, called from the one place combat-action.js already shares for grading both single- and multi-select answers (they'd already drifted once, per the pre-#11 battle-log bug, so this had to not be a second copy): - Correct answer: +100 base - Streak bonus: +25 per consecutive correct beyond the first, capped at +100 (a 5+ streak is the ceiling) - Wrong answer: 0, never negative, streak resets to 0 - A hint used on that question halves its award -- get-hint.js now sets session.pending_q_hint_used on any non-blocked hint grant, which combat-action.js reads and clears when that question is graded - Enemy defeated: +250; surviving to victory adds +1 per HP remaining - difficultyMultiplier defaults to 1x, a hook for issue #9's Easy/ Medium/Hard multipliers rather than combat-action.js reimplementing scoring whenever that lands Score and streak are computed and stored server-side only (session.streak, alongside player.score) -- no client code writes either, so the score can't be edited from the client the way a client-computed one could be. Streak is deliberately reset on every new encounter in start-combat.js/ reset-game.js (a fresh battle's run of correct answers, not one inherited from a fight that already ended), while score is deliberately carried forward across encounters within a session via the existing freshPlayer(existingScore) -- both now called out explicitly in comments rather than left incidental, per the issue's ask. Works for signed-out guests since it all lives in the KV session guests already have. Client: score and streak ride inside combat_state (same pattern as issue #15's action_costs) so no new fetch or call-site plumbing was needed. Added a Score display to the combat HUD with a live "+100" delta annotation and a "Streak x2"+ indicator once it's worth showing. _holoStreak (the existing holo-card shine intensity driver) is now a mirror of the server's streak instead of an independent client-side counter that could drift from it. Found and fixed a real bug while wiring the delta annotation: several call sites (submitQuizAnswer's `finally` block, its multi-select counterpart) re-invoke updateCombatHUD a second time with the exact same already-applied combat_state as a defensive re-render. The #17 Resolve delta got away with a simple value-diff guard against that because Resolve always changes on a graded turn -- score doesn't: a wrong answer is a real, fresh turn that legitimately scores a delta of 0, which a value-diff can't distinguish from "redundant call, nothing changed" and would leave a stale "+100" showing after a miss. Fixed by computing isFreshState once per updateCombatHUD call via object identity against the last-rendered combat_state (a real new server response is always a distinct object; the redundant call reuses window.combatState's existing reference) and gating both the Resolve and score delta annotations on that instead of on value equality. Verified via Playwright: streak-bonus math against real API responses (+100 at streak 1, +125 at streak 2, matching the formula), never a negative delta on a wrong answer, hint-halving via direct API calls (score_delta 50 after a hint on an otherwise-100-point question, resolved against the real corpus for a deterministic correct answer), and that the delta annotation no longer sticks at a stale value across the redundant-re-render turns that exposed the bug above.
Both end screens previously said "Victory!"/"Defeat" and listed questions with a tick or cross -- a receipt, not a result: no score, accuracy, streak, XP, or path to understanding what was missed. combat-action.js computes a run_summary object (score, accuracy, correct/total, best_streak, xp_earned, hints_used, hp_remaining) once outcome is set, extending the existing end-of-run payload rather than requiring a second request. best_streak is new session state, tracked alongside streak and reset per-encounter same as it (start-combat.js, reset-game.js). xp_earned is floor(score / 10) per the scoring issue's own suggested starting divisor. level_results entries now also carry correctAnswer and the question's easy-tier hint (every one of the 800 questions has one) so a missed question's review renders from data already being returned, not a second lookup. Client: one shared renderRunSummary()/handleRunOutcome() replaces three separate copies of the victory/defeat transition logic (in performAction, submitQuizAnswer, and submitQuizAnswerMulti) that had already drifted once -- the single-select path needed its own fix for the 'active'-class toggle that the multi-select path already had. Score is the largest element on the screen (clamp(2.75rem, 9vw, 5.5rem)), above the existing flavour-text heading. Missed questions expand via native <details>/ <summary> (keyboard/screen-reader support from the browser, no custom expand-collapse JS) to reveal the correct answer and hint. renderLevelResults() now runs question/answer/hint text through normalizeMathText (escapes HTML, then adds KaTeX $...$ delimiters, same treatment the quiz modal already gets) instead of interpolating raw corpus text -- unescaped "<" or "&" used to break this markup outright. Added "Play Again (Same Realm)" and "Change Realm" next to the existing "Return to Main Menu". Neither needs a confirmation dialog -- starting a new run from an end screen discards nothing the player hasn't already finished, unlike the nav Home button's mid-run abandonment -- noted explicitly in playAgainSameRealm()'s comment so it doesn't get added reflexively later. Play Again reuses the just-finished run's syllabus_id/ difficulty from window.combatState and calls the existing selectSyllabus() directly; Change Realm shows the syllabus-select screen without a full reset. Both correctly hide the end-screen overlay first (it doesn't clear on its own the way combat-screen's display toggle does). Focus moves onto the summary container (tabindex="-1") when the screen appears, using the same transitionend-or-450ms-fallback pattern as trapFocus's own focus-establishment workaround, since the screen's opacity/visibility change is behind an 0.8s CSS transition. Verified via Playwright: run_summary math against real API responses (a 6-correct-in-a-row victory scored exactly 1280 -- streak-bonus sum 950 + victory bonus 250 + 80 HP remaining, matching the formula by hand); score/accuracy/streak/XP/hints render correctly and match the API response; missed-question review expands to real corpus answer + hint text; focus lands on the summary container after the screen transition; no horizontal overflow at 390px; Play Again returns to combat with score carried forward; Change Realm shows the syllabus grid.
#22) profile.js previously stored one field ({uid, active_game_id}) with its own header comment naming the points economy as deliberately deferred. Extends the schema to lifetime_xp, xp_balance, totals (runs/questions answered/correct), per-realm records (best_score, runs, accuracy inputs, best_streak, keyed by realm name so #9's difficulty tiers can nest under each one later without a migration), and recent_runs capped at 10 (a Firestore document has a 1MB limit; an uncapped array is a slow leak). toFirestoreFields()/fromFirestoreFields() were hand-mapping 2 flat fields -- rewritten as a generic, recursive encoder/decoder (toFirestoreValue/fromFirestoreValue) that handles strings, numbers, booleans, null, arrays (arrayValue), and nested objects (mapValue) for any shape, not just today's. This is what makes the realm-tier-nesting and recent_runs-array requirements possible without hand-writing a converter per field. Verified round-trip correctness (including nested maps/arrays) by mocking fetch and decoding what putProfile() actually sends back through getProfile()'s path -- see backend/test_issue22_firestore_roundtrip.mjs. applyRunToProfile() applies one finished run to a profile in place (pure data manipulation, no I/O) and mergeProfiles() does the guest-to- account merge on first sign-in (sums totals, takes the max of bests, merges recent_runs by recency) rather than an account overwriting a week of guest play. Write timing: combat-action.js writes the profile exactly once, at run end, from the same place run_summary (#21) is computed -- not per answer. XP/score come only from that server-computed run_summary, never the request body, so a client can't post an arbitrary XP total; the Firestore write is wrapped in try/catch so a failure there can't break the run's own response (the player still sees their summary). The client now sends its Firebase ID token on every combat-action call so the server can verify auth.uid matches session.uid before writing. Guests get the identical shape in localStorage (studysaga_guest_profile) via a deliberate line-for-line mirror of applyRunToProfile() in game-simple.js -- there's no shared module system between Pages Functions and the unbundled static frontend, so the duplication is by-hand and called out in a comment on both sides. A new /api/sync-profile endpoint both fetches a signed-in player's profile (so a second device shows the same lifetime totals) and, when passed local_profile, merges it in exactly once right after sign-in -- mergeGuestProfileOnSignIn() only clears the local copy after the server confirms the merge saved, so a network failure leaves it intact to retry rather than silently losing a guest's history. A "My Profile" panel on the main menu (reusing the existing modal/ trapFocus infrastructure) shows lifetime XP, XP balance, total runs, overall and per-realm accuracy, per-realm best score, and the last 10 runs -- reads localStorage directly for guests, calls /api/sync-profile for signed-in players. Security: kept the existing model (client's own ID token, Functions verifies it, no service account). The Firestore rules still give a user full write access to their own document, which profile.js's comment now states explicitly as an accepted tradeoff for a single-player study game with no leaderboards -- the game itself can't be cheated through the normal client since XP is always server-derived from session state, only a player's own private stat display could be edited directly through the Firebase SDK, which only harms their own record-keeping. Revisit before any social comparison feature ships. Verified via Playwright: a full guest run correctly populates localStorage with the right shape and values; the profile panel renders real accumulated data (lifetime XP, per-realm table, recent runs) for guests; recording 12 runs caps recent_runs at exactly 10; /api/sync-profile returns 401 on a missing or invalid token.
XP only counted upward until now. Depends on scoring (#20) and the persistent profile (#22), both shipped earlier this session, and is explicitly gated by its own issue text against trivializing the game before real per-tier enemy scaling (#9) exists -- addressed with a documented, conservative balance target rather than shipping unlimited power creep (see MAX_TOTAL_UPGRADE_LEVELS's comment in game.js): the catalogue's 6 upgrades sum to 23 possible levels, capped at 12 total purchased, so a player can reach roughly half of any one upgrade's ceiling and never max every category simultaneously. Placeholder tuned by inspection, not playtest data -- revisit once #9 ships. Catalogue (game.js UPGRADE_CATALOG, deterministic -- no randomness, no loot boxes, no real money): Neural Capacity (+1 max CAP/level, 5 levels), Resilience (+10 max HP/level, 5), Efficient Recall (+1 Recharge CAP/level, 3), Extra Insight (+1 Simple hint/run/level, 3), Deep Insight (+1 Deep hint/run/level, 2), Focused Strike (+2 Attack damage/level, 5 -- Ability's damage is never upgraded, per the catalogue). Doubling cost curves per upgrade. The real architectural work this issue asked for: action costs/damage, max HP/CAP, and hint budgets were module-level constants (ACTIONS, SIMPLE_HINT_MAX, HARD_HINT_MAX, CONFIG.players.default_kk) read at each use site -- upgrades require them to be per-session. effectiveStats() in game.js resolves a player's purchased levels into actual numbers once, at start-combat time, stored as session.effective_stats and threaded through: freshPlayer() takes it for max_hp/max_cap, actionCosts() for attack damage/recharge gain (both already returned to the client via issue #15's mechanism, so the self-describing buttons automatically show upgraded numbers with no separate client change), and hintsSummary()/get-hint.js's budget check for simple/hard hint maximums. Verified via direct API calls: a player with Resilience x2 + Neural Capacity x1 gets max_hp 130/max_cap 11 (base 110/10); Focused Strike x1 deals 17 real damage in an actual graded turn (base 15); Extra/Deep Insight x1 each raise the hint budget to 4 simple / 2 hard. Where upgrades come from: signed-in players' levels are read server-side from their own Firestore profile via id_token (same accepted tradeoff already documented in profile.js -- this can only affect a player's own run). Guests have no server-side account, so their levels ride in the start-combat request body, same trust model as the rest of the guest profile. New /api/buy-upgrade endpoint validates a purchase entirely against the player's own stored profile -- current level, next cost, the total- level cap, and XP balance are all re-read there, never trusted from the request, so a client can't grant itself an upgrade or spend XP it doesn't have. Both halves of a purchase (XP deduction, level increment) are applied to one in-memory object before a single PATCH write, so a failed write leaves the untouched old profile in place -- atomic by construction, not by a transaction API. Guests buy against their own localStorage balance client-side (buyUpgradeGuest(), mirroring the server logic the same by-hand-sync way applyRunToProfile() already does for issue #22) since there's nothing server-side to validate a guest purchase against. profile.js's schema gains upgrades: { [key]: level }; mergeProfiles() merges it by max-per-key (not summed -- these are permanent levels already paid for once, so a guest-then-signed-in player must not get both sets of levels for the price of one). UI: a card grid matching .syllabus-card's look (name, plain-language effect, level/max, next cost, affordability) reachable from the main menu and both run-summary screens. Unaffordable and maxed cards stay visible with their state shown, never hidden. Buying requires a confirm() dialog -- spending XP is irreversible, unlike issue #21's Play Again/Change Realm which discard nothing and deliberately don't prompt. Verified via Playwright: shop renders all 6 catalogue entries with live cost/affordability; a purchase deducts the correct XP and increments the correct level; the 12-level cap blocks further purchases with the capped-message shown; insufficient XP disables the buy button; /api/buy-upgrade rejects a missing or invalid token.
The corpus has a severe generation-time skew: 88.7% of single-select questions (2641/2976) store the correct answer at options[0]. The server was never shuffling option order before sending a question to the client -- sanitizedOpts was a straight .map() over the stored order -- so every player saw that same skew, live, in production. Reported by the user as "almost all correct answers were choice A." Fixed by shuffling each question's option order once per serve (both places a question is dispatched: the "next question after grading" and "initial question fetch" branches), using the existing shuffle() helper on an index array so the mapping back to the original, unshuffled question.options can be kept. That mapping is stashed as session.pending_option_order and used at grading time to translate the client's submitted answer_index/answer_indices (positions in the shuffled order the player actually saw) back to the corpus's original indices before comparing against the answer key, which is keyed to that original order. Verified via direct API calls across 45 sampled questions (15 games): served option[0] matched the corpus's stored option[0] only 31.1% of the time post-fix (vs. the 88.7% baseline, and in line with chance for mostly-4-to-5-option questions), and all 45 submissions of the corpus-correct answer at its new shuffled position were graded correct -- the translation logic doesn't break grading. This is a data/product-integrity bug, not a regression from this session's other work -- the bias has been live in production since before today. Not yet deployed; needs a manual `wrangler pages deploy` alongside the rest of this session's undeployed work.
…gating #9 (Easy/Medium/Hard difficulty tiers) was only partially done -- the difficulty picker UI and question-tagging existed, but the rest of the issue's scope did not, contrary to an earlier assumption in this session that it was fully shipped. This closes the gap except for the per-question timer, which the issue's own comment thread says is worth splitting into its own issue (no timer of any kind exists anywhere in cf-pages/ today) -- filed separately. Score multiplier: DIFFICULTY_MULTIPLIERS (Easy 1x / Medium 1.5x / Hard 2x) in game.js, wired into scoreForAnswer()'s call site in combat- action.js via session.difficulty (locked in for the encounter at start-combat time). Stated on the difficulty-picker buttons themselves ("2x score") so it's visible before a player commits to a tier, not just reflected in the score afterward. Verified: a correct answer at Hard scored exactly 200 (100 base x 2x). Per-realm-per-tier persistence: profile.js's realm records are now nested by tier (realms.biology.hard.best_score etc.) -- the exact shape issue #22 was deliberately left able to hold without a migration. Each tier record also tracks victories, driving a "Cleared"/"Not cleared" badge per realm x tier in the profile panel. Profiles saved before this nesting existed (a flat record directly on profile.realms[realm]) are migrated to 'medium' the first time they're touched (migrateLegacyRealmRecord(), mirrored client-side for guests) rather than requiring a one-time migration script or silently losing the old data -- verified the migration merges rather than overwrites. Tier-selection gating: /api/syllabi now returns tier_counts and the 15-question floor (MIN_TIER_QUESTIONS in game.js); a difficulty button below the floor is disabled outright instead of silently falling back to the full question pool after a player has already committed to it (that fallback stays as a server-side defensive safety net). All four realms currently sit well above the floor in every tier (lowest is biology/hard at 172), so nothing is disabled in practice today, but the mechanism is real and tested. Docs: added a difficulty tier guidelines section to the README (tier definitions, score multipliers, the question floor) for future question authors, per the issue's acceptance criteria. Verified via Playwright and direct API calls: tier_counts/min_tier_ questions returned correctly; Hard-tier scoring multiplies correctly; per-tier victories/best_score/accuracy track and render correctly in the profile panel with correct Cleared/Not cleared/Not played states; legacy flat realm records migrate into the medium tier and merge (not overwrite) on the next run recorded against them.
Split out of #9's difficulty-modifier scope since no timer, countdown, or deadline logic existed anywhere in cf-pages/ before this. Starting values (tunable in one place, game.js's QUESTION_TIME_LIMIT_MS, not playtested yet): Easy 30s, Medium 20s, Hard 12s. Server: computed once per question serve (both the "next question" and "initial fetch" branches in combat-action.js) from the question's difficulty tag, stored as session.pending_q_deadline, and returned to the client as question.time_limit_ms. At grading time this is enforced as a grace-padded (3s) backstop -- forces isCorrect = false if a real answer somehow arrives after deadline+grace -- consistent with scoring being server-authoritative everywhere else in this codebase (#20) and never trusting client timing for the actual grade. Client: a countdown bar + numeric text in the quiz modal (startQuizTimer()/stopQuizTimer(), cleared on every modal close path so a stray timer can't fire against a question the player already answered). On expiry, auto-submits an empty answer through the exact same submitQuizAnswer()/submitQuizAnswerMulti() path a real click uses -- null/[] never matches the answer key, so no separate "timed_out" signal needs to travel to the server, and CAP is deducted exactly like any other wrong answer (existing "wrong answer still costs CAP" rule, issue #15, applies unchanged). Countdown bar sweep is a CSS transition disabled under prefers-reduced-motion; the numeric seconds-remaining text is the primary signal either way. Found and fixed a message-framing bug while verifying end-to-end: the "Time's up!" message (vs. plain "Incorrect") was keyed only to the grace-padded timedOut check, which the client's own auto-submit almost always lands inside (it fires right at the nominal deadline, which the grace window is specifically there to treat leniently) -- so a genuine timeout would silently read as an ordinary wrong answer. Fixed by keying the message on whether an answer was actually submitted at all (only ever empty via the auto-submit-on-expiry path), keeping timedOut purely as the scoring backstop the two checks were never meant to share. Verified via Playwright: time_limit_ms served correctly per tier (12s for Hard); the countdown UI renders and updates; a real, un-answered 12s Hard-tier question auto-closes the quiz modal, deducts CAP exactly like a wrong answer, and shows "Time's up!" with the correct answer text.
cf-pages/functions/_lib/data.json becomes the single source of truth for the question corpus -- it's what's actually served to players (3,152 questions vs backend's ~800) and already carries bug fixes that never made it back into backend's copy. backend/data.json is now a generated mirror for the local Flask dev server, never edited directly. - merge_corpus_reconciliation.py (one-time): pulled forward 10 questions' better-written hints from backend/data.json before retiring it to a mirror; everything else backend/data.json had that differed (difficulty tags from a classification pass that found zero Hard-tier questions anywhere) was deliberately not imported - sync_corpus.py (ongoing): regenerates backend/data.json from the master; CI's new corpus-drift job runs it and fails the build if the committed mirror doesn't match - validate_corpus.py: added the 3 checks issue #24's acceptance criteria call for that were missing -- duplicate question text within a realm, duplicate option text within a question, and answer_index/answer_indices range + consistency with isCorrect flags - README.md / backend/README.md: documented the master/mirror relationship and indexed the two new scripts Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Merges Eyal's 800-topic tiered corpus (200 topics x 4 subjects) into the production data.json served by cf-pages Functions, adding 2,400 questions across 4 new subjects. Content went through a multi-round independent Gemma re-verification after an initial 47% defect rate was found in what Gemini had scored as clean (dominant issue: hint-tier inversion). Final state: 0 flagged topics, 91.3% of individual hint scores high-quality. Existing biology/math/chemistry/physics content is untouched. Feedback text ships in its verified plain-educational tone rather than the other subjects' combat-flavor voice, to avoid layering new unaudited content on top of what was just verified. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
1 of 3 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This branch is 30 commits ahead of
mainand has diverged significantly in both game features and corpus content. Highlights relevant to review:cf-pages/functions/_lib/data.json). This content went through multi-round independent Gemma re-verification after an initial 47% defect rate was found in what Gemini had scored as clean (dominant issue: hint-tier inversion). Final state: 0 flagged topics, 91.3% of individual hint scores high-quality.main's corpus for the original 4 subjects (biology/math/chemistry/physics) is only 200 questions each (800 total) — much smaller than what's already on this branch (766-798 each, 3,152 total) before the new subjects are even added. That expansion happened across this branch's history and never made it back tomain. Reviewers should be aware this PR brings substantially more than just the 4 new subjects — merging it means adopting this branch's corpus and feature state as the new baseline.Content verification detail (new subjects)
backend/validate_corpus.pypasses: 5,552 questions across 8 realmsbackend/data.jsonmirror regenerated viabackend/sync_corpus.pyto keep CI's corpus-drift check greenTest plan
get-hint.js's exact-text match againstdata.json)🤖 Generated with Claude Code