Shakespearean nickname generation with source-grounded provenance.
Status: 🔵 Research
Live Demo · Methodology · Source Data · Contributing
Fuse the tavern vernacular of Eastcheap with the royal iconography of Agincourt to construct a series of online handles. Each name must be fully parsed: tracing its constituent roots, character ties, Shakespearean plays, semantic registers, morphological synthesis, and period connotations, while explicitly designating the generated term as a neologism.
sack + agincourt → sackagincourt → Sackagincourt1415
The browser is an Early Modern composing room: classical serif text, monospace parameters and provenance, paper white, black ink and dark red. Use FORGE NAME, inspect the paper trail, follow the four-stage formation, and export a JSON record. Replay links preserve the settings and selected impression. Everything runs locally in the browser; there is no generation API or API key.
- Source evidence: 195 indexed words in the current three-play atlas, with speaker identities, printed line labels, speech/word XML IDs and source hashes.
- Semantic affiliation: add-half smoothed log odds contrast selected tavern and royal speaker groups. Royal affinity controls weighted ingredient selection.
- Ranked fusion: compare overlap, concatenation and vowel-group cuts using explicit retention, consonant-cluster and length scores; inspect alternatives.
- Character Markov: order-three conditional generation with real prefix conditioning. Browser records include transitions, probabilities and fallbacks.
- Chronology: distinguish story history, biography, composition context and publication; 1413 is coronation, 1415 is Agincourt, 1623 is the First Folio.
- Reproducibility: seeded generation, structured provenance, Python/Node tests, browser integration checks, source manifests and a documented rebuild process.
This is a literary experiment, not a historical pronunciation model. Speaker groups are editorial interpretations; rare-word associations are uncertain. Evidence attests an ingredient, not the invented name. See methodology.
Python 3.9+; no third-party runtime packages.
git clone https://github.com/styayur/Nickspeare.git
cd Nickspeare
python -m pip install -e .python -m nickspeare generate -n 10 --seed 42
python -m nickspeare generate -n 5 --rule rogue_to_king --style title --year-theme agincourt --royal-affinity 0.85 --json
python -m nickspeare generate -n 5 --rule markov_coin --suffix none --seed 42
python -m nickspeare generate -n 5 --rule markov_epithet --seed 42Python offers rogue_to_king, quote_splice, archaic_coinage, markov_coin
and markov_epithet. Browser offers the first four. --suffix none never adds
a year; number denotes a nonhistorical ornament; both randomly chooses a
historical or numeric suffix. Styles: lower, camel, title, snake, kebab,
upper. Separators transform existing phrase boundaries, not guessed compound roots.
from nickspeare import NicknameGenerator
engine = NicknameGenerator(seed=42, royal_affinity=0.85)
name = engine.generate_one(rule="rogue_to_king", year_theme="agincourt")
print(name.variant)
print(name.as_dict()["provenance"])Python and browser share lexical/evidence data and blend scores but use different
RNGs and transformation rules. Seeds replay within an engine/version, not across
engines. Word Markov trains on individual quotes by default; add --corpus PATH
for Folger XML or line-oriented plain text. Supplemental corpora train the model
but do not silently extend the committed source-evidence index.
python -m nickspeare data
python scripts/build_features.py
python scripts/build_atlas.pyThis rebuilds features and the identical Python/browser atlas from cached Folger
XML for 1H4, 2H4, H5. Full texts stay in ignored .cache/. The source manifest
pins input bytes with SHA-256. Missing source files fail the build.
In RStudio run source("R/setup.R") to install missing xml2 and jsonlite
packages. Then run from a terminal:
Rscript R/extract_features.R --out work/r-featuresThe R pass uses shared exclusions and speaker cohorts and the same smoothed log-odds formula for characteristic words. Its exploratory outputs are kept separate from the published atlas. The Python atlas builder adds evidence records and source hashes. R is optional for nickname generation.
Adding a play is an explicit public workflow, not an undocumented edit to JSON. The current three-play atlas, 195 indexed words, source rights, hashes, exclusions, rebuild commands, and acceptance rules are documented in docs/corpus-extension.md. Full source texts remain ignored; only reviewed evidence and generated atlas data may be committed.
python -m pip install -e '.[dev]'
python -m unittest discover -s tests -v
npm test
python -m buildServe the site with python -m http.server 8765 --directory docs. After
python -m playwright install chromium, run python scripts/check_browser.py
in another terminal. CI runs Python 3.9/3.12/3.14 and headless browser checks.
GitHub Pages publishes main:/docs.
| Path | Responsibility |
|---|---|
nickspeare/generator.py |
Generation, suffixes and structured records |
nickspeare/corpus.py |
Markov engines and TEI parsing |
nickspeare/provenance.py |
Evidence lookup, blend ranking, date explanation |
nickspeare/data/provenance.json |
Packaged evidence atlas |
scripts/build_atlas.py |
Reproducible source evidence build |
R/ |
RStudio dependency setup and extraction |
docs/ |
Static Pages UI, browser engine and shared atlas |
tests/ |
Python and JavaScript regression tests |
Code: AGPL-3.0-or-later, see LICENSE. The cached Folger edition and its derived data retain CC BY-NC 3.0 attribution and restrictions; they are not relicensed by the code license. See DATA_SOURCES.md.
Thanks to the Folger Shakespeare Library and its editors/encoders. Optional corpora include Project Gutenberg and Karpathy’s tiny Shakespeare. Markov techniques are inspired by markov_poem and namemaker.
- Source-grounded nickname generation with per-name provenance on the web demo.
- Extend the corpus and improve the evidence/provenance UI.
- More computational-humanities experiments and comparative provenance views.
- Ungrounded generation or commercial name services.
- GitHub Issues: reproducible bugs and scoped feature proposals.
- GitHub Discussions: not enabled; use Discord for design discussions and early feedback for now.
- Discord: join the community for informal discussion and feedback. It is not an SLA support channel.
- Security: follow SECURITY.md; do not open a public issue for a vulnerability.
- Contributing: CONTRIBUTING.md.
- Releases use
vX.Y.Ztags and publish verified Python distributions withSHA256SUMS.txt; maintainers perform releases.
