Statistical timing analysis of congressional stock-trade disclosures.
Public records in. Permutation tests, FDR correction, and a cryptographic evidence chain in the middle. Exactly one kind of claim out.
The landing page: what the corpus holds, the disclosed volume as a statutory range (never a midpointed number), and the headline result β 2 of 233 tested members, with the distribution of corrected p-values beside it. The one chart is an emphasis chart because the story is a single bin; its palette is CVD-validated, not eyeballed, and every chart on the site ships a table view so no value is hover-only.
Members of Congress must disclose their trades under the STOCK Act. CapitolWatch ingests those disclosures β plus roll-call votes, committees, bills, and market calendars β and asks one narrow, answerable question per member:
Were this member's trades timed unusually close to their chamber's votes, compared with their own trading history re-dealt at random 10,000 times?
For any flagged trade, the system says one of exactly two things:
- π "Disclosed outside the statutory 45-day window β filed N days after the trade." Pure arithmetic on the filing's own dates. The current record holder: 3,698 days late β a 2015 trade disclosed in 2025.
- π "Timing statistically unusual relative to [a specific named vote], corrected p = X." A permutation test with BenjaminiβHochberg false-discovery-rate correction β raw p-values are structurally incapable of reaching the dashboard (the flag store has no column for them).
It never claims intent β in either direction. No member is accused; no member is endorsed. That discipline is mechanically enforced: the test suite scans every user-facing string for verdict vocabulary, accusatory and exculpatory, and fails on a hit.
- π³οΈ 72,856 transactions extracted from 10,651 real filings β the full electronic-PTR era, 2012β2026, both chambers β House PDFs parsed by word geometry across three template generations (the 2013β2017 era has no end-of-table marker and unpadded dates; both fixture-pinned), Senate HTML by the stdlib parser, scanned paper filings routed to human review instead of guessed
- π¬ Member-level permutation test (10,000 resamples of market days within each member's own trading window) β the null model survived two discarded calibration runs that each looked publishable and were each wrong
- π§Ύ One-query evidence chain: every flag β SHA-256 of the original filing bytes, extraction counts, the named vote, test parameters, and the RNG seed
- βΏ Beyond equities: crypto holdings (Bitcoin, Ethereum, Solanaβ¦) resolve to their USD spot pairs and are priced and counted alongside stocks, and renamed tickers (FBβMETA, SQβXYZ) recover pre-rename history β while the timing null is deliberately kept on the equity trading calendar so a 7-day market can't reintroduce a weekend confound
- β Validated four independent ways: synthetic ground truth through the real pipeline, a pinned known-cases regression set, an archived independent dataset (House 2021 agreement: ratio 0.99), and full-corpus re-extraction from original bytes in which 99.66% of electronic filings validate against the count the filing itself states β the 26 that don't are quarantined for review, never silently ingested
- β‘ Caching that's proven, not claimed: full cold build 14,612s β full re-run 1,760s with zero re-downloads and zero re-OCR
- π Plain, stats-first dashboard: a read-only UI β colored data ink on monochrome chrome β whose Analysis view charts yearly purchase/sale trends, top tickers, member leaderboards, a digital-assets-by-token breakdown, and the corrected-p distribution across every tested member: the significant members are the one highlighted bar, the rest cluster at chance
- π§ͺ 95 fixture-based tests, no network, no database β
pytestruns anywhere, including CI
| Filings processed | 10,651 β 7,634 machine-extracted and count-validated, 3,040 paper β review by design, 26 electronic one-offs β review |
| Transactions | 72,856 across 359 members (2012β2026) |
| Vote calendar behind the test | 14,588 roll-call votes, 112thβ119th Congresses β 99.99% of member trades pair with an own-chamber vote window |
| Members tested / excluded | 233 (69,016 trades) / 126 insufficient data β an exclusion, not a result |
| Statistically unusual after BH correction | 2 members (corrected p 0.0233, 0.0349) |
| Disclosure-timing flags / event-timing flags | 12,777 / 881 β all 13,658 with full evidence chains, stale flags from superseded runs retracted by design |
| Market data behind the calendar | 5,584,047 daily closes, 2,450 tickers (incl. 8 crypto USD pairs), trading-day calendar back to 2006 |
| Independent cross-check (House 2021) | ours 5,541 vs archived 5,625 β 0.99 |
Every number is reproducible β docs/METHODOLOGY.md pairs
each with its command.
The test result as the dashboard shows it: 233 members tested, 2 with
statistically unusual timing after BH correction (the highlighted low-p bar), the
rest clustered at chance. Corrected p-values only β the raw ones can't reach this
page. Two members flagged on the shorter 2021β2026 corpus no longer are, for two
different reasons: one gained 1,134 trades that diluted the effect until its raw
p rose 248Γ; the other gained only 14 trades and moved mostly because the
multiple-testing batch grew from 158 to 233. scripts/run_diff.py attributes
every status change to the mechanism that caused it.
Descriptive leaderboards: most disclosed transactions and largest disclosed volume. Volume bars draw the floor of each member's summed statutory ranges; the full minβmax band is always a range, never a midpoint.
House Clerk PDFs Senate eFD HTML Congress.gov + GovTrack Yahoo daily closes
β β β β
βΌ βΌ βΌ βΌ
βββββββββββββββββββββββββββββββββββ ββββββββββββββββββββββ ββββββββββββββββββββ
β Extraction + count validation β β Legislative cache β β Validated price β
β (geometry parse β independent β β (members/committeesβ β cache (poisoning β
β count must agree β or review) β β bills/votes) β β guard) β
ββββββββββββββββββ¬βββββββββββββββββ βββββββββββ¬βββββββββββ ββββββββββ¬ββββββββββ
βΌ βΌ βΌ
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β PostgreSQL β canonical store + audit trail (extraction_runs, stat_runs, flags) β
β entity resolution: nickname-normalized fuzzy match, empirical thresholds, β
β ambiguity guard (ask the two Robert Menendezes) β
ββββββββββββββββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββββββββββββ
βΌ
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Stats engine: per-member permutation, 10k resamples, β
β BH-FDR across the batch, seeded + persisted β
ββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββ
βΌ
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Flag store β corrected p only, one-query evidence β
β chain β read-only FastAPI β React dashboard β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Design rationale, and everything the plan got right or wrong, in
docs/ARCHITECTURE.md.
Prerequisites: Python 3.12 Β· Node 18+ Β· Docker (Compose) Β·
brew install tesseract (or apt install tesseract-ocr) Β· a free
Congress.gov API key
git clone https://github.com/rushilrawat/CapitolWatch.git && cd CapitolWatch
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .env # add your CONGRESS_API_KEY
make up # Postgres 15 via Docker Compose
python -m db.migrate
# project rule: always dry-run a small window before a historical backfill
python -m scripts.backfill --disclosures --dry-run
python -m scripts.backfill --legislative --disclosures --process # hours; resumable
python -m uvicorn api.main:app # read-only API on :8000
cd dashboard && npm install && npm run dev # dashboard on :5173 (proxies /api)make test # 95 tests β fixtures only, no network, no database
make lint # ruff + black --checkEvery stage is independently resumable (manifest dedup, content-hash extraction cache, validated price cache) β an interrupted backfill re-runs safely, and the second pass doubles as the caching proof.
| Check | What it proves | Command |
|---|---|---|
| Synthetic ground truth | Random-timing members through the real pipeline β 0 flags; one injected on-vote-day member β flagged alone | python -m scripts.checkpoint_b_run |
| Known-cases regression | Three publicly reported disclosure-timing cases (532/200/918-day gaps) re-found from our own scrape, every checkpoint | tests/known_cases.yaml |
| Independent cross-check | Per-year counts vs the archived Stock Watcher datasets β House 2021 ratio 0.99 | python -m scripts.crosscheck_stockwatcher |
| Full-corpus re-validation | All 10,654 filings re-extracted from original bytes: 7,634 pass, 2,994 paper β review by design, 26 electronic failures quarantined (99.66% of electronic filings validate) | python -m scripts.checkpoint_a_run --skip-scrape |
| Evidence-chain audit | Sampled flags re-derived end to end β hash, extraction, every stored number | python -m scripts.validate_pipeline --n 12 |
| Run-to-run accountability | Every member whose status changed between two analysis runs, attributed to the mechanism that moved it β own data, own raw p, or the size of the multiple-testing batch | python -m scripts.run_diff --old 7 --new 14 |
The bugs these checks caught β and they caught real ones, repeatedly β are the
subject of docs/ENGINEERING_NOTES.md, each anchored
to its commit hash.
scrapers/ House/Senate disclosure scrapers + Congress.gov/GovTrack clients
extraction/ PDF geometry parser, Senate HTML parser, OCR fallback, count validators
db/ migrations, ingest pipeline (hash-cached, self-healing retries)
market_data/ price client (validated cache), event-window construction
stats_engine/ permutation engine, BH-FDR, seeded + persisted runs
api/ read-only FastAPI (corrected p only β enforced by schema)
dashboard/ React + Vite + Recharts (members, flags, evidence chains, analysis, methodology)
scripts/ backfill runner, checkpoint runners, pipeline validator, cross-check
tests/ 95 fixture-based tests incl. the Rule-1 copy scan (no network)
docs/ methodology, engineering notes, architecture, scalability
In value order β details in
docs/ARCHITECTURE.md β What's next:
- Bill-subject / sector-company linkage β the big statistical upgrade: connect a pharma trade to a pharma bill, not just to a session day
- Prediction-market event sources β legislation-linked markets (Kalshi, Polymarket) as dated, graded salience events (2024+ coverage only)
- Options & derivatives treated distinctly (asymmetric bets carry more signal)
- Hearing transcripts / witness lists as an event source
- A crypto-aware null model β a member-specific 7-day exchangeability set so digital-asset trades are tested on their own calendar, not the equity one (crypto is already ingested and priced; only the null is equity-shaped today)
- Extend the verified rename/override table beyond FBβMETA and SQβXYZ to the remaining delisted symbols; historical committee memberships
Contributions are welcome β CONTRIBUTING.md is short and the
rules that matter are mechanical (the suite enforces them). Please also see the
Code of Conduct and Security Policy. Wrong
numbers are treated as seriously as crashes: if a filing says one thing and
CapitolWatch says another, that's a bug report we want.
MIT. All input data is public record, fetched with rate limits and a
descriptive User-Agent, and every immutable rule this project runs under is public
too β CLAUDE.md. The seven rules exist because this system names
real, identifiable public officials using their own legally mandated disclosures:
statistical timing claims only, ranges never midpoints, corrected p only,
count-validated extraction, polite scraping, fixture-tested stages, and no invented
data β ever.