Skip to content

Latest commit

 

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Paperflow

A zero-cost macroeconomic intelligence platform running entirely on your local machine — 11-module sovereign rating engine (GFCR v6.0), free-data ingestion, and LLM-generated policy narratives.

Python License Engine Data

Live data stack (verified 2026-08)

Provenance statement: the databases shipped in data/warehouse/ were populated by a live fetch run on 2026-08-11 with real World Bank / IMF / OECD / PWT / OWID / ITU / UN DESA data (266,779 rows, zero seed_demo). Full source list, row counts, audit counts and reproduction steps: data/DATA_MANIFEST.md.

paperflow --pipeline (or --fetch then --rate) pulls real data from free public sources into macro_indicators, then runs the GFCR engine. The warehouse ships with a fresh ingest; no demo data is written in live mode (--seed --force seeds reference data only, and --fetch replaces the table with API data).

Source Live rows (2026-08 run) Feeds
World Bank WDI ~129,000 Debt, deficit, CA, reserves, unemployment, demography, CO2, energy, Gini
IMF WEO Datamapper ~73,000 GGXWDG debt, primary balances, GDP, inflation, savings (incl. projections)
Penn World Table 10.01 ~33,000 TFP, GDP/hour, employment, hours
OWID (Our World in Data) ~19,000 CO2 intensity, renewable share
ITU ~11,600 ICT development index
UN DESA 320 (40×8 vintages) Migration stock → per-100k skilled stock
OECD ~250 R&D / GERD
GitHub proxy 40 Developer-density proxy (static)

Honest degradations (by design, all visible in the Data Audit screen):

  • WGI governance indicators are archived from the free WDI endpoint, so the 6 governance inputs fall back to the seeded/static values + estimators.
  • ILO ILOSTAT SDMX is Cloudflare-blocked from most datacenter IPs (and its endpoint moved); core labour indicators are covered by World Bank WDI fallbacks, the rest by documented estimators.
  • WIPO patents have no stable bulk API; curated annual counts from WIPO IP Facts & Figures are used.
  • Every indicator value is recorded in data_audit with source, fallback and imputation flag, so "real" vs "estimated" is always inspectable.

The rating pipeline (--rate) builds the GFCR 11-module panels and writes a full multi-year vintage (the repo ships 2020, 2022, 2024, 2026 rounds) so sparklines and momentum signals see a real history. A single round takes ~2 minutes; scripts/build_multi_year_rounds.py builds any year list.

Architecture

┌─────────────────────────────────────────────────────────────────┐
│                    PAPERFLOW                                     │
├─────────────────────────────────────────────────────────────────┤
│  INGESTION LAYER  →  WAREHOUSE  →  ANALYTICS  →  AI NARRATIVE │
│                                                                  │
│  ┌──────────┐   ┌──────────┐   ┌──────────┐   ┌──────────────┐  │
│  │ Free APIs │ → │ SQLite   │ → │ Rating   │ → │ DeepSeek    │  │
│  │ Scrapers  │   │ + Parquet│   │ Engine   │   │ V4 Flash    │  │
│  │ Bulk CSV  │   │ (Local)  │   │ (Python) │   │ (Local LLM) │  │
│  └──────────┘   └──────────┘   └──────────┘   └──────────────┘  │
│       ↑                                              ↓           │
│  ┌──────────┐                              ┌──────────────┐    │
│  │ Scheduler│                              │ Streamlit UI │    │
│  │ (Cron)   │                              │ (Localhost)  │    │
│  └──────────┘                              └──────────────┘    │
└─────────────────────────────────────────────────────────────────┘

Two Core Purposes

Purpose A: Country Macro Rating — GFCR v6.0

GFC Rating (GFCR v6.0) — an 11-module sovereign capacity engine that replaces the legacy 5-dimension weighted score:

# Module Weight Core indicators
1 Fiscal Stress 0.18 Net debt, deficit, Bohn coefficient, monetary–fiscal regime, liquidity
2 External Vulnerability 0.15 CA balance, net external debt, original sin, reserves, energy
3 Financial Vulnerability 0.15 Banking sector size, household/corporate debt, doom-loop, shadow banking
4 Strategic Capacity 0.12 Governance factor (V-Dem/TI/EIU), crisis response, network centrality, GPR (Markov)
5 Demographic Stress 0.10 Old-age dependency (now & projected), TFR, migration buffer, pension burden
6 Climate Transition Risk 0.08 Physical damage risk, stranded assets, adaptation, policy credibility
7 Technology Risk 0.08 R&D, AI readiness, semiconductors, patents, digital infrastructure
8 Supply Chain Risk 0.07 Food/water/mineral/semi/pharma import dependence, concentration HHI
9 Resilience Buffer −0.07 Employment, labor flexibility, trust/social capital, fiscal space, innovation
10 Dynamical Stress 0.14 Kalman-filtered debt velocity, productivity acceleration, IQ/export/tech trends
11 Interactions — Tensor-product terms: doom loop, aging–fiscal spiral, climate–supply cascade

Methodological core:

  • Johnson SU transform — the empirical S_eff distribution is mapped through a fitted Johnson SU CDF to a 0–100 score (degenerate fits fall back to normal).
  • REML hierarchical normalization — every indicator is decomposed into a global mean + peer-group random effect via statsmodels MixedLM (peer groups: AE / EM / LIDC / FC_overlay), so countries are compared within income peers.
  • Kalman dynamical stress — debt velocity, productivity acceleration and trend terms are smoothed with a Kalman filter over ≥10 years of history (demo dataset seeds 2016–2026 to satisfy MIN_HISTORY).
  • Delta-method uncertainty — module sensitivities × residual stds produce a 90% credible interval per rating (t-distribution, df=10).
  • Offline-friendly — every proprietary indicator (BIS, CPIS, FSB, V-Dem, ND-GAIN, IISS, AI-readiness) has a documented estimator + static seed.

The legacy 5-dimension engine (fiscal 30% / productivity 25% / talent 25% / structural 15% / external 5%) is preserved as engine="legacy" and the old country_ratings table is still written by both engines.

Purpose B: Historical Policy Effectiveness

For any country, detect when major reforms occurred and analyze macro indicators 5 years before vs. 5 years after, with LLM-generated narratives.

Tech Stack

Layer Tool
Language Python 3.11+
DataFrames Polars (not Pandas)
Database SQLite3
HTTP httpx + aiohttp
LLM DeepSeek Coder (local)
UI Streamlit
Caching diskcache
Visualization Plotly
Export xlsxwriter, fpdf2

Quick Start

cd paperflow
python -m venv venv
source venv/bin/activate
pip install -r requirements.txt

# Run the full pipeline (seed + live data fetch + rate) — LIVE data mode
python -m src.cli

# Offline demo mode (synthetic seed data, no API calls):
python -m src.cli --demo

# Rate with the legacy 5-dimension engine instead of GFCR v6.0:
python -m src.cli --rate --engine legacy

# Run the UI
streamlit run src/ui/app.py

# Copy ALL source code to clipboard / stdout:
python -m src.cli --code copy          # or: paperflow --code

# Copy ONLY a section's code (smaller, faster, focused exports):
python -m src.cli --code copy database # DB layer incl. the GFCR database code + schema DDL
python -m src.cli --code copy gfcr     # GFCR v6.0 rating-engine code only
python -m src.cli --code copy database gfcr tests   # any combination
python -m src.cli --code copy list     # list every section

# Build + sync the dedicated GFCR database (data/warehouse/gfcr.db):
python -m src.cli --gfcr-db

⚠️ Live vs demo data: Paperflow does not ship in demo mode by default. The default paperflow run fetches fresh data from all public APIs (World Bank, IMF, OECD, PWT, WIPO, ITU, UN DESA). To work fully offline with the bundled synthetic 2016–2026 demo dataset, run paperflow --demo explicitly.

Data Sources (All Free)

  • World Bank WDI API — 1,400+ indicators
  • IMF WEO API — GDP, debt, inflation
  • OECD SDMX — Productivity, employment
  • UN DESA — Migration stocks
  • Penn World Table 10.01 — TFP, capital stock
  • WIPO — Patent statistics
  • ITU — ICT/digital data
  • GitHub API — Developer population proxy

GFCR Database & Sectioned Code Export

Dedicated GFCR database (data/warehouse/gfcr.db)

The GFCR v6.0 formula stores all of its data in a dedicated SQLite database (src/db/gfcr_db.py):

Table Contents
fiscal_stress … dynamical_stress the 10 component panels fed to the scorers
gfcr_static static indicator seeds (BIS, CPIS, FSB, V-Dem, ND-GAIN, IISS…)
peer_groups AE / EM / LIDC / FC_overlay for hierarchical normalization
countries reference panel (iso3 → income tier)
country_ratings_gfcr formula output per country-year
gfcr_formula_meta the formula itself: weights, tiers, fitted Johnson SU params

It is created automatically by paperflow --rate (engine gfcr) and can be built/synced manually with paperflow --gfcr-db. Given gfcr.db alone you can re-run the GFCR computation formula.

Sectioned code copy

paperflow --code copy can export only the code of one section — useful when working on a single part of the project (smaller clipboard, faster LLM context):

paperflow --code copy            # everything (backward compatible)
paperflow --code copy database   # DB layer: schema, connection, seed, GFCR DB code + DDL
paperflow --code copy gfcr       # GFCR v6.0 engine only
paperflow --code copy ingestion  # API clients only
paperflow --code copy etl llm ui # any combination of sections
paperflow --code copy list       # show all sections

The database section includes the GFCR database file path and its full table DDL, and is self-contained (it also ships config.py + peer_groups.py, the runtime deps of the GFCR store). For just the GFCR database code itself use the dedicated gfcr-db section:

paperflow --code copy gfcr-db       # gfcr_db.py + schema/connection/seed + config + peer_groups + DDL
paperflow --code copy database      # full DB layer (superset of gfcr-db)

Available sections: database, gfcr-db, gfcr, analytics, ingestion, etl, llm, policy, ui, cli, config, tests, scripts, readme (aliases like db/gfdb/engine also work).

Project Structure

paperflow/
├── src/
│   ├── config.py           # Global config, GFCR weights/tiers, thresholds, DB paths
│   ├── db/                 # Schema, connection, seed + gfcr_db.py (dedicated GFCR store)
│   ├── ingestion/          # API clients, data fetchers, scheduler
│   ├── etl/                # Mappings, estimators, builder (legacy + GFCR panels)
│   ├── analytics/          # Rating engines
│   │   ├── gfcr/           # GFCR v6.0: 10 module files, interactions, uncertainty, composite
│   │   ├── composite.py    # Dispatcher: engine='gfcr' | 'legacy'
│   │   ├── composite_legacy.py  # Pre-v6.0 5-dimension engine (preserved)
│   │   ├── hierarchical_norm.py # REML peer normalization + Johnson SU
│   │   └── peer_groups.py  # AE / EM / LIDC / FC_overlay assignments
│   ├── policy/             # Policy event analysis
│   ├── llm/                # DeepSeek narrative generation
│   └── ui/                 # Streamlit dashboard (GFCR module radar)
├── data/                   # Local data warehouse (paperflow.db + gfcr.db)
├── tests/                  # Test suite (incl. tests/test_gfcr_integration.py)
└── scripts/                # Automation scripts

🧩 Related Projects

Project Description
mtui Bloomberg-style keyboard terminal for Paperflow's GFCR v6.0 ratings
papertrail 13F hedge fund cloning & capital-flow intelligence pipeline

📄 License

MIT.

About

Sovereign ratings on real data — GFCR v6.0 11-module engine, live World Bank/IMF/OECD/PWT/OWID/UN ingestion, 2020-2026 rating vintages, Streamlit + ETL scheduler

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages