Financial features as they were known when a decision was made.
Open the online report · 中文说明 · Methodology · Data contract · Interview notes
An invoice dated January can be published in February, arrive in March, and be corrected later. Joining on January's date alone can put future knowledge into an earlier credit decision. PITBridge resolves event time, publication time, ingestion time and observation revisions before exporting a feature snapshot.
git clone https://github.com/dev-belly/PITBridge.git
cd PITBridge
python -m pip install -e .
pitbridge demo --out outputs
pitbridge verify --out outputs
python -m unittest discover -s tests -vOpen outputs/report.html in a browser. The report is self-contained, works offline and includes a filter for changed selections. Python 3.11+; no third-party runtime dependencies.
# Bring your own observations, decisions and feature specs:
pitbridge build --inputs demo/inputs.json --out outputsThe deliberately small synthetic fixture is hand-auditable: 11 observations, 15 decisions, 45 feature lookups across tax, bank and utility sources.
| Saved result | Count | Interpretation |
|---|---|---|
| Unsafe selections using future knowledge | 11 | The event-only baseline selected a record unavailable at decision time. |
| Selections changed by the full contract | 16 | Includes freshness and tombstone effects as well as future knowledge. |
| Selected / missing feature rows | 16 / 29 | Missingness is preserved with a reason, never silently replaced by zero. |
Inspect the comparison, record-level lineage, normalized inputs, and summary. These counts demonstrate the fixture; they are not a population leakage rate.
- Require
event_at <= decision_atandmax(published_at, ingested_at) <= decision_at. - For each entity/source/feature/event, select the highest known revision, even if a lower revision arrives later.
- Apply known tombstones to that event. Do not resurrect its earlier revision.
- Select the latest remaining event within the feature's inclusive freshness window.
- Preserve a row for every decision/spec pair, including
no_history,not_available,deletedandstale.
The production path is a SQLite window-function join. A separate Python enumerator checks its results, including randomized histories and append-future invariance. Read the implementation.
| Artifact | Purpose |
|---|---|
inputs.json |
Canonical, normalized source records, decisions and feature contracts. |
snapshots.csv |
Values, missingness, selected record IDs, revisions and timestamps. |
comparison.csv |
Explicitly unsafe event-only baseline; never used as training features. |
summary.json, report.html |
Derived counts and an offline inspection report. |
manifest.json |
SHA-256 hashes of all five artifacts. |
Verification checks hashes and reruns both algorithms, rejecting a fabricated summary even if its hash was updated. Hashes are not signatures and cannot establish that an external source is truthful.
Rolling sums, observation means and counts resolve availability, revisions and tombstones before aggregating an inclusive event-time window. Every result retains all contributing records; a separate Python temporal enumerator checks membership and values.
pitbridge rolling-demo --out outputs/rolling
pitbridge verify-rolling --out outputs/rolling
pitbridge aggregate --inputs demo/rolling/inputs.json --out outputs/customOnline rolling case · Contract and worked example · Feature rows · Membership
This is a reference implementation for scalar financial observations and explicit rolling aggregates. It does not provide streaming ingestion, access control or label generation. Version numbers and trustworthy availability timestamps must come from the upstream source contract. SQLite is intentionally inspectable; distributed-scale performance has not been benchmarked. Rolling means are observation means, and counts do not establish complete business activity.
PITBridge complements CreditVintage's application-time model evaluation. The repositories are separate components; no integration is claimed. MIT license.
