Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
74 changes: 72 additions & 2 deletions docs/data_developer_guide.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Data Developer Guide

This guide explains how to add new datasets and ETL processes to your catalog repositories.
This guide explains how to create, maintain, add update datasets within a **CFA DataOps** catalog repository. A catalog is a Python package that contains one or more datasets, each with configurable TOML-based ETL workflows. This guide is intended for **dataset developers** who author ETL pipelines, add new dataset versions, and manage catalog content, schemas, and validation logic.

> **Prerequisites**: You need to have a catalog repository created and installed. See [Managing Catalogs](managing_catalogs.md) for setup instructions.

Expand All @@ -23,6 +23,22 @@ The ETL pipeline system is built around:
- Python ETL scripts that handle extraction, transformation and loading
- SQL templates for transformations (optional)
- Schema validation using Pandera
- Catalog repository content (datasets/, reports/, workflows/, etc.)
- datacat; the runtime dataset interface used to inspect, validate, and load dataset versions

## Key directories for developers:

### `datasets/`
contains TOML files defining dataset ETL pipelines, metadata, validation rules, and staging behaviours.

### `workflows/`
contains reusable Python modules or workflow scripts supporting ETL.

### `reports/`
Contains notebook templates or report-genrating logic tied to datasets (optional).

### `catalog_defaults.toml`
Defines common config shared by all datasets in the catalog (e.g. blob paths, validation defaults).

## Update an existing dataset

Expand Down Expand Up @@ -52,7 +68,61 @@ To add a new dataset to your catalog repository:
3. Create a new ETL script in `{your_catalog}/workflows/{workflow_type}/`
4. Add SQL transformation templates if using SQL for transforms (these are [Mako templates](https://www.makotemplates.org/))

### Configuration file
## Versioning Behavior

Dataset versions are typially timestamped (e.g. 2025-10-31). Developers can:

Inspect versions

```python
from cfa.dataops import datacat

datacat.my_project.my_dataset.load.get_versions()
```

Load a version

```python
df = datacat.my_project.mydataset.load.get_dataframe()
```

Load with a version filter

```python
df = datacat.my_project.my_dataset.load.get_dataframe(version=">2024.12.01,<2025.08")
```

See which version would be chosen

```python
v = datacat.my_project.my_dataset.load.resolve_versions(version="latest")
```



## Configuration file

Configuration sections typyically include:

**[extract]**

How raw data is sourced. Common patterns include:
- reading Parquet or CSV from blob storage
- applying schema checks on raw fields
- filtering out malformed input

**[transform]**

Defines transformation logic. Options include:
- SQL expressions (DuckDB or Polars SQL)
- Python functions
- multistage ETL pipelines (split into etl/modules)

**[load]**

Defines how the transformed dataset is written inot versioned storage.
Versions are timestampe-based and automatically assigned when new data is produced.


```toml title="{your_catalog}/datasets/{dataset_name}.toml"
[properties]
Expand Down
15 changes: 12 additions & 3 deletions docs/data_user_guide.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ df = datacat.private.scenarios.covid19vax_trends.load.get_dataframe()

## Accessing Data

When the ETL pipelines are run, the data sources (raw and/or transformed) are stored into Azure Blob Storage. You can access these datasets directly using the `datacat` interface:
Raw and transformed data produced by ETL pipelines are stored in Azure Blob Storage. You can access these datasets directly using the `datacat` interface:

```python
from cfa.dataops import datacat
Expand All @@ -27,12 +27,18 @@ from cfa.dataops import datacat
df = datacat.private.scenarios.covid19vax_trends.load.get_dataframe()

# Get raw data as polars DataFrame
df = datacat.private.scenarios.seroprevalence.extract.get_dataframe(output="polars")
raw_df = datacat.private.scenarios.seroprevalence.extract.get_dataframe(output="polars")

# Get specific version
df = datacat.private.scenarios.covid19vax_trends.load.get_dataframe(
version_df = datacat.private.scenarios.covid19vax_trends.load.get_dataframe(
version_spec="==2025-06-03T17-56-50"
)

# Get raw or transformed data as Polars Lazyframe
lazy_df = datacat.private.scenarios.seropervalence.extract.get_dataframe(output="pl_lazy")

# Get reference datasets
ref_df = datacat.reference.my_reference_dataset.get_dataframe()
```

### Dataset Access Methods
Expand Down Expand Up @@ -138,6 +144,9 @@ vax_df = datacat.private.scenarios.covid19vax_trends.load.get_dataframe()

# Get raw data for analysis
raw_vax = datacat.private.scenarios.covid19vax_trends.extract.get_dataframe()

# Get raw or transformed data as LazyFrame
lazy_vax = datacat.private.scenarios.covid19vax_trends.load.get_dataframe(output="pl_lazy")
```

### Fetching Versions within a Range
Expand Down
74 changes: 74 additions & 0 deletions docs/glossary.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,74 @@
# CFA DataOps Glossary
This glossary provides clear, CDC-context definiitons of key technologies, tools and concepts frequently used in the **cfa-dataops** environment. It is intended to support new developers onboarding into CFA DataOps workflows.


## Azure Blob Storage
**Azure Blob Storage** is Microsoft Azure's cloud object storage solution used for storing large volumes of unstructured data such as CSV files, Parquet datasets, model outputs, logs, and other artifacts.

### Why it matters in cfa-dataops
- It provides secure, scablable storage for ingestion pipelines, cleaned datasets, and analytical outputs used in CFA modeling and analytics
- Many cfa-dataops integration tests rely on Blob Storage access, which requires authenticating with 'az login --identity`
- Enables cloub-based pipelines that mirror production environments, making local-to-cloud reproducibility easier


## Catalog (CFA Catalog)
The **CFA Catalog** is a central structured repository of datasets used by CFA modeling teams. It provides metadata, versioning, provenance, and standardized accessibility, enabling discoverability, reproducibility, and governance.

### Why it matters in cfa-dataops
- Ensures datasets are well-documented and versioned
- Allows analytics teams to locate authoratative ("source of truth") datasets quickly
- Supports publication workflows for modeling and public-facing data products
- Ensure reproducible analytics across CFA teams.


## DuckDB
DuckDB is an in-process OLAP (analytical) database designed for fast, local analytical queries. It runs inside Python and supports fast SQL queries on large data files without requiring a server.

### Why it matters in cfa-dataops
- Supports SQL, making transformations readable and standardized
- Enables reproducible local pipelines before cloud publication
- Ideal for rapid local development and reproducible ETL workflows
- Efficient for working with large CSV/Parquet datasets locally


## Hypothesis
Hypothesis is a property-based testing framework for Python. Instead of manually specifying inputs, Hypothesis automatially generates input data to explore edge cases.

### Why it matters in cfa-dataops
- Helps ensure reliability of ingestion and transformation functions
- Useful for validating data schemas or catalog consistency rules
- Integrated into cfa-dataops testing alongside pytest (unit + property-based tests, unit + randomized checks)


## Polars
Polars is a high-performance DataFrame library for Rust and Python, optimized for tabular data processing.

### Why it matters in cfa-dataops
- Extremely fast for cleaning, filtering, merging, and reshaping datasets
- Offers better performance compared to pandas for large datasets
- Works seamlessly with DuckDB to deliver flexible, efficient ETL patterns
- Offers declarative query patterns and efficient lazy computation


## Pytest
**pytest** is a Python testing framework used to write and execute test suites, including unit tests, integration tests, and property-based tests.

### Why it matters in cfa-dataops
- CFA DataOps uses pytest as its primary test runner, including support for:
- Discovery of test files
- Mocking with pytest-mock
- Coverage reporting
- Property-based tests via Hypothesis
- Unit tests,
- Integration tests
- pytest integrates seamlessly with uv (uv run pytest)
- supports node ID selection for running specific tests.


## UV
`uv` is a fast, modern Python package environment manager designed to replace slower and heavier tools sucha as pip and virtualenv. It ensures reproducible environments and predictable dependency resolution.

### Why it matters in cfa-dataops
- uv provides reliable installs and consistent execution environments across developer machines and CI
- In cfa-dataops, uv is the recommended setup tool for running tests and syncing dependencies (uv sync, uv run pytest)
- It improves the stability of pipelines and reduces environment drift
Loading