Getting Started¶
This guide walks you through installing havn, creating your first project, running a pipeline, and exploring data in the web UI.
Prerequisites¶
- Python 3.10 or later
- pip (Python package manager)
Installation¶
Install from PyPI¶
pip install havn
The web UI frontend is bundled into the package — no separate build step needed.
Install from Source (development)¶
git clone https://github.com/chraltro/havn
cd havn
pip install -e ".[dev]"
cd frontend && npm install && npm run build
Docker¶
docker run -v $(pwd):/project -p 3000:3000 ghcr.io/chraltro/havn serve
Create a New Project¶
Scaffold a new project with havn init:
havn init my-project
cd my-project
This creates the following structure:
my-project/
ingest/ # Python scripts and .dpnb notebooks for data ingestion
earthquakes.dpnb # Sample ingest notebook (USGS earthquake data)
transform/
bronze/ # Light cleanup SQL
earthquakes.sql
silver/ # Business logic SQL
earthquake_events.sql
earthquake_daily.sql
gold/ # Consumption-ready SQL
earthquake_summary.sql
top_earthquakes.sql
region_risk.sql
export/ # Python scripts for exporting data
earthquake_report.py
seeds/ # CSV reference data
magnitude_scale.csv
contracts/ # YAML data quality contracts
quality.yml
notebooks/ # Interactive .dpnb notebooks
explore.dpnb
project.yml # Project configuration
.env # Secrets (never commit this)
.gitignore
warehouse.duckdb # Created after first pipeline run
Run Your First Pipeline¶
The scaffolded project includes a complete earthquake data pipeline. Run it:
havn jobs run full-refresh
This executes the pipeline steps defined in project.yml:
- Ingest -- Fetches earthquake data from the USGS API (falls back to sample data offline)
- Seed -- Loads
seeds/magnitude_scale.csvas a reference table - Transform -- Builds SQL models in dependency order:
bronze->silver->gold - Export -- Generates a summary report
Explore Your Data¶
List Tables¶
havn tables
This shows all tables and views in the warehouse, organized by schema.
Run Queries¶
havn query "SELECT * FROM gold.earthquake_summary LIMIT 10"
Output options:
havn query "SELECT * FROM gold.top_earthquakes" --csv
havn query "SELECT * FROM gold.top_earthquakes" --json
havn query "SELECT COUNT(*) FROM landing.earthquakes" --limit 5
Check Data Quality¶
havn contracts
This runs all YAML contracts from the contracts/ directory and reports pass/fail results.
View Run History¶
havn history
Shows recent pipeline runs with status, duration, and row counts.
Start the Web UI¶
havn serve
This starts the web server on http://localhost:3000 with:
- File Browser -- Edit SQL and Python files with Monaco editor
- Query Panel -- Run ad-hoc SQL queries with autocomplete and
$nameparameters - Table Browser -- Browse schemas, tables, and column profiles
- DAG Viewer -- Interactive dependency graph visualization
- Notebook Runner -- Execute
.dpnbnotebooks interactively - Pipeline Controls -- Run streams and view history
With Authentication¶
havn serve --auth
On first launch with --auth, you will be prompted to create an admin user through the web UI. See Auth for details.
Project Configuration¶
The project.yml file is the central configuration. See Configuration for the full reference. Here is a minimal example:
name: my-project
database:
path: warehouse.duckdb
streams:
full-refresh:
description: "Full pipeline rebuild"
steps:
- seed: [all]
- ingest: [all]
- transform: [all]
- export: [all]
Next Steps¶
- Transforms -- Learn how to write SQL transform models
- Pipelines -- Configure multi-step data pipelines
- Connectors -- Connect to external data sources
- Quality -- Add data quality checks
- CLI Reference -- Full command reference