Connectors¶
Connectors automate data ingestion from external sources. havn includes pre-built connectors for databases, SaaS APIs, file storage, and webhooks. Each connector tests the connection, discovers available resources, generates an ingest script, and updates project.yml.
Available Connectors¶
| Connector | Type | Description |
|---|---|---|
| PostgreSQL | postgres |
PostgreSQL database tables |
| MySQL | mysql |
MySQL/MariaDB database tables |
| CSV Files | csv |
Local or remote CSV files |
| Stripe | stripe |
Stripe payments data |
| Shopify | shopify |
Shopify e-commerce data |
| HubSpot | hubspot |
HubSpot CRM data |
| Google Sheets | google_sheets |
Google Spreadsheets |
| REST API | rest_api |
Generic REST API endpoints |
| S3/GCS | s3_gcs |
Amazon S3 or Google Cloud Storage files |
| Webhook | webhook |
Receive webhook data via HTTP POST |
| Snowflake | snowflake |
Snowflake data warehouse (Arrow transfer) |
| BigQuery | bigquery |
Google BigQuery (Arrow transfer) |
| Redshift | redshift |
Amazon Redshift (via DuckDB postgres extension) |
List all available connectors:
havn connectors available
Setting Up a Connector¶
Interactive Setup¶
Use havn connect to set up a connector interactively:
# PostgreSQL
havn connect postgres --host localhost --database mydb --user admin --password secret
# Stripe
havn connect stripe --api-key sk_live_xxx
# CSV file
havn connect csv --path /data/customers.csv
# Google Sheets
havn connect google-sheets --set spreadsheet_id=1BxiMVs0XRA5nFMdKvBdBZjgmUUqptlbs74OgVE2upms
# With JSON config
havn connect postgres --config '{"host":"db.prod","database":"app","user":"ro","password":"s3cret"}'
# From a config file
havn connect postgres --config ./postgres.json
The setup process:
- Tests the connection -- Verifies credentials and connectivity
- Discovers resources -- Lists available tables, endpoints, or sheets
- Generates an ingest script -- Creates
ingest/connector_<name>.py - Updates project.yml -- Adds the connection and creates a sync stream
- Stores secrets in .env -- Passwords and API keys go to
.env, notproject.yml
Configuration Options¶
havn connect <type> [OPTIONS]
Options:
--name, -n Connection name (default: auto-generated)
--tables, -t Comma-separated tables to sync
--schema, -s Target schema (default: landing)
--schedule Cron schedule for automatic sync
--test Only test the connection
--discover Only list available resources
--config, -c JSON string or file path with params
--set key=value Set individual parameters (repeatable)
Convenience shortcuts for common parameters:
--host Hostname
--port Port number
--database Database name
--user Username
--password Password
--url URL
--api-key API key
--token Access token
--path File or bucket path
Managing Connectors¶
List Configured Connectors¶
havn connectors list
Shows all connectors in project.yml with their type, script path, and status.
Test a Connection¶
havn connectors test prod_postgres
Verifies that the connection still works with the stored credentials.
Sync Data¶
havn connectors sync prod_postgres
Runs the generated ingest script for a connector.
Regenerate Script¶
havn connectors regenerate prod_postgres
Re-discovers resources and regenerates the ingest script. Useful when the connector code is updated or configuration changes.
Remove a Connector¶
havn connectors remove prod_postgres
Deletes the ingest script and removes the connection from project.yml.
Connector Architecture¶
Each connector implements the BaseConnector contract:
test_connection(config)-- Verify the connection worksdiscover(config)-- List available tables/resourcesgenerate_script(config, tables, target_schema)-- Emit a Python ingest script
The generated ingest script is a standard havn Python script that uses the db DuckDB connection. You can customize it after generation.
Secret Handling¶
Connector parameters marked as secret (passwords, API keys, tokens) are:
- Stored in
.envas environment variables (e.g.,PROD_POSTGRES_PASSWORD=...) - Referenced in
project.ymlas${ENV_VAR_NAME}placeholders - Never written to
project.ymlin plaintext
Webhook Connector¶
The webhook connector receives data via HTTP POST and stores it in a landing table:
havn connect webhook --name orders_webhook
Once configured, send data to the webhook endpoint:
curl -X POST http://localhost:3000/api/webhook/orders \
-H "Content-Type: application/json" \
-d '{"order_id": 123, "amount": 99.99}'
Data is stored in landing.<webhook_name>_inbox with columns: id, received_at, payload (JSON).
Warehouse Connectors¶
Install optional dependencies for warehouse connectors:
pip install havn[snowflake] # Snowflake
pip install havn[bigquery] # BigQuery
pip install havn[redshift] # Redshift (uses DuckDB postgres extension, no extra deps)
pip install havn[warehouses] # All warehouse connectors
Snowflake¶
havn connect snowflake --set account=xy12345.us-east-1 --set warehouse=COMPUTE_WH \
--user myuser --password secret --database PROD --tables customers,orders
Supports full and CDC incremental sync. Data is transferred via Arrow (fetch_arrow_all()) for high throughput.
BigQuery¶
havn connect bigquery --set project_id=my-gcp-project --set dataset=analytics \
--set service_account_b64=<base64-encoded-json> --tables events,users
Uses Arrow transfer via to_arrow(). Service account credentials are base64-encoded and stored in .env.
Redshift¶
havn connect redshift --host my-cluster.abc123.us-east-1.redshift.amazonaws.com \
--database analytics --user admin --password secret --tables orders,products
Uses DuckDB's built-in postgres extension for direct connectivity with no additional Python dependencies.
Connector Health¶
Check the last sync status for each connector:
# Via API
curl http://localhost:3000/api/connectors/health
Returns the most recent run status, timestamp, and duration for each ingest script.
Related Pages¶
- CDC -- Incremental sync with change data capture
- Configuration -- Connection configuration in project.yml
- Pipelines -- Running connectors as pipeline steps
- API Reference -- Connector API endpoints