Databricks Lakehouse
Set up the Databricks Lakehouse destination connector.
Sync modes, namespaces and the columns Pipeloom adds are explained once in Connector concepts; workspace variables and custom components in Orchestration.
:::info Direct Load
Starting with version 4.0.0, this destination uses Direct Load architecture. Data is written directly to final tables — raw tables (_airbyte_raw_*) are no longer produced.
If you are upgrading from version 3.x, see the migration guide.
:::
Prerequisites
- A Databricks workspace with Unity Catalog enabled.
- A SQL warehouse or all-purpose compute cluster to run queries against.
- Authentication credentials: an OAuth2 client ID and secret (recommended), or a personal access token.
- Permission to create schemas, tables, and Unity Catalog Volumes in the target catalog. The connector uses Volumes to stage Avro files before loading them into tables.
- Acceptance of the Databricks JDBC ODBC driver license. By using this connector, you agree that it may only be used to connect third-party applications to Apache Spark SQL within a Databricks offering using the ODBC and/or JDBC protocols.
Network access
If you're using Pipeloom and this destination uses IP-based access controls, add Pipeloom's IP addresses to your allowlist.
Step 1: Set up Databricks
Gather the following information from your Databricks workspace.
Server Hostname, HTTP Path, and Port
-
Open your Databricks workspace.
-
Navigate to your SQL warehouse:
-
Open the Connection Details tab:
-
Note the Server Hostname, HTTP Path, and Port values. The default port is
443.
Unity Catalog Name
You also need the name of the Unity Catalog you want to write to. Find this in the Databricks workspace sidebar under Catalog — it is the top-level catalog name, not a schema or table name.
Authentication
OAuth2 (Recommended)
Create a service principal in your Databricks workspace. Generate a client ID and secret, and grant the service principal access to your target catalog and schema.
Personal Access Token
-
In your Databricks workspace, click your profile icon in the top-right corner and go to Settings > Developer > Access tokens > Manage.
-
Click Generate new token. Enter a description and an expiration period (leave blank for no expiration):
Step 2: Set up the Databricks destination in Pipeloom
- Log in to your Pipeloom account.
- In the left navigation bar, click Destinations. In the top-right corner, click + New destination.
- Find and select Databricks Lakehouse from the list of available destinations.
- Enter the Server Hostname, HTTP Path, Port, and Unity Catalog Name from Step 1.
- Select your Authentication method and enter the required credentials.
- Accept the Databricks JDBC driver Terms & Conditions.
- Configure the remaining options:
- Default Schema — The schema where tables are created. Defaults to
default. You can override this per-connection. - CDC deletion mode — Controls how CDC deletions are handled. Hard delete removes the row; soft delete keeps a tombstone record with the
_ab_cdc_deleted_attimestamp. Defaults to hard delete. - Purge Staging Files and Tables — Whether to delete staging Avro files from Unity Catalog Volumes after loading. Leave enabled unless you need to inspect staged files for debugging.
- Default Schema — The schema where tables are created. Defaults to
- Click Set up destination.
Supported sync modes
| Sync mode | Supported? |
|---|---|
| Full Refresh - Overwrite | Yes |
| Full Refresh - Append | Yes |
| Full Refresh - Overwrite + Deduped | Yes |
| Incremental Sync - Append | Yes |
| Incremental Sync - Append + Deduped | Yes |
Output schema
Each stream is written directly to a final table in your configured schema. The table includes your data columns plus the following Pipeloom metadata columns:
| Column | Type | Description |
|---|---|---|
_airbyte_raw_id | STRING | A UUID assigned to each record by Pipeloom. |
_airbyte_extracted_at | TIMESTAMP | When the record was read from the source. |
_airbyte_meta | STRING | JSON object containing sync metadata. |
_airbyte_generation_id | LONG | Tracks refreshes. |
Data type map
| Pipeloom Type | Databricks Type | Notes |
|---|---|---|
string | STRING | |
number | DECIMAL(38, 10) | Max 28 integer digits, 10 fractional |
integer | LONG | 64-bit integer |
boolean | BOOLEAN | |
object | STRING | Serialized as JSON |
array | STRING | Serialized as JSON |
timestamp_with_timezone | TIMESTAMP | Microsecond precision |
timestamp_without_timezone | TIMESTAMP_NTZ | Microsecond precision, no timezone |
time_with_timezone | STRING | No native Databricks equivalent |
time_without_timezone | STRING | No native Databricks equivalent |
date | DATE |
Naming conventions
- Schema and table names are lowercased automatically. Databricks treats them as case-insensitive identifiers.
- Column names preserve the casing from your source data.
- Special characters in identifiers are escaped automatically by the connector.
Namespace support
This destination supports namespaces. The namespace maps to a Databricks schema.