Pipeloom Docs
Connectors

Connector concepts

Sync modes, namespaces, change data capture, refreshes and the columns Pipeloom adds.

The connector pages link here for the ideas every connection shares, so they are explained once.

Sync modes

A connection's sync mode decides two things per stream: how Pipeloom reads from the source (everything, or only what changed) and how it writes to the destination (append, replace, or keep one row per primary key). Not every connector supports every mode; each connector page lists what it supports.

Full refresh overwrite

Every sync reads the whole stream and replaces the destination table with it. Simple, always consistent with the source, and the most expensive on large tables.

Full refresh overwrite deduped

As full refresh overwrite, then keeps one row per primary key. Use it when the source can return duplicates.

Full refresh append

Every sync reads the whole stream and appends it to the destination table. The table keeps a copy of the stream per sync, which is useful for snapshots over time.

Incremental append

Each sync reads only records added or updated since the last one, tracked by a cursor field (such as updated_at), and appends them. An updated record appears once per version.

Incremental append deduped

As incremental append, then keeps only the latest version of each record by primary key, so the table mirrors the source. This is the usual choice for replicating a database.

Namespaces

A namespace is the schema (or dataset, or database) a stream is written into. By default a connection writes into the destination's configured namespace; you can instead mirror the source's namespace or set your own, and add a prefix to every table name.

Change data capture

Database sources (Postgres, MySQL, SQL Server, MongoDB, and others) can read changes from the database's log instead of querying tables. Change data capture sees deletes, puts almost no load on the database, and does not need a cursor column. It needs setup on the database side, described on each connector's page.

Refreshes

A refresh re-reads a stream from the beginning without deleting the connection's state for other streams. Use it after a source's history changes or a destination table is damaged. Depending on the destination, a refresh either replaces the data or appends to it, keeping earlier rows.

Metadata columns

Destinations add a few columns to every table so each row can be traced back to a sync:

ColumnWhat it holds
_airbyte_raw_idA unique id for the record.
_airbyte_extracted_atWhen the record was read from the source.
_airbyte_generation_idWhich generation of the stream wrote it; a refresh starts a new one.
_airbyte_metaChanges made on the way (such as a value that did not fit the column type) and why.

Some destinations also keep an internal schema (airbyte_internal) for staging. Leave it in place.

IP allow list

If your database or API only accepts connections from known addresses, contact support for the addresses Pipeloom's syncs connect from and add them to your allow list.

On this page