Pipeloom Docs
ConnectorsSources

Mongo DB

Set up the Mongo DB source connector.

Sync modes, namespaces and the columns Pipeloom adds are explained once in Connector concepts; workspace variables and custom components in Orchestration.

Pipeloom's certified MongoDB connector offers the following features:

Quick Start

This section provides information about configuring the MongoDB V2 source connector. If you are upgrading from a previous version of the MongoDB V2 source connector, please refer to the upgrade instructions in this document.

New Installation/New Source Connector Configuration

Here is an outline of the minimum required steps to configure a new MongoDB V2 source connector:

  1. Create or discover the configuration of a MongoDB replica set, either hosted in MongoDB Atlas or self-hosted.
  2. Create a new MongoDB source in Pipeloom
  3. (Pipeloom Only) Allow inbound traffic from Pipeloom IPs

Once this is complete, you will be able to select MongoDB as a source for replicating data.

Step 1: Create a dedicated read-only MongoDB user

These steps create a dedicated, read-only user for replicating data. Alternatively, you can use an existing MongoDB user with access to the database.

MongoDB Atlas
  1. Log in to the MongoDB Atlas dashboard.

  2. From the dashboard, click on "Database Access" under "Security"

  3. Click on the "+ ADD NEW DATABASE USER" button.

  4. On the "Add new Database User" modal dialog, choose "Password" for the "Authentication Method".

  5. In the "Password Authentication" section, set the username to READ_ONLY_USER in the first text box and set a password in the second text box.

  6. Under "Database User Privileges", click on "Select one built-in role for this user" under "Built-in Role" and choose "Only read any database".

  7. Under "Database User Privileges", navigate to "Specific Privileges", then click "Add Specific Privilege" and add readAnyDatabase.

:::info Starting in version v2.0.0, change data capture now supports monitoring the entire cluster, not just a single database. This allows you to sync multiple collections across different databases using a single source.

The readAnyDatabase privilege is required for this expanded access. Without it, the connection will fail with an authorization error. :::

  1. Enable "Restrict Access to Specific Clusters/Federated Database instances" and enable only those clusters/database that you wish to replicate.

  2. Click on "Add User" at the bottom to save the user.

Self Hosted

These instructions assume that the MongoDB shell is installed. To install the MongoDB shell, please follow these instructions.

  1. From a terminal window, launch the MongoDB shell:
> mongosh <connection string to cluster> --username <user with admin permissions>;
  1. Switch to the admin database:
test> use admin
switched to db admin
  1. Create the READ_ONLY_USER user with the read role:
admin> db.createUser({user: "READ_ONLY_USER", pwd: "READ_ONLY_PASSWORD", roles: [{role: "read", db: "TARGET_DATABASE"}]})

:::note Replace READ_ONLY_PASSWORD with a password of your choice and TARGET_DATABASE with the name of the database to be replicated. :::

  1. Next, enable authentication, if not already enabled. Start by editing the /etc/mongodb.conf by adding/editing these specific keys:
net:
  bindIp: 0.0.0.0

security:
  authorization: enabled

:::note Setting the bindIp key to 0.0.0.0 will allow connections to database from any IP address. Setting the security.authorization key to enabled will enable security and only allow authenticated users to access the database. :::

Step 2: Discover the MongoDB cluster connection string

These steps outline how to discover the connection string of your MongoDB instance.

MongoDB Atlas

Atlas is MongoDB's cloud-hosted offering. Below are the steps to discover the connection configuration for a MongoDB Atlas-hosted replica set cluster:

  1. Log in to the MongoDB Atlas dashboard.

  2. From the dashboard, click on the "Connect" button of the source cluster.

  3. On the "Connect to <cluster name>" modal dialog, select "Shell" under the "Access your data through tools" section.

  4. Copy the connection string from the entry labeled "2. Run your connection string in your command line" on the modal dialog, removing/avoiding the quotation marks.

Self Hosted Cluster

Self-hosted clusters are MongoDB instances that are hosted outside of MongoDB Atlas. Below are the steps to discover the connection string for a MongoDB self-hosted replica set cluster.

  1. Refer to the MongoDB connection string documentation for instructions on discovering a self-hosted deployment connection string.

Step 3: Configure the Pipeloom MongoDB Source

To configure the Pipeloom MongoDB source, use the database credentials and connection string from steps 1 and 2, respectively. The source will test the connection to the MongoDB instance upon creation.

IP allow list

If you use Pipeloom and your organization restricts access to specific IPs, add the Pipeloom IP addresses to your allow list.

Replication Methods

The MongoDB source utilizes change data capture (CDC) as a reliable way to keep your data up to date. In addition, MongoDB source now allows for syncing in a full refresh mode.

CDC

Pipeloom utilizes the change streams feature of a MongoDB replica set to incrementally capture inserts, updates and deletes using a replication plugin. To learn more how Pipeloom implements CDC, refer to Change Data Capture (CDC).

Full Refresh

The Full Refresh sync mode added in v1.3.0 allows for reading the entire contents of a collection repeatedly. The MongoDB source connector is using checkpointing in Full Refresh read so a sync job that failed for network error for example, Rather than starting over it will continue its full refresh read from a last known point.

Schema Enforcement

By default, the MongoDB V2 source connector enforces a schema. This means that while setting up a connector it will sample a configurable number of documents and will create a set of fields to sync. From that set of fields, an admin can then deselect specific fields from the Replication screen to filter them out from the sync.

When the schema enforced option is disabled, MongoDB collections are read in schema-less mode which doesn't assume documents share the same structure. This allows for greater flexibility in reading data that is unstructured or vary a lot in between documents in a single collection. When schema is not enforced, each document will generate a record that only contains the following top-level fields:

{
  "_id": <document id>,
  "data": {<a JSON containing the entire set of fields found in document>}
}

The contents of data will vary according to the contents of each document read from MongoDB. Unlike in Schema enforced mode, the same field can vary in type between documents. For example, field "xyz" may be a String on one document and a Date on another. As a result no field will be omitted and no document will be rejected. When Schema is not enforced there is no way to deselect fields as all fields are read for every document.

Array Type Normalization

Most destinations can safely cast between primitive types (i.e. integer to string), but they cannot reconcile composite/compound types like an object versus an array. Without normalization, records with these structural mismatches would be written as NULL in the destination. To ensure data consistency, in schema-enforced mode our implementation automatically converts any non-array field to an array when the schema expects an array.

// Document 1 - Array (matches schema)
{ "reviews": [{"rating": "A"}, {"rating": "B"}] }

// Document 2 - Single object (auto-wrapped)
{ "reviews": {"rating": "A"} }  →  { "reviews": [{"rating": "A"}] }

:::info Important This normalization is intended for occasional type inconsistencies. If your collection has fields that frequently change types, consider using schemaless mode instead. :::

Limitations & Troubleshooting

To see connector limitations, or troubleshoot your MongoDB connector, see more in our MongoDB troubleshooting guide.

Schema discovery performance impact

:::warning Schema discovery runs heavy aggregation queries against your MongoDB cluster in parallel across all collections. On production clusters with many collections or large documents, this can cause significant resource pressure, including degraded query performance, replication lag, or in extreme cases, cluster instability. :::

Because MongoDB collections are schemaless, no sample size can guarantee a complete or stable schema — new fields can appear in documents at any time. When schema enforcement is enabled, the connector's Discover phase runs $sample aggregation pipelines against every collection in the configured databases in parallel. On clusters with many collections or large documents, these concurrent queries can cause significant resource pressure on your MongoDB deployment, including degraded performance or cluster instability.

For recommended approaches and other alternatives to protect your production cluster, see Schema discovery performance impact in the troubleshooting guide.

MongoDB CDC Limitations

MongoDB has a 16MB maximum document size limit for BSON documents. During CDC syncs, change stream events can exceed this limit when documents are large, causing a BSONObjectTooLarge error. For details on resolving this error, see the MongoDB CDC Limitations section in the troubleshooting guide.

Configuration Parameters

Parameter NameDescription
Cluster TypeThe type of the MongoDB cluster (MongoDB Atlas replica set or self-hosted replica set).
Connection StringThe connection string of the source MongoDB cluster. For Atlas hosted clusters, see the quick start guide for steps to find the connection string. For self-hosted clusters, refer to the MongoDB connection string documentation for more information.
Database NamesThe names of the MongoDB databases that contain the source collection(s) to sync. Allows specifying multiple databases to discover and sync collections from.
UsernameThe username which is used to access the database. Required for MongoDB Atlas clusters.
PasswordThe password associated with this username. Required for MongoDB Atlas clusters.
Authentication Source(MongoDB Atlas clusters only) Specifies the database that the supplied credentials should be validated against. Defaults to admin. See the MongoDB documentation for more details.
Schema EnforcedControls whether schema is discovered and enforced. See discussion in Schema Enforcement.
Initial Waiting Time in Seconds (Advanced)The amount of time the connector will wait when it launches to determine if there is new data to sync or not. Defaults to 300 seconds. Valid range: 120 seconds to 1200 seconds.
Size of the queue (Advanced)The size of the internal queue. This may interfere with memory consumption and efficiency of the connector, please be careful.
Discovery Sample Size (Advanced)The maximum number of documents to sample when attempting to discover the unique fields for a collection. Default is 10,000 with a valid range of 1,000 to 100,000. See the MongoDB sampling method for more details.
Update Capture Mode (Advanced)Determines how Pipeloom retrieves update events during CDC (Incremental) syncs.

Lookup (default): fetches the document's latest available state when the update event is processed. If a document is updated multiple times in rapid succession, Pipeloom may capture the newest available version for multiple events instead of the document state immediately after each individual update. As a result, some intermediate document states may not be captured (see Troubleshooting: Lookup vs Post Image).

Post Image: uses MongoDB's built-in change stream post-images to capture the document state immediately after each change. IMPORTANT: Post Image requires MongoDB 6.0+ and collections must be configured to return pre and post images. Failure to do so may lead to data loss.
Document discovery timeout in seconds (Advanced)The amount of time the connector will wait when discovering documents. Defaults to 600 seconds. Valid range: 5 seconds to 1200 seconds.
Invalid CDC position behavior (Advanced)Determines whether Pipeloom should fail or re-sync data in case of a stale or invalid cursor value in the WAL. If "Fail sync" is chosen, you will need to manually reset the connection before syncing data again. If "Re-sync data" is chosen, Pipeloom will automatically trigger a full refresh, which increases cloud costs and may miss incremental changes that occurred while the cursor was invalid. Default is "Fail sync".
Initial Load Timeout in Hours (Advanced)The amount of time an initial load is allowed to continue before catching up on CDC logs. Default is 8 hours. Valid range: 4 to 24 hours.

For more information regarding configuration parameters, please see MongoDb Documentation.

On this page