Pipeloom Docs
ConnectorsSources

Shopify

Set up the Shopify source connector.

Sync modes, namespaces and the columns Pipeloom adds are explained once in Connector concepts; workspace variables and custom components in Orchestration.

This page contains the setup guide and reference information for the Shopify source connector.

Prerequisites

Setup guide

This connector supports OAuth2.0 and API Password (for private applications) authentication methods.

Set up Shopify

For Pipeloom:

  1. Navigate to the Pipeloom dashboard.
  2. Click Sources and then click + New source.
  3. On the Set up the source page, select Shopify from the Source type dropdown.
  4. Enter a name for the Shopify connector.

Create a custom app

Authentication to the Shopify API requires a custom application. Follow these instructions to create a custom app and find your Admin API Access Token.

  1. Log in to your Shopify account.
  2. In the dashboard, navigate to Settings > App and sales channels > Develop apps > Create an app.
  3. Select a name for your new app.
  4. Select Configure Admin API scopes.
  5. Grant access to the following list of scopes. Only select scopes prefixed with read_, not write_ (e.g. read_locations,read_price_rules, etc ).
  6. Click Install app to give this app access to your data.
  7. Once installed, go to API Credentials to copy the Admin API Access Token. You are now ready to set up the source in Pipeloom!

Connect using API Password

  1. Enter a Source name.
  2. Enter your Shopify Store name. You can find this in your URL when logged in to Shopify or within the Store details section of your Settings.
  3. For API Password, enter your custom application's Admin API access token.
  4. (Optional) You may set a Replication Start Date as the starting point for your data replication. Any data created before this date will not be synced. Please note that this defaults to January 1st, 2020.
  5. Click Set up source and wait for the connection test to complete.

Custom app scopes

Add the following scopes to your custom app to ensure Pipeloom can sync all available data. For more information on access scopes, see the Shopify docs.

  • read_analytics
  • read_assigned_fulfillment_orders
  • read_content
  • read_customers
  • read_discounts
  • read_draft_orders
  • read_fulfillments
  • read_gdpr_data_request
  • read_gift_cards
  • read_inventory
  • read_legal_policies
  • read_locations
  • read_locales
  • read_marketing_events
  • read_markets
  • read_merchant_managed_fulfillment_orders
  • read_online_store_pages
  • read_order_edits
  • read_orders
  • read_price_rules
  • read_product_listings
  • read_products
  • read_publications
  • read_reports
  • read_resource_feedbacks
  • read_script_tags
  • read_shipping
  • read_shopify_payments_accounts
  • read_shopify_payments_bank_accounts
  • read_shopify_payments_disputes
  • read_shopify_payments_payouts
  • read_themes
  • read_third_party_fulfillment_orders
  • read_translations

Supported sync modes

The Shopify source connector supports the following sync modes:

  • Full Refresh
  • Incremental

This source syncs data using the Shopify REST API, the Shopify GraphQL API, and the Shopify GraphQL BULK API. Streams labeled "(GraphQL)" in the list below use the GraphQL or BULK API; unlabeled streams use the REST API.

Supported Streams

Entity-Relationship Diagram (ERD)

Countries and market-driven shipping

Shopify is moving merchant shipping configuration from delivery profiles to Markets. Once a shop is on market-driven shipping, the deliveryProfiles API that backs the Countries stream returns a frozen snapshot that no longer reflects changes made by the merchant.

The connector checks shop.features.marketDrivenShipping when the Countries or Market Countries stream starts syncing:

  • Shops on legacy shipping: Countries syncs as before; Market Countries emits no records.
  • Shops on market-driven shipping: Countries keeps emitting the snapshot of the shipping configuration as it was at migration time (and logs a warning), so previously synced data is not wiped; Market Countries emits one record per market region (a whole country, or a country subdivision such as a US state) with the market's current shipping options.

Market Countries differs from Countries in shape: a country that belongs to several markets appears once per market, whole-country regions carry no province list, and shipping_enabled / shipping_options are null when the market inherits its shipping configuration from a parent market. Markets of every type and status are included, so filter on market_type = REGION and market_status = ACTIVE for buyer-facing shipping countries.

Shopify's rollout starts on October 1, 2026 (merchants who opt in, plus shops that Shopify moves automatically once their installed apps are confirmed compatible) and is planned to cover all shops by July 1, 2027, so both stream behaviors coexist during that period. Nothing changes in your connection automatically and no global switch is required: keep Countries if you need the legacy snapshot, and enable Market Countries for shops that have migrated.

Enabling Market Countries for a migrated shop

  1. Make sure the connector has the read_markets scope. With OAuth2.0, re-authenticate the source on its Settings page. With API password, add read_markets to your custom app's Admin API scopes and save the source again. Without the scope, the stream is not in the catalog and the source logs a warning when the schema is discovered.
  2. Refresh the source schema on the connection so Market Countries appears in the stream list.
  3. Enable the Market Countries stream. You can leave Countries enabled; it continues to hold the last legacy snapshot.
  4. Run a sync and check that market_countries contains one record per market region (market_handle, code, subdivision_code, shipping_options). On a shop that has not migrated yet, the stream stays empty until Shopify migrates the shop; no action is needed.

Capturing deleted records

The connector captures deletions for records in the Articles, Blogs, CustomCollections, Orders, Pages, and PriceRules streams. When a record is deleted, the connector outputs a record with the ID of that record and the deleted_at, deleted_message, and deleted_description fields filled out. No other fields are filled out for the deleted records.

For deleted products, use the dedicated Deleted Products stream, which queries the Shopify GraphQL Events API for product deletion events. This stream returns records with id (the product ID), deleted_at, deleted_message, and shop_url fields.

Check the following Shopify documentation for more information about retrieving deleted records.

Syncing discount codes

The connector provides two streams for discount codes:

StreamMethodBest for
Discount CodesBulk Operations (async)Stores where every discount has fewer than ~100 redeem codes. High throughput, low API cost.
Discount Codes SyncSynchronous cursor-paginated GraphQLStores that have one or more discounts with more than ~100 redeem codes. Guarantees every code is returned.

Shopify's Bulk Operations API silently truncates the nested codes connection at approximately 100 records per parent discount, even with groupObjects: true. If a discount has 500 codes, the bulk stream will only return ~100 of them without any error or warning.

Discount Codes Sync avoids this limitation by querying parent discounts with cursor pagination, then explicitly paging through each parent's child codes (up to 250 per page). The trade-off is higher per-record API cost and lower throughput compared to the bulk stream.

Recommendation: Enable the Discount Codes Sync stream if you have any discounts with more than ~100 redeem codes. You may run both streams simultaneously — they produce records with the same schema, so you can deduplicate downstream by id.

Marketing Attribution data

Data related to marketing attribution can be found across a few different streams. Sync these streams to understand marketing performance:

  • Customer Journey Summary (firstVisit.source, firstVisit.sourcetype)
  • Orders (referring_site, source_name)
  • Abandoned Checkouts (referring_site, source_name)

Features

FeatureSupported?(Yes/No)
Full Refresh SyncYes
Incremental - Append SyncYes
NamespacesNo

IP allow list

If you use Pipeloom and your organization restricts access to specific IPs, add the Pipeloom IP addresses to your allow list.

Data type map

Integration TypePipeloom Type
stringstring
numbernumber
arrayarray
objectobject
booleanboolean

Limitations & Troubleshooting

Expand to see details about Shopify connector limitations and troubleshooting

Connector limitations

Rate limiting

Shopify has some rate limit restrictions. Typically, there should not be issues with throttling or exceeding the rate limits but, in some edge cases, you may encounter the following warning message:

"Caught retryable error '<some_error> or null' after <some_number> tries.
Waiting <some_number> seconds then retrying..."

This is expected when the connector hits a 429 - Rate Limit Exceeded HTTP Error. The sync operation will continue successfully after a short backoff period.

For all Shopify GraphQL Bulk API requests, Shopify's bulk operation limitations apply. Different request types can have different limitations.

Missing records after bulk stream checkpointing

Version 3.3.3 fixes an issue where some incremental GraphQL Bulk streams could skip parent records after a bulk job checkpointed mid-output. The issue affected these streams:

  • metafield_customers
  • metafield_products
  • metafield_product_images
  • product_images
  • customer_address

If you synced one of these streams on an earlier connector version and suspect missing historical records, clear the affected stream and run a sync to backfill data. Clearing a stream deletes the data Pipeloom wrote for that stream in your destination. For more information, see Clearing your data.

BULK job checkpoint collisions

Incremental GraphQL Bulk streams checkpoint a bulk job once it has collected BULK Job checkpoint (rows collected) lines, then start the next job from the newest cursor value they saw. If a single cursor value holds more lines than that threshold, the next job cannot advance past it, and the sync fails with:

The stream: `<stream_name>` checkpoint collision is detected. Try to increase the
`BULK Job checkpoint (rows collected)` to the bigger value. The stream will be
synced again during the next sync attempt.

This most often affects the metafield streams and discount_codes, where many child rows hang off one parent record, and it usually follows a bulk change in your store that stamped a large set of records with the same updated_at value.

First, check whether it cleared itself. Look at whether any sync has succeeded since the collision:

  • A later sync succeeded. The next attempt got past the cluster on its own. No action needed, though the failed attempts still counted toward your usage.
  • No sync has succeeded since. The stream is blocked and will not recover without a configuration change. Recent records are not reaching your destination, so treat it as urgent.

To unblock a stream that is not recovering, raise BULK Job checkpoint (rows collected) in your source configuration. The value must exceed the number of lines sharing that one cursor value, and the maximum is 1,000,000. Increase it in steps rather than jumping straight to the maximum: each bulk job then covers more data, so syncs take longer and use more temporary disk space while results are sorted.

Two things that do not help:

  • Lowering GraphQL BULK Date Range in Days. A collision means the sync could not get past a single cursor value, and no date range can subdivide a single value.
  • Clearing the stream. The cluster is in your source data, so a fresh sync reaches it again.

If the stream still collides at 1,000,000, or if raising the value does not change the behavior, contact Pipeloom Support — the affected cursor value is in the sync logs, on the line reading Stream <stream_name>, continue from checkpoint:.

Troubleshooting

  • If you encounter access errors while using OAuth2.0 authentication, please make sure you've followed this Shopify Article to request the access to the client's store first. Once the access is granted, you should be able to proceed with OAuth2.0 authentication.
  • If you receive a "The BULK job couldn't be created at this time, since another job is running." error, please check your operation's progress with the Shopify GraphQL BULK api.
  • If you receive a "checkpoint collision is detected" error for a stream, see BULK job checkpoint collisions above to tell a self-clearing failure from a blocked stream.
  • If you need to cancel a Shopify GraphQL BULKjob, please follow these steps. You will need the current in-progress job ID to cancel.
  • Check out common troubleshooting issues for the Shopify source connector on our Pipeloom Forum here.

On this page