GitHub
Set up the GitHub source connector.
Sync modes, namespaces and the columns Pipeloom adds are explained once in Connector concepts; workspace variables and custom components in Orchestration.
This page contains the setup guide and reference information for the GitHub source connector.
Prerequisites
- List of GitHub Repositories (and access for them in case they are private)
For Pipeloom:
- Personal Access Token (see Permissions and scopes)
Setup guide
Step 1: Set up GitHub
Create a GitHub Account.
Pipeloom additional setup steps
Log into GitHub and then generate a personal access token. To load balance your API quota consumption across multiple API tokens, input multiple tokens separated with ,.
Step 2: Set up the GitHub connector in Pipeloom
- For Pipeloom: Authenticate with Personal Access Token. To generate a personal access token, log into GitHub and then generate a personal access token. Enter your GitHub personal access token. To load balance your API quota consumption across multiple API tokens, input multiple tokens separated with
,.
-
GitHub Repositories - Enter a list of GitHub organizations/repositories, e.g.
airbytehq/airbytefor single repository,airbytehq/airbyte airbytehq/another-repofor multiple repositories. If you want to specify the organization to receive data from all its repositories, then you should specify it according to the following example:airbytehq/*.:::caution Repositories with the wrong name or repositories that do not exist or have the wrong name format will be skipped with
WARNmessage in the logs. ::: -
Start date (Optional) - The date from which you'd like to replicate data for streams. For streams which support this configuration, only data generated on or after the start date will be replicated.
-
These streams will only sync records generated on or after the Start Date:
comments,commit_comment_reactions,commit_comments,commits,deployments,events,issue_comment_reactions,issue_events,issue_milestones,issue_reactions,issues,project_cards,project_columns,projects,pull_request_comment_reactions,pull_requests,pull_request_stats,releases,review_comments,reviews,stargazers,workflow_runs,workflows. -
The Start Date does not apply to the streams below and all data will be synced for these streams:
assignees,branches,collaborators,issue_labels,organizations,pull_request_commits,repositories,tags,teams,users
-
-
Branch (Optional) - List of GitHub repository branches to pull commits from, e.g.
airbytehq/airbyte/master. If no branches are specified for a repository, the default branch will be pulled. (e.g.airbytehq/airbyte/master airbytehq/airbyte/my-branch). -
API URL (Optional) - If you use a self-hosted GitHub instance, enter its API URL, for example
https://github.company.org. Leave empty to usehttps://api.github.com/. -
Max Waiting Time (in minutes) (Optional) - Maximum time the connector waits when every configured API token is rate-limited before it fails the sync. The default is 120 minutes, which covers GitHub's 60-minute rate limit reset window plus margin. You can set any value between 1 and 240 minutes. If you provide multiple personal access tokens, the connector rotates through them first, and only waits after every token is exhausted.
-
Number of Concurrent Threads (Optional) - How many partitions the connector reads in parallel. The default is 4 and the maximum is 25. While the declarative migration is in progress this setting applies only to the streams already moved to the declarative manifest (
repositories,assignees,branches,collaborators,issue_labels,tags,organizations,teams,users,events,pull_requests,commit_comments,issue_milestones,stargazers,projects,issue_events,deployments,workflows,comments,issues,review_comments,pull_request_commits,project_columns,project_cards,team_members,team_memberships,issue_timeline_events,commit_comment_reactions,issue_comment_reactions,commits,contributor_activity,workflow_runs,workflow_jobs); the remaining streams are still read one at a time. Raising it speeds up those streams but increases the risk of hitting GitHub's secondary rate limits.
For Pipeloom:
- Navigate to the Pipeloom dashboard.
- Click Sources and then click + New source.
- On the Set up the source page, select GitHub from the Source type dropdown.
- Enter a name for the GitHub connector.
Supported sync modes
The GitHub source connector supports the following sync modes:
- Full Refresh - Overwrite
- Full Refresh - Append
- Incremental Sync - Append
- Incremental Sync - Append + Deduped
Supported Streams
This connector outputs the following full refresh streams:
- Assignees
- Branches
- Contributor Activity
- Collaborators
- Issue labels
- Organizations
- Pull request commits
- Tags
- TeamMembers
- TeamMemberships
- Teams
- Users
- Issue timeline events
This connector outputs the following incremental streams:
- Comments
- Commit comment reactions
- Commit comments
- Commits
- Deployments
- Events
- Issue comment reactions
- Issue events
- Issue milestones
- Issue reactions
- Issues
- Project (Classic) cards
- Project (Classic) columns
- Projects (Classic)
- ProjectsV2
- Pull request comment reactions
- Pull request stats
- Pull requests
- Releases
- Repositories
- Review comments
- Reviews
- Stargazers
- WorkflowJobs
- WorkflowRuns
- Workflows
Entity-Relationship Diagram (ERD)
Notes
-
Only 4 streams (
comments,commits,issuesandreview comments) from the listed above streams are pure incremental meaning that they:- read only new records;
- output only new records.
-
Streams
workflow_runsandworkflow_jobsare almost pure incremental: -
Other 19 incremental streams are also incremental but with one difference, they:
- read all records;
- output only new records.
Consider this behavior when using these incremental streams, because it may affect your API call limits.
-
Sometimes for large streams specifying very distant
start_datein the past may result in keep on getting error from GitHub instead of records (respectiveWARNlog message will be outputted). In this case Specifying more recentstart_datemay help. The "Start date" configuration option does not apply to the streams below, because the GitHub API does not include dates which can be used for filtering:assigneesbranchescollaboratorsissue_labelsorganizationspull_request_commitstagsteamsusers
-
Adding a repository or organization to a connection that has already synced does not backfill its history. See Adding repositories or organizations to an existing connection.
IP allow list
If you use Pipeloom and your organization restricts access to specific IPs, add the Pipeloom IP addresses to your allow list.
Limitations & Troubleshooting
Expand to see details about GitHub connector limitations and troubleshooting.
Connector limitations
Rate limiting
You can use a personal access token to make API requests. Additionally, you can authorize a GitHub App or OAuth app, which can then make API requests on your behalf. All of these requests count towards your personal rate limit of 5,000 requests per hour (15,000 requests per hour if the app is owned by a GitHub Enterprise Cloud organization).
:::info
REST API and GraphQL API rate limits are counted separately. The REST API uses a request-based limit, while the GraphQL API uses a point-based limit where each query costs a calculated number of points. Streams that use the GraphQL API include pull_request_stats, reviews, pull_request_comment_reactions, issue_reactions, releases, and projects_v2.
:::
:::tip In the event that limits are reached before all streams have been read, it is recommended to take the following actions:
- Utilize Incremental sync mode.
- Set a higher sync interval.
- Divide the sync into separate connections with a smaller number of streams.
- Provide multiple personal access tokens in the Personal Access Tokens field, separated by commas. The connector rotates through all tokens and only waits once every token's rate limit is exhausted.
:::
When every configured token is rate-limited, the connector waits for the limit to reset rather than failing immediately. The wait is capped by the Max Waiting Time (in minutes) configuration option (default: 120 minutes, maximum: 240 minutes), and that cap applies both when the connector already knows its tokens are spent and when GitHub rejects a request and asks it to wait. Rate-limit exhaustion is classified as a transient error, so Pipeloom will retry the sync according to your connection's retry behavior if the connector does exceed this wait. If you supply several tokens, a rate-limited request switches to another token instead of waiting, whatever this setting is.
Test connection is exempt from the wait: it answers within seconds with the rate-limit message instead of sleeping, so a rate-limited token never leaves the connection setup page hanging.
Refer to GitHub article Rate limits for the REST API.
Adding repositories or organizations to an existing connection
Widening the GitHub Repositories field on a connection that has already synced does not backfill what the new repositories or organizations did in the past. The connector keeps one sync position per repository or organization, and a newly added one starts from the connection's current overall position instead of from your configured Start date. Anything created or last updated before that position is never emitted, and the sync reports no warning or error.
For example, a connection syncing docker/* since 2026-01-01 that you widen to docker/*, airbytehq/* will pick up the airbytehq repositories updated after 2026-01-01, but not the ones whose last update is older than that.
To pull the full history of a newly added repository or organization, clear the affected streams (or refresh the connection) after saving the new value, then sync. Each stream then re-reads from the beginning of the range it supports — your configured Start date for streams that honor it, and everything available for the streams listed above that do not.
This currently affects the repositories stream. Other streams still fall back to the Start date for a repository they have not seen before; they will follow the rule above as they move to the connector's declarative implementation.
GitHub Enterprise Server with rate limiting disabled
GitHub Enterprise Server ships with HTTP API rate limiting turned off, and an instance in that state answers GET /rate_limit with 404 Rate limiting is not enabled.. The connector reads that as "this instance does not track quotas" and continues without quota tracking — it no longer treats it as a failed connection. Requests are still authenticated, multiple tokens are still used in turn, and any rate limiting the instance does enforce (secondary rate limits are a separate GHES setting) is still honored through the usual retry and backoff.
Releases stream asset limit
The Releases stream uses the GitHub GraphQL API and fetches up to 100 assets per release. Releases with more than 100 assets will only include the first 100. Sub-pagination for release assets is not currently supported.
Unreadable repositories and repeated server errors
When a repository cannot be read for a reason that is specific to that repository, the connector logs a message and moves on to the next one, so a single bad repository does not fail the whole stream. This covers a repository that was deleted or renamed (404), one your token cannot access (403), one with no commits yet (409), and one where the feature backing the stream is turned off — for example the issue_labels stream on a repository with Issues disabled (410).
Repeated server errors are treated differently as of version 2.2.0. If GitHub keeps answering 502 Bad Gateway or 504 Gateway Timeout after the connector has retried, the stream now fails instead of skipping the affected repository or organization and continuing. Earlier versions logged a warning, skipped it, and reported the sync as successful — which meant a sync could complete with records silently missing. Affected streams so far: repositories (since 2.2.0), assignees, branches, collaborators, issue_labels, tags, organizations, teams, users (since 2.3.0), events, pull_requests, commit_comments, issue_milestones, stargazers, projects, issue_events, deployments, workflows, comments, issues, review_comments (since 2.4.0), pull_request_commits, project_columns, project_cards, team_members, team_memberships, issue_timeline_events, commit_comment_reactions, issue_comment_reactions (since 2.5.0), and commits, contributor_activity, workflow_runs, workflow_jobs (since 2.6.0). If a sync starts failing on one of these after upgrading, it is worth checking whether that repository or organization was already being skipped before.
The same streams also stopped resuming mid-stream. A failed attempt used to pick up from the repository or organization it stopped at; it now restarts the stream from the first one. Nothing is lost — the retry re-reads what it already read — but a stream that fails late costs its earlier partitions' requests again against your rate limit. Same versions as above. For issues this also means state is now saved once per repository rather than once per page, so an interrupted sync restarts the repository it was reading instead of resuming part-way through it.
contributor_activity fails after five 202 answers
GitHub answers 202 Accepted on the contributor statistics endpoint while it computes the statistics. The connector retries five times, 90 seconds apart, as before. As of version 2.6.0 it fails the stream when GitHub is still computing after that, where earlier versions logged a message and skipped the repository with the sync reported as successful.
One repeated record per parent on project_columns, project_cards, commit_comment_reactions and issue_comment_reactions
As of version 2.5.0 these four streams keep their cursor per parent record (per project, column or comment) and re-emit, on each incremental sync, the one record per parent whose cursor equals the saved value, for the same reason as above. Existing connections keep their saved cursors: the nested state the previous versions wrote is converted on the first sync after upgrading. As of version 2.6.0 the same applies to commits (one cursor per branch), workflow_runs and workflow_jobs.
One repeated record per repository on comments, issues and review_comments
As of version 2.4.0, each incremental sync of these three streams re-emits the single record whose updated_at is exactly the timestamp the previous sync stopped at. GitHub's since filter is inclusive, so that record is always on the first page, and the connector no longer filters it out locally. The record is unchanged, and destinations using Incremental | Append + Deduped collapse it on the primary key; on Incremental | Append you will see one extra row per repository per sync.
One repeated record per repository on the semi-incremental streams
As of version 2.4.0, each incremental sync of events, pull_requests, commit_comments, issue_milestones, stargazers, projects, issue_events, deployments, workflows re-emits the single record per repository whose cursor value (updated_at, created_at or starred_at) is exactly the timestamp the previous sync stopped at. These streams read GitHub's listing and filter it locally; the connector used to keep only records strictly newer than the saved cursor and now keeps the boundary record too. The record is unchanged, and destinations using Incremental | Append + Deduped collapse it on the primary key; on Incremental | Append you will see one extra row per repository per sync.
Two smaller changes in the same release:
pull_requestsnow fillsbase.repo_idwith the id of the base repository. Earlier versions always emittednullthere because of a bug in the record transformation;head.repo_idis unchanged.pull_requestsalways lists a repository newest-first. Earlier versions listed it oldest-first on the very first sync of a connection, which let a failed first sync resume part-way through a repository; a failed attempt now restarts the repository, as it already did on every later sync.
Permissions and scopes
If you use OAuth authentication method, the OAuth2.0 application requests the next list of scopes: repo, read:org, read:repo_hook, read:user, read:discussion, read:project, workflow. For personal access token you need to manually select needed scopes.
Your token should have at least the repo scope. Depending on which streams you want to sync, the user generating the token needs more permissions:
- For syncing Collaborators, the user which generates the personal access token must be a collaborator. To become a collaborator, they must be invited by an owner. If there are no collaborators, no records will be synced. Read more about access permissions here.
- Syncing Teams is only available to authenticated members of a team's organization. Personal user accounts and repositories belonging to them don't have access to Teams features. In this case no records will be synced.
- To sync the Projects stream, the repository must have the Projects feature enabled.
Troubleshooting
- Check out common troubleshooting issues for the GitHub source connector on our Pipeloom Forum