# GitHub Backup Repo (`onescales/github-backup-repo`) Actor

Back up public GitHub repositories in bulk. Paste repo URLs and get a ZIP of each main branch, stamped with the exact commit SHA, plus one ZIP of the whole batch.

- **URL**: https://apify.com/onescales/github-backup-repo.md
- **Developed by:** [One Scales](https://apify.com/onescales) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: 5.00 out of 5 stars

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## GitHub Backup Repo

***

### Introduction

**Back up any public GitHub repository to a ZIP file in one click.** Paste one repository URL or a hundred, and get a downloadable ZIP of each repository's main branch, plus a single ZIP of the whole batch at the end. Download links are private and expire after 24 hours by default, or any time you choose up to about a month.

Each backup is stamped with the exact commit SHA it captured, along with the file count, size and a UTC timestamp. That gives you a dated record of exactly what the code looked like when you took the backup. If a repository has no `main` branch (older projects often use `master`), its default branch is backed up instead, so nothing is skipped.

The files are the archives GitHub itself produces, stored byte for byte. Nothing is modified or repackaged.

**Back up only repositories you own or are authorized to copy.** Being public on GitHub does not make code free to copy. Each repository's license and GitHub's Terms of Service still apply, and you are responsible for how you use what you download.

***

[![GitHub Backup Repo — video walkthrough](https://img.youtube.com/vi/zYphyKPuzh4/maxresdefault.jpg)](https://www.youtube.com/watch?v=zYphyKPuzh4)

### Features

- **Bulk backups** — back up one repository or hundreds in a single run, each with its own download link.
- **Commit-stamped** — every backup records the exact commit SHA, file count, size and time it was taken.
- **Main branch, with a fallback** — backs up `main`, and falls back to the default branch when a repository has no `main`.
- **Expiring download links** — links last 24 hours by default; set anywhere from 1 hour to 730 hours (about one month). The final row downloads every backup in one ZIP.
- **Accepts any GitHub link** — paste repository pages, `.git` clone URLs, SSH addresses or deep links into a folder.

***

### Use Cases

- **Off-site copy of your code** — keep a backup of your own public repositories outside GitHub, in case of account lockout, accidental deletion or a bad force-push.
- **Scheduled snapshots** — run it on a schedule to keep a dated history of your project's main branch.
- **Organization-wide backups** — back up every public repository your company or team owns in one run.
- **Release records** — capture the exact code you shipped, tied to a commit SHA.
- **Compliance and records** — keep a time-stamped archive of your own code for legal, licensing or audit needs.
- **Migration prep** — pull down your repositories before moving them to another host.
- **Client handover** — give a client a dated ZIP of the code you built for them.
- **Portfolio archiving** — keep a permanent copy of your personal projects.

***

### Input

Configure the actor in the UI, or pass the same fields via the API.

| Field | Type | Required | Description |
|-------|------|----------|-------------|
| **Repository URLs** (`repoUrls`) | Array of strings | Yes | Public GitHub repository URLs, one per line — e.g. `https://github.com/owner/repo`. Also accepts `github.com/owner/repo`, `.git` clone URLs, `git@github.com:owner/repo.git`, and links to any page inside the repository. |
| **Download Link Expiry (hours)** (`linkExpiryHours`) | Integer | No — defaults to `24` | How long the download links keep working. Minimum `1` hour, maximum `730` hours (about one month). |
| **Proxy Configuration** (`proxyConfiguration`) | Object | No — defaults to residential | Proxy used to download the repositories, which reduces rate-limiting and blocking. The backups themselves are unaffected. |

```json
{
  "repoUrls": [
    "https://github.com/onescales/token-proof-of-work"
  ],
  "linkExpiryHours": 24,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": ["RESIDENTIAL"]
  }
}
```

***

### Output

One row per repository, in the order you entered them, each with a download link. **The final row is the bulk download**: one link to a ZIP of every repository that was backed up successfully. Its `fileCount` is the number of repositories it holds, and `sizeBytes` is their combined size. Export the table as CSV, JSON, or Excel.

| Field | Type | Description |
|-------|------|-------------|
| `repoUrl` | string | The repository URL, normalized to `https://github.com/owner/repo` |
| `owner` | string | Repository owner (user or organization) |
| `repo` | string | Repository name |
| `branch` | string | `main`, or `default` when the repository has no main branch |
| `commitSha` | string | The exact commit captured in the backup |
| `fileName` | string | Name of the backup ZIP file |
| `fileCount` | integer | Number of files in the backup (repositories on the final row) |
| `sizeBytes` | integer | Size of the backup ZIP in bytes |
| `downloadUrl` | string | Download link for the backup, valid for `linkExpiryHours` (24 hours by default) |
| `backedUpAt` | string | When the backup was taken (ISO 8601, UTC) |
| `status` | string | `success`, or `error: <reason>` |

```json
[
  {
    "repoUrl": "https://github.com/onescales/token-proof-of-work",
    "owner": "onescales",
    "repo": "token-proof-of-work",
    "branch": "main",
    "commitSha": "808667a07d3988decfb81ac8a6260e30e55917ea",
    "fileName": "onescales-token-proof-of-work-main.zip",
    "fileCount": 5,
    "sizeBytes": 10421,
    "downloadUrl": "https://api.apify.com/v2/key-value-stores/.../records?signature=...&prefix=repo-00000-onescales-token-proof-of-work-main.zip",
    "backedUpAt": "2026-10-01T02:23:03.832Z",
    "status": "success"
  },
  {
    "repoUrl": "ALL REPOSITORIES (ZIP)",
    "owner": "N/A",
    "repo": "N/A",
    "branch": "N/A",
    "commitSha": "N/A",
    "fileName": "all-repos.zip",
    "fileCount": 1,
    "sizeBytes": 10421,
    "downloadUrl": "https://api.apify.com/v2/key-value-stores/.../records?signature=...&prefix=repo-",
    "backedUpAt": "2026-10-01T02:23:04.202Z",
    "status": "success"
  }
]
```

**Good to know**

- **Your own repositories only.** Use it for repositories you own or have permission to copy, and respect each repository's license.
- **Public repositories only.** A private, deleted or misspelled repository returns `error: repository not found or not public`.
- **What is in a backup:** the files on the branch at that moment, exactly as GitHub's "Download ZIP" gives them. It does not include git history, issues, pull requests, wikis, releases, or Git LFS file contents (LFS files appear as small pointer files).
- **Download links expire.** Each link works for `linkExpiryHours` hours from when it was created (24 by default, 1 to 730 allowed), then returns an access error. To download after that, open the run's storage in the Apify Console, for as long as your plan keeps run data. Your plan's data retention also limits links: a link set to 730 hours stops working early if the run's storage is deleted first.
- **Each link downloads a ZIP that contains the backup ZIP.** That is how Apify serves expiring links. Unzip once to get the repository archive (e.g. `onescales-token-proof-of-work-main.zip`), and the bulk link gives you every repository archive in one ZIP.
- **Links are scoped to the run, not the file.** Anyone who has a link and knows how to edit it can reach the other backups from the same run until it expires. Share links only with people you would share the whole run with.
- **Size limit:** repositories whose archive is over 2 GB are skipped with an error status. Archives are streamed to disk, never held in memory.
- **Retries:** every download and file write is retried up to 3 times (2s, 8s, 20s), including when GitHub rate-limits the run, with a fresh proxy IP on each attempt. Permanent failures such as a 404 fail fast.
- **Speed and memory:** repositories are backed up one at a time with a 2-second pause between them, to stay gentle on GitHub. In tests that works out to about 2–2.5 seconds per small repository: roughly 25–30 repositories a minute, so 100 take about 4 minutes and 500 about 20 minutes. Large repositories add their own download time. Memory stays flat at about 200 MB however large the repositories are, even for a 500 MB repository, so 512 MB is plenty, and more memory does not make it faster.
- **Proxy traffic:** downloads go through a residential proxy by default, and residential traffic is billed per GB on your Apify account. For very large repositories, switch to datacenter proxies or turn the proxy off in the input.
- **Pricing:** pay-per-event. You only pay for rows actually produced. Before it starts, the actor reads your spending limit and remaining account balance, backs up exactly as many repositories as the budget covers, and never goes over it.

***

### Support

**Need help or want a feature added?**

Contact Support at **<https://docs.google.com/forms/d/e/1FAIpQLSfsKyzZ3nRED7mML47I4LAfNh_mBwkuFMp1FgYYJ4AkDRgaRw/viewform?usp=dialog>** — Fill out this quick form.

We respond quickly and are happy to add new fields or custom integrations.

***

### Tags

github backup, backup github repository, github repo backup, download github repo, github zip download, download repository as zip, github archive, archive github repository, bulk github download, github repo downloader, github source code download, clone github repo, github mirror, repository snapshot, code snapshot, main branch backup, master branch download, default branch download, commit sha, open source backup, open source archiving, dependency backup, vendor code audit, software escrow, code preservation, git backup, git repository archive, source code archive, code backup tool, public repository backup, scheduled github backup, automated repo backup, disaster recovery, compliance archive, code audit trail, research dataset source code, offline code access, repository migration, github organization backup, developer tools, devops backup, zip archive, bulk download, expiring download links, temporary download link, apify actor

# Actor input Schema

## `repoUrls` (type: `array`):

Enter one or more public GitHub repository URLs, one per line — for example https://github.com/owner/repo. The main branch of each is downloaded; if a repository has no main branch, its default branch is used instead. Private repositories are not supported.

## `linkExpiryHours` (type: `integer`):

How long the download links stay valid, in hours. Defaults to 24 hours. Minimum 1 hour, maximum 730 hours (about one month). After that the links stop working. You can still download the files from the run's storage in the Apify Console for as long as your plan keeps run data.

## `proxyConfiguration` (type: `object`):

Proxy settings for downloading the repositories. Residential proxies reduce the chance of GitHub rate-limiting or blocking the run. Residential proxy traffic is billed per GB, so for very large repositories you can switch to datacenter proxies or turn the proxy off. The backups themselves are unaffected.

## Actor input object example

```json
{
  "repoUrls": [
    "https://github.com/onescales/token-proof-of-work"
  ],
  "linkExpiryHours": 24,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `repoBackups` (type: `string`):

Every repository with its branch, commit SHA, file count, size, and download URL. The last row links to one ZIP of every backup.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "repoUrls": [
        "https://github.com/onescales/token-proof-of-work"
    ],
    "linkExpiryHours": 24,
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("onescales/github-backup-repo").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "repoUrls": ["https://github.com/onescales/token-proof-of-work"],
    "linkExpiryHours": 24,
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("onescales/github-backup-repo").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "repoUrls": [
    "https://github.com/onescales/token-proof-of-work"
  ],
  "linkExpiryHours": 24,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call onescales/github-backup-repo --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,onescales/github-backup-repo"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/eTMiXlcnT32yCa8xu/builds/d9HYHdtuqYANO2Jtc/openapi.json
