# Instagram Content Scraper (`bgfc97/instagram-content-scraper`) Actor

Scrape Instagram posts and reels by profile, hashtag or URL: caption, likes, comments, views, media, owner and date. Public profile mode needs no login.

- **URL**: https://apify.com/bgfc97/instagram-content-scraper.md
- **Developed by:** [Bruno](https://apify.com/bgfc97) (community)
- **Categories:** Social media
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$2.00 / 1,000 item scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Instagram Content Scraper (Posts, Reels, Hashtags, Comments)

Scrape **public Instagram content** — no login required. Point it at a profile, a hashtag, or a
single post/reel URL and get back clean, structured data: one row per content item, plus optional
comments.

It calls Instagram's **own public web endpoints** (the same ones instagram.com uses to render
pages) through **Apify Proxy residential IPs** with browser-grade TLS/header fingerprints, rotating
IP/session on every retry. When residential keeps getting blocked it can escalate to the Apify
**UNBLOCKER** group.

> This actor is about **content**. If you want profile stats (bio, followers, verified, etc.), use a
> profile scraper instead — this one returns the posts/reels themselves.

### What it does

| Mode | Give it | You get |
|------|---------|---------|
| **Profile posts** | `@username`, `username`, or a profile URL | the profile's most recent posts & reels |
| **Hashtag posts** | `#travel` or `travel` | the hashtag's recent (or top) posts |
| **Single post/reel** | `https://www.instagram.com/p/CODE/` or `/reel/CODE/` | that one item, in full detail |

In **auto** mode (default) each line is classified on its own, so you can mix profiles, hashtags and
URLs in a single run.

### Input

```json
{
  "queries": ["nasa", "#travel", "https://www.instagram.com/reel/C8Qh1p_p3Kp/"],
  "mode": "auto",
  "maxItems": 30,
  "hashtagSection": "recent",
  "includeComments": false,
  "maxComments": 24,
  "sessionCookie": "",
  "proxyConfiguration": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"] },
  "useUnblockerFallback": true
}
```

- **queries** (required) — profiles, hashtags and/or post/reel URLs, one per line.
- **mode** — `auto` (detect per line), or force `user` / `hashtag` / `url`.
- **maxItems** — max posts per profile/hashtag (a URL always returns 1).
- **hashtagSection** — `recent` or `top`.
- **includeComments** / **maxComments** — also pull comments (author, text, likes, date) per item.
- **sessionCookie** *(optional, advanced)* — your **own** Instagram cookie (`sessionid=...; csrftoken=...`)
  to unlock more data and reduce blocking. Never use someone else's credentials.
- **proxyConfiguration** — residential strongly recommended.
- **useUnblockerFallback** — escalate to the UNBLOCKER group as a last resort.

### Output (one item per content)

```json
{
  "query": "nasa",
  "mode": "user",
  "username": "nasa",
  "id": "3419735883757493278",
  "shortcode": "C8Qh1p_p3Kp",
  "url": "https://www.instagram.com/p/C8Qh1p_p3Kp/",
  "type": "carousel",
  "is_video": false,
  "caption": "…",
  "likes": 123456,
  "comments_count": 789,
  "views": null,
  "taken_at": "2024-06-20T15:04:05.000Z",
  "thumbnail": "https://…",
  "media_url": "https://…",
  "dimensions": { "width": 1080, "height": 1350 },
  "carousel_media": [ { "is_video": false, "url": "https://…", "display_url": "https://…" } ],
  "owner_username": "nasa",
  "owner_id": "528817151",
  "owner_full_name": "NASA",
  "comments": [
    { "id": "…", "username": "someone", "text": "🔥", "likes": 12, "created_at": "2024-06-20T16:00:00.000Z" }
  ]
}
```

`type` is one of `image`, `video`, `carousel`, `reel`. `views` is populated for videos/reels.
`comments` appears only when **includeComments** is on.

### What works without login vs. what needs a cookie

Instagram has locked down most of its web API for logged-out visitors. Here is the honest state:

| Mode | Without login (public) | With `sessionCookie` |
|------|------------------------|----------------------|
| **Profile posts (@user)** | ✅ **Full data** — the profile's recent posts/reels with likes, comment counts, caption, media, owner, carousel, dimensions. Paginates for more. | ✅ Same, more stable + deeper pagination |
| **Single post/reel URL** | ✅ **Partial data** from the public post embed — shortcode, type, caption, comment count, thumbnail/media URL, owner. (Likes and exact timestamp aren't in the public embed.) | ✅ **Full data** incl. likes, timestamp, comments |
| **Hashtag posts** | ❌ Instagram redirects logged-out hashtag requests to its login page | ✅ Works with a valid cookie |
| **Comments** | ⚠️ Limited/none without login | ✅ Full comments (author, text, likes, date) |

The actor detects blocking and, for a logged-out run, returns a clear message telling you when a
`sessionCookie` is required (hashtags, full single-post detail, comments).

### Notes & honest limitations

- **Residential proxies are essential** (datacenter/no-proxy requests are almost always blocked).
  The actor rotates IP/session per attempt and escalates to `UNBLOCKER` as a last resort.
- It primes an anonymous `csrftoken` automatically, but Instagram still requires a logged-in session
  for hashtags, full single-post detail (its GraphQL returns 401 logged-out), and bulk comments — so
  supply your **own** `sessionCookie` to unlock those.
- Profile pagination and comment threads are best-effort and may stop early when logged out.
- Only public content is returned. Private profiles/posts require a valid `sessionCookie`, and even
  then only what that account can see.
- Respect Instagram's Terms and applicable laws; scrape responsibly.

# Actor input Schema

## `queries` (type: `array`):

What to scrape. Each line can be: a profile (@username, username, or a profile URL like https://www.instagram.com/nasa/) to get its recent posts & reels; a hashtag (#travel or travel) to get recent posts of that tag; or a single post/reel URL (https://www.instagram.com/p/CODE/ or https://www.instagram.com/reel/CODE/). In 'auto' mode each line's type is detected from its value. You can mix many lines in one run.

## `mode` (type: `string`):

How to interpret every line in 'Profiles, hashtags or post/reel URLs'. 'auto' detects per line (URL -> single post/reel, leading # -> hashtag, otherwise -> profile). Or force one type: 'user' (profile posts), 'hashtag', or 'url' (single post/reel).

## `maxItems` (type: `integer`):

Maximum number of content items to return for each profile or hashtag query (a single post/reel URL always returns exactly 1). The actor reads Instagram one page at a time and paginates (best-effort) until this many items are collected or no further page is available.

## `hashtagSection` (type: `string`):

For hashtag queries, whether to return the hashtag's 'recent' feed (most recent posts) or its 'top' posts (Instagram's most-engaged selection).

## `includeComments` (type: `boolean`):

If enabled, for every returned post/reel the actor also fetches its comments (author @, text, likes, date). This makes extra requests per item, so it is slower and more likely to hit rate limits on large runs.

## `maxComments` (type: `integer`):

When 'Also scrape comments' is on, the maximum number of comments to collect per post/reel. Instagram returns comments in pages; the actor paginates (best-effort) up to this number.

## `sessionCookie` (type: `string`):

OPTIONAL. Your own Instagram cookie string (e.g. "sessionid=...; csrftoken=...") copied from a logged-in browser. Providing it unlocks more data and greatly reduces blocking. Use ONLY a cookie from an account you own and accept the risk to that account — never share credentials. Leave empty to scrape purely public data.

## `proxyConfiguration` (type: `object`):

Apify Proxy settings, with a fresh session per retry attempt. Residential proxies are strongly recommended — Instagram blocks almost all datacenter IPs and no-proxy requests.

## `useUnblockerFallback` (type: `boolean`):

If enabled and residential proxies keep getting blocked for a query, the actor retries the last attempts through the Apify UNBLOCKER group, which defeats heavier anti-bot / JS challenges but is slower and more expensive. Used only as a last resort.

## `timeoutSecs` (type: `integer`):

Timeout per HTTP request to Instagram, in seconds (5-120).

## Actor input object example

```json
{
  "queries": [
    "nasa",
    "#travel",
    "https://www.instagram.com/reel/C8Qh1p_p3Kp/"
  ],
  "mode": "auto",
  "maxItems": 30,
  "hashtagSection": "recent",
  "includeComments": false,
  "maxComments": 24,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  },
  "useUnblockerFallback": true,
  "timeoutSecs": 40
}
```

# Actor output Schema

## `dataset` (type: `string`):

Scrape PUBLIC Instagram content, no login required: recent posts & reels of a profile (@user), recent posts of a hashtag, or a single post/reel URL. Returns one row per content item with id, shortcode, URL, type (image/video/carousel/reel), caption, likes, comment count, views, ISO date, thumbnail/media URL and owner. Optionally scrapes comments (author, text, likes, date). Uses Instagram's own public web endpoints through Apify Proxy residential IPs with browser-grade TLS/headers, and can escalate to the UNBLOCKER group when Instagram blocks residential.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "nasa",
        "#travel",
        "https://www.instagram.com/reel/C8Qh1p_p3Kp/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("bgfc97/instagram-content-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "queries": [
        "nasa",
        "#travel",
        "https://www.instagram.com/reel/C8Qh1p_p3Kp/",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("bgfc97/instagram-content-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "nasa",
    "#travel",
    "https://www.instagram.com/reel/C8Qh1p_p3Kp/"
  ]
}' |
apify call bgfc97/instagram-content-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,bgfc97/instagram-content-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/attK0Mmbp2UMCT3PH/builds/4Y6a9VnxH1mpmPVCz/openapi.json
