# Instagram Post Scraper (`fertech/instagram-post-scraper`) Actor

- **URL**: https://apify.com/fertech/instagram-post-scraper.md
- **Developed by:** [Fertech](https://apify.com/fertech) (community)
- **Stats:** 3 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.49 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Instagram Post & Reel Scraper

Give it Instagram post, reel or IGTV URLs. Get back structured data with
**exact** like and comment counts — not the rounded "102K" the page shows.

No login, no cookies, no session tokens.

**One flat rate, whatever Apify plan you are on. No plan ladder — see the
pricing section of this listing for the current rate.**

### Why this one

**Exact numbers, not rounded ones.** Instagram's page says "24.8K". The
payload behind it says 24,815 — and that is what this Actor reads.

**No guessing.** Where Instagram publishes no number, the field is `null`:
never a substituted `0`, never a different metric wearing the wrong name.
`videoPlayCount` stays empty rather than being filled with the view count,
because plays and views are not the same thing.

### What you get

Feed it a list of Instagram post, reel or IGTV URLs, get one structured
record per post.

- **Engagement**: exact likes, exact comments, video views
- **Post**: caption, hashtags, mentions, type, product type
- **Media**: display dimensions, cover image, duration
- **Owner**: username, id
- **Music**: track metadata when Instagram publishes it

### Input

```json
{
  "directUrls": [
    "https://www.instagram.com/reel/CxAbc123DeF/",
    "https://www.instagram.com/p/CyDef456GhI/"
  ],
  "maxAttemptsPerUrl": 3
}
```

Share links with `?igsh=`, `?igsi=` or `?stkn=` tracking parameters work as
given — the parameters are stripped, so the same post submitted in two forms
is fetched, delivered and charged once.

URLs prefixed with the owner's username — the form Instagram serves when you
open a post from a profile, e.g.
`https://www.instagram.com/<username>/p/<code>/` or
`https://www.instagram.com/<username>/reel/<code>/` — are also accepted; the
username is dropped and the post is normalised to its bare `/p/` or `/reel/`
form.

### Output

One record per URL, with every field below always present. Field names,
types and shapes are exactly what the Actor emits; the values are
illustrative:

```json
{
  "inputUrl": "https://www.instagram.com/reel/CxAbc123DeF",
  "id": "3123456789012345678",
  "type": "Video",
  "shortCode": "CxAbc123DeF",
  "caption": "sunset over the harbour 🌅 #travel #photography #goldenhour",
  "hashtags": ["travel", "photography", "goldenhour"],
  "mentions": [],
  "url": "https://www.instagram.com/p/CxAbc123DeF/",
  "commentsCount": 312,
  "firstComment": "",
  "latestComments": [],
  "dimensionsWidth": 640,
  "dimensionsHeight": 1136,
  "originalWidth": null,
  "originalHeight": null,
  "displayUrl": "https://scontent.cdninstagram.com/v/t51.71878-15/000000000_0000000000000000_0000000000000000000_n.jpg?_nc_cat=…&oe=…",
  "images": [],
  "videoUrl": "",
  "alt": "",
  "likesCount": 24815,
  "videoViewCount": 247282,
  "timestamp": null,
  "childPosts": [],
  "ownerFullName": "",
  "ownerUsername": "example_user",
  "ownerId": "1234567890",
  "productType": "clips",
  "videoDuration": 18.7,
  "musicInfo": { "uses_original_audio": true },
  "isCommentsDisabled": false,
  "videoPlayCount": null,
  "paidPartnership": false
}
```

`displayUrl` is shortened above for readability. Real output carries the
full signed CDN URL — several hundred characters, and **valid for a limited
time**, so download the image rather than storing the link.

**Nine fields are always empty.** They are listed in the table below and
explained under [What this Actor does not return](#what-this-actor-does-not-return).

#### What the values mean

| Field | Notes |
|---|---|
| `likesCount`, `commentsCount` | Exact numbers from the page payload — not the rounded "102K" the page displays. |
| `dimensionsWidth` / `Height` | The display-size cover image (640×1136 above). |
| `originalWidth` / `Height` | Always `null` — the source resolution is not published on the surface this Actor reads. `dimensions*` above is the display size. |
| `videoDuration` | Seconds, fractional. Present for videos and reels only. |
| `type` | `Image`, `Video`, `Sidecar` (carousel), or `Unknown`. |
| `caption` | Raw caption text. `hashtags` and `mentions` are parsed out of it. |
| `alt` | Always `""` — the accessibility description is not published. |
| `musicInfo` | Track metadata (artist, song, audio id) when Instagram publishes it. When it publishes none, `{ "uses_original_audio": null }` — the Actor does not guess. |
| `firstComment`, `latestComments` | Always `""` and `[]` — comment bodies are not published. `commentsCount` is still the true total. |
| `timestamp` | Always `null` — the post's publication date is not published. |
| `videoUrl` | Always `""` — no playable media URL is published. `displayUrl` (the cover image) is. |
| `ownerFullName` | Always `""` — only `ownerUsername` and `ownerId` are published. |
| `isCommentsDisabled` | Always `false`. Not published, so this is a default rather than a reading. |
| `images`, `childPosts` | Always `[]` — carousel children are not expanded. |
| `videoViewCount` | A video's view count. `null` for images and carousels — Instagram publishes no count for those. |
| `videoPlayCount` | Always `null` — see below. |
| `paidPartnership` | Always `false` — not present in the payload. |

#### Views are not plays

`videoViewCount` is Instagram's **view** count. It is not the same as a
**play** count, which counts replays — for one reel Instagram reported
429,413 plays against 247,282 views. Instagram publishes no play count on
any public surface, so `videoPlayCount` is always `null` and this Actor will
not substitute the view count for it. A number in the wrong field is worse
than an empty one.

**A number that is missing is `null`, never `0`.** A zero here always means a
real zero. If Instagram did not publish a figure, this Actor says so rather
than reporting a plausible-looking number you cannot distinguish from data.

### Errors

A URL that cannot be scraped still produces a record, carrying `error` and
`errorDescription`:

| `error` | Meaning | Retried |
|---|---|---|
| `invalid-url` | Not a post, reel or IGTV URL | no |
| `not-found` | Deleted, private, or never existed | no |
| `unsupported-page` | The response was not an Instagram post page at all | no — re-fetching returns the same page |
| `network` | Connection, TLS, timeout or proxy failure | yes, up to `maxAttemptsPerUrl` |
| `blocked` | Instagram served a page but withheld the post data from every IP tried, or returned HTTP 401/403/429 | yes, up to `maxAttemptsPerUrl`; the session rotates once it has failed enough times |

Instagram intermittently serves a real page with the post data stripped out,
depending on which IP asks. It is transient: the Actor detects it, switches
to a different IP and retries, so it rarely reaches your dataset. If you do
see `blocked` records, raising `maxAttemptsPerUrl` gives each URL more IPs
to try.

### Pricing: every submitted URL is charged once

See the pricing section of this listing for the current rate.

**One URL in, one charge out.**

Posts get deleted, accounts go private, and Instagram sometimes withholds a
post's data. When a URL can't be scraped you get an error record instead of
a post record, so you always know what happened to it — and it costs the
same as a delivered one.

Retries and duplicates are not charged on top. You pay once per URL you
submit, however many attempts it takes.

### Requirements

**Residential proxies are recommended.** Instagram rate-limits repeated
requests from one IP, and sometimes serves a real page with the post data
stripped out depending on which IP asks.

The Actor uses Apify residential proxies by default, so there is nothing to
configure — no proxy picker in the input form, though an API caller can
still override it.

Residential proxies are **not included in the Apify Free plan.** On the Free
plan this Actor stops with a clear message rather than spending your credit
on requests that mostly cannot succeed.

### Limitations

- Single post URLs only. Profile, hashtag and location feeds are not
  supported.
- Carousel child media is not expanded.

#### What this Actor does not return

Nine fields are always empty, because the surface this Actor reads does not
publish them: `timestamp`, `videoUrl`, `originalWidth`, `originalHeight`,
`firstComment`, `latestComments`, `ownerFullName`, `alt` and
`isCommentsDisabled`.

They are present in every record, empty, rather than dropped — so existing
integrations keep parsing. If you need the post date or a playable media
URL, this Actor cannot give them to you.

### Reading the run summary

Each run writes a summary to the log and to the key-value store, under
`RUN_SUMMARY`:

```
requested 200 | delivered 187 (93.5%) | unpaid 0 | notfound 8 | failed 5 | blocked-responses 42 | callbackErrors 0
```

| Field | What it counts |
|---|---|
| `requested` | URLs you submitted |
| `delivered` | URLs that produced a post record |
| `notfound` | Deleted, private or non-existent posts |
| `failed` | URLs that ended in any other error record |
| `blocked-responses` | *Responses* where Instagram withheld the data — not URLs |
| `unpaid` | URLs processed after the charge budget ran out |
| `callbackErrors` | Rows that could not be written. Should always be 0 |

**These do not sum to `requested`**, and that is deliberate. Every URL ends
as exactly one of `delivered`, `notfound` or `failed`. `blocked-responses`
sits alongside them: one post withheld twice and then fetched adds 1 to
`delivered` and 2 to `blocked-responses`.

### Using the API

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")

run = client.actor("YOUR_USERNAME/instagram-post-scraper").call(run_input={
    "directUrls": ["https://www.instagram.com/reel/CxAbc123DeF/"],
})

for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item["likesCount"], item["ownerUsername"])
```

### Is scraping Instagram legal?

This Actor collects only publicly available data — information anyone can
see without logging in. It does not access private accounts or login-walled
content. You remain responsible for how you use the data, particularly
regarding personal information and the GDPR. If you plan to process personal
data, seek your own legal advice first.

### Issues and requests

Found a bug, or want profiles, hashtags or full comment threads supported?
Open an issue on the Actor's Issues tab — feature demand genuinely drives
what gets built next.

# Changelog

This Actor's version history is a separate document: https://apify.com/fertech/instagram-post-scraper/changelog.md

# Actor input Schema

## `directUrls` (type: `array`):

One or more Instagram post, reel or IGTV URLs — for example https://www.instagram.com/reel/CxAbc123DeF/. The username-prefixed forms (https://www.instagram.com/<username>/reel/<code>/) and share links with ?igsh= or ?stkn= tracking parameters also work. Profile, hashtag and location URLs are not supported.

## `maxAttemptsPerUrl` (type: `integer`):

How many times to retry a URL that fails at the network or proxy level, each on a new proxy session. Pages that load but carry no post data are not retried, because re-fetching returns the same page.

## Actor input object example

```json
{
  "maxAttemptsPerUrl": 3
}
```

# Actor output Schema

## `results` (type: `string`):

All scraped post records and any error records.

## `runSummary` (type: `string`):

Delivery counts for the run: requested, delivered, unpaid, notfound, failed.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("fertech/instagram-post-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("fertech/instagram-post-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call fertech/instagram-post-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,fertech/instagram-post-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/nBsqRAp8E6vRQxycf/builds/bxZ3TEnnQFgCdrVzV/openapi.json
