# Website Profile Scraper (`reapx/instagram-following-scraper`) Actor

Compare profiles from the open web without opening them one by one. From web page URLs, each web record retains profile link, names, profile links, and response status.

- **URL**: https://apify.com/reapx/instagram-following-scraper.md
- **Developed by:** [ReapX](https://apify.com/reapx) (community)
- **Categories:** SEO tools, Developer tools, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $70.00 / 1,000 followed accounts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Website Profile Scraper

Compare profiles from the open web without opening them one by one. From web page URLs, each web record retains profile link, names, profile links, and response status.

![Website Profile Scraper interface](https://reapx.dev/assets/products/website-profile-scraper/readme.png)

### The record

The dataset schema names every field before the run. The first working set is `requestedUrl`, `finalUrl`, `route`, `scrapedAt`, `wireBytes`, `decodedBytes`, `responseStatus`, `pagesChecked`, `profileUrl`, `entityType`, `name`, and `jobTitle`. Dataset views keep related fields together without changing the underlying row.

#### Captured row

```json
{
  "entityType": "Person",
  "responseStatus": 200
}
```

### Input

Website Profile Scraper accepts source URLs. Run controls stay in the same form.

| Field | What it controls | Starting value |
| --- | --- | --- |
| `startUrls` | Paste exact URLs from the open web, one per line. | `["https://www.kottke.org","https://about.gitlab.com"]` |
| `maxItems` | Stop after this many dataset rows. | `20` |
| `maxSeconds` | Stop after this many seconds and keep completed rows. | `180` |

#### Example input

```json
{
  "startUrls": [
    "https://www.kottke.org",
    "https://about.gitlab.com"
  ],
  "maxItems": 3,
  "maxSeconds": 180
}
```

### Price

$100 per 1,000 followed accounts on the Free plan. Other Apify plans use the rates shown in the Pricing tab.

![Website Profile Scraper input and result demonstration](https://reapx.dev/assets/products/website-profile-scraper/demo.svg)

### Console, API, schedules, and exports

Runs can begin in Apify Console, from a saved task, or through the Actor API. A schedule can reuse the same input. Completed rows remain in the run dataset for API retrieval and Apify dataset exports.

```text
POST https://api.apify.com/v2/acts/B60RE4TVT4GPgcyAI/runs
GET  https://api.apify.com/v2/datasets/{datasetId}/items
```

### Saved tasks

Twenty task pages cover distinct research, comparison, operations, automation, and export jobs. The opening set is:

- **Web profile single-page capture**: Review one web record around `name`, `entityType`, `jobTitle`, and `affiliation`. The saved task uses `startUrls`, `maxItems`, and `maxSeconds` and opens the `overview` view. Configured in Website Profile Scraper.
- **Web profile page comparison**: Compare web profiles using `requestedUrl`, `finalUrl`, `route`, and `scrapedAt`. The `identity` view keeps the differences close together. Configured in Website Profile Scraper.
- **Web profile URL list**: Process a saved web input queue with `startUrls`, `maxItems`, and `maxSeconds`. Source and identifier fields remain visible in the `source` view. Configured in Website Profile Scraper.
- **Web profile canonical index**: Index web profiles by `name`, `scrapedAt`, `wireBytes`, and `decodedBytes`. The saved input keeps the same matching keys from run to run. Configured in Website Profile Scraper.
- **Web profile metadata register**: Assemble a focused web register centered on `requestedUrl`, `finalUrl`, `route`, and `scrapedAt`. The task keeps `startUrls`, `maxItems`, and `maxSeconds` visible for later review. Configured in Website Profile Scraper.
- **Web profile response benchmark**: Compare a second field set across web profiles, led by `requestedUrl`, `finalUrl`, `route`, and `scrapedAt`. Results open as the `proof` table. Configured in Website Profile Scraper.

### Related products

- [Website Email Scraper](https://apify.com/reapx/shopify-email-scraper)
- [Website Data Scraper](https://apify.com/reapx/upwork-job-scraper)
- [Website Lead Scraper](https://apify.com/reapx/notion-scraper)
- [Web Event Data Scraper](https://apify.com/reapx/maps-scraper)
- [Website Contact Scraper](https://apify.com/reapx/linkedin-contact-scraper)

### Support

Include the Actor ID, run ID, saved task name, and affected input when reporting an issue. That is enough to locate the run and its dataset.

# Actor input Schema

## `startUrls` (type: `array`):

Paste exact URLs from the open web, one per line.

## `maxItems` (type: `integer`):

Stop after this many dataset rows.

## `maxSeconds` (type: `integer`):

Stop after this many seconds and keep completed rows.

## Actor input object example

```json
{
  "startUrls": [
    "https://www.kottke.org",
    "https://about.gitlab.com"
  ],
  "maxItems": 20,
  "maxSeconds": 180
}
```

# Actor output Schema

## `results` (type: `string`):

Rows returned by this run.

## `json` (type: `string`):

Clean JSON records from this run.

## `csv` (type: `string`):

CSV export from this run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://www.kottke.org",
        "https://about.gitlab.com"
    ],
    "maxItems": 20,
    "maxSeconds": 180
};

// Run the Actor and wait for it to finish
const run = await client.actor("reapx/instagram-following-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [
        "https://www.kottke.org",
        "https://about.gitlab.com",
    ],
    "maxItems": 20,
    "maxSeconds": 180,
}

# Run the Actor and wait for it to finish
run = client.actor("reapx/instagram-following-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://www.kottke.org",
    "https://about.gitlab.com"
  ],
  "maxItems": 20,
  "maxSeconds": 180
}' |
apify call reapx/instagram-following-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,reapx/instagram-following-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/B60RE4TVT4GPgcyAI/builds/EP7MndC3i4a0fi2tP/openapi.json
