# X (Twitter) Profile Scraper — Accounts by Handle (`thenetaji/x-profile-scraper`) Actor

Turn a list of X handles into a clean table of accounts: who each one says it is, how big its audience is, when it joined, and whether it carries a verification mark today. Profile links and @handles work too, and no X login or account is involved.

- **URL**: https://apify.com/thenetaji/x-profile-scraper.md
- **Developed by:** [The Netaji](https://apify.com/thenetaji) (community)
- **Categories:** Social media, Automation, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.28 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## X (Twitter) Profile Scraper

Turn a list of X handles into a clean table of accounts. Each row says who the account claims to be,
how big its audience is, when it joined, where it points people, and whether it carries a
verification mark today.

No X login, cookie or session is involved at any point in a run. Nothing has to be connected, and no
account of anyone's is put at risk to read a public profile.

Bare handles, @handles and pasted profile links are all accepted, and a single list may mix them.

### Accepted input

`handles` is required and takes one or more accounts. `maxItems` caps how many rows are saved and
defaults to 100; setting it to `0` removes the cap and runs the whole list.

```json
{
  "handles": ["nasa", "@SpaceX", "https://x.com/ISS_Research"],
  "maxItems": 100
}
```

### Response fields

One row per account that resolves.

| Field | Contents |
|---|---|
| `user_id` | X's numeric identifier for the account, as a string |
| `handle`, `requested_handle` | X's casing of the handle, and the value the run asked with |
| `name`, `bio` | Display name and bio |
| `profile_url` | Link to the profile |
| `location` | Free text the account entered; not geocoded and not an address |
| `website` | The link in the profile header, as the account stated it |
| `joined_at` | Account creation time, in X's own date format |
| `followers_count`, `following_count` | Audience size, and how many accounts it follows |
| `tweet_count`, `media_count`, `listed_count` | Posts, posts carrying media, and public lists including the account |
| `favourites_count` | Posts the account has liked, in X's own spelling |
| `verified`, `is_blue_verified` | Legacy check and current mark |
| `protected` | Whether posts are restricted to approved followers |
| `profile_image_url`, `profile_banner_url` | Avatar and header image |
| `professional` | X's business-account object, whole; null on ordinary accounts |

```json
{
  "user_id": "11348282",
  "handle": "NASA",
  "requested_handle": "nasa",
  "name": "NASA",
  "profile_url": "https://x.com/NASA",
  "bio": "Making the seemingly impossible, possible. ✨",
  "location": "Pale Blue Dot",
  "website": "https://t.co/9NkQJKAVks",
  "joined_at": "Wed Dec 19 20:20:32 +0000 2007",
  "followers_count": 92309202,
  "following_count": 119,
  "tweet_count": 74339,
  "verified": false,
  "is_blue_verified": true
}
```

### Questions that come up

**Why do `verified` and `is_blue_verified` disagree?**

Because they are two different questions, and on most currently verified accounts the answers differ.
On the account sampled above, `verified` is `false` while `is_blue_verified` is `true`. `verified` is
the pre-2023 legacy check; `is_blue_verified` is the current mark, and it is the field to read for a
single answer. I publish both rather than merging them into one boolean, because a merge would
destroy the distinction and reading `verified` alone returns the wrong answer for most accounts that
carry a mark today.

**Which column joins back to the input list?**

`requested_handle`. X returns handles in its own casing, so a request for `nasa` produces `NASA` in
the `handle` field. `requested_handle` echoes the value supplied, reduced from an @handle or a pasted
profile link, so a join against an input list never depends on two sides agreeing about
normalisation. For anything tracked over time the key is `user_id`, which survives a rename; the
handle does not.

**Is `location` a real place?**

No. It is free text the account typed into the profile field, and nothing geocodes it, here or
upstream. `Pale Blue Dot` is a genuine value from the account sampled above. Treat the column as a
string, not as an address, and expect jokes, place names and nulls in the same column.

**Does `tweet_count` predict how many posts an export will return?**

No, and using it as a budget is the most common way to plan a run badly. It is how many posts X
states the account has made. A timeline walk I measured reached 122 posts over eight pages and was
still being offered more, so the reachable depth is a floor rather than a number this field predicts.
`tweet_count` describes an account; it does not size an export.

**What comes back for a protected account?**

Its public record, in full, with `protected` set to `true` on the row. A profile is public even when
the posts behind it are not. The timeline is a separate matter and is not readable, since this Actor
holds no X account and follows nobody.

**What happens to a handle that is wrong or retired?**

A handle outside X's own rule of 1 to 15 letters, digits or underscores is reported in the run log
and skipped before a request is spent on it. A handle that does not resolve to any account is
likewise reported and skipped, because X answers an unknown account as a normal empty response rather
than as an error and there is nothing to retry. The remaining accounts in the list are still
collected. Duplicates are collapsed before any request is made, so `nasa`, `@NASA` and
`https://x.com/nasa` in one list are one account and one charge.

**Why are `profile_banner_url` and `professional` sometimes null?**

Because the account never set a header image, and because the account is not a business account.
Fields absent from the upstream response are returned as null rather than omitted, so every row has
the same shape and a column never disappears mid-export. `professional` is republished whole rather
than flattened, since its shape varies by account type.

**Does `profile_image_url` give the full-size avatar?**

It gives the URL exactly as X states it, which usually carries a `_normal` size marker. It is not
rewritten to a larger variant, because that would be a guess about a CDN path rather than a value
anyone measured. The same restraint applies to `website`, which is often an X short link on accounts
that set it recently and is not followed or expanded here.

**Can follower and following lists, search results, or an account's media tab be collected?**

No, and not by any setting. Each of those needs a logged-in X account, and this Actor holds none.
That is the trade for an Actor that asks for no login, no cookie and no session, and it is a measured
boundary rather than something still to be built. The counts themselves are published:
`followers_count`, `following_count`, `media_count` and `listed_count` all arrive on the row, and a
post's media arrives with the post itself, inside that post's `entities`.

**Are `followers_count` and the other counts stable between runs?**

No. They are live counters, and two reads minutes apart will differ. That is drift rather than a
fault. `joined_at` is the only fixed value on the row, and it is left in X's own date format rather
than rewritten to ISO 8601, because rewriting a timestamp is how a timezone bug gets into data that
did not have one.

If a field returns null where a value is clearly present on x.com, the Actor's Issues tab is the
fastest way to reach me; a handle in the report is usually enough to reproduce it.

### Related Actors

[X (Twitter) Tweets Scraper](https://apify.com/thenetaji/x-tweets-scraper) exports what these
accounts posted, taking the same handles this Actor takes and paging each timeline by cursor.
[X (Twitter) Post Scraper](https://apify.com/thenetaji/x-post-scraper) takes post links directly and
returns one row per post, which is the shorter route when specific posts are already known.

# Actor input Schema

## `handles` (type: `array`):

One or more X accounts. Bare handles (`nasa`), @handles (`@nasa`) and profile links (`https://x.com/nasa`) are all accepted, and a single list may mix them. A handle outside X's own rule of 1-15 letters, digits or underscores is reported in the run log and skipped rather than ending the run.

## `maxItems` (type: `integer`):

Rows to save across all accounts in this run. Set 0 for no limit.

## Actor input object example

```json
{
  "handles": [
    "nasa",
    "https://x.com/SpaceX"
  ],
  "maxItems": 50
}
```

# Actor output Schema

## `dataset` (type: `string`):

All records scraped by this run

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "handles": [
        "nasa"
    ],
    "maxItems": 50
};

// Run the Actor and wait for it to finish
const run = await client.actor("thenetaji/x-profile-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "handles": ["nasa"],
    "maxItems": 50,
}

# Run the Actor and wait for it to finish
run = client.actor("thenetaji/x-profile-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "handles": [
    "nasa"
  ],
  "maxItems": 50
}' |
apify call thenetaji/x-profile-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,thenetaji/x-profile-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/k60H3TNDcR272tWLX/builds/NW31AAtqqbLOx8JGU/openapi.json
