# X (Twitter) Post Scraper — Full Data From a Link (`thenetaji/x-post-scraper`) Actor

Paste X post links and get the numbers behind them: what the post says, when it went out, and how much it was liked, shared, bookmarked and seen. A retweet or quote brings the original back in full. No X login or account is involved.

- **URL**: https://apify.com/thenetaji/x-post-scraper.md
- **Developed by:** [The Netaji](https://apify.com/thenetaji) (community)
- **Categories:** Social media, Automation, Marketing
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.85 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## X (Twitter) Post Scraper

Paste X post links and get the numbers behind them. Each link returns one row: what the post says,
when it went out, and how much it was liked, retweeted, replied to, quoted, bookmarked and seen, with
the author and the author's follower count on the same row. A retweet or a quote brings the original
post back in full alongside it, so the text being amplified is never lost.

No X login, cookie or session is involved at any point in a run. Nothing has to be connected, and no
account of anyone's is put at risk to read a public post.

### Accepted input

`tweet_ids` is required and takes one or more posts. Post links and the numeric identifier inside
them are both accepted, and a single list may mix them. `enrichParentPost` attaches the post that
each reply is answering, and is off by default.

```json
{
  "tweet_ids": [
    "https://x.com/nasa/status/2087966752512561302",
    "2087966752512561302"
  ]
}
```

### Response fields

One row per post that still exists.

| Field | Contents |
|---|---|
| `tweet_id`, `tweet_url` | The post's identifier, and a link to it |
| `requested_tweet_id` | The identifier the run asked for, reduced from a pasted link |
| `text` | The post body; `is_retweet` and `is_long_post` qualify it |
| `is_long_post` | `true` when `text` is a long post's complete body rather than its first 280 characters |
| `posted_at`, `lang` | Publish time in X's own date format, and X's language guess |
| `author_id`, `author_handle`, `author_name` | The account that posted it |
| `author_followers_count` | That account's follower count |
| `author_verified`, `author_is_blue_verified` | Legacy check and current mark |
| `author_profile_image_url` | The author's avatar |
| `like_count`, `retweet_count`, `reply_count`, `quote_count` | Engagement counts |
| `bookmark_count`, `view_count` | Bookmarks and impressions, both as numbers |
| `is_retweet`, `retweeted_tweet` | Whether the row is a retweet, and the original in full when it is |
| `is_quote`, `quoted_tweet` | Whether the post quotes another, and that post in full |
| `conversation_id`, `reply_to_tweet_id`, `reply_to_handle` | Position within a thread |
| `entities` | X's own `hashtags`, `urls`, `user_mentions`, `symbols` and `media` structure |
| `possibly_sensitive` | X's sensitive-content flag |
| `reply_to_tweet` | The post this row replies to; present only when `enrichParentPost` is on |

```json
{
  "tweet_id": "2087966752512561302",
  "requested_tweet_id": "2087966752512561302",
  "tweet_url": "https://x.com/NASA/status/2087966752512561302",
  "text": "Yesterday's total solar eclipse in parts of Europe – and the partial eclipse in parts of North America – gave us spectacular views. @NASAHQPhoto was on location to capture the sights. See more: https://t.co/IQ6yWeVbXl",
  "posted_at": "Thu Aug 13 18:17:21 +0000 2026",
  "lang": "en",
  "author_handle": "NASA",
  "author_followers_count": 92309471,
  "like_count": 16970,
  "reply_count": 258,
  "view_count": 1792803,
  "bookmark_count": 1180,
  "is_retweet": false,
  "is_quote": false
}
```

### Questions that come up

**Is there a limit on how many posts one run takes?**

There is no `maxItems` here, because the length of the list is the size of the run: each entry
produces at most one row, and one request. That is the difference from a timeline export, where a
single request returns a page of posts and a cap is the only thing keeping a run finite.

**Which link forms are accepted?**

A post link is reduced to the identifier inside it, in both spellings X has used for the path segment
(`/status/` and the legacy `/statuses/`), and on both domains it answers on, so old `twitter.com`
links work unchanged. A bare numeric identifier is taken as it stands. Anything that is neither is
reported in the run log and skipped before a request is spent on it, and the rest of the list still
runs. Duplicates are collapsed before any request is made, so a link and the identifier inside it in
the same list are one post and one charge.

Typing identifiers out of URLs should rarely be necessary. Every row the Tweets Scraper saves
publishes `tweet_id`, as does every `retweeted_tweet` and `quoted_tweet` nested inside one, so a
timeline export feeds this Actor directly with nothing constructed and nothing guessed.

**What happens to a post that has been deleted?**

It is reported in the run log and skipped, producing no row and no charge, and the remaining posts in
the list are still collected. A post that never existed, or that sits on a protected account, reads
the same way: X answers with an empty response rather than an error, so there is nothing to retry.

**Why does `text` start with `RT @` on some rows?**

Because the row is a retweet, and that truncated form is X's own. The original post is published
whole on the same row under `retweeted_tweet`, with its own author, engagement counts and untruncated
text, so nothing is lost. Quote posts behave the same way: `text` is the commentary and
`quoted_tweet` is the post being commented on. Both halves are published rather than one being
substituted for the other, because they answer different questions: the outer record is what the
account did, and the nested one is what it amplified.

**Are long posts cut off at 280 characters?**

No. When `is_long_post` is `true`, `text` is the complete body. X's ordinary post field would return
a 4,000-character post as its first 280 characters, with no ellipsis and no marker, which is a
truncation that reads exactly like a complete post; that is the failure the flag exists to prevent.

**Where are the images and videos on a post?**

Inside `entities`, in X's own structure, alongside `hashtags`, `urls`, `user_mentions` and `symbols`,
each carrying the character offsets into `text` where it appears. Anything already written against a
Twitter payload reads those keys as they stand. There is deliberately no flattened column of media
links beside them: the shape varies by attachment type, and publishing one column would mean guessing
at key names I have not measured.

**Can the replies under a post be collected, or search results, follower lists or a media tab?**

No, and not by any setting. Each of those needs a logged-in X account, and this Actor holds none.
That is the trade for an Actor that asks for no login, no cookie and no session, and it is a measured
boundary rather than something still to be built. `reply_count` still reports how many replies X
counts on a post. A thread is therefore walkable upward and not downward: `reply_to_tweet_id` is an
ordinary post identifier, so feeding it back into `tweet_ids` fetches the parent, one hop at a time.

**What does `enrichParentPost` actually cost?**

One extra request per reply, billed once per reply whose parent is returned. Posts that are not
replies are untouched and cost nothing, and a parent that has been deleted or sitting on a protected
account is not billed. The same hop can be taken by hand, since `reply_to_tweet_id` is a plain post
identifier; the option exists to do it inside one run rather than two.

**Why is `view_count` empty on some posts?**

Because X published no view count for them, which is normal on older posts and on some reply-scoped
ones. Null means that no count was stated and is never collapsed to `0`, so an average over the
column is not quietly dragged towards zero by posts that simply predate the metric. Where it is
present it reports impressions, which typically run one to two orders of magnitude above the like
count. Views and bookmarks are both numbers rather than text, so sorting on them orders `54752` above
`9` instead of the reverse.

**Are missing fields dropped from a row?**

No. Every row has the same shape, and a field absent upstream is `null` rather than omitted. The
three flags a reader has to branch on are the exception in the other direction: `is_retweet`,
`is_quote` and `is_long_post` are always `true` or `false`, never null, so a check on one never has
to handle a third case.

**Do these rows line up with a timeline export?**

Exactly. The row shape is identical to the one
[X (Twitter) Tweets Scraper](https://apify.com/thenetaji/x-tweets-scraper) produces, so a dataset
from one stacks with a dataset from the other without a second parser. The only difference is the
echo column: timeline rows carry `requested_handle`, and rows here carry `requested_tweet_id`. The
nested `retweeted_tweet`, `quoted_tweet` and `reply_to_tweet` records are a level below that. They
carry the same information as a row, under X's own key names and with the author kept as one nested
object rather than spread across `author_*` columns, because flattening a twin of every post would
double the width of every row to serve the minority that carries one.

If a field returns null where a value is clearly present on x.com, the Actor's Issues tab is the
fastest way to reach me; a post identifier in the report is usually enough to reproduce it.

### Related Actors

[X (Twitter) Tweets Scraper](https://apify.com/thenetaji/x-tweets-scraper) exports a whole account's
posts from its handle and returns this same row shape. It is the cheaper route per row, because one
request returns a page of posts rather than one; this Actor is the direct route when the posts are
already known. [X (Twitter) Profile Scraper](https://apify.com/thenetaji/x-profile-scraper) returns
the accounts behind these posts, including the bio, join date and the full set of counts X publishes.

# Actor input Schema

## `tweet_ids` (type: `array`):

One or more X posts. Post links (`https://x.com/nasa/status/2087966752512561302`) and the numeric ID inside them are both accepted, and a single list may mix them. Every row the Tweets Scraper saves carries that ID, so a timeline export is directly consumable here.

## `enrichParentPost` (type: `boolean`):

Fetches the post that each reply is answering and attaches it in full, with its own text, author and engagement counts. Rows that are not replies are unaffected and cost nothing. One extra request per reply, billed per reply whose parent is returned; a parent that has been deleted, or that sits on a protected account, is not billed.

## Actor input object example

```json
{
  "tweet_ids": [
    "https://x.com/nasa/status/2087966752512561302"
  ],
  "enrichParentPost": false
}
```

# Actor output Schema

## `dataset` (type: `string`):

All records scraped by this run

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "tweet_ids": [
        "https://x.com/nasa/status/2087966752512561302"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("thenetaji/x-post-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "tweet_ids": ["https://x.com/nasa/status/2087966752512561302"] }

# Run the Actor and wait for it to finish
run = client.actor("thenetaji/x-post-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "tweet_ids": [
    "https://x.com/nasa/status/2087966752512561302"
  ]
}' |
apify call thenetaji/x-post-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,thenetaji/x-post-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/KOfUT08E9op874Kh9/builds/SX10COrGpwX9HYZol/openapi.json
