# X (Twitter) Tweet Scraper — browserless (`arthurvianna/x-tweet-scraper`) Actor

Scrapes public tweets from X without a browser engine, using guest-token GraphQL over plain HTTP. Free runs return up to 10 results.

- **URL**: https://apify.com/arthurvianna/x-tweet-scraper.md
- **Developed by:** [Arthur Vianna](https://apify.com/arthurvianna) (community)
- **Categories:** Lead generation, Social media
- **Stats:** 1 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## x-tweet-scraper

A browserless X (Twitter) scraper, shipped as an Apify Actor.

No browser engine, no account credentials, no purchased tweet data. It talks to X's own
GraphQL API over plain HTTP with guest tokens, the same surface a logged-out visitor gets,
and it ships with a free-tier limit that lives on the server and cannot be lifted by
editing the input.

Two things about it are worth your time. The first is that **X gates keyword search for
guest tokens**, which quietly invalidates the obvious design — so the Actor takes a
different route to the same data, and §1–§2 show the measurements that forced it. The
second is the **free-tier gate** in §5, which is harder than it looks: the run executes on
the customer's own Apify account, holding the customer's own token, writing to the
customer's own storage. Almost every natural place to put the limit is somewhere they
control.

|         |                                                    |
| ------- | -------------------------------------------------- |
| Store   | https://apify.com/arthurvianna/x-tweet-scraper     |
| Source  | https://github.com/ArthurVianna96/x-tweet-scraper  |
| Console | https://console.apify.com/actors/PNZugrwspnmMj70at |

```bash
npm install
npm test                     # 221 tests, offline, ~0.5s
npm run probe                # re-derive the endpoint capability matrix (~20s, no credentials)
npm run start:dev            # run the Actor locally
```

Every measurement quoted below is reproducible. `npm run probe` re-runs the endpoint
probe, `src/tools/benchmark.ts` re-runs the performance numbers, and every run writes its
own diagnostics to `OUTPUT` so you can contradict the claims with the Actor itself. The
raw evidence, with dates, is in [`docs/README-data-source.md`](docs/README-data-source.md).

#### Which surfaces are implemented

The brief (§2a) asks us to say this plainly, so:

| Surface                | Operation                           | Status                                      |
| ---------------------- | ----------------------------------- | ------------------------------------------- |
| **Tweets by author**   | `UserTweets`                        | ✅ required — the extraction engine         |
| **Single tweet by id** | `TweetResultByRestId`               | ✅ required — `tweetIds`, one request each  |
| **Profile by handle**  | `UserByScreenName`                  | ✅ required — returns §5's `author` in full |
| **Free-text search**   | `SearchTimeline` is `404` to guests | ⚠️ **stretch, served a different way** (§3) |

All three required surfaces are guest-reachable, browserless, and at the §5 schema.

`searchTerms` works, but not through X: X's search timeline is walled to guests, so the
keywords are answered by seeding account discovery from a public web index and filtering
natively (§3). That is the "equivalent public HTTP source you justify" route rather than
the programmatic-auth route, and its honest limitation is that **recall is seed-bounded** —
you get matching tweets from accounts that discuss the topic, not every tweet on X (§10).
It is not rejected as unsupported; it is implemented, measured, and scoped.

***

### Contents

1. [The finding everything rests on](#1-the-finding-everything-rests-on)
2. [Architecture: separating *who* from *what*](#2-architecture-separating-who-from-what)
3. [Keyword search, without X's search](#3-keyword-search-without-xs-search)
4. [Scraping X without a browser](#4-scraping-x-without-a-browser)
5. [The free tier you cannot edit away](#5-the-free-tier-you-cannot-edit-away)
6. [Speed and cost, measured](#6-speed-and-cost-measured)
7. [Output contract, and the calls behind it](#7-output-contract-and-the-calls-behind-it)
8. [Running it](#8-running-it)
9. [robots.txt, ToS, and what we would tell a client](#9-robotstxt-tos-and-what-we-would-tell-a-client)
10. [What it cannot do](#10-what-it-cannot-do)
11. [Decisions and trade-offs](#11-decisions-and-trade-offs)

***

### 1. The finding everything rests on

The brief says X's search timeline is auth-walled to guests, and asks us to work out
**which operations are guest-reachable today** and scope the feature set around them. So
the interesting question is not whether `SearchTimeline` is closed — it is closed — but
where exactly the wall runs.

`SearchTimeline` returns `404` with a zero-length body to a guest token, from every host,
method and product variant we tried. The guest tokens themselves are perfectly healthy;
the gate is on the *operation*.

Rather than take that on trust, `npm run probe` re-derives the whole matrix in about 20
seconds without credentials. The method turns on one useful detail: **X distinguishes
refusal from rejection by status code.**

| Status                          | What it means                                                                |
| ------------------------------- | ---------------------------------------------------------------------------- |
| `404`, zero-length body         | operation **gated** — refused before X even validated the request            |
| `422 GRAPHQL_VALIDATION_FAILED` | operation **permitted** — it reached validation and our variables were wrong |
| `200`                           | operation **permitted**                                                      |

So a `404` here is not a typo in a path or a stale `queryId`. Operations that *are*
permitted answer with a descriptive `422` on the same token in the same second.

Of 23 operations probed on 2026-08-17, **five are open to guests**:

| Open to guests        | What it gives us                                                  |
| --------------------- | ----------------------------------------------------------------- |
| `UserTweets`          | **the extraction engine** — full tweet objects + cursors          |
| `UserByScreenName`    | the profile surface — handle → `userId` and §5's `author`         |
| `TweetResultByRestId` | the by-id surface — one fully hydrated tweet per `tweetIds` entry |
| `GenericTimelineById` | unused                                                            |
| `TrendHistory`        | trend metadata, no tweets — *and newly permitted*                 |

Everything else is `404`: `SearchTimeline`, `ListSearchTimeline`, `ExplorePage`,
`TrendRelevantUsers`, `Followers`, `Following`, `SimilarPosts`, `TweetDetail`, `UserMedia`,
`UserOriginalsTimeline` and every other narrow timeline variant. That permitted set is not arbitrary — it is exactly
what a logged-out browser can render: **one profile, or one tweet.** Search is not on it,
and neither is anything adjacent to it. It also maps one-to-one onto the three required
surfaces, which is the point: the doors that are open are exactly the ones the brief asks
us to build on.

`TrendHistory` is the interesting entry. It was gated on 2026-08-14 and permitted on
2026-08-17. This surface is undocumented and it moves, which is the whole argument for
resolving `queryId`s at runtime instead of shipping them (§4).

We looked for a way around the gate before accepting it. Every attempt failed, including
`SearchTimeline` via `api.x.com` / `twitter.com` / POST / `product=Top`, X's legacy iPhone
bearer, `/i/api/2/search/adaptive.json`, the `cdn.syndication.twimg.com` timelines, and
`x.com/hashtag/<tag>` logged out. Requesting `/search` with a Googlebot user-agent also
returns `404` — X verifies crawlers by reverse DNS rather than by UA string, and we did
not attempt to defeat that. The full list, with status codes, is in
[the appendix](docs/README-data-source.md#3-workarounds-ruled-out).

**Free-text search is not reachable from the guest surface.** The three required surfaces
are. Everything below follows from that.

***

### 2. Architecture: separating *who* from *what*

Search is closed, but profiles and timelines are wide open. So the problem splits in two:
work out **which accounts** to read, then **read their tweets**. Only the first half ever
leaves X, and it happens once.

```
  tweetIds ───────────────────────────►  TweetResultByRestId  ─┐   one request per id
                                                               │
  searchTerms / hashtags ─┐                                    │
                          ├─►  DiscoveryStrategy (port) ─┐     │
  fromUsers ──────────────┘     ├─ DirectHandleDiscovery │     │
                                └─ SeededTopicDiscovery  │     │
                                                         ▼     │
                                      UserByScreenName → userId│  ← 100% native X from here
                                                         │     │
                                                         ▼     │
                                      UserTweets (cursor-paged)│
                                                         │     │
                                                         ▼     ▼
                    normalize → filter → dedupe → ResultSink → dataset
                                             ▲
                                    free-tier cap enforced here
```

Both sources are lazy generators feeding one sink, which is what lets the cap stop the
*fetching* on either surface (§5.3).

That split is what keeps the compromise contained. Discovery sits behind a port, so the
one part of the problem X refuses to help with is isolated in a single swappable adapter —
and the other 95% of the system neither knows nor cares which strategy ran.

**`UserTweets` returns complete tweet objects** — full text, all six metrics, entities,
media, author. Nothing needs a second hydration call. The cost unit is *one request ≈ 20
tweets*, not one request per tweet, and that single fact is why the native path is fast
(§6).

**Growing the account set is free.** Mentions and retweeted authors are already sitting in
pages we have paid for, so the frontier expands at zero extra request cost. It is
depth-limited (default 1) because mentions from a topical account are not all topical, and
precision decays quickly.

### 3. Keyword search, without X's search

`searchTerms` is the one surface X does not hand us. The brief scopes it as a **stretch,
not a requirement** — implement the author/id/profile paths and say honestly that search
is walled. It is implemented here anyway, through a public web index rather than through
X, which the brief names as the "equivalent public HTTP source you justify" route.

The trick is *what question you ask the index*.

#### Why we ask a search engine about people, not posts

The obvious version of external discovery is to search the web for tweet URLs and hydrate
them by id. We built that, measured it, and threw it away.

For a keyword query the freshest indexed tweet was **36 days old** (median 181). A hashtag
query returned **zero** tweet URLs — only profile pages. You cannot serve `sortBy: latest`
honestly out of a web index.

The same measurement is what rescues the *handle* variant. Search engines index X
**profiles** well and X **posts** slowly. So we ask them only the question they can
actually answer — *who talks about this?* — and get recency from X itself. The lookup runs
**once per run at cold start**, is **skipped entirely** when `fromUsers` is supplied, and
sits behind the port so it can be swapped for a static roster without touching extraction.

#### Discovery is a cascade, not a dependency

Search engines fight automation, and they do not fight it consistently. Measured
2026-08-17 from one IP after roughly a dozen queries: DuckDuckGo started answering `202`
with a 14 KB anti-bot challenge, while Brave answered `200` with usable results on the
same query in the same minute. Mojeek and Ecosia returned `403`. Bing returned `200` but
wraps every result in a base64 redirect, so no handles survive its HTML.

So `SeededTopicDiscovery` walks an ordered list — DuckDuckGo's HTML and Lite endpoints,
then Brave, then Startpage — cheapest first, since this traffic crosses the same paid
proxy as everything else. The first engine that yields handles wins. A blocked engine falls through. A throwing engine does not fail the run. All
engines blocked yields no seeds and a logged warning, not a crash.

#### What it costs, and what it cannot do

Once the handles are in hand the run is 100% native X: profiles resolve to ids, timelines
page normally, and the keyword is applied as a filter over the normalized text. Mentions
and retweeted authors found along the way widen the frontier for free.

The honest limit is **recall**. You get matching tweets from accounts that discuss the
topic, not every tweet on X, and no amount of engineering closes that gap from the guest
surface. Selectivity is low as a result — a measured keyword run matched 1.6% of what it
fetched, against 84% on the author path — which is why keyword runs cost ≈$7.49 per 1k
results against ≈$0.41 (§6), and why the Actor carries a request budget.

Supplying `fromUsers` removes the dependency entirely: no search engine, no third party,
nothing but X.

#### What we considered instead

| Option                                                       | Why not                                                                                                                                                                                                                                                                                           |
| ------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Paid tweet-search API** (e.g. a pay-per-event Apify actor) | Keyword results immediately — and the X-specific extraction becomes the vendor's work, not ours. Worth knowing: `apidojo-io/twitter-scraper-lite`, the widely-referenced "X scraper", makes **no HTTP calls to X at all**. It is an `apify_client` wrapper around paid actor `nfp1fpt5gUlBwPcor`. |
| **X official API v2 recent search**                          | Legitimate, but needs a paid app bearer — a hardcoded server-side credential, which is exactly the mechanism the brief's §3 rules out.                                                                                                                                                            |
| **Search index → tweet IDs → hydrate**                       | Measured and rejected: 36-day-old freshest result, zero hashtag coverage.                                                                                                                                                                                                                         |
| **Logged-in account pool**                                   | Prohibited by §3 of the brief.                                                                                                                                                                                                                                                                    |

***

### 4. Scraping X without a browser

No browser engine is installed and none is needed. The entire client is
[`got-scraping`](https://github.com/apify/got-scraping) behind a one-function port:

```ts
export type HttpClient = (req: HttpRequest) => Promise<HttpResponse>;
```

Everything that leaves the process goes through that function. It is why 221 tests run
offline in ~0.5 s with no module mocking anywhere: a test hands the constructor a canned
responder and asserts. The port is deliberately status-code-transparent — a `404` or `429`
is a *response*, not an exception — because the error taxonomy cannot classify what the
transport has already thrown away.

Three pieces of protocol knowledge make the browserless path actually work.

#### 1. Resolve `queryId`s at runtime, never ship them

X's GraphQL endpoints are keyed by an opaque `queryId` that rotates with every frontend
deploy. We watched three distinct bundle hashes in three days —
`main.e4aca26a.js` → `main.4f5b42da.js` → `main.b07c4c6a.js` — one of those changes
landing inside a single afternoon. Hardcoding them guarantees a silent `404` on some
future Tuesday.

So the Actor fetches `x.com/explore` (not `x.com/`, which serves a server-rendered login
wall with no app bundle), extracts `main.<hash>.js`, and parses the webpack modules
carrying `{queryId, operationName, metadata:{featureSwitches, fieldToggles}}`. Cached for
the run, refreshed exactly once on an unexpected `404`, fatal after that.

#### 2. Extract by path, never by type

Tweets are read at exactly
`instructions[].entries[].content.itemContent.tweet_results.result` — plus the module
variant one level deeper — and never recursed into.

The tempting alternative is to walk the payload and collect every `__typename === "Tweet"`.
It is wrong, and expensively so: a tweet's `retweeted_status_result` and
`quoted_status_result` are structurally identical to a top-level tweet, so recursion turns
a 20-entry page into **32 items**, the same content emitted two and three times.

Pinned entries are skipped for a related reason. A pinned tweet is an arbitrarily old
tweet served at the top of a timeline, and emitting it silently corrupts `sortBy: latest`.

#### 3. X's own stop signal is advisory

`TimelineTerminateTimeline` sounds like it means "stop paging". It does not: X emits it on
*every* page of a paginated timeline while the bottom cursor keeps returning fresh tweets.
Obey it and `@apify` truncates from 92 tweets to 19 — a **79% recall loss**, and nothing
errors. The logs look perfectly healthy.

That is the failure mode worth designing against here: not a crash, but silent data loss
no alert fires on. So the Actor stops on structural signals only — no bottom cursor, a
cursor that did not advance, an empty page, or a page containing nothing new.

Guests also see two response modes, and the extractor handles both: paginated (~17–20
tweets per page plus a cursor) and single-snapshot (~98–100 tweets, no cursor,
non-chronological). Per-account measurements are
[in the appendix](docs/README-data-source.md#5-timeline-behaviour-measured-2026-08-17).

#### Normalization: the text pipeline, in order

```
1. SELECT SOURCE OBJECT   isRetweet ? legacy.retweeted_status_result.result : self
2. SELECT TEXT FIELD      note_tweet…result.text  ??  legacy.full_text
3. EXPAND t.co            → entities[].expanded_url  (urls *and* media)
4. DECODE HTML ENTITIES   &amp; &lt; &gt;            ← last
```

Every step is there because something measurably broke without it.

- **Step 1 — read the original, not the wrapper.** A retweet's own `legacy.full_text` is
  the `"RT @handle: …"` wrapper. Worse, so are its *metrics*: the wrapper in our fixture
  reports `favorite_count: 0` against the original's `13`. Read the wrapper and
  `minLikes: 1` silently discards every retweet in the run. Metrics come from the
  original; identity, author and timestamp stay the retweet's own.

- **Step 2 — long-form text lives in `note_tweet`.** And the trap here is a good one: the
  truncated version can be *longer* in raw characters than the complete one. Measured 302
  vs 283 on one `@apify` post, because X appends a 23-character t.co pointer to the text
  it cut off (`display_text_range` ends at 278). "Take whichever string is longer" emits
  the truncated text. When `note_tweet` supplies the text, entities must come from its
  `entity_set` too, or you are describing one string with another's offsets.

- **Step 3 — expand media links, not just URL links.** The trailing photo/video link is a
  t.co as well, but it lives in `entities.media[]`, not `entities.urls[]`. Expand only the
  latter and you leave a bare `https://t.co/…` in the text of every tweet with media —
  which is most of them.

- **Step 4 — decode last, and never index.** X's `indices` are offsets into the *raw*
  string, so decoding `&amp;` (5 chars) to `&` (1 char) shifts every later offset by four.
  We sidestep the arithmetic entirely by replacing t.co tokens as strings. That also dodges
  a second trap: X computes indices in Unicode code points while JavaScript slices in
  UTF-16 code units, so one emoji early in a tweet desynchronises them.

Two schema notes that most published scrapers still get wrong: **`legacy.followers_count`
no longer exists** — follower data moved to `user_results.result.relationship_counts` —
and `screen_name`/`name` now live under `core`, not `legacy`. More
[in the appendix](docs/README-data-source.md#6-schema-notes).

***

### 5. The free tier you cannot edit away

**The requirement:** unverified users get at most 10 results per run, and the Actor must
*stop fetching and pushing* at 10 regardless of what the input says. Client-side limits do
not count as protection, and environment variables are not trusted.

What makes this genuinely hard is where the code runs. An Apify Actor executes **on the
customer's account, under the customer's token, writing to the customer's storage**. Most
of the obvious places to put a limit are places they own.

#### 4.1 Identity comes from the credential, not the claim

```ts
const me = await new ApifyClient({ token: process.env.APIFY_TOKEN }).user('me').get();
```

Not `APIFY_USER_ID`. Per Apify's own docs that variable is "ID of the user who started the
Actor" — an environment variable, which the brief explicitly names as untrusted.

The token is strictly stronger, because it is a credential **the authority validates**.
Asking the platform "whose token is this?" is self-validating: a forged token is either
rejected (→ fail closed → free) or genuinely someone else's (→ correctly returns *their*
entitlement). There is no third outcome.

#### 4.2 The entitlement store is public on purpose

Here is the trap that breaks the obvious design. **Inside a run, `APIFY_TOKEN` belongs to
the runner.** `Actor.getValue()`, `Actor.apifyClient` and the default key-value store are
all authenticated *as them*. A private store on our account is simply unreachable from
inside the run that needs to read it.

So the store is **public-read, and the authority is write access** — which stays ours:

- a runner's own token can read it, so no shared secret is needed at read time;
- public reads cannot change a verdict;
- keys are `HMAC-SHA256(runnerUserId)`, so a world-readable store leaks no customer IDs —
  it is a page of hashes that names nobody;
- **blast radius:** a leaked Apify API token would let an attacker write to the store and
  grant themselves paid. A leaked HMAC key would only let them compute a key in a store
  that was already public. We put the authority behind the credential whose leak costs
  less.

`ENTITLEMENTS_HMAC_KEY` **must be marked Secret** in the Apify Console. A published
Actor's non-secret environment variables are publicly visible on its detail page, so a
plain env var would publish the authority itself.

#### 4.3 One chokepoint, consumed lazily

```ts
async push(item: T): Promise<boolean> {
  if (this.pushed >= this.opts.cap) return false;  // check…
  this.pushed++;                                   // …and increment, with no await between
  await this.opts.push(item);
  return this.pushed < this.opts.cap;
}
```

```ts
for await (const tweet of crawl(seeds)) {
  // lazy: pages cursors only when pulled
  if (!matches(tweet, filters)) continue;
  if (!(await sink.push(tweet))) break; // unwinds the whole generator chain
}
```

`break` stops the consumer, which stops the crawl, which stops cursor paging — and
`mergeConcurrent` returns every in-flight account chain, so none keeps paging in the
background. **A free user who asks for 1000 results costs us one page.** That is asserted
directly: the test drives a 100-page source and asserts exactly one page was fetched and
the generator was closed.

The check-then-increment ordering is load-bearing under concurrency. With N account chains
feeding one sink, any `await` between the check and the increment lets every worker pass
the check on the same value. A test fires 50 concurrent pushes at a slow sink and asserts
10; swapping those two lines fails it with 50, which we verified by mutation.

**But the cap alone does not bound cost.** That is the part worth dwelling on, because it
is not obvious and it is not small.

The sink stops the run at 10 **matches**. A low-selectivity search may never reach 10 — it
exhausts the account frontier first, and the gate never engages at all. Measured on the
shipped Actor: a free run with `searchTerms: ["web scraping"]` fetched **7,287 tweets
across 284 pages, spending 328 requests and 10 guest tokens, to deliver 9 results**. The
cap was working correctly the entire time. It simply had nothing to stop.

So an unverified run is bounded on **both** axes: results by the cap, and requests by an
allowance proportional to what it may return — 10 requests per permitted result, so 100
for a 10-result cap. That is roughly 2,000 tweets scanned: a generous sample, and two
orders of magnitude below "however many accounts exist". A paid run keeps its configured
budget. Re-running that same scenario afterwards: **29 requests, 14 pages, 3 tokens, 10
results.**

`maxResults` is also clamped up front, as an optimisation — but the clamp is not the
protection, and the input schema deliberately carries **no `"maximum": 10`**. A limit
expressed in the input is exactly the client-side artifact the brief rejects, and it would
break paying users.

#### 4.4 Fail closed, and say which kind of closed

```ts
const isPaid = entitlement?.paid === true; // ✅ undefined → false → free
// const isPaid = entitlement?.paid !== false;  ❌ undefined → true → unlimited
```

Every path resolves to free: a throw, a `null` record, a malformed record, a `paid` field
that is the *string* `"true"`. Zod validates the record so `undefined` can never reach a
boolean.

The verdict then distinguishes two cases that share a cap but not a meaning:

```jsonc
{ "limited": true, "reason": "free_tier",               "cap": 10 }  // verified free
{ "limited": true, "reason": "entitlement_unavailable", "cap": 10 }  // could not verify
```

The second is the **alertable** one — it may be capping a paying customer because of our
own outage — so it is emitted as a warning rather than an info line.

#### 4.5 The bypass that state persistence creates

The run's key-value store belongs to the runner, so a persisted counter is
attacker-writable:

```
1. Free run pushes 10, persists { pushed: 10 }.
2. User PUTs { pushed: 0 } with their own token.
3. User resurrects the run → resumes at 0 → 10 more into the same dataset.
4. Repeat.
```

A naive resume hands out 20 on the first resurrect with no tampering at all. The fix is to
floor the counter on an authority the user cannot lower:

```ts
const pushed = Math.max(persisted.pushed ?? 0, dataset.itemCount);
```

They cannot reduce `itemCount` without deleting the results they were trying to
accumulate. Entitlement is also re-resolved on resume and never cached in runner-writable
storage.

The principle generalises: *persisted state is fine for cursors — worst case the user
re-scrapes and pays for it. It is not fine for the counter that enforces the cap.*

#### 4.6 Anti-fork: the honest answer

- **Layer 0 — Distribution.** Production: private repo. *(This submission is public by
  requirement.)*
- **Layer 1 — Platform.** `Settings → Hide source files from Actor detail`, or the Store
  republishes the source regardless of where git lives.
- **Layer 2 — Credential.** The HMAC key as a Secret env var: absent from the repo, absent
  from the public env listing, unreadable via API or Console.
- **Layer 3 — Authority.** The verdict comes from a store only our token can write. A fork
  cannot *impersonate* a paid user — it can only *delete the check*.
- **Layer 4 — The honest limit.** Deletion is unpreventable. A *permission* check asks a
  server a question and can be removed; only a *capability* the server supplies cannot. A
  capability moat needs something expensive to reproduce or perishable — and guest tokens
  are freely mintable while queryIds are freely extractable. **This task has no genuine
  capability moat.** The real moat is the Store listing, maintained queryId resolution, and
  operational upkeep.

Where that reasoning ends is instructive. `apidojo-io/twitter-scraper-lite` is a public
repo with no HTTP calls to X at all — a thin wrapper around a paid actor. It is perfectly
fork-proof, because forking it gets you nothing. We deliberately did not go there: the
brief asks for a browserless extractor, not a billing wrapper.

**This gate is fully tamper-proof and explicitly not fork-proof, and for this deliverable
that is the right trade.**

***

### 6. Speed and cost, measured

The benchmark is time to 100 items from a single high-volume author, residential proxy,
paid user. One property sets the ceiling: **a single author is a single cursor chain, and
a cursor chain cannot be parallelised** — page 2 needs page 1's cursor. So this measures
sequential paging rather than concurrency.

Measured on the Apify platform, and published as a distribution rather than a best sample,
because the brief re-runs it. Our timer is stricter than the brief's: it includes the ~4 s
cold start the brief excludes.

| Author                      | n   | Wall clock                 | Grade A   |
| --------------------------- | --- | -------------------------- | --------- |
| `@elonmusk` (snapshot mode) | 5   | 3.7–12.0 s                 | **5 / 5** |
| `@apify` (paginated mode)   | 8   | 10.1–85.1 s, median 33.2 s | 4 / 8     |

Every run was clean — zero `429`s, zero errors, zero duplicates — and on the paginated
account every run did identical work: 12 pages, 13 requests, 46% selectivity. The work is
constant; only the clock moves.

**Reliably Grade A on a snapshot-mode author, and a coin flip on a paginated one.** The
variance is residential-proxy latency multiplied across 12 sequential fetches: a good run
paged at ~1.1 s, matching the same account with no proxy at all, and a bad one at ~7 s.

Page count is the part we control, and it follows from selectivity. Counting retweets on
the same account raises selectivity to 69%, shortens the chain to 8 pages and halves the
median to 14.8 s (7/8 Grade A) — but the tail survives, so this attributes the cost rather
than fixing it. The brief leaves `includeRetweets` at `false`, so **4 / 8 is the number
that answers the brief.**

We also tried retiring slow exit nodes to escape the tail. It measured no better and was
reverted; a replacement token is minted through the same slow pool.

There is no guest-reachable way to page more efficiently: `UserOriginalsTimeline`, which
would let us ask X for originals only, is one of the `404`s in §1. On a retweet-heavy
account we pay for what we discard.

#### Cost

Per 1,000 results at Apify list prices (residential proxy $12.50/GB, compute unit $0.40).
The constants live in `src/actor/summary.ts` so you can correct them for your plan.

|           | Author path          | Seeded              |
| --------- | -------------------- | ------------------- |
| Proxy     | 0.032 GB → **$0.40** | 0.59 GB → **$7.39** |
| Compute   | 0.008 CU → $0.003    | 0.26 CU → $0.10     |
| **Total** | **≈ $0.41 / 1k**     | **≈ $7.49 / 1k**    |

The extrapolation is linear and part of the cost is not. Cold start is paid once whatever
the run size, so a small run overstates the figure — the same path reports ≈$2.19/1k on a
10-item run and ≈$0.41/1k on a 100-item one. `bytesTransferred` is published beside it.

The 18× gap between the two paths is the cost of the stretch surface: extraction from
known handles is cheap, and keyword matching is not, because recall is seed-bounded and
selectivity is low. An early run against weaker seeds measured 0.09% selectivity and an
implied $128/1k, which is why the Actor has a request budget (`maxRequests`, default 500).
When it stops a run the summary says so.

Every run writes its own diagnostics to `OUTPUT`, so any claim here can be re-derived:

```jsonc
{
  "requested": 100,
  "fetched": 119,
  "pushed": 100,
  "limited": false,
  "reason": null,
  "cap": null,
  "hydratedById": { "requested": 0, "hydrated": 0, "missing": 0 },
  "discoveryStrategy": "direct",
  "seedsResolved": 10,
  "accountsCrawled": 6,
  "pagesFetched": 6,
  "filteredOut": 19,
  "selectivity": 0.8403,
  "duplicatesDropped": 0,
  "accountsSkipped": { "protected": 0, "suspended": 0, "notFound": 0 },
  "budgetExhausted": false,
  "tokensConsumed": 1,
  "xRequests": 12,
  "totalRequests": 14,
  "bytesTransferred": 3200000,
  "errors": { "429": 0, "403": 0, "404": 0, "5xx": 0, "timeout": 0, "other": 0 },
  "estimatedCostPer1kResults": { "proxyGB": 0.032, "computeUnits": 0.008, "usd": 0.41 },
  "wallClockMs": 2902,
}
```

#### Why the design looks like this

The binding constraint is the **request budget**, not latency. X grants roughly 50 requests
per 15 minutes per guest token, a page is ~20 tweets, and a page costs ~700–900 ms.

That is why a *pool* of session triples is the core design rather than an optimisation, and
why the response to a `429` is to rotate rather than to sleep: a fresh token costs ~200 ms,
and waiting fifteen minutes for a free resource is the largest throughput mistake available
here. Better still, triples retire at ≤5 remaining requests, so the `429` is never taken at
all — both benchmark runs report zero.

Parallelism runs **across accounts**, because a cursor chain cannot be parallelised
internally: page 2 needs page 1's cursor.

> If `SearchTimeline` were available, the A-grade path would be time-window sharding —
> split `since`/`until` into N sub-ranges, run N independent cursor chains in parallel, and
> merge through the global seen-set. A single cursor chain is inherently sequential, so
> sharding the query is the only way to parallelise search paging. It is the same reasoning
> that makes this design parallelise across accounts rather than across pages.

***

### 7. Output contract, and the calls behind it

Every item conforms exactly to the required shape. **Missing values are `null` — never
omitted, never `undefined`** — so the dataset keeps a stable column set. Timestamps are
ISO-8601 UTC, counts are integers, and IDs are strings, never JS numbers (they exceed
`Number.MAX_SAFE_INTEGER`).

```jsonc
{
  "id": "2088525549626867786",
  "url": "https://x.com/apify/status/2088525549626867786",
  "text": "…",
  "lang": "en",
  "createdAt": "2026-08-15T07:17:48.000Z",
  "conversationId": "2088525549626867786",
  "isReply": false,
  "isRetweet": true,
  "isQuote": false,
  "inReplyToId": null,
  "quotedTweetId": null,
  "author": {
    "id": "3510729917",
    "username": "apify",
    "name": "Apify",
    "verified": false,
    "followers": 11840,
    "following": 296,
  },
  "metrics": {
    "likes": 13,
    "retweets": 1,
    "replies": 0,
    "quotes": 0,
    "bookmarks": 1,
    "views": 713,
  },
  "entities": {
    "hashtags": [],
    "mentions": ["apify"],
    "urls": [],
    "media": [{ "type": "photo", "url": "…", "thumbnail": "…" }],
  },
  "source": "Twitter Web App",
  "scrapedAt": "2026-08-17T10:15:00.000Z",
}
```

The brief leaves a number of behaviours undefined. Each one is decided, tested, and written
down here, so documented behaviour can be diffed against actual behaviour:

| Ruling                         | Decision                                                   | Why                                                                                                                                                                                                        |
| ------------------------------ | ---------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `mediaType: images`            | matches if **≥1 photo**, other content allowed             | "tweets with images" is the natural reading; "photos only" surprises                                                                                                                                       |
| `animated_gif`                 | grouped under `video`                                      | X stores GIFs as MP4, and the enum has no `gif` value                                                                                                                                                      |
| `mediaType: links`             | ≥1 entry in `entities.urls`                                | —                                                                                                                                                                                                          |
| `mediaType: text_only`         | **no media AND no links**                                  | `links` is its own enum value, so allowing links here makes the enum incoherent                                                                                                                            |
| `onlyVerified`                 | `is_blue_verified \|\| verification.verified`              | X conflates paid Blue with legacy verification and the schema has one boolean. Never key off `verified_type`: `@grok` returns `verified_type: "Business"` with `verified: false` while being blue-verified |
| `includeReplies` default       | `false`                                                    | the brief specifies the default only for retweets; we default both and say so                                                                                                                              |
| `sortBy: latest`               | descending Snowflake ID                                    | IDs are monotonic, so ID order **is** chronological order                                                                                                                                                  |
| `sortBy: top`                  | descending `likes + retweets` **within the collected set** | X's relevance ranking is not reproducible from the guest surface — approximation, declared                                                                                                                 |
| `since` / `until`              | **inclusive**, applied to the Snowflake                    | a bare date in `until` covers the whole day to `23:59:59.999`; treating it as midnight silently drops a day                                                                                                |
| `tweetIds` vs `includeReplies` | an explicit id opts into replies and retweets              | naming a tweet by id *is* the selection; those defaults exist to shape a timeline sweep, and applying them here would silently drop the exact tweet asked for                                              |
| `hashtags`                     | post-filter, never a target                                | brief §4: it constrains the timelines a target produced, so a hashtags-only run has nothing to fetch from and is rejected                                                                                  |
| Multiple values in one filter  | OR                                                         | `hashtags: ["a","b"]` means a **or** b                                                                                                                                                                     |
| Multiple filters               | AND                                                        | and an unspecified filter is **no constraint**, never a narrowing one                                                                                                                                      |
| Missing metric vs `minLikes`   | counts as 0                                                | absent data is not evidence of engagement                                                                                                                                                                  |
| Retweet metrics                | the **original's**                                         | the wrapper's counters are structurally zero (§3)                                                                                                                                                          |

Results are buffered and written in one batch at the end, because `sortBy` is a property of
the whole result set and cannot be honoured by an append-only stream. The buffer is bounded
by the cap and checkpointed on migration.

***

### 8. Running it

#### The fastest way to verify this

**On the platform.** The Actor is published:
[apify.com/arthurvianna/x-tweet-scraper](https://apify.com/arthurvianna/x-tweet-scraper).
Run it there and the gate is live — free accounts get 10 results, and the run summary in
`OUTPUT` says why.

**Locally, with no account at all.** A fresh clone runs the whole thing, and the gate is
visible immediately because it fails closed without an entitlements store:

```bash
git clone https://github.com/ArthurVianna96/x-tweet-scraper && cd x-tweet-scraper
npm install && npm test              # 221 tests, offline, ~0.5s
mkdir -p storage/key_value_stores/default
echo '{"fromUsers":["apify"],"maxResults":1000}' \
  > storage/key_value_stores/default/INPUT.json
npm run start:dev
```

To exercise the by-id surface instead, swap the input for a list of ids — no accounts
needed, one request each:

```bash
echo '{"tweetIds":["2089366645768643034","1"],"maxResults":1000}' \
  > storage/key_value_stores/default/INPUT.json
npm run start:dev
```

The summary reports `hydratedById: { requested: 2, hydrated: 1, missing: 1 }` — the second
id does not exist, and is counted rather than fatal.

That first run asks for 1000 and returns **10**, logging `reason: "entitlement_unavailable"` —
the fail-closed path. To see the *verified*-free path (`reason: "free_tier"`) and the paid
path you need an entitlements store: either run the deployed Actor above, or stand up your
own ([`docs/DEPLOYMENT.md`](docs/DEPLOYMENT.md)).

#### Deployment status

Deployed and verified on Apify on 2026-08-17:

| Check                                       | Result                                                                                   |
| ------------------------------------------- | ---------------------------------------------------------------------------------------- |
| Paid run, `maxResults: 15`                  | 15 items, `limited: false`                                                               |
| Paid run, `maxResults: 1000` on one account | 233 items — the account's full reachable timeline                                        |
| **Free run, `maxResults: 1000`**            | **10 items**, `reason: "free_tier"`, 3 requests, 1 token                                 |
| Secret env vars on the Actor                | both `isSecret: true` — the HMAC key is not on the detail page                           |
| Entitlement propagation                     | a `grant`/`revoke` is visible to the next run immediately; no caching observed over 60 s |

The free run is the one that matters. **3 requests and 1 guest token for a 1000-result
request** means the cap stopped the *crawl*, not just the output.

#### Deploying your own

Standing up your own copy takes two secrets and one store setting.
[`docs/DEPLOYMENT.md`](docs/DEPLOYMENT.md) is the runbook: configuring the gate,
provisioning paid customers, and the one step that fails silently.

That step is worth naming here, because it is the sort of bug that never shows up in your
own testing. The entitlements store must be **public-read** (`generalAccess`, not
`isPublic`). Get it wrong and every customer's run falls back to
`entitlement_unavailable` — capped at 10, paying or not — while your own test runs look
perfect throughout, because your token can read your own private store. The bug is
invisible from the inside and total from the outside.

#### Running locally

Input is read from a file, not from arguments. Create it once:

```bash
npm install
cp .env.example .env          # optional — without it the gate caps at 10
mkdir -p storage/key_value_stores/default
cat > storage/key_value_stores/default/INPUT.json <<'JSON'
{ "fromUsers": ["apify", "naval"], "maxResults": 25, "sortBy": "latest",
  "proxyConfiguration": { "useApifyProxy": false } }
JSON
```

Then either runner works:

```bash
npm run start:dev     # tsx, reads .env — fastest loop, no Apify login
npx apify-cli run     # the platform's own runner; reads `apify secrets` instead of .env
```

Results land as one file per item in `storage/datasets/default/`, and the run summary in
`storage/key_value_stores/default/OUTPUT.json`.

**Both runners purge the default dataset and key-value store on start** (keeping `INPUT`),
so each local run is clean and there is no stale-state trap. The side effect is that the
resume protection in §5.5 is invisible locally — it needs storage to survive. To watch it
work, disable the purge and run twice:

```bash
APIFY_PURGE_ON_START=0 npm run start:dev   # first run fills the dataset
APIFY_PURGE_ON_START=0 npm run start:dev   # → "resuming", fetched 0, pushed 0
```

The second run reports `fetched: 0` — it did not pull a single page, because the cap was
already spent. To see the resurrect-and-reset bypass fail, forge the counter the way a
runner who owns this storage could, then run again:

```bash
python3 - <<'PY'
import json; p='storage/key_value_stores/default/CRAWL_STATE.json'
d=json.load(open(p)); d['pushed']=0; json.dump(d,open(p,'w'))
PY
APIFY_PURGE_ON_START=0 npm run start:dev   # still "alreadyPushed: 25" — floored on itemCount
```

A residential proxy is recommended (`proxyConfiguration: {"useApifyProxy": true,
"apifyProxyGroups": ["RESIDENTIAL"]}`). X rate-limits per token and per IP, and the seed
lookup is more likely to be challenged from a datacenter IP.

#### Development

|                                                               |                                          |
| ------------------------------------------------------------- | ---------------------------------------- |
| `npm test`                                                    | 221 tests, offline, no platform          |
| `npm run typecheck` / `npm run lint`                          | strict TS, ESLint                        |
| `npm run probe`                                               | re-derive the endpoint capability matrix |
| `npx tsx src/tools/capture-fixtures.ts <handles…>`            | refresh committed fixtures from live X   |
| `npx tsx src/tools/benchmark.ts native \| seeded <arg>`       | reproduce §6                             |
| `npm run entitlement -- <key\|grant\|revoke\|check> <userId>` | provision a paid customer                |

Layering is one-way — `actor → adapters → domain` — and `domain/` imports nothing from the
other two. Every seam is constructor injection; there is no module mocking anywhere in the
suite. See `CLAUDE.md` for the full conventions.

***

### 9. robots.txt, ToS, and what we would tell a client

`https://x.com/robots.txt` sets `User-agent: * → Disallow: /`. The `Allow:` rules for
`/search`, `/hashtag/*` and `/i/api/` apply only to the named `Googlebot`/`Bingbot` group.
An automated client that is not a verified search-engine crawler is therefore outside what
robots.txt permits, and X's Terms separately restrict automated access.

We tested whether that crawler allowance was user-agent-gated. It is not: requesting
`/search` with a Googlebot user-agent returns `404`, because X verifies crawler identity by
reverse DNS. We did not attempt to defeat that.

This Actor collects **public data only**, uses no account credentials, and holds a
conservative request rate. Before running it for a client in production we would raise
three things:

1. robots.txt does not permit it, and a commercial agreement or X's licensed API is the
   compliant path at scale;
2. GDPR/CCPA obligations — tweets and author profiles are personal data, so a lawful basis
   and a retention policy are required;
3. guest-token access is undocumented and can be withdrawn without notice, as
   `SearchTimeline` itself demonstrates.

***

### 10. What it cannot do

Stated plainly, because a limitation you find in the README is cheaper than one you find in
production.

- **`searchTerms` recall is seed-bounded.** This is the stretch surface (§2a), and the one
  place we cannot match what a logged-in client could do. You get matching tweets from
  accounts that discuss the topic, not every tweet on X — a hard ceiling of the guest
  surface, not an implementation shortcut. The run summary reports how many accounts were
  seeded and crawled, so callers can judge coverage for themselves. The three required
  surfaces have no such ceiling.
- **`sortBy: top` is an approximation** — engagement ranking within the collected set. X's
  own relevance ranking is not reproducible from the guest surface.
- **On a paginated author the benchmark is proxy-bound, and Grade A is not guaranteed.**
  Measured 4 of 8 runs under 30 s, median 33.2 s, on constant work (§6). A single cursor
  chain is sequential by construction, so residential-proxy latency multiplies across every
  page. We tried rotating away from slow nodes and measured no improvement. Shortening the
  chain does help — counting retweets takes it to 7 of 8 — but the tail stays, because the
  proxy sets the variance and only the page count is ours to influence.
- **Fixtures cannot detect X changing its response shape.** The suite is offline and
  deterministic by choice, and the price of that choice is that these tests stay green
  while production breaks. In a production system you close this with a scheduled,
  non-blocking contract test against the live API that alerts on drift;
  `src/tools/capture-fixtures.ts` makes re-capturing a one-command job. Out of scope here,
  and stated rather than papered over.
- **The seed lookup depends on third-party search engines**, which challenge automated
  traffic (§3). The cascade mitigates it; `fromUsers` removes the dependency entirely.
- **Buffered output** trades peak memory for correct ordering. At the scale this Actor
  targets — hundreds to low thousands of results — that is the right trade. A
  hundred-thousand-result run would want a spill-to-disk merge sort instead.

Delivered from the bonus list (§11): **`searchTerms` via a justified public HTTP source**
rather than left unsupported; incremental/resumable scraping keyed on stored cursors (§5.5);
global deduplication and a seen-set across overlapping targets — required anyway, given
nested retweets, snowball overlap and ids that also appear in a crawled timeline; graceful
handling of protected / suspended / deleted accounts and dead tweet ids, neither of which
fails a run; and cost-per-1k reporting. Not delivered: the finish webhook.

***

### 11. Decisions and trade-offs

1. **Build on the doors that are open, and say which they are.** The three required
   surfaces — author, id, profile — are guest-reachable and implemented natively. Search is
   not, and is scoped as a stretch served a different way rather than half-worked into the
   required paths.
2. **Separate discovery from extraction, and put discovery behind a port.** The one part of
   the problem X does not permit is isolated in a single swappable adapter, and the other
   95% of the system is native and unaffected by that choice.
3. **Ask a web index about profiles, not posts.** Driven by measurement — 36-day-old
   freshest indexed tweet, zero hashtag coverage — not by preference.
4. **Refuse to solve the hard part by buying it.** A paid search API would have made keyword
   search work immediately, and made the extraction someone else's work.
5. **Resolve `queryId`s at runtime.** Three bundle hashes in three days; hardcoding
   guarantees a silent failure on some future Tuesday.
6. **Extract by path; never recurse.** Recursion double-counts nested originals — 32 items
   from a 20-entry page.
7. **Distrust X's own stop signal.** `TimelineTerminateTimeline` costs 79% of recall on a
   paginated account if believed.
8. **Pin the session triple; rotate on `429`; retire before the limit.** Coherence is what
   abuse detection looks for, and proactive retirement is worth more than any backoff.
9. **Enforce the cap at a single lazy chokepoint.** It stops fetching, not just pushing —
   and the seam is constructor injection, so the required test needs no network.
10. **Derive identity from the token, not the environment; fail closed on every path.**
11. **Floor the resume counter on `dataset.itemCount`.** Persisted state is fine for cursors
    and unacceptable for the counter that enforces the cap.
12. **Budget requests explicitly.** Low selectivity is normal, and without a budget the cost
    of a run is bounded only by how many accounts exist.
13. **Publish the diagnostics that would let a reviewer contradict us.** Selectivity, tokens,
    bytes, `429` count and cost are all in `OUTPUT`.

# Actor input Schema

## `fromUsers` (type: `array`):

Handles without '@'. When supplied, the Actor makes no third-party call at all: these accounts are crawled directly on X. This is the fastest and most deterministic path — clear it to discover accounts from `searchTerms`/`hashtags` instead.

## `tweetIds` (type: `array`):

Specific tweet ids to hydrate, one request each. This is a target in its own right: a run may supply only tweetIds and no accounts. An id that names a deleted or unreachable tweet is counted and skipped, never fatal.

## `searchTerms` (type: `array`):

Keywords to match against tweet text (case-insensitive). Multiple terms are OR'd. X's search timeline is closed to guest tokens, so this is the stretch surface: with no `fromUsers` supplied these keywords also seed account discovery via one search-engine lookup. Recall is seed-bounded — see the README.

## `hashtags` (type: `array`):

Without '#'. A post-filter over the timelines a target produced, matched case-insensitively against the tweet's hashtag entities — not a target on its own.

## `since` (type: `string`):

ISO date or datetime, e.g. 2026-08-01. Inclusive. Applied to the tweet's Snowflake ID, so it is exact.

## `until` (type: `string`):

ISO date or datetime. Inclusive — a bare date covers the whole day up to 23:59:59.999 UTC.

## `language` (type: `string`):

ISO-639-1 code, matched against X's own language detection (e.g. en, pt, es).

## `minLikes` (type: `integer`):

Inclusive floor.

## `minRetweets` (type: `integer`):

Inclusive floor.

## `minReplies` (type: `integer`):

Inclusive floor.

## `onlyVerified` (type: `boolean`):

X conflates paid Blue with legacy verification, so this matches either. Leave unset for no constraint — it never excludes verified authors.

## `mediaType` (type: `string`):

images = has at least one photo; video = has video or GIF (X stores GIFs as MP4); links = has at least one URL; text\_only = no media and no links.

## `includeReplies` (type: `boolean`):

Excluded by default.

## `includeRetweets` (type: `boolean`):

Excluded by default. When included, a retweet carries the original's full text and metrics — the retweet wrapper's counters are structurally zero.

## `sortBy` (type: `string`):

latest = newest first, exact (tweet IDs are monotonic). top = ranked by likes + retweets within the collected set; X's own relevance ranking is not reproducible from the logged-out surface.

## `maxResults` (type: `integer`):

How many tweets to return. Free runs are capped at 10 results regardless of this value; the cap is enforced server-side. Paid runs receive the full amount.

## `proxyConfiguration` (type: `object`):

Residential proxies are recommended: X rate-limits per guest token and per IP, and the Actor pins one token to one IP for the life of a session.

## `maxConcurrency` (type: `integer`):

Advanced. A single timeline's cursor chain is sequential, so parallelism happens across accounts.

## `maxAccounts` (type: `integer`):

Advanced. Upper bound on how many accounts a single run will crawl, including accounts found by seed expansion.

## `maxPagesPerAccount` (type: `integer`):

Advanced. One page is ~20 tweets on a paginated timeline.

## `expansionDepth` (type: `integer`):

Advanced. 0 disables snowball expansion. Mentions and retweeted authors are harvested from pages already fetched, so expansion costs no extra requests — but precision decays with depth.

## `maxSessions` (type: `integer`):

Advanced. Each token is worth ~50 requests per 15 minutes. Tokens are free and instant to mint.

## `maxRequests` (type: `integer`):

Advanced. Hard ceiling on requests to X for the whole run — the cost control. Low-selectivity keyword filters can otherwise fetch thousands of tweets for a handful of matches.

## Actor input object example

```json
{
  "fromUsers": [
    "apify",
    "naval"
  ],
  "includeReplies": false,
  "includeRetweets": false,
  "sortBy": "latest",
  "maxResults": 100,
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "maxConcurrency": 4,
  "maxAccounts": 50,
  "maxPagesPerAccount": 25,
  "expansionDepth": 1,
  "maxSessions": 5,
  "maxRequests": 500
}
```

# Actor output Schema

## `tweets` (type: `string`):

One item per tweet, conforming exactly to the documented schema. Absent values are null, never omitted.

## `runSummary` (type: `string`):

Requested vs fetched vs pushed, the free-tier verdict (limited / reason / cap), selectivity, guest tokens consumed, error counts and estimated cost per 1k results.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "fromUsers": [
        "apify",
        "naval"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("arthurvianna/x-tweet-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "fromUsers": [
        "apify",
        "naval",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("arthurvianna/x-tweet-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "fromUsers": [
    "apify",
    "naval"
  ]
}' |
apify call arthurvianna/x-tweet-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,arthurvianna/x-tweet-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/PNZugrwspnmMj70at/builds/wj4iT8pgxQKYZrlWc/openapi.json
