# X Community Notes Scraper - Fact Checks on Posts (`thirdwatch/x-community-notes-scraper`) Actor

Export X (Twitter) Community Notes from X's own public dataset. Filter fact-check notes by post, keyword, classification, media, date, and whether the note is actually showing on X. No login or cookies required.

- **URL**: https://apify.com/thirdwatch/x-community-notes-scraper.md
- **Developed by:** [Thirdwatch](https://apify.com/thirdwatch) (community)
- **Categories:** Social media, Education
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## X Community Notes Scraper

Export X (Twitter) Community Notes: the crowd-written fact checks that appear under posts. Filter by post, keyword, classification, media, date, and whether a note is actually showing on X.

### Why this one is different

Almost every X scraper depends on an endpoint that X can close, and most of them broke when it did. This Actor reads the Community Notes corpus that **X publishes itself**, as dated snapshots under `ton.twimg.com/birdwatch-public-data`. No login, no cookies, no proxy, no rate limit, and nothing that stops working when X changes its API.

### What you get

One row per note:

| Field | Description |
| --- | --- |
| `summary` | The note text shown to readers |
| `tweetId`, `tweetUrl` | The post the note annotates |
| `classification` | `MISINFORMED_OR_POTENTIALLY_MISLEADING` or `NOT_MISLEADING` |
| `isMisleading` | Boolean shortcut for the above |
| `reasons` | Why it was flagged: factual error, missing context, manipulated media, satire, outdated, and so on |
| `trustworthySources` | Whether the note cites sources |
| `isMediaNote` | Note attached to an image or video rather than post text |
| `isCollaborativeNote` | Written collaboratively (the newer note format) |
| `createdAt` | When the note was written, ISO 8601 |
| `noteId`, `noteUrl` | The note itself |
| `noteAuthorParticipantId` | Pseudonymous contributor id, as published by X |
| `status`, `isShowingOnX`, `statusLabel` | Present when **Include helpfulness status** is on |
| `snapshotDate` | Which daily snapshot the row came from |

### Written notes vs notes people actually see

Most notes never appear on X. A note only displays once contributors who usually disagree with each other both rate it helpful, so the corpus is far larger than what readers ever see.

If you care about the notes that actually shipped, turn on **Include helpfulness status** and **Only notes showing on X**. That joins a larger file, so the run takes longer, which is why it is off by default.

### Recent notes vs the full archive

By default the Actor reads the recent slice, a few megabytes covering roughly the last several months. Turn on **Search the historical archive** to include every note back to 2021, about 3 million of them. That downloads a ~480 MB archive, so expect several minutes.

The Actor streams it to disk rather than into memory and stops as soon as your `maxResults` is met, so a filtered search stays cheap even against the full corpus.

### Example uses

**Has this post been fact-checked?** Put post URLs or IDs into **Post IDs or URLs**. You get any notes attached to them, or an empty result if there are none.

**Track misinformation on a topic.** Set **Text contains** to a term and **Classification** to *Misleading only*.

**Study manipulated media.** Turn on **Media notes only** with the historical archive, since media notes are largely a pre-2026 population.

**Monitor what is actually being shown.** Combine **Only notes showing on X** with **Written on or after** to see recently displayed notes.

### Input

| Field | Description | Default |
| --- | --- | --- |
| `maxResults` | Notes to publish | 1000 |
| `classification` | All, misleading only, or not misleading only | all |
| `searchText` | Keep notes whose text contains this phrase | — |
| `tweetIds` | Only notes on these posts, URLs or bare IDs | — |
| `sinceDate` | Drop notes written before this date | — |
| `mediaNotesOnly` | Only notes on images or video | false |
| `includeStatus` | Join current helpfulness status | false |
| `onlyShowingOnX` | Only notes displayed publicly, implies the join | false |
| `includeHistorical` | Search the ~480 MB archive back to 2021 | false |
| `snapshotDate` | Use a specific daily snapshot | latest |

### Notes on the data

- X publishes with a lag and skips some days. With no `snapshotDate` the Actor walks back up to 14 days to find the newest snapshot that exists, and reports which one it used in `snapshotDate`.
- The dataset gives a pseudonymous contributor id, not a handle. That is X's own privacy design and cannot be resolved back to an account.
- `tweetUrl` uses X's handle-free `/i/status/` permalink, because the dataset does not carry the post author's handle.
- The legacy reason flags are all zero on 2026 collaborative notes. They are populated on older notes, so use `includeHistorical` when you need them.

### Output fields

- One row per result with stable keys.

### Use cases

- Marketers monitoring social presence and engagement
- Researchers building social-media datasets
- Teams tracking public figures and trends

### Related Actors

- [Douyin Scraper](https://apify.com/thirdwatch/douyin-scraper) — Thirdwatch
- [Facebook Ad Library Scraper - Active Ads & Creatives](https://apify.com/thirdwatch/fb-ad-library-scraper) — Thirdwatch
- [Facebook Marketplace Scraper - Listings, Prices & Alerts](https://apify.com/thirdwatch/facebook-marketplace-scraper) — Thirdwatch
- [Facebook Scraper - Page Posts, Group Posts, Comments & Pages](https://apify.com/thirdwatch/facebook-search-scraper) — Thirdwatch

Last verified: 2026-09

More scrapers at [thirdwatch.dev](https://thirdwatch.dev).

# Actor input Schema

## `maxResults` (type: `integer`):

How many notes to publish.

## `classification` (type: `string`):

Keep only misleading or only not-misleading notes.

## `searchText` (type: `string`):

Keep notes whose text contains this phrase (case-insensitive).

## `tweetIds` (type: `array`):

Only return notes attached to these posts. Accepts x.com URLs or bare post IDs. Use this to check whether specific posts have been fact-checked.

## `sinceDate` (type: `string`):

Drop notes written before this date (YYYY-MM-DD).

## `mediaNotesOnly` (type: `boolean`):

Keep only notes attached to an image or video rather than post text.

## `includeStatus` (type: `boolean`):

Join each note to its current rating status, so you can tell which notes are actually showing on X. Downloads a large file, so runs take longer.

## `onlyShowingOnX` (type: `boolean`):

Keep only notes rated helpful enough to display publicly. Turns on the status join automatically.

## `includeHistorical` (type: `boolean`):

Also search every note back to 2021 (about 3 million). This downloads a ~480 MB archive, so runs take several minutes. Leave off to search only recent notes, which is a few MB.

## `snapshotDate` (type: `string`):

Use a specific daily snapshot (YYYY-MM-DD). Leave empty for the most recent one X has published.

## Actor input object example

```json
{
  "maxResults": 1000,
  "classification": "all",
  "tweetIds": [],
  "mediaNotesOnly": false,
  "includeStatus": false,
  "onlyShowingOnX": false,
  "includeHistorical": false
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxResults": 1000
};

// Run the Actor and wait for it to finish
const run = await client.actor("thirdwatch/x-community-notes-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "maxResults": 1000 }

# Run the Actor and wait for it to finish
run = client.actor("thirdwatch/x-community-notes-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxResults": 1000
}' |
apify call thirdwatch/x-community-notes-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,thirdwatch/x-community-notes-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/2lNPANiHHpxO0egNS/builds/jU3KZIiOue9aAouMH/openapi.json
