X Community Notes Scraper - Fact Checks on Posts
Pricing
from $2.00 / 1,000 results
X Community Notes Scraper - Fact Checks on Posts
Export X (Twitter) Community Notes from X's own public dataset. Filter fact-check notes by post, keyword, classification, media, date, and whether the note is actually showing on X. No login or cookies required.
Pricing
from $2.00 / 1,000 results
Rating
0.0
(0)
Developer
Thirdwatch
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 hours ago
Last modified
Categories
Share
X Community Notes Scraper
Export X (Twitter) Community Notes: the crowd-written fact checks that appear under posts. Filter by post, keyword, classification, media, date, and whether a note is actually showing on X.
Why this one is different
Almost every X scraper depends on an endpoint that X can close, and most of them broke when it did. This Actor reads the Community Notes corpus that X publishes itself, as dated snapshots under ton.twimg.com/birdwatch-public-data. No login, no cookies, no proxy, no rate limit, and nothing that stops working when X changes its API.
What you get
One row per note:
| Field | Description |
|---|---|
summary | The note text shown to readers |
tweetId, tweetUrl | The post the note annotates |
classification | MISINFORMED_OR_POTENTIALLY_MISLEADING or NOT_MISLEADING |
isMisleading | Boolean shortcut for the above |
reasons | Why it was flagged: factual error, missing context, manipulated media, satire, outdated, and so on |
trustworthySources | Whether the note cites sources |
isMediaNote | Note attached to an image or video rather than post text |
isCollaborativeNote | Written collaboratively (the newer note format) |
createdAt | When the note was written, ISO 8601 |
noteId, noteUrl | The note itself |
noteAuthorParticipantId | Pseudonymous contributor id, as published by X |
status, isShowingOnX, statusLabel | Present when Include helpfulness status is on |
snapshotDate | Which daily snapshot the row came from |
Written notes vs notes people actually see
Most notes never appear on X. A note only displays once contributors who usually disagree with each other both rate it helpful, so the corpus is far larger than what readers ever see.
If you care about the notes that actually shipped, turn on Include helpfulness status and Only notes showing on X. That joins a larger file, so the run takes longer, which is why it is off by default.
Recent notes vs the full archive
By default the Actor reads the recent slice, a few megabytes covering roughly the last several months. Turn on Search the historical archive to include every note back to 2021, about 3 million of them. That downloads a ~480 MB archive, so expect several minutes.
The Actor streams it to disk rather than into memory and stops as soon as your maxResults is met, so a filtered search stays cheap even against the full corpus.
Example uses
Has this post been fact-checked? Put post URLs or IDs into Post IDs or URLs. You get any notes attached to them, or an empty result if there are none.
Track misinformation on a topic. Set Text contains to a term and Classification to Misleading only.
Study manipulated media. Turn on Media notes only with the historical archive, since media notes are largely a pre-2026 population.
Monitor what is actually being shown. Combine Only notes showing on X with Written on or after to see recently displayed notes.
Input
| Field | Description | Default |
|---|---|---|
maxResults | Notes to publish | 1000 |
classification | All, misleading only, or not misleading only | all |
searchText | Keep notes whose text contains this phrase | — |
tweetIds | Only notes on these posts, URLs or bare IDs | — |
sinceDate | Drop notes written before this date | — |
mediaNotesOnly | Only notes on images or video | false |
includeStatus | Join current helpfulness status | false |
onlyShowingOnX | Only notes displayed publicly, implies the join | false |
includeHistorical | Search the ~480 MB archive back to 2021 | false |
snapshotDate | Use a specific daily snapshot | latest |
Notes on the data
- X publishes with a lag and skips some days. With no
snapshotDatethe Actor walks back up to 14 days to find the newest snapshot that exists, and reports which one it used insnapshotDate. - The dataset gives a pseudonymous contributor id, not a handle. That is X's own privacy design and cannot be resolved back to an account.
tweetUrluses X's handle-free/i/status/permalink, because the dataset does not carry the post author's handle.- The legacy reason flags are all zero on 2026 collaborative notes. They are populated on older notes, so use
includeHistoricalwhen you need them.
Output fields
- One row per result with stable keys.
Use cases
- Marketers monitoring social presence and engagement
- Researchers building social-media datasets
- Teams tracking public figures and trends
Related Actors
- Douyin Scraper — Thirdwatch
- Facebook Ad Library Scraper - Active Ads & Creatives — Thirdwatch
- Facebook Marketplace Scraper - Listings, Prices & Alerts — Thirdwatch
- Facebook Scraper - Page Posts, Group Posts, Comments & Pages — Thirdwatch
Last verified: 2026-09
More scrapers at thirdwatch.dev.