X Community Notes Scraper - Fact Checks on Posts avatar

X Community Notes Scraper - Fact Checks on Posts

Pricing

from $2.00 / 1,000 results

Go to Apify Store
X Community Notes Scraper - Fact Checks on Posts

X Community Notes Scraper - Fact Checks on Posts

Export X (Twitter) Community Notes from X's own public dataset. Filter fact-check notes by post, keyword, classification, media, date, and whether the note is actually showing on X. No login or cookies required.

Pricing

from $2.00 / 1,000 results

Rating

0.0

(0)

Developer

Thirdwatch

Thirdwatch

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 hours ago

Last modified

Share

X Community Notes Scraper

Export X (Twitter) Community Notes: the crowd-written fact checks that appear under posts. Filter by post, keyword, classification, media, date, and whether a note is actually showing on X.

Why this one is different

Almost every X scraper depends on an endpoint that X can close, and most of them broke when it did. This Actor reads the Community Notes corpus that X publishes itself, as dated snapshots under ton.twimg.com/birdwatch-public-data. No login, no cookies, no proxy, no rate limit, and nothing that stops working when X changes its API.

What you get

One row per note:

FieldDescription
summaryThe note text shown to readers
tweetId, tweetUrlThe post the note annotates
classificationMISINFORMED_OR_POTENTIALLY_MISLEADING or NOT_MISLEADING
isMisleadingBoolean shortcut for the above
reasonsWhy it was flagged: factual error, missing context, manipulated media, satire, outdated, and so on
trustworthySourcesWhether the note cites sources
isMediaNoteNote attached to an image or video rather than post text
isCollaborativeNoteWritten collaboratively (the newer note format)
createdAtWhen the note was written, ISO 8601
noteId, noteUrlThe note itself
noteAuthorParticipantIdPseudonymous contributor id, as published by X
status, isShowingOnX, statusLabelPresent when Include helpfulness status is on
snapshotDateWhich daily snapshot the row came from

Written notes vs notes people actually see

Most notes never appear on X. A note only displays once contributors who usually disagree with each other both rate it helpful, so the corpus is far larger than what readers ever see.

If you care about the notes that actually shipped, turn on Include helpfulness status and Only notes showing on X. That joins a larger file, so the run takes longer, which is why it is off by default.

Recent notes vs the full archive

By default the Actor reads the recent slice, a few megabytes covering roughly the last several months. Turn on Search the historical archive to include every note back to 2021, about 3 million of them. That downloads a ~480 MB archive, so expect several minutes.

The Actor streams it to disk rather than into memory and stops as soon as your maxResults is met, so a filtered search stays cheap even against the full corpus.

Example uses

Has this post been fact-checked? Put post URLs or IDs into Post IDs or URLs. You get any notes attached to them, or an empty result if there are none.

Track misinformation on a topic. Set Text contains to a term and Classification to Misleading only.

Study manipulated media. Turn on Media notes only with the historical archive, since media notes are largely a pre-2026 population.

Monitor what is actually being shown. Combine Only notes showing on X with Written on or after to see recently displayed notes.

Input

FieldDescriptionDefault
maxResultsNotes to publish1000
classificationAll, misleading only, or not misleading onlyall
searchTextKeep notes whose text contains this phrase—
tweetIdsOnly notes on these posts, URLs or bare IDs—
sinceDateDrop notes written before this date—
mediaNotesOnlyOnly notes on images or videofalse
includeStatusJoin current helpfulness statusfalse
onlyShowingOnXOnly notes displayed publicly, implies the joinfalse
includeHistoricalSearch the ~480 MB archive back to 2021false
snapshotDateUse a specific daily snapshotlatest

Notes on the data

  • X publishes with a lag and skips some days. With no snapshotDate the Actor walks back up to 14 days to find the newest snapshot that exists, and reports which one it used in snapshotDate.
  • The dataset gives a pseudonymous contributor id, not a handle. That is X's own privacy design and cannot be resolved back to an account.
  • tweetUrl uses X's handle-free /i/status/ permalink, because the dataset does not carry the post author's handle.
  • The legacy reason flags are all zero on 2026 collaborative notes. They are populated on older notes, so use includeHistorical when you need them.

Output fields

  • One row per result with stable keys.

Use cases

  • Marketers monitoring social presence and engagement
  • Researchers building social-media datasets
  • Teams tracking public figures and trends

Last verified: 2026-09

More scrapers at thirdwatch.dev.