Substack Notes Scraper avatar

Substack Notes Scraper

Pricing

$0.01 / 1,000 notes

Go to Apify Store
Substack Notes Scraper

Substack Notes Scraper

Collect public Substack Notes from topic phrases or public source URLs. Save note text, stable IDs, public links, author details, times, engagement, media, and discovery data.

Pricing

$0.01 / 1,000 notes

Rating

0.0

(0)

Developer

Maxime Dupré

Maxime Dupré

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Share

📝 Track public Substack Notes

Researchers, marketers, and newsletter teams can collect public Substack Notes without copying them by hand. This Actor returns one structured row per saved Note with text, stable identity, a public URL, source details, times, author and publication context, visible engagement, media, and links.

Use it to:

📦 Structured Note data

Each saved row describes one public Substack Note. It includes the visible Note text, stable Note ID, canonical URL, source type, publication and collection times, author and publication context, visible engagement, media, outbound links, and thread fields. If the same Note appears again through another phrase or submitted URL, the first eligible match is saved and later matches are ignored. discoveredFrom shows only the phrase or URL that first produced the row. Missing public values stay null when the source does not show them, and arrays can be empty.

▶️ Run a public Notes collection

Choose a target

Select Topic search for phrases, or Public source URLs for public profile, feed, or Note pages.

Set a limit

maxItems is an optional positive limit for unique Notes. Leave it empty to return all available results until the source is exhausted.

Filter dates

dateField checks publication time or change time. startDate and endDate are inclusive UTC dates.

Skip known Notes

Add stable Note IDs in excludeNoteIds when you already have them.

Public search ranking, availability, rate limits, and source changes can affect what is returned. This Actor does not collect newsletter posts, full article bodies, or private or subscriber-only content.

⚙️ Input

Input fields

FieldTypeWhat it does
targetstringRequired. Choose topicSearch for topic phrases or sourceUrls for public source URLs.
topicPhrasesarray of stringsAdds one or more phrases for Topic search. This field is ignored when target is sourceUrls.
searchOrderstringChooses relevance or recent order for Topic search. It is ignored when target is sourceUrls.
sourceUrlsarray of objectsAdds public Substack profile, feed, or Note URLs. This field is ignored when target is topicSearch.
sourceUrls[].urlstringA public Substack profile, feed, or Note URL to read.
dateFieldstringChooses whether dates check published time or changed time.
startDatestringOptional inclusive UTC start date in YYYY-MM-DD format.
endDatestringOptional inclusive UTC end date in YYYY-MM-DD format.
excludeNoteIdsarray of stringsSkips the listed stable Note IDs.
maxItemsintegerOptional positive limit for unique Notes. Leave it empty to return all available results until the source is exhausted.

Example input

This is the public input from a successful current-beta default-input run:

{
"target": "topicSearch",
"topicPhrases": [
"technology"
],
"searchOrder": "relevance",
"dateField": "published",
"maxItems": 100
}

🧾 Output

Every saved row uses the same shape for topic, profile, feed, and Note sources. Values can be null when a public source does not show them. Arrays can be empty.

Output fields

FieldTypeWhat it does
noteIdstringStable public ID for the Note.
textstringVisible text of the public Note.
noteUrlstringCanonical public URL for the Note.
sourceTypestringSays how the Note was found: topicSearch, profile, feed, or note.
discoveredFromobjectGroups the first topic phrase or public source URL that produced the Note.
discoveredFrom.typestringSays whether the first discovery value was a topicPhrase or sourceUrl.
discoveredFrom.valuestringThe first topic phrase or public source URL that produced the Note.
publishedAtstringUTC time when the Note was published.
editedAtstring or nullUTC time when the Note was last edited, when public edit time is available.
collectedAtstringUTC time when the Actor collected the Note.
authorobjectGroups public author identity and profile details.
author.idstring or nullPublic author ID, when available.
author.namestring or nullVisible author name, when available.
author.handlestring or nullVisible author handle, when available.
author.profileUrlstring or nullPublic author profile URL, when available.
author.photoUrlstring or nullPublic author photo URL, when available.
author.biostring or nullVisible author bio, when available.
publicationobject or nullGroups the public publication linked to the Note, when available.
publication.idstring or nullPublic publication ID, when available.
publication.namestring or nullVisible publication name, when available.
publication.urlstring or nullPublic publication URL, when available.
engagementobjectGroups visible counts and reaction details.
engagement.likesnumber or nullVisible like or reaction count, when available.
engagement.restacksnumber or nullVisible restack count, when available.
engagement.repliesnumber or nullVisible reply count, when available.
engagement.totalnumber or nullTotal visible engagement count, when available.
engagement.reactionsarray of objects or nullLists visible reactions and their counts, when available.
engagement.reactions[].emojistringThe visible reaction or emoji label.
engagement.reactions[].countnumberVisible count for that reaction.
mediaarray of objectsLists media attached to the Note. It can be empty.
media[].urlstringPublic URL for attached media.
media[].typestring or nullMedia type, when exposed.
media[].widthnumber or nullMedia width in pixels, when available.
media[].heightnumber or nullMedia height in pixels, when available.
outboundLinksarray of stringsPublic URLs found in the Note text. It can be empty.
threadobjectGroups reply status and public parent or root Note relationships.
thread.isReplyboolean or nullSays whether the Note is a reply, when available.
thread.parentNoteIdstring or nullPublic ID of the direct parent Note, when available.
thread.parentUrlstring or nullPublic URL of the direct parent Note, when available.
thread.rootNoteIdstring or nullPublic ID of the root Note, when available.
thread.rootUrlstring or nullPublic URL of the root Note, when available.

Example row

This is a genuine, unshortened row from a successful current-beta profile run:

{
"noteId": "321255318",
"text": "Also as we are going to kickstart our new training series - AI Avatar Training and Social Media Domination.\n\nWe are doing a limited time 50% off on our Substack plan.\n\nAnd you will be having full zoom access and replay and all our previous ai courses.\n\nDetails here:\n\nhttps://sifuyik.substack.com/p/24-hours-only-100-hours-of-ai-training",
"noteUrl": "https://substack.com/@sifuyik/note/c-321255318",
"sourceType": "profile",
"discoveredFrom": {
"type": "sourceUrl",
"value": "https://substack.com/@sifuyik"
},
"publishedAt": "2026-08-24T01:32:51.983Z",
"editedAt": null,
"collectedAt": "2026-08-25T16:34:50.254Z",
"author": {
"id": "73134922",
"name": "Sifu Yik Chan",
"handle": "sifuyik",
"profileUrl": "https://substack.com/@sifuyik",
"photoUrl": "https://substack-post-media.s3.amazonaws.com/public/images/658267de-b7ce-4827-a4a4-bf4aec4d9ddc_866x866.jpeg",
"bio": "AI Tips and News"
},
"publication": {
"id": "7223942",
"name": "Sifu Yik's Substack",
"url": "https://sifuyik"
},
"engagement": {
"likes": 2,
"restacks": 0,
"replies": 0,
"total": 2,
"reactions": [
{
"emoji": "❤",
"count": 2
}
]
},
"media": [
{
"url": "https://substack-post-media.s3.amazonaws.com/public/images/831fada1-9358-40eb-9a73-f000a1804f6e_2560x3200.png",
"type": "image",
"width": 2560,
"height": 3200
}
],
"outboundLinks": [
"https://sifuyik.substack.com/p/24-hours-only-100-hours-of-ai-training"
],
"thread": {
"isReply": false,
"parentNoteId": null,
"parentUrl": null,
"rootNoteId": null,
"rootUrl": null
}
}

💳 Pricing

Pay per event. The primary event is Note at $0.00001 for each public Note saved to your dataset. The total follows the number of Notes saved.

🔌 Integrations

Start the Actor in Apify Console or through the standard Actor API. Read the dataset through its API link or export it after the run.

❓ FAQ

What does this Actor collect?

It collects public Substack Notes from topic searches and public profile, feed, or Note sources. It does not collect newsletter posts, full article bodies, or private and subscriber-only content.

Can I use more than one topic phrase?

Yes. Add several topic phrases to one Topic search run. They use the same search order, date filters, exclusions, and note limit.

Can I use a public profile, feed, or Note URL?

Yes. Choose Public source URLs and add one or more public Substack URLs. The URL can point to a profile, feed, or Note page.

What happens when the same Note matches twice?

The first eligible match is saved. Later matches from another phrase or submitted URL are ignored, and discoveredFrom keeps the first matching value.

Can I filter Notes by a date range?

Yes. Choose publication time or change time, then add an inclusive UTC start date, end date, or both.

Does this return every available Note?

No result set can promise full coverage. Public availability, search ranking, rate limits, and source changes can affect which Notes are returned.

What happens when a public field is missing?

The Actor keeps the missing value as null when the schema allows it. It does not invent a name, count, URL, or time.

Do I need a Substack account or API key?

No. The Actor reads anonymous public Substack surfaces and does not require authenticated access.

📝 Changelog

0.0: Initial release

🆘 Support

For issues, questions, or feature requests, file a ticket and I'll fix or implement it in less than 24h 🫡

Made with ❤️ by Maxime Dupré