Discourse Forum Scraper – Topics, Posts, Replies & Categories avatar

Discourse Forum Scraper – Topics, Posts, Replies & Categories

Pricing

from $1.00 / 1,000 dataset items

Go to Apify Store
Discourse Forum Scraper – Topics, Posts, Replies & Categories

Discourse Forum Scraper – Topics, Posts, Replies & Categories

Scrape any Discourse forum — community.openai.com, community.n8n.io, discuss.pytorch.org, community.home-assistant.io and thousands more. Get rows for categories, topics and posts: title, URL, author, views, likes, tags, accepted-answer flag, and post bodies as HTML, Markdown and plain text.

Pricing

from $1.00 / 1,000 dataset items

Rating

0.0

(0)

Developer

R.L.

R.L.

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

0

Monthly active users

8 days ago

Last modified

Categories

Share

Discourse Forum Scraper — Topics, Posts, Replies & Categories from Any Community

What does this Discourse forum scraper do?

Discourse Forum Scraper extracts categories, topics and posts from any Discourse-powered community forum — no per-site configuration needed. Discourse powers most modern developer and product communities, and instead of parsing HTML this actor talks to the built-in JSON API every Discourse site exposes (/latest.json, /c/{category}.json, /search.json, /t/{id}.json, ...), which makes it fast, immune to theme and layout changes, and generic across sites. Point it at a forum (see the list below), choose what to crawl, and you get one row per category, topic and post.

Run it on the Apify platform for scheduling, API access, webhooks, and proxy rotation without managing any infrastructure yourself.

Why scrape a Discourse forum?

  • Community research & sentiment analysis — pull recurring questions, complaints, and feature requests from a product's support forum.
  • Competitive intelligence — monitor what users say about a competitor's product in its official community.
  • Support/knowledge-base mining — build a searchable dataset of solved threads and accepted answers.
  • Content/dataset building — collect real-world Q&A pairs or discussion threads for training or analysis.

Which Discourse forums does it work on?

Any of them. The actor checks /site.json — which every Discourse install serves — and refuses a host that isn't Discourse, so there is no per-site configuration and nothing to maintain when a forum changes its theme. These were each confirmed to answer with a valid Discourse /site.json:

ForumWhat it covers
community.openai.comOpenAI API and ChatGPT developer community
community.n8n.ion8n automation: questions, jobs, showcases
community.home-assistant.ioHome Assistant smart-home support
discuss.pytorch.orgPyTorch usage and troubleshooting
discuss.streamlit.ioStreamlit apps and deployment
discuss.elastic.coElasticsearch, Kibana, Logstash
forums.developer.nvidia.comCUDA, Jetson, drivers
forum.arduino.ccArduino hardware and sketches
forums.docker.comDocker Desktop and Engine
community.grafana.comGrafana dashboards and data sources
users.rust-lang.orgRust language help
forum.obsidian.mdObsidian plugins and workflows
forum.rclone.orgrclone configuration and bug reports
community.smartthings.comSmartThings devices and automations
discourse.ubuntu.comUbuntu announcements and docs discussion
community.wanikani.comWaniKani Japanese-learning community
meta.discourse.orgDiscourse's own forum about Discourse

Self-hosted and private-company Discourse instances work the same way, as long as the content you point at is public.

How do I run it?

  1. Set Forum URL to the base URL of any Discourse forum (e.g. https://community.n8n.io).
  2. Pick a Mode:
    • categories — list every category/section on the forum, so you can find the right categorySlug without guessing (run this first if you don't already know it).
    • latest / top — the site-wide topic list.
    • category — one category (set Category slug/path to a category item's category_slug_for_scraping from a categories run, e.g. questions/12).
    • search — Discourse's search endpoint (set Search query; supports Discourse search filters like order:latest #category).
    • topicUrls — scrape only the specific topic URLs you provide.
  3. Toggle Include full posts to also fetch every reply in each topic (off = topic metadata only, much faster).
  4. Set Max topics / Max posts per topic to bound the run, and hit Start.

What inputs does it take?

See the Input tab for the full schema. Key fields:

FieldDescription
forumUrlBase URL of the Discourse forum
modecategories, latest, top, category, search, or topicUrls
categorySlugCategory slug/path (mode=category)
searchQuerySearch query (mode=search)
topicUrlsList of direct topic URLs (mode=topicUrls)
includePostsFetch full post stream per topic
maxTopics / maxPostsPerTopicLimits (0 = unlimited)
requestDelaySecs / maxConcurrencyPoliteness/rate-limit controls
proxyConfigurationApify Proxy settings

What does a result row look like?

Up to three item shapes land in the same dataset, distinguished by type:

{
"type": "category",
"forum": "https://community.n8n.io",
"category_id": 13,
"title": "Jobs",
"slug": "jobs",
"url": "https://community.n8n.io/c/jobs/13",
"parent_category_id": null,
"topic_count": 615,
"category_slug_for_scraping": "jobs/13"
}
{
"type": "topic",
"forum": "https://community.n8n.io",
"topic_id": 302663,
"title": "AI Assistant on self-hosted n8n: early setup instructions",
"slug": "ai-assistant-on-self-hosted-n8n-early-setup-instructions",
"url": "https://community.n8n.io/t/ai-assistant-on-self-hosted-n8n-early-setup-instructions/302663",
"posts_count": 4,
"views": 512,
"like_count": 3,
"created_at": "2026-01-05T10:00:00.000Z"
}
{
"type": "post",
"topic_id": 302663,
"post_id": 565721,
"post_number": 1,
"username": "Ophir_Prusak",
"cooked_html": "<p>Here's how to set it up:</p><ol><li>Install the <strong>self-hosted</strong> package</li><li>Restart n8n</li></ol>",
"content_markdown": "Here's how to set it up:\n\n1. Install the **self-hosted** package\n2. Restart n8n",
"content_text": "Here's how to set it up: Install the self-hosted package Restart n8n",
"url": "https://community.n8n.io/t/ai-assistant-on-self-hosted-n8n-early-setup-instructions/302663/1"
}

cooked_html is Discourse's raw rendered post body. content_markdown and content_text are cleaned conversions of it (via markdownify and BeautifulSoup) — use these for LLM input, search indexing, or anywhere raw HTML is unwanted.

Download the dataset as JSON, CSV, Excel, or HTML from the Storage tab.

Which fields does each item type contain?

type: "category" — one row per category, from mode: categories:

FieldDescription
forumBase URL of the forum the row came from
category_idDiscourse's numeric category ID
titleCategory name
slugCategory slug as it appears in the URL
urlLink to the category
parent_category_idParent category's ID, or null for a top-level category
posts_countPosts in the category, as Discourse reports it
topic_countTopics in the category
category_slug_for_scrapingslug/id string to paste straight into categorySlug

type: "topic" — one row per topic:

FieldDescription
forum, topic_id, slug, urlWhere the topic lives
titleTopic title
category_idCategory the topic sits in
tagsArray of the topic's tags
created_at / last_posted_atISO timestamps of the first and latest post
posts_countPosts in the topic, including the first one
reply_countReplies as Discourse counts them
viewsView count
like_countLikes on the topic
pinned, closed, archivedTopic state flags
has_accepted_answertrue when a reply is marked as the solution

type: "post" — one row per post, when Include full posts is on:

FieldDescription
forum, topic_id, topic_titleThread the post belongs to
post_id, post_number, urlThe post's own ID, its position in the thread, and a direct link
username / nameAuthor's forum handle and display name
created_at / updated_atISO timestamps
reply_to_post_numberPost number this one replies to, or null for a top-level reply
reply_countReplies to this post
like_countLikes on the post
cooked_htmlRendered HTML of the post body
content_markdownPost body as clean Markdown
content_textPost body as clean plain text

What do the Discourse terms mean?

  • cooked — Discourse's own name for a post body after its Markdown has been rendered to HTML. cooked_html is that HTML verbatim; content_markdown and content_text are conversions of it, so quotes, code blocks and lists survive as text.
  • post stream — the ordered list of post IDs in a topic. The actor walks it in chunks of 20 so a 500-reply thread is fetched completely, not just the first page.
  • post_number — a post's 1-based position in its topic. 1 is the original post; everything above it is a reply.
  • accepted answer — a reply marked as the solution on forums that run Discourse's Solved plugin. has_accepted_answer is true on those topics and absent or null on forums without the plugin.
  • slug — the human-readable part of a Discourse URL. category_slug_for_scraping pairs it with the numeric ID (jobs/13), which is the form categorySlug wants.

How do I turn a forum into research? (worked example)

You can run this actor end-to-end through the Apify MCP server to answer research questions without writing scraping code. Example: "What kind of freelance work do n8n users want to buy in 2026?"

  1. Discover the category — call the actor with mode: "categories" on https://community.n8n.io. This lists every category with a ready-to-use category_slug_for_scraping, e.g. "jobs/13" for the Jobs category.
  2. Scrape it with posts — call the actor again with mode: "category", categorySlug: "jobs/13", includePosts: true, maxTopics: 80. This returns each job-post topic plus its replies, with content_text already cleaned of HTML.
  3. Analyze — pull content_text across all post items and count keyword mentions (e.g. "AI agent", "CRM", "voice AI", "RAG", "self-hosted") to see what buyers are actually asking for, and compare "hiring"/"looking for" vs. "for hire"/"available" topic titles to gauge demand vs. supply.

A real run against community.n8n.io/c/jobs (61 topics, 204 rows) found: AI agent building was the top request (72 mentions), usually combined with CRM integration (92), document/RAG/OCR processing (60/38/22), and a growing voice AI niche (Vapi + Twilio + n8n stacks, 31+23 mentions). Buy-side posts outnumbered for-hire posts roughly 2:1, and self-hosted n8n (21 mentions) was a common hard requirement. Rates ranged $15–50/hr or $250–1,500 flat, with a shift toward long-term retainers over one-off builds.

This same discover → scrape → analyze flow works for any Discourse forum's category (support backlogs, feature requests, showcase/built-with sections, etc.), not just job boards.

How much does it cost?

This actor is pay per event: $1.00 per 1,000 dataset items, and nothing else. Every row counts once — a category row, a topic row, or a post row.

So the cost follows the shape of the crawl, not the size of the forum:

RunRowsCost
mode: categories on a forum with 40 categories40$0.04
100 topics, Include full posts off100$0.10
100 topics averaging 8 posts each, posts on100 topics + 800 posts$0.90

maxTopics caps the topics and maxPostsPerTopic caps the posts per topic, so you can set a ceiling before you start. There is no browser and no rendering — just JSON requests — so runs are quick and the compute cost underneath is small. If you set a maximum cost per run on the Apify side, the actor stops pushing rows once that limit is reached instead of running past it.

Tips for a clean crawl

  • Leave includePosts off for a quick topic-index crawl, then re-run with topicUrls on just the topics you care about.
  • Lower maxConcurrency / raise requestDelaySecs on smaller or self-hosted forums to stay polite.
  • Use mode: search with Discourse's own filters (order:latest, #category, after:2026-01-01) when you want a slice of a big forum rather than its whole topic list.

FAQ

Do I need an account, an API key or a login?

No. Every Discourse forum serves a JSON twin of each page (/latest.json, /c/{slug}.json, /search.json, /t/{slug}/{id}.json) and the actor reads those anonymously. The flip side: categories that require a login are invisible to it, because they are invisible to an anonymous visitor.

How do I find the right category slug?

Run mode: categories once. Each category row carries category_slug_for_scraping (e.g. jobs/13) — paste that into categorySlug with mode: category. No guessing from URLs.

Does it get every reply in a long thread?

Yes. It reads the topic's full post stream and then fetches the posts that weren't inlined, 20 IDs per request, so a 500-reply thread comes out complete. Set maxPostsPerTopic if you only want the first N posts of each thread in thread order.

Can I scrape just a few specific threads?

Yes — mode: topicUrls with the topic URLs. Each distinct host is validated once against /site.json, and a URL on a host that isn't Discourse is skipped with a warning instead of failing the run.

Can I run it on a schedule?

Yes. A scheduled latest or search run is a simple way to watch a community for new threads. The actor does not remember previous runs, so rows repeat across runs — de-duplicate on topic_id and post_id if you append them to your own store.

Will it get me rate-limited or blocked?

It is built to avoid that: requestDelaySecs (default 1 second) and maxConcurrency (default 3) throttle it, and an HTTP 429 triggers a back-off of 5, then 10, then 15 seconds before retrying. Lower the concurrency further on small or self-hosted forums, and enable Apify Proxy for large crawls.

What happens if a page fails, or the site isn't Discourse?

A host whose /site.json isn't Discourse fails the run immediately with a message saying so, rather than producing empty results. After that, failures are contained: a request is retried three times and then skipped with an error in the log, a 404 is logged and skipped, and a topic that throws is skipped with a warning while the rest of the crawl continues.

Can I feed the posts straight into an LLM?

That is what content_markdown and content_text are for — the same post body without the HTML, so a thread can go into a prompt, an embedding or a search index as-is. cooked_html stays in the row for when you need the original markup.

Good to know

  • Only scrape publicly accessible forum content, and respect each forum's Terms of Service and robots.txt.
  • This actor is not affiliated with Discourse or with any of the forums named above.
  • Found a Discourse site this doesn't work on? Open an issue in the Issues tab.

Community research toolkit

Part of the Community research toolkit — research online communities and creators across forums, link-in-bio pages, and social platforms:

Did you find this useful?

⭐ Rate this actor on Apify! Your feedback helps other users find it and helps us keep improving it.