Discourse Forum Scraper – Topics, Posts, Replies & Categories
Pricing
from $1.00 / 1,000 dataset items
Discourse Forum Scraper – Topics, Posts, Replies & Categories
Scrape any Discourse forum — community.openai.com, community.n8n.io, discuss.pytorch.org, community.home-assistant.io and thousands more. Get rows for categories, topics and posts: title, URL, author, views, likes, tags, accepted-answer flag, and post bodies as HTML, Markdown and plain text.
Pricing
from $1.00 / 1,000 dataset items
Rating
0.0
(0)
Developer
R.L.
Maintained by CommunityActor stats
0
Bookmarked
3
Total users
0
Monthly active users
8 days ago
Last modified
Categories
Share
Discourse Forum Scraper — Topics, Posts, Replies & Categories from Any Community
What does this Discourse forum scraper do?
Discourse Forum Scraper extracts categories, topics and posts from any Discourse-powered community forum — no per-site configuration needed. Discourse powers most modern developer and product communities, and instead of parsing HTML this actor talks to the built-in JSON API every Discourse site exposes (/latest.json, /c/{category}.json, /search.json, /t/{id}.json, ...), which makes it fast, immune to theme and layout changes, and generic across sites. Point it at a forum (see the list below), choose what to crawl, and you get one row per category, topic and post.
Run it on the Apify platform for scheduling, API access, webhooks, and proxy rotation without managing any infrastructure yourself.
Why scrape a Discourse forum?
- Community research & sentiment analysis — pull recurring questions, complaints, and feature requests from a product's support forum.
- Competitive intelligence — monitor what users say about a competitor's product in its official community.
- Support/knowledge-base mining — build a searchable dataset of solved threads and accepted answers.
- Content/dataset building — collect real-world Q&A pairs or discussion threads for training or analysis.
Which Discourse forums does it work on?
Any of them. The actor checks /site.json — which every Discourse install serves — and refuses a host that isn't Discourse, so there is no per-site configuration and nothing to maintain when a forum changes its theme. These were each confirmed to answer with a valid Discourse /site.json:
| Forum | What it covers |
|---|---|
| community.openai.com | OpenAI API and ChatGPT developer community |
| community.n8n.io | n8n automation: questions, jobs, showcases |
| community.home-assistant.io | Home Assistant smart-home support |
| discuss.pytorch.org | PyTorch usage and troubleshooting |
| discuss.streamlit.io | Streamlit apps and deployment |
| discuss.elastic.co | Elasticsearch, Kibana, Logstash |
| forums.developer.nvidia.com | CUDA, Jetson, drivers |
| forum.arduino.cc | Arduino hardware and sketches |
| forums.docker.com | Docker Desktop and Engine |
| community.grafana.com | Grafana dashboards and data sources |
| users.rust-lang.org | Rust language help |
| forum.obsidian.md | Obsidian plugins and workflows |
| forum.rclone.org | rclone configuration and bug reports |
| community.smartthings.com | SmartThings devices and automations |
| discourse.ubuntu.com | Ubuntu announcements and docs discussion |
| community.wanikani.com | WaniKani Japanese-learning community |
| meta.discourse.org | Discourse's own forum about Discourse |
Self-hosted and private-company Discourse instances work the same way, as long as the content you point at is public.
How do I run it?
- Set Forum URL to the base URL of any Discourse forum (e.g.
https://community.n8n.io). - Pick a Mode:
categories— list every category/section on the forum, so you can find the rightcategorySlugwithout guessing (run this first if you don't already know it).latest/top— the site-wide topic list.category— one category (set Category slug/path to a category item'scategory_slug_for_scrapingfrom acategoriesrun, e.g.questions/12).search— Discourse's search endpoint (set Search query; supports Discourse search filters likeorder:latest #category).topicUrls— scrape only the specific topic URLs you provide.
- Toggle Include full posts to also fetch every reply in each topic (off = topic metadata only, much faster).
- Set Max topics / Max posts per topic to bound the run, and hit Start.
What inputs does it take?
See the Input tab for the full schema. Key fields:
| Field | Description |
|---|---|
forumUrl | Base URL of the Discourse forum |
mode | categories, latest, top, category, search, or topicUrls |
categorySlug | Category slug/path (mode=category) |
searchQuery | Search query (mode=search) |
topicUrls | List of direct topic URLs (mode=topicUrls) |
includePosts | Fetch full post stream per topic |
maxTopics / maxPostsPerTopic | Limits (0 = unlimited) |
requestDelaySecs / maxConcurrency | Politeness/rate-limit controls |
proxyConfiguration | Apify Proxy settings |
What does a result row look like?
Up to three item shapes land in the same dataset, distinguished by type:
{"type": "category","forum": "https://community.n8n.io","category_id": 13,"title": "Jobs","slug": "jobs","url": "https://community.n8n.io/c/jobs/13","parent_category_id": null,"topic_count": 615,"category_slug_for_scraping": "jobs/13"}
{"type": "topic","forum": "https://community.n8n.io","topic_id": 302663,"title": "AI Assistant on self-hosted n8n: early setup instructions","slug": "ai-assistant-on-self-hosted-n8n-early-setup-instructions","url": "https://community.n8n.io/t/ai-assistant-on-self-hosted-n8n-early-setup-instructions/302663","posts_count": 4,"views": 512,"like_count": 3,"created_at": "2026-01-05T10:00:00.000Z"}
{"type": "post","topic_id": 302663,"post_id": 565721,"post_number": 1,"username": "Ophir_Prusak","cooked_html": "<p>Here's how to set it up:</p><ol><li>Install the <strong>self-hosted</strong> package</li><li>Restart n8n</li></ol>","content_markdown": "Here's how to set it up:\n\n1. Install the **self-hosted** package\n2. Restart n8n","content_text": "Here's how to set it up: Install the self-hosted package Restart n8n","url": "https://community.n8n.io/t/ai-assistant-on-self-hosted-n8n-early-setup-instructions/302663/1"}
cooked_html is Discourse's raw rendered post body. content_markdown and content_text are cleaned conversions of it (via markdownify and BeautifulSoup) — use these for LLM input, search indexing, or anywhere raw HTML is unwanted.
Download the dataset as JSON, CSV, Excel, or HTML from the Storage tab.
Which fields does each item type contain?
type: "category" — one row per category, from mode: categories:
| Field | Description |
|---|---|
forum | Base URL of the forum the row came from |
category_id | Discourse's numeric category ID |
title | Category name |
slug | Category slug as it appears in the URL |
url | Link to the category |
parent_category_id | Parent category's ID, or null for a top-level category |
posts_count | Posts in the category, as Discourse reports it |
topic_count | Topics in the category |
category_slug_for_scraping | slug/id string to paste straight into categorySlug |
type: "topic" — one row per topic:
| Field | Description |
|---|---|
forum, topic_id, slug, url | Where the topic lives |
title | Topic title |
category_id | Category the topic sits in |
tags | Array of the topic's tags |
created_at / last_posted_at | ISO timestamps of the first and latest post |
posts_count | Posts in the topic, including the first one |
reply_count | Replies as Discourse counts them |
views | View count |
like_count | Likes on the topic |
pinned, closed, archived | Topic state flags |
has_accepted_answer | true when a reply is marked as the solution |
type: "post" — one row per post, when Include full posts is on:
| Field | Description |
|---|---|
forum, topic_id, topic_title | Thread the post belongs to |
post_id, post_number, url | The post's own ID, its position in the thread, and a direct link |
username / name | Author's forum handle and display name |
created_at / updated_at | ISO timestamps |
reply_to_post_number | Post number this one replies to, or null for a top-level reply |
reply_count | Replies to this post |
like_count | Likes on the post |
cooked_html | Rendered HTML of the post body |
content_markdown | Post body as clean Markdown |
content_text | Post body as clean plain text |
What do the Discourse terms mean?
- cooked — Discourse's own name for a post body after its Markdown has been rendered to HTML.
cooked_htmlis that HTML verbatim;content_markdownandcontent_textare conversions of it, so quotes, code blocks and lists survive as text. - post stream — the ordered list of post IDs in a topic. The actor walks it in chunks of 20 so a 500-reply thread is fetched completely, not just the first page.
- post_number — a post's 1-based position in its topic.
1is the original post; everything above it is a reply. - accepted answer — a reply marked as the solution on forums that run Discourse's Solved plugin.
has_accepted_answeristrueon those topics and absent ornullon forums without the plugin. - slug — the human-readable part of a Discourse URL.
category_slug_for_scrapingpairs it with the numeric ID (jobs/13), which is the formcategorySlugwants.
How do I turn a forum into research? (worked example)
You can run this actor end-to-end through the Apify MCP server to answer research questions without writing scraping code. Example: "What kind of freelance work do n8n users want to buy in 2026?"
- Discover the category — call the actor with
mode: "categories"onhttps://community.n8n.io. This lists every category with a ready-to-usecategory_slug_for_scraping, e.g."jobs/13"for theJobscategory. - Scrape it with posts — call the actor again with
mode: "category",categorySlug: "jobs/13",includePosts: true,maxTopics: 80. This returns each job-post topic plus its replies, withcontent_textalready cleaned of HTML. - Analyze — pull
content_textacross allpostitems and count keyword mentions (e.g. "AI agent", "CRM", "voice AI", "RAG", "self-hosted") to see what buyers are actually asking for, and compare "hiring"/"looking for" vs. "for hire"/"available" topic titles to gauge demand vs. supply.
A real run against community.n8n.io/c/jobs (61 topics, 204 rows) found: AI agent building was the top request (72 mentions), usually combined with CRM integration (92), document/RAG/OCR processing (60/38/22), and a growing voice AI niche (Vapi + Twilio + n8n stacks, 31+23 mentions). Buy-side posts outnumbered for-hire posts roughly 2:1, and self-hosted n8n (21 mentions) was a common hard requirement. Rates ranged $15–50/hr or $250–1,500 flat, with a shift toward long-term retainers over one-off builds.
This same discover → scrape → analyze flow works for any Discourse forum's category (support backlogs, feature requests, showcase/built-with sections, etc.), not just job boards.
How much does it cost?
This actor is pay per event: $1.00 per 1,000 dataset items, and nothing else. Every row counts once — a category row, a topic row, or a post row.
So the cost follows the shape of the crawl, not the size of the forum:
| Run | Rows | Cost |
|---|---|---|
mode: categories on a forum with 40 categories | 40 | $0.04 |
| 100 topics, Include full posts off | 100 | $0.10 |
| 100 topics averaging 8 posts each, posts on | 100 topics + 800 posts | $0.90 |
maxTopics caps the topics and maxPostsPerTopic caps the posts per topic, so you can set a ceiling before you start. There is no browser and no rendering — just JSON requests — so runs are quick and the compute cost underneath is small. If you set a maximum cost per run on the Apify side, the actor stops pushing rows once that limit is reached instead of running past it.
Tips for a clean crawl
- Leave
includePostsoff for a quick topic-index crawl, then re-run withtopicUrlson just the topics you care about. - Lower
maxConcurrency/ raiserequestDelaySecson smaller or self-hosted forums to stay polite. - Use
mode: searchwith Discourse's own filters (order:latest,#category,after:2026-01-01) when you want a slice of a big forum rather than its whole topic list.
FAQ
Do I need an account, an API key or a login?
No. Every Discourse forum serves a JSON twin of each page (/latest.json, /c/{slug}.json, /search.json, /t/{slug}/{id}.json) and the actor reads those anonymously. The flip side: categories that require a login are invisible to it, because they are invisible to an anonymous visitor.
How do I find the right category slug?
Run mode: categories once. Each category row carries category_slug_for_scraping (e.g. jobs/13) — paste that into categorySlug with mode: category. No guessing from URLs.
Does it get every reply in a long thread?
Yes. It reads the topic's full post stream and then fetches the posts that weren't inlined, 20 IDs per request, so a 500-reply thread comes out complete. Set maxPostsPerTopic if you only want the first N posts of each thread in thread order.
Can I scrape just a few specific threads?
Yes — mode: topicUrls with the topic URLs. Each distinct host is validated once against /site.json, and a URL on a host that isn't Discourse is skipped with a warning instead of failing the run.
Can I run it on a schedule?
Yes. A scheduled latest or search run is a simple way to watch a community for new threads. The actor does not remember previous runs, so rows repeat across runs — de-duplicate on topic_id and post_id if you append them to your own store.
Will it get me rate-limited or blocked?
It is built to avoid that: requestDelaySecs (default 1 second) and maxConcurrency (default 3) throttle it, and an HTTP 429 triggers a back-off of 5, then 10, then 15 seconds before retrying. Lower the concurrency further on small or self-hosted forums, and enable Apify Proxy for large crawls.
What happens if a page fails, or the site isn't Discourse?
A host whose /site.json isn't Discourse fails the run immediately with a message saying so, rather than producing empty results. After that, failures are contained: a request is retried three times and then skipped with an error in the log, a 404 is logged and skipped, and a topic that throws is skipped with a warning while the rest of the crawl continues.
Can I feed the posts straight into an LLM?
That is what content_markdown and content_text are for — the same post body without the HTML, so a thread can go into a prompt, an embedding or a search index as-is. cooked_html stays in the row for when you need the original markup.
Good to know
- Only scrape publicly accessible forum content, and respect each forum's Terms of Service and
robots.txt. - This actor is not affiliated with Discourse or with any of the forums named above.
- Found a Discourse site this doesn't work on? Open an issue in the Issues tab.
Community research toolkit
Part of the Community research toolkit — research online communities and creators across forums, link-in-bio pages, and social platforms:
- Reddit API Scraper — Scrape Reddit posts, comments, search results, subreddits, and user profiles.
- Linktree Profile Scraper — Scrape profile info and all links from Linktree pages.
- phpBB Forum Scraper — Generic scraper for phpBB-powered forums.
- Hacker News Scraper — Scrape Hacker News by search, user, listing, or thread.
- Social Blade Scraper — YouTube, TikTok, Instagram, Twitch Stats — Social Blade stats and growth projections for creators.
Did you find this useful?
⭐ Rate this actor on Apify! Your feedback helps other users find it and helps us keep improving it.