Discourse Forum API — Topics, Posts & Categories avatar

Discourse Forum API — Topics, Posts & Categories

Pricing

from $0.30 / 1,000 topics

Go to Apify Store
Discourse Forum API — Topics, Posts & Categories

Discourse Forum API — Topics, Posts & Categories

Topics, posts and categories from any Discourse community as clean rows, from the public JSON Discourse serves every visitor: latest, new, top, a category, a tag, a search or listed topics, with post text as Markdown or plain text. Usernames only, never real names. No key, no login.

Pricing

from $0.30 / 1,000 topics

Rating

0.0

(0)

Developer

Insight Solutions

Insight Solutions

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

9 hours ago

Last modified

Share

Point it at any Discourse community — meta.discourse.org, community.openai.com, discuss.python.org, forum.obsidian.md, forums.docker.com, discourse.mozilla.org, forum.gitlab.com and thousands more — and get its topics, posts and categories as clean rows in one schema. Latest, new, top by period, one category, one tag, a search query, or the topics you list; each topic's posts as Markdown, plain text or HTML.

It reads exactly what Discourse serves to any anonymous visitor: every list and topic page is also available as JSON by adding .json to its address. No API key, no login, no cookie, no browser. Private categories are not readable anonymously, so they are not read — and never attempted.

At a glance

Input — this is the Store prefill; paste it and run:

{ "forumUrl": "https://meta.discourse.org", "mode": "latest", "maxTopics": 20, "includePosts": true,
"maxPostsPerTopic": 10, "postFormat": "markdown", "includeCategories": true, "maxConcurrency": 3,
"maxRunSecs": 240, "proxyConfiguration": { "useApifyProxy": true } }

The 20 most recently active topics on meta.discourse.org with up to 10 posts each, plus the forum's 46 categories and its about-page stats — about 24 requests.

Output — one topic row per topic; the fields you will use most are title, url, categoryName, tags, postsCount, views, likeCount and lastPostedAt. Then one post row per post, with username, content and likeCount (full list under Output reference). Anything that could not be read — a forum behind bot protection, a host that is not Discourse, a topic that does not exist — comes back as a free diagnostic row (ok: false, errorType, error) instead of a charge.

Price — $0.50 per 1,000 topics + $0.20 per 1,000 posts (+ $0.001 per run) on the FREE tier; categories, forum stats, diagnostics and empty runs free; no API key, no login, no browser, limited permissions, works over the Apify MCP server (mcp.apify.com) and with x402 agentic payments.

From code — client.actor("insight.solutions/discourse-forum-api").call(run_input={…}) with apify-client, or

POST https://api.apify.com/v2/acts/insight.solutions~discourse-forum-api/run-sync-get-dataset-items
.


What you get

A topic (from the prefill's run; empty columns left out here):

{
"ok": true,
"rowType": "topic",
"input": "latest",
"source": "meta.discourse.org",
"sourceUrl": "https://meta.discourse.org/t/413448.json",
"topicId": 413448,
"title": "Could usernames be included in user_badge webhook payload?",
"url": "https://meta.discourse.org/t/could-usernames-be-included-in-user-badge-webhook-payload/413448",
"categoryId": 6,
"categoryName": "Support",
"categorySlug": "support",
"tags": ["badges", "webhooks"],
"createdAt": "2026-09-28T02:25:16.465Z",
"lastPostedAt": "2026-09-30T18:14:14.584Z",
"postsCount": 15,
"replyCount": 11,
"views": 156,
"likeCount": 6,
"pinned": false,
"closed": false,
"hasAcceptedAnswer": false,
"excerpt": "Any chance the user_badge webhooks could include username along with user_id? Currently the payload for a user_badge event looks like this: …",
"originalPosterUsername": "burke",
"lastPosterUsername": "putty",
"posterUsernames": ["burke", "itsbhanusharma", "NateDhaliwal", "RGJ", "putty"],
"participantCount": 5,
"wordCount": 1095,
"rank": 2
}

A post — the third reply in that topic, which quotes the first:

{
"rowType": "post",
"topicId": 413448,
"postId": 2044613,
"postNumber": 3,
"url": "https://meta.discourse.org/t/could-usernames-be-included-in-user-badge-webhook-payload/413448/3",
"username": "itsbhanusharma",
"trustLevel": 3,
"isStaff": false,
"createdAt": "2026-09-28T13:29:35.464Z",
"likeCount": 0,
"reads": 26,
"content": "> **burke:**\n> I’m creating an external service that I was hoping to trigger when specific badge(s) are granted; however, it needs the username and badge webhooks only returns user IDs.\n\nFwiw, fetching username from user id is just one API call away, unless you’re manually granting thousands of badges in an instant, you should be fine with the default API limits.",
"contentText": "> burke:\n> I’m creating an external service … you should be fine with the default API limits.",
"wordCount": 62,
"replyToPostNumber": null,
"quoteCount": 1,
"isAcceptedAnswer": false
}

Plus, free: one category row per category (categoryName, categorySlug, description, topicCount, postCount, topicsWeek, parentCategoryId, subcategoryIds, …) and one forum row (forumTitle, discourseVersion and the about page's stats: topics_count, posts_count, users_count, active_users_7_days, …).

The Markdown keeps paragraphs, links (made absolute), lists, headings, tables, quotes (attributed by username), fenced code blocks with their language, images at full size and emoji as :shortcodes:. Link previews become one titled link.


Quick start

The latest topics, text only, no posts — one request per 30 topics:

{ "forumUrl": "https://discuss.python.org", "mode": "latest", "maxTopics": 300, "includePosts": false }

Everything in one category, with every post:

{ "forumUrl": "https://meta.discourse.org", "mode": "category", "category": "support", "maxTopics": 100, "maxPostsPerTopic": 0 }

A search, Discourse's own operators passed through:

{ "forumUrl": "https://meta.discourse.org", "mode": "search", "query": "webhook #support after:2026-01-01", "maxPostsPerTopic": 20 }

Monitoring a tag — schedule it daily; since keeps only topics with a post in the last day and stops the walk at the first page that is entirely older:

{ "forumUrl": "https://community.openai.com", "mode": "tag", "tag": "devday-2026", "since": "24h", "maxTopics": 200, "includeCategories": false }

Specific topics, by id or URL:

{ "forumUrl": "https://meta.discourse.org", "mode": "topics", "topicIds": ["https://meta.discourse.org/t/could-usernames-be-included-in-user-badge-webhook-payload/413448/15", "1"] }

Works with any Discourse community

Seven forums were read through the Apify datacenter proxy while this Actor was built, and all answered the same JSON: meta.discourse.org, community.openai.com, discuss.python.org, forum.obsidian.md, forums.docker.com, discourse.mozilla.org and forum.gitlab.com. A subfolder install (https://example.org/forum) works too — give its full address.

Some forums put a bot-protection challenge in front of everything (a Cloudflare "Just a moment…" page; community.cloudflare.com and community.home-assistant.io did in the capture). You get one free blocked row naming the host after a single retry from a fresh IP, and nothing else — the challenge is never worked around. A site that is not Discourse at all gets one free not-discourse row ("example.com does not look like a Discourse forum: /latest.json returned 404 text/html").


Input

FieldTypeDefaultWhat it does
forumUrlstringhttps://meta.discourse.orgThe community's address. https:// is added, a trailing slash or a pasted topic URL is trimmed back to the forum, a subfolder is kept.
modeselectlatestlatest · new (newest created) · top · category · tag · search · topics
periodselectmonthlyFor top: all, yearly, quarterly, monthly, weekly, daily
categorystring""For category: support, support/6, support/self-hosting, 6 or the category URL
tagstring""For tag: ai (or #ai, or the tag URL)
querystring""For search: passed through with its operators (order:latest, #category, tags:ai, after:2026-01-01, in:title)
topicIdsstring[][]For topics: ids, /t/<slug>/<id> paths or full topic URLs
maxTopicsinteger100Topic rows per run; list modes page until they reach it. 0 = no cap (the time budget still applies)
includePostsbooleantrueRead each topic and return its posts
maxPostsPerTopicinteger50Post rows per topic, first post first. 0 = all
postFormatselectmarkdownmarkdown · text · html — what content holds. contentText is always plain text
keywordsstring[][]Keep topics whose title or excerpt — or, with posts on, post text — contains any of these
sincestring—Keep topics with a post on or after this (created_at in new mode): an ISO date-time, a date or a window like 24h / 7d
includeCategoriesbooleantrueOne free category row per category
maxConcurrencyinteger3Topics fetched at once, each slot pacing itself to one request per 300 ms
maxRunSecsinteger240The run's time budget, 30–3600
proxyConfigurationobject{ "useApifyProxy": true }Datacenter by default

A run with no valid forumUrl, or a mode whose field is empty or unreadable, fails before any request with a free invalid-input row, and costs nothing.


Output reference

Every row has the same 77 columns (null where a column does not apply), so the dataset exports as one table.

Every row: ok, rowType (topic · post · category · forum · diagnostic), input (the mode and its key: latest, top:monthly, category:support, tag:ai, search:…, topic:413448), error, errorType, scrapedAt, source (the forum's host), sourceUrl (the JSON URL the row came from).

topic: topicId, title, slug, url, categoryId, categoryName, categorySlug, tags, createdAt, lastPostedAt, bumpedAt, postsCount, replyCount, views, likeCount, opLikeCount, pinned, closed, archived, visible, hasAcceptedAnswer, hasSummary, excerpt, imageUrl, featuredLink, originalPosterUsername, lastPosterUsername, posterUsernames, participantCount and wordCount (when the topic itself was read), matchSnippet (search), rank (1-based, in the forum's own order), locale.

post: postId, topicId, topicTitle, topicUrl, postNumber, url, username, userTitle, isStaff, trustLevel, createdAt, updatedAt, content, contentText, contentHtml (only with postFormat: html), wordCount, replyToPostNumber, replyCount, likeCount, reads, score, quoteCount, incomingLinkCount, links ({ url, title, internal, clicks }), isAcceptedAnswer, isWiki, version. Staff action notices ("pinned this topic", "closed this topic") are not posts anyone wrote and are skipped.

category (free): categoryId, categoryName, categorySlug, url, description, parentCategoryId, topicCount, postCount, topicsWeek, topicsMonth, topicsYear (top-level categories), readRestricted, color, position, subcategoryIds.

forum (free): forumTitle, forumDescription, discourseVersion, stats.

diagnostic (free): errorType is one of not-discourse, blocked, not-found, rate-limited, http, timeout, deadline, budget, no-results, upstream-format, invalid-input; error says what happened in a sentence.


Usernames, not names

Every user object Discourse sends carries the public username and a name field (a real name, when the person filled one in), an avatar, and on posts a display_username that repeats the name. This Actor returns usernames only — the handle a person chose to post under. It never returns name, display_username or avatar URLs; quotes inside posts are attributed from the quote's own username attribute and their avatar images are removed, even from postFormat: html. Anything shaped like an e-mail address in any text column is replaced with [email hidden], and the forum's contact address on its about page is never read. The test suite walks every key and string of every row a run produces to hold all of this in place.


What you are never charged for

  • Every category row and the forum row.
  • Every diagnostic row: a forum behind bot protection, a host that is not Discourse, a topic, tag or category that does not exist, a rate limit, a timeout.
  • Topics dropped by since or keywords — filters run before the charge.
  • Staff action notices inside a topic.
  • A run that returns nothing at all: it finishes SUCCEEDED with zero results and a "0 results. N diagnostic row(s) explain why. Nothing was charged." status message, and bills nothing, start fee included.

Pricing

EventWhat it isFREEBRONZESILVERGOLD
actor-startOnce per run, only after the run has returned a topic$0.001$0.001$0.001$0.001
topicOne topic row$0.0005$0.0005$0.0004$0.0003
postOne post row$0.0002$0.0002$0.00016$0.00012

$0.50 per 1,000 topics + $0.20 per 1,000 posts (+ $0.001 per run).

RunCost
The prefill — 20 topics, up to 10 posts each (191 posts in the test fixtures)$0.0492
100 topics with their first 20 posts each: 100 × 0.0005 + 2,000 × 0.0002 + 0.001$0.451
1,000 topics, includePosts: false$0.501
A daily tag monitor finding 15 topics with 10 posts each$0.0385 a day
A forum behind bot protection, an unknown tag, an empty search$0.00

Charging is charge-after-push: rows are in your dataset before the event is recorded, and a topic is pushed together with its posts. ACTOR_MAX_TOTAL_CHARGE_USD is respected — each topic and its posts are planned against what is left before either is pushed, the start fee kept in reserve — and when it is reached the run stops fetching, adds a free budget row and finishes SUCCEEDED.


Proxy and politeness

The default is { "useApifyProxy": true } — Apify's datacenter pool. meta.discourse.org returned byte-identical JSON through it and with no proxy at all, and six other forums answered through it.

Discourse limits anonymous readers per IP (on a default install, 200 requests a minute and 50 in ten seconds). Each topic slot uses its own proxy session and waits 300 ms between its requests; list pages are 250–600 ms apart; with no proxy at all, the slots share one clock at one request per 350 ms. A 429, a 403, a challenge page or Discourse's own rate-limit answer is retried once from a fresh session (after a short wait for a rate limit); a second refusal becomes a free row, and rows already returned are kept.


Use it from an AI agent, or from code

One JSON object in, one flat array out — the shape agent runtimes want. The Actor runs with limited permissions, uses pay-per-event pricing and never enters Standby, so it works over the Apify MCP server and with x402 agentic payments. The Integrations tab pushes results to Slack, a webhook, Zapier, Make, Google Sheets, Snowflake or BigQuery.

curl -X POST "https://api.apify.com/v2/acts/insight.solutions~discourse-forum-api/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"forumUrl":"https://meta.discourse.org","mode":"search","query":"webhook order:latest","maxTopics":50,"maxPostsPerTopic":5}'
# pip install apify-client
from apify_client import ApifyClient
client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("insight.solutions/discourse-forum-api").call(run_input={
"forumUrl": "https://discuss.python.org",
"mode": "top",
"period": "weekly",
"maxTopics": 50,
"maxPostsPerTopic": 10,
"postFormat": "markdown",
})
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
if row["rowType"] == "post":
print(row["topicTitle"], row["postNumber"], row["username"], row["content"][:200], sep=" | ")

FAQ

Can it read private categories or private messages? No. It reads what an anonymous visitor sees, and the anonymous JSON simply does not contain private categories. It never logs in, never asks for a key and never sends a cookie, so it never tries. A topic that is deleted or hidden from anonymous visitors comes back as a free row — not-found when the forum answers 404, blocked when it answers 403.

What about rate limits? Discourse's limits are per IP and generous for reading: the default pacing (three slots, 300 ms apart, each on its own proxy session) stays well inside them. If a forum is stricter, a refused request is retried once from a fresh IP and then becomes a free rate-limited row — lower maxConcurrency and run again.

How does paging work? List pages hold 30 topics (50 on top). Discourse says whether there is another page (more_topics_url) and the Actor follows exactly what it says until maxTopics, your budget, the time budget, the end of the list, or — with since — the first page that is entirely older. Search pages hold 50 results and are followed while Discourse reports more. A topic carries its first 20 posts; the rest are fetched 20 at a time, only as many as maxPostsPerTopic needs.

Why is excerpt often empty? Many forums only send an excerpt for pinned topics in their lists. With posts on, the full first post is in the post rows.

Why do some topics have no posterUsernames in search mode? Search results do not carry the poster list; with includePosts on, each topic is read and the list is filled from its participants.

Does keywords search the whole forum? No — it filters what a list returns. To search a forum, use mode: "search".


Limitations

  • The upstream JSON may change with Discourse versions and plugins. It has been stable for years and seven forums answered the same shape, but an old install or an unusual plugin can differ; a shape the Actor does not recognise comes back as a free upstream-format row.
  • Forums behind a bot-protection challenge cannot be read; you get a free blocked row.
  • A 403 from a forum's own JSON (for example a forum that requires login for everything) is reported as blocked too — the two cannot be told apart from the status alone.
  • topicsWeek / topicsMonth / topicsYear are published for top-level categories only.
  • Posts are converted from Discourse's rendered HTML; rich embeds (polls, calendar events, custom plugin blocks) come through as their text, not their structure.
  • Very long posts are cut at 50,000 characters with a trailing "…".

Our other Actors

Every Insight Solutions Actor is pay-per-result with no browser, no login and no API key, and every one of them returns free diagnostic rows instead of billing for failures. Prices are per 1,000 results.

Video, audio & social

News, documents & the web

Business, finance & jobs

Apps & games