# Discourse Forum API — Topics, Posts & Categories (`insight.solutions/discourse-forum-api`) Actor

Topics, posts and categories from any Discourse community as clean rows, from the public JSON Discourse serves every visitor: latest, new, top, a category, a tag, a search or listed topics, with post text as Markdown or plain text. Usernames only, never real names. No key, no login.

- **URL**: https://apify.com/insight.solutions/discourse-forum-api.md
- **Developed by:** [Insight Solutions](https://apify.com/insight.solutions) (community)
- **Categories:** Social media, Developer tools, Other
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.30 / 1,000 topics

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Discourse Forum API — Topics, Posts & Categories

Point it at any **Discourse** community — meta.discourse.org, community.openai.com,
discuss.python.org, forum.obsidian.md, forums.docker.com, discourse.mozilla.org,
forum.gitlab.com and thousands more — and get its **topics, posts and categories
as clean rows** in one schema. Latest, new, top by period, one category, one tag,
a search query, or the topics you list; each topic's posts as Markdown, plain text
or HTML.

It reads exactly what Discourse serves to any anonymous visitor: every list and
topic page is also available as JSON by adding `.json` to its address. **No API
key, no login, no cookie, no browser.** Private categories are not readable
anonymously, so they are not read — and never attempted.

### At a glance

**Input** — this is the Store prefill; paste it and run:

```json
{ "forumUrl": "https://meta.discourse.org", "mode": "latest", "maxTopics": 20, "includePosts": true,
  "maxPostsPerTopic": 10, "postFormat": "markdown", "includeCategories": true, "maxConcurrency": 3,
  "maxRunSecs": 240, "proxyConfiguration": { "useApifyProxy": true } }
```

The 20 most recently active topics on meta.discourse.org with up to 10 posts each,
plus the forum's 46 categories and its about-page stats — about 24 requests.

**Output** — one `topic` row per topic; the fields you will use most are `title`,
`url`, `categoryName`, `tags`, `postsCount`, `views`, `likeCount` and
`lastPostedAt`. Then one `post` row per post, with `username`, `content` and
`likeCount` (full list under *Output reference*). Anything that could not be read —
a forum behind bot protection, a host that is not Discourse, a topic that does not
exist — comes back as a free diagnostic row (`ok: false`, `errorType`, `error`)
instead of a charge.

**Price** — $0.50 per 1,000 topics + $0.20 per 1,000 posts (+ $0.001 per run) on
the FREE tier; categories, forum stats, diagnostics and empty runs free; no API
key, no login, no browser, limited permissions, works over the Apify MCP server
(`mcp.apify.com`) and with x402 agentic payments.

**From code** —
`client.actor("insight.solutions/discourse-forum-api").call(run_input={…})` with
`apify-client`, or `POST
https://api.apify.com/v2/acts/insight.solutions~discourse-forum-api/run-sync-get-dataset-items`.

***

### What you get

A **topic** (from the prefill's run; empty columns left out here):

```json
{
  "ok": true,
  "rowType": "topic",
  "input": "latest",
  "source": "meta.discourse.org",
  "sourceUrl": "https://meta.discourse.org/t/413448.json",
  "topicId": 413448,
  "title": "Could usernames be included in user_badge webhook payload?",
  "url": "https://meta.discourse.org/t/could-usernames-be-included-in-user-badge-webhook-payload/413448",
  "categoryId": 6,
  "categoryName": "Support",
  "categorySlug": "support",
  "tags": ["badges", "webhooks"],
  "createdAt": "2026-09-28T02:25:16.465Z",
  "lastPostedAt": "2026-09-30T18:14:14.584Z",
  "postsCount": 15,
  "replyCount": 11,
  "views": 156,
  "likeCount": 6,
  "pinned": false,
  "closed": false,
  "hasAcceptedAnswer": false,
  "excerpt": "Any chance the user_badge webhooks could include username along with user_id? Currently the payload for a user_badge event looks like this: …",
  "originalPosterUsername": "burke",
  "lastPosterUsername": "putty",
  "posterUsernames": ["burke", "itsbhanusharma", "NateDhaliwal", "RGJ", "putty"],
  "participantCount": 5,
  "wordCount": 1095,
  "rank": 2
}
```

A **post** — the third reply in that topic, which quotes the first:

```json
{
  "rowType": "post",
  "topicId": 413448,
  "postId": 2044613,
  "postNumber": 3,
  "url": "https://meta.discourse.org/t/could-usernames-be-included-in-user-badge-webhook-payload/413448/3",
  "username": "itsbhanusharma",
  "trustLevel": 3,
  "isStaff": false,
  "createdAt": "2026-09-28T13:29:35.464Z",
  "likeCount": 0,
  "reads": 26,
  "content": "> **burke:**\n> I’m creating an external service that I was hoping to trigger when specific badge(s) are granted; however, it needs the username and badge webhooks only returns user IDs.\n\nFwiw, fetching username from user id is just one API call away, unless you’re manually granting thousands of badges in an instant, you should be fine with the default API limits.",
  "contentText": "> burke:\n> I’m creating an external service … you should be fine with the default API limits.",
  "wordCount": 62,
  "replyToPostNumber": null,
  "quoteCount": 1,
  "isAcceptedAnswer": false
}
```

Plus, free: one **`category`** row per category (`categoryName`, `categorySlug`,
`description`, `topicCount`, `postCount`, `topicsWeek`, `parentCategoryId`,
`subcategoryIds`, …) and one **`forum`** row (`forumTitle`, `discourseVersion`
and the about page's `stats`: `topics_count`, `posts_count`, `users_count`,
`active_users_7_days`, …).

The Markdown keeps paragraphs, links (made absolute), lists, headings, tables,
quotes (attributed by username), fenced code blocks with their language, images
at full size and emoji as `:shortcodes:`. Link previews become one titled link.

***

### Quick start

**The latest topics, text only, no posts** — one request per 30 topics:

```json
{ "forumUrl": "https://discuss.python.org", "mode": "latest", "maxTopics": 300, "includePosts": false }
```

**Everything in one category**, with every post:

```json
{ "forumUrl": "https://meta.discourse.org", "mode": "category", "category": "support", "maxTopics": 100, "maxPostsPerTopic": 0 }
```

**A search**, Discourse's own operators passed through:

```json
{ "forumUrl": "https://meta.discourse.org", "mode": "search", "query": "webhook #support after:2026-01-01", "maxPostsPerTopic": 20 }
```

**Monitoring a tag** — schedule it daily; `since` keeps only topics with a post in
the last day and stops the walk at the first page that is entirely older:

```json
{ "forumUrl": "https://community.openai.com", "mode": "tag", "tag": "devday-2026", "since": "24h", "maxTopics": 200, "includeCategories": false }
```

**Specific topics**, by id or URL:

```json
{ "forumUrl": "https://meta.discourse.org", "mode": "topics", "topicIds": ["https://meta.discourse.org/t/could-usernames-be-included-in-user-badge-webhook-payload/413448/15", "1"] }
```

***

### Works with any Discourse community

Seven forums were read through the Apify datacenter proxy while this Actor was
built, and all answered the same JSON: **meta.discourse.org**,
**community.openai.com**, **discuss.python.org**, **forum.obsidian.md**,
**forums.docker.com**, **discourse.mozilla.org** and **forum.gitlab.com**. A
subfolder install (`https://example.org/forum`) works too — give its full address.

Some forums put a bot-protection challenge in front of everything (a Cloudflare
"Just a moment…" page; community.cloudflare.com and community.home-assistant.io
did in the capture). You get **one free `blocked` row naming the host** after a
single retry from a fresh IP, and nothing else — the challenge is never worked
around. A site that is not Discourse at all gets one free `not-discourse` row
("example.com does not look like a Discourse forum: /latest.json returned 404
text/html").

***

### Input

| Field | Type | Default | What it does |
|---|---|---|---|
| `forumUrl` | string | `https://meta.discourse.org` | The community's address. `https://` is added, a trailing slash or a pasted topic URL is trimmed back to the forum, a subfolder is kept. |
| `mode` | select | `latest` | `latest` · `new` (newest created) · `top` · `category` · `tag` · `search` · `topics` |
| `period` | select | `monthly` | For `top`: `all`, `yearly`, `quarterly`, `monthly`, `weekly`, `daily` |
| `category` | string | `""` | For `category`: `support`, `support/6`, `support/self-hosting`, `6` or the category URL |
| `tag` | string | `""` | For `tag`: `ai` (or `#ai`, or the tag URL) |
| `query` | string | `""` | For `search`: passed through with its operators (`order:latest`, `#category`, `tags:ai`, `after:2026-01-01`, `in:title`) |
| `topicIds` | string\[] | `[]` | For `topics`: ids, `/t/<slug>/<id>` paths or full topic URLs |
| `maxTopics` | integer | `100` | Topic rows per run; list modes page until they reach it. `0` = no cap (the time budget still applies) |
| `includePosts` | boolean | `true` | Read each topic and return its posts |
| `maxPostsPerTopic` | integer | `50` | Post rows per topic, first post first. `0` = all |
| `postFormat` | select | `markdown` | `markdown` · `text` · `html` — what `content` holds. `contentText` is always plain text |
| `keywords` | string\[] | `[]` | Keep topics whose title or excerpt — or, with posts on, post text — contains any of these |
| `since` | string | — | Keep topics with a post on or after this (`created_at` in `new` mode): an ISO date-time, a date or a window like `24h` / `7d` |
| `includeCategories` | boolean | `true` | One free `category` row per category |
| `maxConcurrency` | integer | `3` | Topics fetched at once, each slot pacing itself to one request per 300 ms |
| `maxRunSecs` | integer | `240` | The run's time budget, 30–3600 |
| `proxyConfiguration` | object | `{ "useApifyProxy": true }` | Datacenter by default |

A run with no valid `forumUrl`, or a mode whose field is empty or unreadable,
**fails before any request** with a free `invalid-input` row, and costs nothing.

***

### Output reference

Every row has the same 77 columns (null where a column does not apply), so the
dataset exports as one table.

**Every row:** `ok`, `rowType` (`topic` · `post` · `category` · `forum` ·
`diagnostic`), `input` (the mode and its key: `latest`, `top:monthly`,
`category:support`, `tag:ai`, `search:…`, `topic:413448`), `error`, `errorType`,
`scrapedAt`, `source` (the forum's host), `sourceUrl` (the JSON URL the row came
from).

**`topic`:** `topicId`, `title`, `slug`, `url`, `categoryId`, `categoryName`,
`categorySlug`, `tags`, `createdAt`, `lastPostedAt`, `bumpedAt`, `postsCount`,
`replyCount`, `views`, `likeCount`, `opLikeCount`, `pinned`, `closed`,
`archived`, `visible`, `hasAcceptedAnswer`, `hasSummary`, `excerpt`, `imageUrl`,
`featuredLink`, `originalPosterUsername`, `lastPosterUsername`,
`posterUsernames`, `participantCount` and `wordCount` (when the topic itself was
read), `matchSnippet` (search), `rank` (1-based, in the forum's own order),
`locale`.

**`post`:** `postId`, `topicId`, `topicTitle`, `topicUrl`, `postNumber`, `url`,
`username`, `userTitle`, `isStaff`, `trustLevel`, `createdAt`, `updatedAt`,
`content`, `contentText`, `contentHtml` (only with `postFormat: html`),
`wordCount`, `replyToPostNumber`, `replyCount`, `likeCount`, `reads`, `score`,
`quoteCount`, `incomingLinkCount`, `links` (`{ url, title, internal, clicks }`),
`isAcceptedAnswer`, `isWiki`, `version`. Staff action notices ("pinned this
topic", "closed this topic") are not posts anyone wrote and are skipped.

**`category`** (free): `categoryId`, `categoryName`, `categorySlug`, `url`,
`description`, `parentCategoryId`, `topicCount`, `postCount`, `topicsWeek`,
`topicsMonth`, `topicsYear` (top-level categories), `readRestricted`, `color`,
`position`, `subcategoryIds`.

**`forum`** (free): `forumTitle`, `forumDescription`, `discourseVersion`,
`stats`.

**`diagnostic`** (free): `errorType` is one of `not-discourse`, `blocked`,
`not-found`, `rate-limited`, `http`, `timeout`, `deadline`, `budget`,
`no-results`, `upstream-format`, `invalid-input`; `error` says what happened in a
sentence.

***

### Usernames, not names

Every user object Discourse sends carries the public `username` **and** a `name`
field (a real name, when the person filled one in), an avatar, and on posts a
`display_username` that repeats the name. This Actor returns **usernames only** —
the handle a person chose to post under. It never returns `name`,
`display_username` or avatar URLs; quotes inside posts are attributed from the
quote's own username attribute and their avatar images are removed, even from
`postFormat: html`. Anything shaped like an e-mail address in any text column is
replaced with `[email hidden]`, and the forum's contact address on its about page
is never read. The test suite walks every key and string of every row a run
produces to hold all of this in place.

***

### What you are never charged for

- Every `category` row and the `forum` row.
- Every `diagnostic` row: a forum behind bot protection, a host that is not
  Discourse, a topic, tag or category that does not exist, a rate limit, a
  timeout.
- Topics dropped by `since` or `keywords` — filters run before the charge.
- Staff action notices inside a topic.
- A run that returns nothing at all: it finishes **SUCCEEDED with zero results**
  and a "0 results. N diagnostic row(s) explain why. Nothing was charged." status
  message, and bills nothing, **start fee included**.

***

### Pricing

| Event | What it is | FREE | BRONZE | SILVER | GOLD |
|---|---|---|---|---|---|
| `actor-start` | Once per run, only after the run has returned a topic | $0.001 | $0.001 | $0.001 | $0.001 |
| **`topic`** | One topic row | **$0.0005** | **$0.0005** | **$0.0004** | **$0.0003** |
| **`post`** | One post row | **$0.0002** | **$0.0002** | **$0.00016** | **$0.00012** |

**$0.50 per 1,000 topics + $0.20 per 1,000 posts (+ $0.001 per run).**

| Run | Cost |
|---|---|
| The prefill — 20 topics, up to 10 posts each (191 posts in the test fixtures) | **$0.0492** |
| 100 topics with their first 20 posts each: 100 × 0.0005 + 2,000 × 0.0002 + 0.001 | **$0.451** |
| 1,000 topics, `includePosts: false` | **$0.501** |
| A daily tag monitor finding 15 topics with 10 posts each | **$0.0385** a day |
| A forum behind bot protection, an unknown tag, an empty search | **$0.00** |

Charging is charge-after-push: rows are in your dataset before the event is
recorded, and a topic is pushed together with its posts. `ACTOR_MAX_TOTAL_CHARGE_USD`
is respected — each topic and its posts are planned against what is left
before either is pushed, the start fee kept in reserve — and when it is reached
the run stops fetching, adds a free `budget` row and finishes SUCCEEDED.

***

### Proxy and politeness

The default is `{ "useApifyProxy": true }` — Apify's **datacenter** pool.
meta.discourse.org returned byte-identical JSON through it and with no proxy at
all, and six other forums answered through it.

Discourse limits anonymous readers per IP (on a default install, 200 requests a
minute and 50 in ten seconds). Each topic slot uses its own proxy session and
waits 300 ms between its requests; list pages are 250–600 ms apart; with no proxy
at all, the slots share one clock at one request per 350 ms. A 429, a 403, a
challenge page or Discourse's own rate-limit answer is retried **once** from a
fresh session (after a short wait for a rate limit); a second refusal becomes a
free row, and rows already returned are kept.

***

### Use it from an AI agent, or from code

One JSON object in, one flat array out — the shape agent runtimes want. The Actor runs with **limited permissions**, uses **pay-per-event** pricing and never enters Standby, so it works over the Apify MCP server and with x402 agentic payments. The **Integrations** tab pushes results to Slack, a webhook, Zapier, Make, Google Sheets, Snowflake or BigQuery.

```bash
curl -X POST "https://api.apify.com/v2/acts/insight.solutions~discourse-forum-api/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"forumUrl":"https://meta.discourse.org","mode":"search","query":"webhook order:latest","maxTopics":50,"maxPostsPerTopic":5}'
```

```python
## pip install apify-client
from apify_client import ApifyClient

client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("insight.solutions/discourse-forum-api").call(run_input={
    "forumUrl": "https://discuss.python.org",
    "mode": "top",
    "period": "weekly",
    "maxTopics": 50,
    "maxPostsPerTopic": 10,
    "postFormat": "markdown",
})

for row in client.dataset(run["defaultDatasetId"]).iterate_items():
    if row["rowType"] == "post":
        print(row["topicTitle"], row["postNumber"], row["username"], row["content"][:200], sep=" | ")
```

***

### FAQ

**Can it read private categories or private messages?** No. It reads what an
anonymous visitor sees, and the anonymous JSON simply does not contain private
categories. It never logs in, never asks for a key and never sends a cookie, so it
never tries. A topic that is deleted or hidden from anonymous visitors comes back
as a free row — `not-found` when the forum answers 404, `blocked` when it answers
403\.

**What about rate limits?** Discourse's limits are per IP and generous for
reading: the default pacing (three slots, 300 ms apart, each on its own proxy
session) stays well inside them. If a forum is stricter, a refused request is
retried once from a fresh IP and then becomes a free `rate-limited` row — lower
`maxConcurrency` and run again.

**How does paging work?** List pages hold 30 topics (50 on `top`). Discourse says
whether there is another page (`more_topics_url`) and the Actor follows exactly
what it says until `maxTopics`, your budget, the time budget, the end of the list,
or — with `since` — the first page that is entirely older. Search pages hold 50
results and are followed while Discourse reports more. A topic carries its first
20 posts; the rest are fetched 20 at a time, only as many as `maxPostsPerTopic`
needs.

**Why is `excerpt` often empty?** Many forums only send an excerpt for pinned
topics in their lists. With posts on, the full first post is in the `post` rows.

**Why do some topics have no `posterUsernames` in search mode?** Search results
do not carry the poster list; with `includePosts` on, each topic is read and the
list is filled from its participants.

**Does `keywords` search the whole forum?** No — it filters what a list returns.
To search a forum, use `mode: "search"`.

***

### Limitations

- **The upstream JSON may change with Discourse versions and plugins.** It has
  been stable for years and seven forums answered the same shape, but an old
  install or an unusual plugin can differ; a shape the Actor does not recognise
  comes back as a free `upstream-format` row.
- Forums behind a bot-protection challenge cannot be read; you get a free
  `blocked` row.
- A 403 from a forum's own JSON (for example a forum that requires login for
  everything) is reported as `blocked` too — the two cannot be told apart from the
  status alone.
- `topicsWeek` / `topicsMonth` / `topicsYear` are published for top-level
  categories only.
- Posts are converted from Discourse's rendered HTML; rich embeds (polls, calendar
  events, custom plugin blocks) come through as their text, not their structure.
- Very long posts are cut at 50,000 characters with a trailing "…".

### Our other Actors

Every Insight Solutions Actor is pay-per-result with no browser, no login and no API key, and every one of them returns free diagnostic rows instead of billing for failures. Prices are per 1,000 results.

**Video, audio & social**

- [YouTube Transcript API](https://apify.com/insight.solutions/youtube-transcript-api) — captions as timed segments, text, SRT or VTT, with language fallback and translation.
- [YouTube Comments API](https://apify.com/insight.solutions/youtube-comments-api) — comments and replies with likes, pinned and hearted flags, newest or top sort.
- [YouTube Channel API](https://apify.com/insight.solutions/youtube-channel-api) — a channel's videos, Shorts and live streams, plus YouTube search.
- [Podcast Search, Episodes & Charts API](https://apify.com/insight.solutions/podcast-api) — Apple Podcasts search, charts and full episode feeds.
- [Bluesky Scraper](https://apify.com/insight.solutions/bluesky-scraper) — profiles, posts, followers and follows from the public AT Protocol API.
- [Telegram Channel Scraper](https://apify.com/insight.solutions/telegram-channel-scraper) — posts, views and channel stats from public Telegram channels.
- [Substack Scraper](https://apify.com/insight.solutions/substack-scraper) — posts with full free text, comments and publication profiles.
- [Hacker News API](https://apify.com/insight.solutions/hacker-news-api) — stories, comments, users, front page and a structured "Who is hiring?" parser from the official HN APIs.

**News, documents & the web**

- [Google News Search, Topics & Real Article URLs](https://apify.com/insight.solutions/google-news-api) — news search and topic feeds with the publisher's real URL decoded.
- [Website to Markdown — Content Extractor for LLMs & RAG](https://apify.com/insight.solutions/website-content-extractor) — any site as clean Markdown, text and heading-aware chunks.
- [Internet Archive API](https://apify.com/insight.solutions/internet-archive-api) — archive.org search, item metadata, files and reviews.
- [Wayback Machine Toolkit](https://apify.com/insight.solutions/wayback-toolkit) — archived URL inventories, snapshots and text diffs between dates.
- [Website Technology Detector](https://apify.com/insight.solutions/website-tech-detector) — the tech stack behind any site, with the evidence for each detection.
- [Domain Intelligence API](https://apify.com/insight.solutions/domain-intelligence-api) — DNS, RDAP registration, TLS certificate and HTTP facts in one row per domain.
- [SEO Page Audit](https://apify.com/insight.solutions/seo-page-audit) — sitemap crawl with on-page checks, structured data and broken-link reports.
- [Keyword Suggestions API](https://apify.com/insight.solutions/keyword-suggestions-api) — Google, YouTube, Bing, Amazon and eBay autocomplete with alphabet and question expansions.
- [Website Contact Extractor](https://apify.com/insight.solutions/website-contact-extractor) — emails, phone numbers and social profiles from any list of websites.
- [Web Search Results API](https://apify.com/insight.solutions/web-search-api) — Bing and DuckDuckGo organic results with snippets, no key, no browser.
- [Company Enrichment API](https://apify.com/insight.solutions/company-enrichment-api) — a domain in, a company profile out: firmographics, contacts, tech stack, DNS and hiring signal.
- [Company Dossier API](https://apify.com/insight.solutions/company-dossier-api) — one company in, twelve sections out: profile, tech, contacts, DNS, open roles, news, SEC filings, federal awards, recalls, YC batch and apps.
- [Press Releases API](https://apify.com/insight.solutions/press-releases-api) — GlobeNewswire and PR Newswire releases plus any newsroom feed, by keyword, company, ticker or subject.
- [Federal Register API](https://apify.com/insight.solutions/federal-register-api) — rules, proposed rules, notices and the Public Inspection desk with dockets, comment deadlines and CFR references.
- [Academic Papers Search API](https://apify.com/insight.solutions/academic-papers-api) — OpenAlex, Crossref, arXiv and PubMed in one row per paper: abstract, citations, open-access PDF, authors and venue.
- [RSS & Atom Feed Monitor](https://apify.com/insight.solutions/rss-feed-monitor) — any RSS, Atom or JSON feed (or an OPML file) in, only the new items out, with keyword filters and a webhook.
- [Website Change Monitor](https://apify.com/insight.solutions/website-change-monitor) — watch any pages, diff the text between runs, get change rows with added/removed lines, keyword alerts and a webhook.
- [Wikipedia & Wikidata API](https://apify.com/insight.solutions/wikipedia-api) — article text, search, daily pageviews and Wikidata entity facts, any language edition.

**Business, finance & jobs**

- [Congress & Insider Trades API](https://apify.com/insight.solutions/congress-insider-trades-api) — STOCK Act periodic transaction reports and SEC Form 4 insider trades in one schema.
- [Federal Contracts, Grants & Lobbying API](https://apify.com/insight.solutions/federal-contracts-grants-api) — SAM.gov opportunities, USAspending awards, Grants.gov notices and Senate lobbying filings in one schema.
- [SEC EDGAR API](https://apify.com/insight.solutions/sec-edgar-api) — filings, XBRL financials and full-text search by ticker or CIK.
- [Clinical Trials & FDA API](https://apify.com/insight.solutions/clinical-trials-fda-api) — ClinicalTrials.gov studies plus openFDA recalls, labels, approvals, 510(k)s and adverse-event reports.
- [Product & Vehicle Recalls API](https://apify.com/insight.solutions/product-recalls-api) — CPSC, NHTSA, FDA and USDA recalls, vehicle complaints and ratings, plus a VIN decoder.
- [Y Combinator Companies, Batches & Founders](https://apify.com/insight.solutions/yc-companies-directory) — the YC directory with founders and social links, filterable by batch, industry and hiring status.
- [Career Site Jobs API](https://apify.com/insight.solutions/ats-jobs-api) — jobs straight from Greenhouse, Lever, Ashby, Workable and 10+ other ATS career sites.
- [New Job Postings Monitor](https://apify.com/insight.solutions/job-postings-monitor) — new, closed and changed postings on the career sites you watch.
- [Hiring Signals API — Open Roles & Hiring Surge by Company](https://apify.com/insight.solutions/hiring-signals-api) — one row per company per run: open roles, what opened and closed, department and seniority breakdowns, and a hiring-surge flag.
- [Remote Jobs API](https://apify.com/insight.solutions/remote-jobs-api) — RemoteOK, Remotive, We Work Remotely, Himalayas, Jobicy and more in one schema, deduplicated.
- [Shopify Products API](https://apify.com/insight.solutions/shopify-products-api) — any Shopify store's catalogue, variants, prices and stock signals.
- [Shopify Store Monitor](https://apify.com/insight.solutions/shopify-store-monitor) — price drops, sales, restocks, sell-outs and new products on any Shopify store, one row per change.
- [Public Tenders API](https://apify.com/insight.solutions/public-tenders-api) — EU TED, UK Find a Tender and Contracts Finder notices by keyword, CPV code, country, stage and deadline.
- [Nonprofit & IRS 990 Lookup API](https://apify.com/insight.solutions/nonprofit-990-api) — search US nonprofits and get EIN, NTEE code and multi-year Form 990 financials.
- [OpenStreetMap Places API](https://apify.com/insight.solutions/osm-places-api) — businesses and points of interest by category and area from OpenStreetMap: name, address, coordinates, website, phone, opening hours.

**Apps & games**

- [App Store & Google Play Reviews API](https://apify.com/insight.solutions/app-reviews-api) — reviews from both stores with ratings, versions and developer replies.
- [App Store Top Charts & App Search API](https://apify.com/insight.solutions/app-charts-api) — Apple top charts by country and genre, plus app search and details.
- [App Store Keyword Rank Tracker](https://apify.com/insight.solutions/app-store-keyword-rank-tracker) — where any app ranks for any keyword on the App Store and Google Play, with rank changes and ASO suggestions.
- [Steam Reviews API](https://apify.com/insight.solutions/steam-reviews-api) — Steam reviews with playtime, helpfulness and game details.
- [Steam Game Data API](https://apify.com/insight.solutions/steam-store-stats-api) — prices, tags, review scores, live player counts and top charts.

# Actor input Schema

## `forumUrl` (type: `string`):

The Discourse community to read, e.g. `https://meta.discourse.org`, `community.openai.com` or `https://example.org/forum` for a subfolder install. `https://` is added when missing, and a pasted topic or list URL is trimmed back to the forum's own address.

## `mode` (type: `string`):

Which topics to read. `latest` is the forum's front list (recently active first); `new` is newest-created first; `top` is the most active in a period; `category`, `tag` and `search` need the matching field below; `topics` reads the topic ids or URLs you list.

## `period` (type: `string`):

For `top`: the window Discourse ranks over.

## `category` (type: `string`):

For `category`: a slug (`support`), slug/id (`support/6`), parent/child slugs (`support/self-hosting`), an id (`6`) or the category's URL. A bare slug is looked up in the forum's own category list; if two subcategories share it you get a free row naming both.

## `tag` (type: `string`):

For `tag`: the tag name, e.g. `ai` (a leading `#` or a tag URL is fine). A tag that does not exist — or has no topics anonymous visitors can see — comes back as a free `not-found` row.

## `query` (type: `string`):

For `search`: passed to the forum's own search, operators included — `order:latest`, `#category`, `tags:ai`, `after:2026-01-01`, `in:title`. Each result is a topic, with the matching post's snippet in `matchSnippet`.

## `topicIds` (type: `array`):

For `topics`: numeric ids, `/t/<slug>/<id>` paths or full topic URLs (`https://…/t/some-topic/413448/15` reads topic 413448). Entries that are neither come back as free `invalid-input` rows.

## `maxTopics` (type: `integer`):

The most topic rows this run returns. List modes page until they reach it. 0 means no cap — the time budget is then the only bound.

## `includePosts` (type: `boolean`):

Read each topic (`/t/<id>.json`) and return its posts as `post` rows, with the text as Markdown, plain text or HTML. Off returns topic rows only — one request per list page instead of one per topic.

## `maxPostsPerTopic` (type: `integer`):

Cap on post rows per topic, first post first. 0 means every post. The topic itself carries 20; more are fetched 20 at a time.

## `postFormat` (type: `string`):

What `content` holds. `markdown` keeps paragraphs, links, lists, quotes (attributed by username) and code blocks; `text` is plain text; `html` is Discourse's rendered HTML, minus quote avatars. `contentText` is always the plain text.

## `keywords` (type: `array`):

Keep only topics whose title or excerpt — or, with posts on, whose post text — contains one of these (case-insensitive). Applied before anything is charged.

## `since` (type: `string`):

Keep topics active on or after this: `last_posted_at` (or `created_at` in `new` mode) at or after an ISO date-time (`2026-09-01T00:00:00Z`), a date (`2026-09-01`) or a window (`24h`, `7d`). A list walk stops at the first page that is entirely older. A window means the same thing on every run of a schedule.

## `includeCategories` (type: `boolean`):

Return one free `category` row per category the forum shows anonymous visitors: name, slug, description, topic and post counts, parent and subcategories.

## `maxConcurrency` (type: `integer`):

Topics fetched at once. Each slot waits 300 ms between its requests and uses its own proxy session; Discourse limits anonymous readers per IP, so raising this far rarely helps.

## `maxRunSecs` (type: `integer`):

Stop after this many seconds. Rows already returned are kept and a free `deadline` row says the budget ran out.

## `proxyConfiguration` (type: `object`):

Discourse's public JSON answered from the Apify datacenter proxy and from no proxy at all, byte-identical, so datacenter is the default. A forum behind a bot-protection challenge comes back as one free `blocked` row; the Actor does not try to get past it.

## Actor input object example

```json
{
  "forumUrl": "https://meta.discourse.org",
  "mode": "latest",
  "period": "monthly",
  "category": "",
  "tag": "",
  "query": "",
  "topicIds": [],
  "maxTopics": 20,
  "includePosts": true,
  "maxPostsPerTopic": 10,
  "postFormat": "markdown",
  "keywords": [],
  "includeCategories": true,
  "maxConcurrency": 3,
  "maxRunSecs": 240,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `results` (type: `string`):

Topic, post, category and forum rows from the Discourse forum's public JSON, in one schema, plus free diagnostic rows. Delivered as JSON items in the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "forumUrl": "https://meta.discourse.org",
    "mode": "latest",
    "maxTopics": 20,
    "includePosts": true,
    "maxPostsPerTopic": 10,
    "postFormat": "markdown",
    "includeCategories": true,
    "maxConcurrency": 3,
    "maxRunSecs": 240,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("insight.solutions/discourse-forum-api").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "forumUrl": "https://meta.discourse.org",
    "mode": "latest",
    "maxTopics": 20,
    "includePosts": True,
    "maxPostsPerTopic": 10,
    "postFormat": "markdown",
    "includeCategories": True,
    "maxConcurrency": 3,
    "maxRunSecs": 240,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("insight.solutions/discourse-forum-api").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "forumUrl": "https://meta.discourse.org",
  "mode": "latest",
  "maxTopics": 20,
  "includePosts": true,
  "maxPostsPerTopic": 10,
  "postFormat": "markdown",
  "includeCategories": true,
  "maxConcurrency": 3,
  "maxRunSecs": 240,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call insight.solutions/discourse-forum-api --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,insight.solutions/discourse-forum-api"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/FnToK4jzBQmJcvKn5/builds/bQDCF29xw62EIhcAq/openapi.json
