# Truthsocial Scraper (`dtrungtin/truthsocial-scraper`) Actor

Scrapes the latest posts truthsocial.com. By default it
saves the **3 newest posts of @realDonaldTrump**; give it any other usernames or profile URLs and a different
number of posts. Each post with its text, date, media, link preview, reply/ReTruth/like
counts, hashtags and mentions.

- **URL**: https://apify.com/dtrungtin/truthsocial-scraper.md
- **Developed by:** [Tin](https://apify.com/dtrungtin) (community)
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $10.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Truthsocial Scraper

Scrapes the latest posts ("Truths") of public [Truth Social](https://truthsocial.com) accounts. By default it
saves the **3 newest posts of @realDonaldTrump**; give it any other usernames or profile URLs and a different
number of posts. Each post becomes one dataset item with its text, date, media, link preview, reply/ReTruth/like
counts, hashtags and mentions.

Truth Social is built on Mastodon, so public profiles are served by a Mastodon-compatible JSON API
(`/api/v1/accounts/lookup` and `/api/v1/accounts/{id}/statuses`). The Actor reads that API instead of rendering
pages in a browser, which is fast and cheap. Requests use browser-like headers and TLS fingerprints
(got-scraping), because the site is protected by Cloudflare.

### Input

| Field | Type | Description |
| --- | --- | --- |
| `usernames` | array | Required. Usernames or profile URLs: `realDonaldTrump`, `@realDonaldTrump`, `https://truthsocial.com/@realDonaldTrump` (a post URL works too). Duplicates are removed; invalid entries are skipped with a warning. |
| `maxPosts` | integer | Newest posts to save per account, 1 to 1000. Default `3`. |
| `excludeReplies` | boolean | Skip replies to other posts. Default `true`, so you get the account's own top-level posts. |
| `excludeReTruths` | boolean | Skip ReTruths (re-posts). Default `false`. |
| `proxyConfiguration` | object | Optional proxy. If runs are blocked, use `{ "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"], "apifyProxyCountry": "US" }`. |
| `accessToken` | string | Optional, secret. Only needed if Truth Social stops serving public posts to visitors who are not logged in (see below). |

Example input:

```json
{
  "usernames": ["realDonaldTrump", "https://truthsocial.com/@TeamTrump"],
  "maxPosts": 3,
  "excludeReplies": true,
  "excludeReTruths": false,
  "proxyConfiguration": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"], "apifyProxyCountry": "US" }
}
```

### Output

One item per post, newest first per account (`position` 1 is the newest). The dataset's "Latest posts" view
shows the main fields as a table.

```json
{
  "position": 1,
  "username": "realDonaldTrump",
  "displayName": "Donald J. Trump",
  "accountId": "107780257626128497",
  "profileUrl": "https://truthsocial.com/@realDonaldTrump",
  "id": "115000000000001000",
  "url": "https://truthsocial.com/@realDonaldTrump/115000000000001000",
  "createdAt": "2026-09-27T03:00:00.000Z",
  "text": "First paragraph of the post.\nSecond line https://example.com/article\n\nSecond paragraph #MAGA",
  "contentHtml": "<p>First paragraph of the post.<br>Second line <a href=\"https://example.com/article\">...</a></p><p>Second paragraph #MAGA</p>",
  "isReply": false,
  "inReplyToId": null,
  "isReTruth": false,
  "reTruthOf": null,
  "quoteOf": null,
  "repliesCount": 100,
  "reTruthsCount": 200,
  "likesCount": 300,
  "media": [
    { "type": "image", "url": "https://static-assets-1.truthsocial.com/...jpg", "previewUrl": "https://static-assets-1.truthsocial.com/...jpg", "description": null }
  ],
  "link": { "url": "https://example.com/article", "title": "Article title", "description": "Article summary", "imageUrl": "https://example.com/a.jpg" },
  "mentions": [],
  "hashtags": ["MAGA"],
  "language": "en",
  "sensitive": false,
  "spoilerText": null,
  "visibility": "public",
  "scrapedAt": "2026-09-27T04:13:40.736Z"
}
```

(The values above illustrate the format; they are not a real post.)

Notes on the fields:

- `text` is the post as plain text: paragraphs separated by a blank line, links written out in full.
  `contentHtml` is the original HTML.
- For a **ReTruth** (`isReTruth: true`), `text`, `media`, `link` and the counts describe the re-posted original,
  and `reTruthOf` names its author, URL, date and text. `createdAt` is when the account re-posted it.
- For a quote post, `quoteOf` holds the quoted post (or just its `id` when the API gives no details).
- `likesCount` is Mastodon's "favourites" count, shown as likes on Truth Social.

The run's status message summarises the result, e.g. "Saved 3 post(s) from 1 account(s)." and lists accounts that
were not found or could not be scraped.

### Dataset schema and views

`.actor/dataset_schema.json` describes every output field with its type (the `fields` JSON Schema above is the
exact contract: all fields are always present, and fields that can be empty are `null` rather than missing).
The Apify platform checks each saved item against this schema, and the Actor converts API values to the declared
types before saving, so an unexpected value from Truth Social is stored as `null` instead of failing the run.

The dataset offers three views in Apify Console and the API (`?view=<name>`):

| View | Shows |
| --- | --- |
| `overview` (Latest posts) | Account, position, date, text, ReTruth flag, reply/ReTruth/like counts, URL. |
| `engagement` (Engagement) | Likes, ReTruths and replies per post, plus reply and ReTruth flags. |
| `media` (Media and links) | Attachments, link preview, hashtags and mentions. |

### Output schema and run summary

`.actor/output_schema.json` tells Apify Console and the API where a run's results are. The run's **Output** tab
(and the `output` of a run started via API) shows these links:

| Output | Link |
| --- | --- |
| `latestPosts` | Dataset items in the "Latest posts" view (`?view=overview`). |
| `engagement` | Dataset items in the "Engagement" view (`?view=engagement`). |
| `mediaAndLinks` | Dataset items in the "Media and links" view (`?view=media`). |
| `allFields` | All dataset items with every field. |
| `summary` | The run summary, key-value store record `OUTPUT`. |

At the end of every run, the Actor saves the summary as the `OUTPUT` record, also when the run fails:

```json
{
  "status": "SUCCEEDED",
  "message": "Saved 2 post(s) from 1 account(s). Not found: @nobody. Failed: @blockedUser.",
  "postsSaved": 2,
  "accounts": [
    { "username": "realDonaldTrump", "status": "OK", "postsSaved": 2, "error": null },
    { "username": "nobody", "status": "NOT_FOUND", "postsSaved": 0, "error": null },
    { "username": "blockedUser", "status": "FAILED", "postsSaved": 0, "error": "truthsocial.com refused the request to /api/v1/accounts/lookup: Cloudflare bot protection (HTTP 403)." }
  ],
  "skippedInputs": [{ "value": "bad user!", "reason": "\"bad user!\" is not a valid username (letters, digits and _ only)" }],
  "settings": { "maxPosts": 2, "excludeReplies": true, "excludeReTruths": false, "proxy": false, "accessToken": false },
  "startedAt": "2026-09-27T04:50:00.000Z",
  "finishedAt": "2026-09-27T04:50:03.000Z"
}
```

`status` is `SUCCEEDED` or `FAILED`; `accounts[].status` is `OK`, `NOT_FOUND` or `FAILED`. `settings.accessToken`
only says whether a token was given; the token itself is never saved. When the input is invalid, `settings` is
`null` and `message` explains the problem.

### Memory

The Actor runs with **256 MB** by default and is capped at a **maximum of 2 GB** (2048 MB), set in
`.actor/actor.json` (`defaultMemoryMbytes`, `maxMemoryMbytes`). Runs cannot be started with more than 2 GB.
The Actor only makes JSON requests, so 256 MB is enough even for 1000 posts per account; raise it (up to 2 GB)
only when scraping many accounts in one run, which also gives the run more CPU.

### Blocking, proxies and login

- **Cloudflare:** Truth Social often blocks datacenter IP addresses, including Apify's default servers. The Actor
  recognises a block (Cloudflare challenge, HTTP 403/429, an HTML page instead of data), retries up to 3 times with
  back-off and, with a proxy configured, a new proxy session each time. If it stays blocked, the run fails with a
  message saying what to change. The most reliable setting is **Apify Proxy, RESIDENTIAL group, country US**.
- **Rate limits:** HTTP 429 answers are retried after the time the site asks for (at most 60 s).
- **Login:** public profiles are readable without an account. If Truth Social ever requires login for them, the run
  fails with "login required" and asks for `accessToken`. To get one: log in at truthsocial.com in your browser,
  open the developer tools (F12) > Network, reload, click any request to `truthsocial.com/api/...` and copy the
  value after `Bearer ` in its `Authorization` request header. The token belongs to your account; the input field
  is stored encrypted and never logged.
- If some accounts fail and others succeed, the run still succeeds with the posts it got; the status message lists
  the failed accounts. It fails only when no posts could be saved.

### Run locally

```bash
npm install
apify run --purge          # uses the prefill input: the 3 latest posts of @realDonaldTrump
```

Without the Apify CLI:

```bash
mkdir -p storage/key_value_stores/default
echo '{ "usernames": ["realDonaldTrump"], "maxPosts": 3 }' > storage/key_value_stores/default/INPUT.json
APIFY_LOCAL_STORAGE_DIR=./storage node src/main.js
```

Results are written to `storage/datasets/default/`. For offline development you can point the Actor at a mock of
the API with the environment variable `TRUTHSOCIAL_BASE_URL` (for example `http://127.0.0.1:18765`).

### Known limitations

- **Not tested against the live site yet.** This version was verified against a local mock of Truth Social's
  Mastodon-compatible API (pagination, filters, ReTruths, quotes, media, blocks, rate limits, login errors), not
  against truthsocial.com itself. If Truth Social's API differs, the first real run will show it in the log.
- Whether plain HTTP requests get through Cloudflare depends on the IP address. Expect to need a residential proxy
  on the Apify platform; residential proxy traffic is billed by Apify separately.
- Only public accounts and public posts are available. Private or suspended accounts are reported as not found.
- Engagement counts are what the API returns at scrape time; Truth Social may round or cache them.
- Truth Social may list a pinned post first; the Actor sorts by date, so pinned posts only appear if they are among
  the newest.

### Changelog

#### 0.3

- Output schema (`.actor/output_schema.json`, linked from `actor.json`): the run's Output tab links the three dataset
  views, all dataset fields and the run summary.
- New `OUTPUT` record with a run summary (status, posts saved per account, accounts not found or failed, skipped
  inputs, settings used), saved on success and on failure.
- Memory limits (2 GB max, 256 MB default) and the dataset schema from 0.2 are unchanged.

#### 0.2

- Maximum run memory capped at 2 GB (`maxMemoryMbytes: 2048`); default run memory set to 256 MB
  (`defaultMemoryMbytes: 256`) so the default always stays below the cap.
- Complete dataset schema: every output field is declared with its type and description, and all fields are
  required (nullable where a value can be missing). Two new views, "Engagement" and "Media and links", next to
  "Latest posts".
- Item values are converted to the schema types before saving (IDs as strings, counts as non-negative integers,
  empty strings as `null`), so the platform's schema check cannot reject an item because of an unexpected API value.
- No change to the input or to which posts are scraped.

#### 0.1

- First version: latest posts of public Truth Social accounts via the Mastodon-compatible API, with reply and
  ReTruth filters, optional proxy and access token, and block/rate-limit handling.

# Actor input Schema

## `usernames` (type: `array`):

Truth Social usernames or profile URLs, e.g. realDonaldTrump, @realDonaldTrump or https://truthsocial.com/@realDonaldTrump.

## `maxPosts` (type: `integer`):

How many of the newest posts to save for each account, newest first.

## `excludeReplies` (type: `boolean`):

Skip posts that are replies to other posts, so only the account's own top-level posts are saved.

## `excludeReTruths` (type: `boolean`):

Skip ReTruths (re-posts of other people's posts). When kept, a ReTruth's text, media and counts describe the original post.

## `proxyConfiguration` (type: `object`):

Truth Social is protected by Cloudflare, which often blocks datacenter IPs. If runs are blocked, use Apify Proxy with the RESIDENTIAL group and country US.

## `accessToken` (type: `string`):

Only needed if Truth Social stops serving public posts without login. Bearer token of your own Truth Social account; see the README for how to copy it. Stored encrypted.

## Actor input object example

```json
{
  "usernames": [
    "realDonaldTrump"
  ],
  "maxPosts": 3,
  "excludeReplies": true,
  "excludeReTruths": false
}
```

# Actor output Schema

## `latestPosts` (type: `string`):

Newest posts per account: date, text, counts and URL (dataset view "overview").

## `engagement` (type: `string`):

Likes, ReTruths and replies per post (dataset view "engagement").

## `mediaAndLinks` (type: `string`):

Attachments, link previews, hashtags and mentions (dataset view "media").

## `allFields` (type: `string`):

Every field of every post, as described by the dataset schema.

## `summary` (type: `string`):

OUTPUT record: status, posts saved per account, accounts not found or failed, settings used.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "usernames": [
        "realDonaldTrump"
    ],
    "maxPosts": 3
};

// Run the Actor and wait for it to finish
const run = await client.actor("dtrungtin/truthsocial-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "usernames": ["realDonaldTrump"],
    "maxPosts": 3,
}

# Run the Actor and wait for it to finish
run = client.actor("dtrungtin/truthsocial-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "usernames": [
    "realDonaldTrump"
  ],
  "maxPosts": 3
}' |
apify call dtrungtin/truthsocial-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,dtrungtin/truthsocial-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ziHpFmkMLGhP8hGZm/builds/Yfzbv7bNbg40Fn1F4/openapi.json
