# Social Comment Classifier — buying intent & questions · $0.5/1k (`leoworks/social-comment-classifier`) Actor

Classify Instagram, TikTok, Facebook and YouTube comments from any comment scraper's dataset: purchase intent, questions, complaints, requests, praise, spam and sentiment, with probabilities. Comment sentiment analysis, no prompts, no LLM key.

- **URL**: https://apify.com/leoworks/social-comment-classifier.md
- **Developed by:** [Leoworks](https://apify.com/leoworks) (community)
- **Categories:** AI, Social media, Agents
- **Stats:** 3 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.43 / 1,000 comment classifications

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Social Comment Classifier — purchase intent, questions & sentiment

**For brands, social media managers, agencies and AI agents that already scrape comments** — from an Instagram, TikTok, Facebook or YouTube comment scraper, any other Apify dataset, or pasted as text — and need every comment labelled by what it is, for $0.50 per 1,000 comments, with no prompt writing or LLM key.

- **Comment type** — `purchase_intent` · `question` · `complaint` · `request` · `praise` · `spam` · `other` (one main type, with a probability for every type)
- **Sentiment** — positive / neutral / negative
- **Needs reply** — whether the brand or creator should answer (questions, problems, "how do I buy?")
- **Your own labels** — up to 10 yes/no criteria in plain language (e.g. "mentions shipping or delivery", "asks about a discount code")

**Use it to:** find buyers in your comments ("price?", "link?", "when is the restock?") · answer unanswered questions · catch complaints before they spread · hide spam and fake-shop promotion · collect product requests · Instagram, TikTok and YouTube comment sentiment analysis at scale · compare comment mix across posts or competitors.

### Output sample

Real rows from run `PgvMClRhIRV9gKmp7` (2026-10-09): public comments from the four comment scrapers below, with one custom label ("mentions shipping or delivery").

| comment | source | type (probability) | sentiment | needs reply | custom: shipping |
|---|---|---|---|---|---|
| When will you restock Generation G Fuzz? | Instagram | **purchase_intent** (1.00) | neutral | yes | no |
| Link | Facebook (live shopping) | **purchase_intent** (0.56) | neutral | no | no |
| How much do they hold? | Instagram | **question** (1.00) | neutral | yes | no |
| Temu, you guys are up over here making tiktoks. When im still am waiting for my package… | TikTok | **complaint** (1.00) | negative | yes | **yes** |
| I would like a refund for the 2 weeks the towers were down in my area. | Facebook | **complaint** (0.94) | negative | yes | no |
| can y'all make the resurfacing retinol serum in a bigger bottle ? | TikTok | **request** (1.00) | neutral | yes | no |
| obsessed with every single one of these! | Instagram | **praise** (1.00) | positive | no | no |
| For me personally… I do buy quality fakes/knockoffs at (shop name) | YouTube | **spam** (0.72) | neutral | no | no |

Each row also keeps the comment ID and post fields you choose (`id`, `cid`, `postUrl`, `videoWebUrl`, …), and in full mode `typeScores` with the probability of every type.

### Input example

The form default — two pasted comments, no dataset needed (about $0.001, 2 seconds):

```json
{
  "texts": [
    "Where can I buy this in Canada? Need it!!",
    "Ordered 3 weeks ago and still nothing. Is anyone answering messages?"
  ]
}
```

To classify a comment scraper's output, pass its dataset instead — the text, post title and ID fields are detected automatically:

```json
{
  "datasetId": "YOUR_COMMENTS_DATASET_ID",
  "customLabels": ["mentions shipping or delivery"]
}
```

### Pricing

Pay only for classified comments — no subscription.

| Event | Price | When |
|---|---|---|
| `comment-judged` | $0.0005 | One comment classified (type such as purchase intent, question or complaint; sentiment; needs-reply flag and any custom labels). |

That is **$0.50 per 1,000 comments**. **First run with the form defaults: about $0.001** (2 comments, 2 seconds).

**Cost examples**

| Comments | Cost |
|---|---|
| 100 | $0.05 |
| 1,000 | $0.50 |
| 10,000 | $5.00 |
| 100,000 | $50.00 |

With the free $5 monthly Apify credit you can classify about **10,000 comments**.

Items without comment text are skipped and **not charged**. Comments that fail after retries are reported with an `error` field and **not charged**. If you set a maximum cost per run, the Actor stops cleanly when it is reached.

### Works with

**The most used comment scrapers on Apify Store** — run one, then pass its dataset to this Actor. Field detection was checked against real output of each (2026-10-09, 20 comments per scraper).

| Platform | Scraper on Apify Store | Text field | Post title used | IDs kept |
|---|---|---|---|---|
| Instagram | [Instagram Comments Scraper (apify)](https://apify.com/apify/instagram-comment-scraper) | `text` | — | `id`, `commentUrl`, `postUrl` |
| TikTok | [TikTok Comments Scraper (clockworks)](https://apify.com/clockworks/tiktok-comments-scraper) | `text` | — | `cid`, `videoWebUrl` |
| Facebook | [Facebook Comments Scraper (apify)](https://apify.com/apify/facebook-comments-scraper) | `text` | `postTitle` | `commentId`, `id`, `commentUrl`, `facebookUrl` |
| YouTube | [YouTube Comments Scraper (streamers)](https://apify.com/streamers/youtube-comments-scraper) | `comment` | `title` (video title) | `cid`, `videoId` |

**Also:** the Apify API and JavaScript/Python clients · Apify Schedules (e.g. classify new comments every morning) · Claude, Cursor and Claude Code through the Apify MCP server (next section) · for product reviews, our [Korean](https://apify.com/leoworks/korean-review-classifier), [Japanese](https://apify.com/leoworks/japanese-review-classifier) and [AliExpress](https://apify.com/leoworks/aliexpress-reviews-classifier) review classifiers.

### Use with Claude, Cursor or Claude Code (MCP)

Add the Apify MCP server with this Actor as a tool and ask your agent in plain language — for example *"Which of these comments are from people who want to buy? …"* or *"Classify dataset abc123 from my TikTok comments run and list the unanswered questions."* The agent calls the tool `leoworks--social-comment-classifier` and reads the labels with `get-dataset-items`.

Claude Desktop or Cursor (`mcp.json`):

```json
{
  "mcpServers": {
    "apify": {
      "url": "https://mcp.apify.com?tools=leoworks/social-comment-classifier",
      "headers": { "Authorization": "Bearer YOUR_APIFY_TOKEN" }
    }
  }
}
```

Claude Code: `claude mcp add --transport http apify "https://mcp.apify.com?tools=leoworks/social-comment-classifier" --header "Authorization: Bearer YOUR_APIFY_TOKEN"`. Leave out the header to sign in with OAuth in the browser instead. Your Apify token is in Console → Settings → API & Integrations. We verified this setup with the Apify MCP server (v0.17.4) on 2026-10-09: the agent classified a pasted comment in 6 seconds end to end (run `XTdACdD0TchL5eQ34`).

### Output (one row per comment)

```json
{
  "cid": "7694457506550088461",
  "videoWebUrl": "https://www.tiktok.com/@temu/video/7694358584565583117",
  "text": "Temu, you guys are up over here making tiktoks. When im still am waiting for my package like, give me my freaking package, please.",
  "labels": {
    "type": { "label": "complaint", "probability": 1, "confidence": 1 },
    "sentiment": { "label": "negative", "probability": 1, "confidence": 1 },
    "needsReply": { "probability": 0.84, "matched": true },
    "custom": [{ "label": "mentions shipping or delivery", "probability": 0.98, "matched": true }],
    "typeScores": [
      { "label": "complaint", "probability": 1 },
      { "label": "spam", "probability": 0 }
    ]
  }
}
```

`typeScores` lists all seven types (shortened here). **Minimal mode** returns label keys only (`"type": "complaint", "sentiment": "negative", "needsReply": true`).

**How types are chosen:** each comment gets one main type. When a comment fits several, the first in this order wins: spam → purchase_intent → complaint → question → request → praise → other. So "Love it! Where can I buy it in the UK?" is `purchase_intent`, and "It stopped working after a week, any tips?" is `complaint`. Use `typeScores` or **Needs reply** when you want every comment that asks something.

#### Summary by post (REPORT)

Each run also saves a **REPORT** record (Output tab → *Summary by post*) at no extra charge: for every post, the comment type mix, sentiment shares, the share of comments that need a reply and up to 3 example comments for purchase intent, questions, complaints and requests — plus the same for all comments together. Posts are grouped by `summaryGroupField` (detected automatically from fields such as `postUrl`, `videoWebUrl`, `facebookUrl` or `videoId` when empty).

```json
{
  "groupField": "videoWebUrl",
  "groups": [{
    "group": "https://www.tiktok.com/@brand/video/1",
    "comments": 120,
    "classified": 118,
    "types": [{ "label": "praise", "count": 51, "share": 0.432 }, { "label": "purchase_intent", "count": 22, "share": 0.186 }],
    "sentiment": { "positive": 0.55, "neutral": 0.36, "negative": 0.09 },
    "needsReplyRate": 0.31,
    "examples": { "purchase_intent": ["price?", "Do you ship to Canada?"], "question": ["Is it waterproof?"] }
  }]
}
```

### Accuracy

Measured on 205 hand-labelled public English comments from Instagram, TikTok, Facebook and YouTube (beauty brands, a mobile carrier, a retailer's live shopping and product review videos; 2026-10-09). The questions were adjusted on one part, then checked once on a separate part that was not used for adjusting.

| Set | Comments | Main type | Main type (either label where two fit) | Sentiment |
|---|---|---|---|---|
| **Check set — not used for adjusting** | 100 | **90%** | 95% | **87%** |
| Adjusting set | 105 | 88% | 93% | 84% |

On the check set, complaints were found 97% of the time, questions 100%, spam 86% and purchase intent 71% (7 comments). Short comments without context ("No", "Vibes") are the hardest. Probabilities are calibrated — use `typeScores` and **Label threshold** to trade coverage for certainty. Automated labels can be wrong; check samples before making big decisions.

### Limits

| Item | Limit |
|---|---|
| Comment length | First 4,000 characters are used (post title: 300) |
| Custom labels | Up to 10, each up to 200 characters |
| Dataset size | Any — datasets are read in pages of 1,000 |
| Language | English (measured). Other languages are accepted and usually work, but accuracy has not been measured — treat them as **beta** |
| Replies | Each row is classified on its own. Replies nested inside a comment (e.g. an Instagram `replies` array) are not classified — use the scraper's option to output replies as rows |
| Speed | About 100 comments in 3–4 seconds |
| Data | Only the comment text, the post title and your custom labels are sent for classification; usernames and profile fields are not used or output |

### FAQ

**Which AI makes the judgments?** Jev, TypeSafe's decision model (version `jev-1.13.0`, pinned). Jev answers each label with a calibrated probability instead of generated text, so the same input gets the same answer from run to run. Only the comment text, the post title and your custom labels are sent to Jev; usernames and other fields are not.

**Does it scrape Instagram, TikTok, Facebook or YouTube?** No. It classifies comments you already have — run a comment scraper first (see **Works with**) or paste texts.

**Does it write replies?** No. It only assigns labels with probabilities — fast, cheap and consistent. Use **Needs reply** to decide which comments to answer.

**Why `request` as well as `complaint`?** Product requests ("please make a bigger bottle") are feedback for product teams, not problems for support. They are kept apart so each team gets its own list.

### Disclaimer

> Independent tool — not affiliated with, endorsed by or sponsored by Meta (Instagram, Facebook), TikTok, Google (YouTube), or by the authors of the scrapers listed above. Names are used only to describe compatible data sources. The comments shown are public comments used as examples.

### Reviews and support

If this Actor saved you time, a short review on Apify Store helps others find it. Questions or a dataset whose fields are not detected? Open an issue in the **Issues** tab — we answer within a day.

### Changelog

See the Changelog tab.

# Changelog

This Actor's version history is a separate document: https://apify.com/leoworks/social-comment-classifier/changelog.md

# Actor input Schema

## `datasetId` (type: `string`):

An Apify dataset containing comments — for example the output of an Instagram, TikTok, Facebook or YouTube comment scraper. Use this OR “Comment texts”.

## `texts` (type: `array`):

Paste comment texts directly (one per line). Use this OR “Comments dataset”.

## `textField` (type: `string`):

Field holding the comment text. Leave empty to auto-detect (text, comment, content, body, …).

## `contextField` (type: `string`):

Field holding the title or caption of the post the comment was written under (helps with short comments like “link?”). Leave empty to auto-detect (postTitle, title, …).

## `idFields` (type: `array`):

Fields copied unchanged from each input item to the output so you can join results back (e.g. id, cid, postUrl). Leave empty to auto-pick the comment ID and post URL.

## `customLabels` (type: `array`):

Up to 10 extra yes/no labels (each up to 200 characters) written in plain language, e.g. “mentions shipping to Canada”, “asks about a discount code”. Each gets a probability.

## `labelThreshold` (type: `number`):

Minimum probability (0–1) for the needs-reply flag and custom labels to be marked as matched. Raise it for fewer, surer matches.

## `outputMode` (type: `string`):

Minimal mode returns only label keys — smaller and easier to aggregate.

## `summaryGroupField` (type: `string`):

Field in your dataset that identifies the post (e.g. `postUrl`, `videoWebUrl`, `videoId`). The run then saves a REPORT record with, per post: comment type mix, sentiment shares, share needing a reply and example purchase-intent, question, complaint and request comments. Leave empty to detect it automatically (overall summary only if none is found). No extra charge.

## `maxItems` (type: `integer`):

Classify at most this many comments (0 = all).

## `maxConcurrency` (type: `integer`):

Comments classified in parallel.

## `healthCheck` (type: `boolean`):

Internal: fail the run when results look degraded (used by the developer's scheduled checks).

## Actor input object example

```json
{
  "texts": [
    "Where can I buy this in Canada? Need it!!",
    "Ordered 3 weeks ago and still nothing. Is anyone answering messages?"
  ],
  "idFields": [],
  "customLabels": [],
  "labelThreshold": 0.5,
  "outputMode": "full",
  "maxItems": 0,
  "maxConcurrency": 10,
  "healthCheck": false
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

## `report` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "texts": [
        "Where can I buy this in Canada? Need it!!",
        "Ordered 3 weeks ago and still nothing. Is anyone answering messages?"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("leoworks/social-comment-classifier").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "texts": [
        "Where can I buy this in Canada? Need it!!",
        "Ordered 3 weeks ago and still nothing. Is anyone answering messages?",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("leoworks/social-comment-classifier").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "texts": [
    "Where can I buy this in Canada? Need it!!",
    "Ordered 3 weeks ago and still nothing. Is anyone answering messages?"
  ]
}' |
apify call leoworks/social-comment-classifier --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,leoworks/social-comment-classifier"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/bRcr9i2Hfer5vY5rH/builds/zKu5AJTieXSsPH8cO/openapi.json
