# Facebook Hashtag Search Scraper (`memo23/facebook-hashtag-search-scraper`) Actor

Search Facebook posts by hashtag. Supports any language including Thai. No browser or login required.

- **URL**: https://apify.com/memo23/facebook-hashtag-search-scraper.md
- **Developed by:** [Muhamed Didovic](https://apify.com/memo23) (community)
- **Categories:** Social media
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 posts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Facebook Hashtag Search Scraper

#### How it works

![How Facebook Hashtag Search Scraper works](https://raw.githubusercontent.com/muhamed-didovic/muhamed-didovic.github.io/main/assets/how-it-works-facebook-hashtag.png)

**Find every public Facebook post that mentions a hashtag — across reels, photos, videos, and posts — without a browser, without a login.** Combines Facebook's own hashtag search with deep multi-engine SERP discovery (DuckDuckGo, Brave, Mojeek, Yandex), solves DuckDuckGo's anti-bot JavaScript challenge in pure Node.js, and delivers structured rows ready for analytics.

*Pure HTTP. No Puppeteer. No Playwright. $3 per 1,000 posts, and `maxItems` is a hard cap.*

***

### Why Use This Scraper?

- **Multi-engine SERP discovery** — runs Facebook's hashtag search AND fans out across DuckDuckGo, Brave, Mojeek, and Yandex in parallel.
- **Per-post-type fan-out** — `Any` automatically queries `site:facebook.com/photo`, `/videos`, and `/reel` separately for the deepest possible coverage.
- **DuckDuckGo bot-bypass** — extracts the `dp` token from the landing page and solves DDG's JavaScript anti-bot challenge in pure Node (linkedom + `vm.runInContext`). No Chrome required.
- **Hashtag-presence quality filter** — drops false positives where search engines surfaced loosely-matched URLs whose body doesn't actually contain the searched hashtag.
- **Auto-region detection** — Thai-script queries route to DDG's `kl=th-th` index; Japanese, Cyrillic, and Arabic auto-route to their matching regions.
- **Pure HTTP — no browser** — runs in seconds, not minutes; minimal RAM; cheap on Apify compute.
- **Pin-shape (post-shape) output** — every dataset row is a Facebook post object: message, author, engagement, attachments. Same shape regardless of which engine surfaced it.

***

### Overview

The **Facebook Hashtag Search Scraper** is a hashtag-first scraper for Facebook public posts. You provide a hashtag (or a `facebook.com/hashtag/{tag}` URL); the actor returns every public Facebook post — across reels, videos, photos, and standard posts — that contains that hashtag.

The dataset shape is **post rows**. Whether a result was discovered via Facebook's own hashtag GraphQL flow or via DuckDuckGo/Brave/Mojeek/Yandex site-search, every row has the same structure: post text, author info, engagement metrics, attachments, hashtags, and permalink URL.

This actor is designed for marketers monitoring sponsored brand hashtags (e.g. `#ทุเรียนยิ้มฟินเว่อ`, `#workout`), agencies verifying influencer deliverables, researchers building hashtag-conversation datasets, and analysts mapping competitive hashtag activity.

***

### Supported Inputs

Two ways to specify hashtags:

#### Plain hashtag strings

The simplest path. Provide one or more hashtags with or without `#`, in any language (Latin, Thai, Cyrillic, Arabic, CJK):

```json
{
  "hashtags": ["workout", "#travel", "ทุเรียนยิ้มฟินเว่อ"]
}
```

#### Hashtag URLs

Paste full Facebook hashtag URLs — the actor parses out the tag automatically:

```json
{
  "startUrls": [
    { "url": "https://www.facebook.com/hashtag/workout" },
    { "url": "https://www.facebook.com/hashtag/travel" }
  ]
}
```

You can mix both — duplicates are deduped.

#### Pattern rules

| Pattern | Supported? |
|---|---|
| `https://www.facebook.com/hashtag/{tag}` | ✅ yes |
| `https://m.facebook.com/hashtag/{tag}` | ✅ yes |
| Percent-encoded Unicode (`/hashtag/%E0%B8%97...`) | ✅ yes |
| `https://www.facebook.com/groups/...` | ❌ unsupported (use the Facebook Group Posts scraper) |
| `https://www.facebook.com/{user}/posts/...` | ❌ unsupported (single-post URLs aren't hashtag searches) |
| Hosts outside `*.facebook.com` | ❌ unsupported |

Non-hashtag URLs in `startUrls` are skipped with a warning rather than failing the run.

***

### Use Cases

| Audience | What they get |
|---|---|
| **Marketing teams** | Track every public post under your campaign hashtag — spot influencer reach and engagement-rich participants. |
| **Brand agencies** | Verify deliverables on hashtag-required sponsored posts (campaign tags like `#YourBrandChallenge`). |
| **SEO / content teams** | Map hashtag conversation depth, top author handles, and engagement-by-media-type. |
| **Researchers** | Build longitudinal datasets of public hashtag activity in any language including Thai, Arabic, Japanese. |
| **Competitive analysts** | Snapshot a competitor's branded-hashtag footprint across post types. |
| **Brand teams** | Monitor unsanctioned use of trademarked hashtags. |

***

### How It Works

1. **You provide hashtags** (or hashtag URLs) and a post type (`any`, `photos`, `videos`, or `reels`).
2. **Actor fans out across discovery sources** — Facebook's hashtag GraphQL flow + DuckDuckGo + Brave Search + Mojeek + Yandex, one query per concrete post type.
3. **Anti-bot bypass** — DuckDuckGo's JavaScript challenge is solved in pure Node.js using `linkedom` + `vm.runInContext` to compute the `jsa` token and resolve the next-page URL.
4. **Per-URL enrichment** — every discovered Facebook URL is fetched and parsed for full post data (message, author, reactions, comments, shares, media, hashtags).
5. **Quality filter** — posts whose page body doesn't actually contain the searched hashtag are dropped (false-positive control).
6. **Export** — structured JSON, CSV, or Excel.

***

### Input Configuration

The scraper requires a simple JSON input. Real fields only:

| Field | Type | Description | Default |
|-------|------|-------------|---------|
| `hashtags` | Array<string> | Hashtags to search (with or without `#`). | `[]` |
| `startUrls` | Array<{url}> | Facebook hashtag URLs (alternative or supplement to `hashtags`). | `[]` |
| `postType` | Enum | `any`, `photos`, `videos`, `reels`. `any` fans out to all three. | `"any"` |
| `timeFilter` | Enum | SERP date filter: `any`, `day`, `week`, `month`, `year`. | `"any"` |
| `maxItems` | Integer | Maximum number of posts in the dataset. | `50` |
| `minDelay` | Integer | Min seconds between paginated GraphQL requests. | `5` |
| `maxDelay` | Integer | Max seconds between paginated GraphQL requests (random jitter). | `10` |
| `proxy` | Object | Proxy configuration. **Highly recommended** to use Apify Residential to avoid blocks. | Apify Residential |

#### Example input — single niche hashtag

```json
{
  "hashtags": ["ทุเรียนยิ้มฟินเว่อ"],
  "postType": "any",
  "timeFilter": "any",
  "maxItems": 50
}
```

#### Example input — popular hashtag, reels only

```json
{
  "hashtags": ["workout"],
  "postType": "reels",
  "timeFilter": "any",
  "maxItems": 500
}
```

#### Example input — mixed strings + URLs

```json
{
  "hashtags": ["travel"],
  "startUrls": [
    { "url": "https://www.facebook.com/hashtag/food" }
  ],
  "postType": "any"
}
```

***

### Output Overview

Every dataset item is a **Facebook post row** — same shape regardless of which discovery source surfaced it (FB hashtag search, DDG, Brave, Mojeek, or Yandex). Posts whose body doesn't contain the searched hashtag are filtered out before export.

You can export the dataset as:

- **JSON** — full nested objects (attachments, engagement breakdown, etc.)
- **CSV / Excel** — flat rows; nested arrays/objects are stringified

Field availability varies slightly by source — Facebook GraphQL-discovered rows have richer engagement breakdowns; SERP-discovered rows always include core fields (id, url, message, author, hashtags) and engagement counts.

***

### Output Samples

*First row from a real `ทุเรียนยิ้มฟินเว่อ` run, shortened for readability:*

```json
{
  "id": "1286256033720319",
  "url": "https://www.facebook.com/lotussgofreshth/posts/pfbid09R4R3A9pb16334EcKmoxzbey1QXvEtB2vj5PBbF1a7bYKaXrR3xNu2hBU3Qn8pZYl",
  "message": "ทุเรียนหมอนทองโลตัส โก เฟรช อร่อย หวานมัน ไม่ต้องลุ้น ยิ้มเลย\nซื้อยกลูก ราคาดี สะดวกใกล้บ้าน\n🗓️โปรโมชั่นตั้งแต่วันที่ 7 พ.ค. 69 – 13 พ.ค. 69\n…\n#ทุเรียนโลตัส #Letsdurian #เทศกาลทุเรียน #เล็ทส์ดูเรียน #ทุเรียนยิ้มฟินเว่อ #ทุเรียน #โลตัสโกเฟรช #lotussGoFresh",
  "externalUrls": [
    "https://l.lotuss.com/Kteov",
    "https://l.lotuss.com/8r7xj"
  ],
  "hashtags": [
    "#ทุเรียนโลตัส", "#Letsdurian", "#เทศกาลทุเรียน",
    "#เล็ทส์ดูเรียน", "#ทุเรียนยิ้มฟินเว่อ", "#ทุเรียน",
    "#โลตัสโกเฟรช", "#lotussGoFresh"
  ],
  "authorName": "Lotus's go fresh",
  "authorId": "100070078029083",
  "authorUrl": "https://www.facebook.com/lotussgofreshth",
  "creationTime": 1778140860,
  "date": "2026-05-07T08:01:00.000Z",
  "elapsedMinutes": 2174,
  "elapsedSeconds": 130460,
  "privacyScope": "Unknown",
  "attachments": [
    {
      "type": "Photo",
      "url": "https://scontent-bcn1-1.xx.fbcdn.net/v/t39.30808-6/690601217_1285566873789235_2626861004532548182_n.jpg",
      "thumbnailUrl": "https://scontent-bcn1-1.xx.fbcdn.net/v/t39.30808-6/690601217_1285566873789235_2626861004532548182_n.jpg",
      "accessibility": "May be an image of text that says 'LET'S DURIAN ที่โลตัส โก เฟรช (สาขาเล็ก) อร่อย หวานมัน ไม่ต้องลุ้น ยิ้มเลย'",
      "mediaId": "1286256000386989"
    }
  ],
  "mediaCount": 5,
  "reactionCount": 601,
  "commentCount": 33,
  "shareCount": 23,
  "engagement": {
    "reactions": {
      "total": 601,
      "topReactions": [
        { "type": "Like", "count": 591 },
        { "type": "Love", "count": 5 },
        { "type": "Wow", "count": 4 },
        { "type": "Haha", "count": 1 }
      ]
    }
  },
  "locationId": null,
  "hashtag": "ทุเรียนยิ้มฟินเว่อ"
}
```

***

### Key Output Fields

Grouped by category:

#### Post Core

- `id` — Facebook post ID
- `url` — permalink URL
- `message` — post text
- `creationTime` — Unix timestamp (seconds)
- `date` — ISO 8601 string
- `elapsedMinutes` / `elapsedSeconds` — age at scrape time
- `privacyScope` — `"Public"`, `"Unknown"`, etc.
- `hashtag` — the hashtag the search was for
- `hashtags[]` — every hashtag found in the post body (Unicode-aware)
- `externalUrls[]` — external links in the post

#### Author

- `authorName` — display name
- `authorId` — Facebook numeric ID
- `authorUrl` — author profile URL

#### Engagement

- `reactionCount` / `commentCount` / `shareCount`
- `engagement.reactions.total`
- `engagement.reactions.topReactions[]` — array of `{ type, count }` (Like, Love, Haha, Wow, Sad, Angry)

#### Media (Attachments)

- `attachments[]` — each entry: `{ type, url, thumbnailUrl, accessibility, mediaId }`
- `mediaCount` — convenience count

#### Geo

- `locationId` — currently always `null` (not yet extracted; planned)

***

### Pricing

Pay per event, no subscription:

| Event | Price |
|---|---|
| Post (one row in the dataset) | $0.003 ($3 per 1,000) |
| Run start | $0.005 per GB of run memory |

`maxItems` is a hard cap: the run never pushes, or charges for, a post past it. If you set a spending limit on the run, the scraper stops at the number of posts that limit covers. Residential proxy is included and does not show up as a separate cost.

***

### FAQ

**How many posts can I expect per hashtag?**
Depends on the hashtag's footprint:

- **Popular generic hashtags** (`#workout`, `#travel`, `#food`): 100–300+ posts.
- **Niche brand campaigns** (e.g. `#ทุเรียนยิ้มฟินเว่อ`): 10–25 posts. This is a content limit, not a code limit — the scraper paginates to offset 430+ on DuckDuckGo and exhausts the index.

**Why do I sometimes get fewer results than I expect?**
Three reasons: (a) the hashtag genuinely doesn't have many indexed posts on the open web; (b) DuckDuckGo / Brave's index varies by proxy geography (request from a Thailand IP for Thai content for best coverage); (c) very recent posts may not be indexed yet.

**Does this scraper need a Facebook login?**
No. Everything runs against public endpoints — Facebook's logged-out hashtag GraphQL flow plus public web-search engines. No account, no login, no ban risk.

**Is this scraper fast?**
Pure HTTP, so no browser start-up. Most of a run is waiting on the search engines and on each post page: two hashtags with `maxItems: 20` took about three minutes on the platform. A niche tag with few posts finishes sooner.

**How does it bypass DuckDuckGo's anti-bot?**
DDG serves a JavaScript anti-bot challenge that requires DOM-aware execution (malformed-HTML `innerHTML.length` probes). The scraper solves it in pure Node.js using `linkedom` (lightweight DOM emulator) + `vm.runInContext` — no Chrome required.

**What's the difference between `Any` and selecting a specific post type?**
`Any` runs THREE separate SERP queries per source (`site:facebook.com/photo`, `/videos`, `/reel`) — much deeper coverage than a bare `site:facebook.com` query. Specific types narrow to that single type but go just as deep.

**Why does the scraper sometimes return slightly different counts on consecutive runs?**
Search engines vary their results per IP/region/time. Run-to-run variance of ±10–20% on the same hashtag is normal.

**Can I scrape private hashtags?**
There is no such thing — hashtags are public on Facebook by design. The scraper only accesses public posts that contain the hashtag.

**Can I mix multiple hashtags in one run?**
Yes — pass multiple hashtags in `hashtags[]` (or as `startUrls`). Each runs through the full discovery pipeline; `maxItems` is the global cap across all of them.

***

### Support

- For issues or feature requests, please use the [Issues](https://console.apify.com/actors/wNBB9C0B0zXZRLE0Y/issues) section of this actor.
- If you need customization or have questions, feel free to contact the author:
  - Author's website: <https://muhamed-didovic.github.io/>
  - Email: <muhamed.didovic@gmail.com>

### Additional Services

- Request customization or whole dataset: <muhamed.didovic@gmail.com>
- If you need anything else scraped, or this actor customized, email: <muhamed.didovic@gmail.com>
- For API services of this scraper (no Apify fee, just usage fee for the API), contact: <muhamed.didovic@gmail.com>

### Explore More Scrapers

If you found this Facebook Hashtag Search Scraper useful, be sure to check out our other powerful scrapers and actors at [memo23's Apify profile](https://apify.com/memo23). We offer a wide range of tools to enhance your web scraping and automation needs across various platforms and use cases.

***

### ⚠️ Disclaimer

This Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by Meta Platforms, Inc. or any of its subsidiaries. All trademarks (Facebook, Meta) mentioned are the property of their respective owners.

The scraper accesses only publicly available Facebook hashtag search pages and public post pages — no authenticated endpoints, paid features, or content behind the Facebook login wall. Discovery uses public web-search engines (DuckDuckGo, Brave, Mojeek, Yandex) querying the public web. Users are responsible for ensuring their use complies with Facebook's Terms of Service, applicable data-protection law (GDPR, CCPA, PDPA, etc.), and any contractual obligations of their own organization.

***

### SEO Keywords

facebook hashtag scraper, scrape facebook hashtag, facebook hashtag API, facebook hashtag search, Apify facebook hashtag, facebook reels scraper, facebook posts scraper, facebook brand campaign monitoring, hashtag tracking scraper, facebook influencer scraper, facebook engagement data, facebook brand mentions scraper, hashtag analytics, facebook content monitoring, facebook reels API, social media data extraction, brand campaign reach scraper, multilingual hashtag scraper, thai hashtag scraper, facebook search scraper

# Actor input Schema

## `hashtags` (type: `array`):

List of hashtags (with or without #) to search on Facebook, e.g. 'travel' or '#สวัสดี'. Supports any language. Provide either this or `startUrls` (or both).

## `startUrls` (type: `array`):

Facebook hashtag URLs (e.g. https://www.facebook.com/hashtag/travel). Merged with `hashtags` and deduped. Non-hashtag URLs are skipped with a warning.

## `postType` (type: `string`):

Narrows discovery to a specific Facebook post type. `Any` returns all post types. Specific types only return URLs of that type.

## `timeFilter` (type: `string`):

Date filter applied to web-search discovery. Use `startDate` if you also want to date-filter posts found via Facebook's own hashtag search.

## `maxItems` (type: `integer`):

Maximum number of posts to push to the dataset per run.

## `minDelay` (type: `integer`):

Minimum seconds to wait between paginated GraphQL requests. Lower values risk rate-limiting / blocked responses.

## `maxDelay` (type: `integer`):

Maximum seconds to wait between paginated GraphQL requests. Each delay is randomised between min and max.

## `proxy` (type: `object`):

Apify proxy settings used by the scraper.

## Actor input object example

```json
{
  "hashtags": [
    "travel"
  ],
  "startUrls": [],
  "postType": "any",
  "timeFilter": "any",
  "maxItems": 50,
  "minDelay": 5,
  "maxDelay": 10,
  "proxy": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "hashtags": [
        "travel"
    ],
    "startUrls": [],
    "proxy": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("memo23/facebook-hashtag-search-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "hashtags": ["travel"],
    "startUrls": [],
    "proxy": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("memo23/facebook-hashtag-search-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "hashtags": [
    "travel"
  ],
  "startUrls": [],
  "proxy": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call memo23/facebook-hashtag-search-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,memo23/facebook-hashtag-search-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/wNBB9C0B0zXZRLE0Y/builds/hbhauJiHuC0sbbgvy/openapi.json
