# Tumblr Scraper: Blogs, Tags, Search & Posts (`thescrapelab/tumblr-blog-tag-search-scraper`) Actor

Export public Tumblr posts and blog profiles from blogs, tags, search, and post URLs. Includes media metadata, reblog context, and optional repeat-run change monitoring. No Tumblr login or API key.

- **URL**: https://apify.com/thescrapelab/tumblr-blog-tag-search-scraper.md
- **Developed by:** [Inus Grobler](https://apify.com/thescrapelab) (community)
- **Categories:** Social media
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.25 / 1,000 tumblr posts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Tumblr Scraper: Blogs, Tags, Search & Posts

Collect public Tumblr posts from blogs, individual post URLs, blog-specific tag pages, global tags, and search results. Get text, dates, tags, note counts, media URLs, links, and reblog context in an exportable dataset. Optional monitoring compares posts with the last complete run for the same source set.

No Tumblr account, Tumblr API key, or login cookies are required. Global tag and search discovery uses the public Tumblr site in a browser; blog archives and post pages use public HTML. Blogs that redirect to a custom domain are discovered through their public `www.tumblr.com` page when available. The scraper does not send direct requests to Tumblr's API.

### Use cases

- **Brand and topic monitoring:** find public posts from Tumblr search and tags, then compare repeat runs for new or changed posts.
- **Creator and blog research:** export a public blog's recent posts, tags, media links, and available note counts.
- **Content research:** collect public posts from several blogs, tags, and searches in one dataset for analysis or reporting.
- **Post lookup:** extract one public post URL and its available text, media metadata, links, and reblog context.

### Quick start in Apify Console

1. Open the **Input** tab and enter a blog handle, public Tumblr URL, `#tag`, or `search:phrase` in **Tumblr sources**.
2. Set **Maximum post rows**. Start with 10 posts to check that the source is public and available.
3. Click **Start**. Open the **Output** tab to view or export the dataset. Check **Run summary** if a source returns fewer posts than expected.

The Console example requests up to 10 posts across the public Tumblr Staff blog, `#photography`, and `search:independent artists`, demonstrating all three source types. The Actor works with no Tumblr credentials. An Apify token is needed only if you choose to run it programmatically through the Apify API.

### Input: choose Tumblr sources

Enter any mix of public Tumblr URLs, blog handles, `#tags`, and `search:phrases` in **Tumblr sources**. Duplicate sources are removed.

Supported examples:

- `https://staff.tumblr.com/` — public blog and older archive pages
- `https://staff.tumblr.com/tagged/tumblr%20premium` — posts under one blog's tag
- `https://staff.tumblr.com/post/828009069026721792` — one public post
- `https://www.tumblr.com/tagged/photography` — global tag discovery
- `https://www.tumblr.com/search/photography` — global search discovery
- `staff` or `@staff` — blog handle
- `#photography` — global tag
- `search:independent artists` — search phrase

Raise **Maximum post rows** for larger exports; the Actor adjusts blog pages and discovery scrolls to that limit. With multiple sources, it shares the result budget across sources so the first source does not consume the whole run. **Monitoring mode** defaults to an independent snapshot; choose `compare` for repeat-run changes. Enable **Apify Proxy** if direct requests are rate-limited.

```json
{
  "startUrls": ["staff", "#photography", "search:independent artists"],
  "maxItems": 50
}
```

### Output dataset

The default dataset contains one row per unique post and a blog profile row for each explicitly selected blog or blog tag. Post rows include:

- `postId`, `postUrl`, `blogName`, `blogUrl`, `publishedAt`
- `title`, `text`, available `html`, `postType`, `tags`
- `noteCount`, `media` metadata, outbound `links`, `isReblog`, `rootPostUrl`
- `sourceUrls`, `contentCompleteness`, `scrapedAt`
- `changeType` and `noteDelta` when monitoring is enabled

```json
{
  "recordType": "post",
  "postId": "828009069026721792",
  "postUrl": "https://staff.tumblr.com/post/828009069026721792",
  "blogName": "staff",
  "publishedAt": "2026-09-17T09:16:22.000Z",
  "title": "Premium just got better",
  "text": "Your feedback matters. Tumblr Premium now includes...",
  "postType": "text",
  "tags": ["tumblr premium", "new features"],
  "noteCount": 2948,
  "media": [{"type": "image", "url": "https://64.media.tumblr.com/..."}],
  "contentCompleteness": "full"
}
```

`RUN_SUMMARY` in the run's key-value store reports counts, per-source status, and any access or depth limit. Overall status is `bounded` when an item, page, scroll, source share, or spending limit was reached; `partial` indicates a source or extraction gap. A source's `publicMirrorUrl` identifies fallback discovery through Tumblr's public page for a custom-domain blog. A result marked `preview` came from a public discovery card because the full post page was unavailable. Posts are saved as they are found, so a timed-out run can still leave partial results in the dataset. A browser scroll limit does not mean all historical global tag or search posts were collected.

### Pricing

Pay $1.25 per 1,000 saved post rows ($0.00125 each), plus $0.00005 per Actor start per GB of memory. The default 2 GB setting charges two start events ($0.00010 total). Blog profile rows are free, and platform usage is included. The 10-post Console example costs $0.01260 if all 10 posts are saved, or $0.01255 if you select 1 GB for a blog-only run. Use **Maximum cost per run** to cap spending; Apify may stop the run when the cap is reached, and posts already saved remain in the dataset. Allow room above the expected post charges if you want the run to finish normally.

### Compare repeat runs

Set **Monitoring mode** to `compare`. The Actor automatically keeps a separate baseline for each source set. The first successful run marks observed posts `new`. Later runs mark them `new`, `changed`, or `unchanged`; `noteDelta` reports engagement change separately. The saved baseline contains compact post hashes and note counts, not post bodies. It advances after complete or bounded runs, but not after source failures, parse gaps, or spending limits. The Actor does not create schedules.

### Run through the Apify API

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("thescrapelab/tumblr-blog-tag-search-scraper").call(run_input={
    "startUrls": ["staff"],
    "maxItems": 10,
})
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
    if row["recordType"] == "post":
        print(row["postUrl"], row.get("noteCount"))
```

### Public access limits

Tumblr blogs hidden from visitors without an account, login-only posts, and blocked pages cannot be collected. This Actor does not download media files or enumerate individual likes, reblogs, or replies. HTML themes and Tumblr's discovery UI can change; per-source status and content-completeness fields make those limits visible. Use modest limits for recurring monitoring, and respect the content owner's rights when reusing text or images.

### Troubleshooting and support

- **No posts or fewer than requested:** check that the blog or post is public, then open **Run summary** for the source's status and stop reason. Tag and search pages may expose only a limited window of results.
- **Rate-limited or blocked:** enable **Apify Proxy**, lower the post limit, and allow time between repeated runs. Tumblr can still restrict access.
- **Preview instead of full content:** `contentCompleteness: "preview"` means the public discovery card was available but the full post page was not.
- **Run marked ABORTED at its cost cap:** saved rows remain available. Increase **Maximum cost per run** or request fewer posts for the next run.
- **Monitoring did not advance:** check **Run summary** for a source failure, extraction gap, or spending limit. The baseline advances only after a complete or bounded run.

For help with a specific run, share its Apify run ID and the source status shown in **Run summary** through the Actor's **Issues** tab. Do not include private credentials in an issue.

# Actor input Schema

## `startUrls` (type: `array`):

Enter a public Tumblr URL, blog handle (staff or @staff), #tag, or search:phrase. Mix sources freely; duplicates are removed.

## `maxItems` (type: `integer`):

Global cap on unique post rows, excluding free blog profile rows. Tag and search discovery uses the default 2 GB memory setting.

## `monitoringMode` (type: `string`):

Snapshot each run, or compare posts with the previous complete run for the same sources. The baseline is managed automatically.

## `proxyConfiguration` (type: `object`):

Optional. Enable Apify Proxy if Tumblr rate-limits or blocks direct cloud requests.

## Actor input object example

```json
{
  "startUrls": [
    "staff"
  ],
  "maxItems": 10,
  "monitoringMode": "snapshot",
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "staff"
    ],
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("thescrapelab/tumblr-blog-tag-search-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": ["staff"],
    "maxItems": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("thescrapelab/tumblr-blog-tag-search-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "staff"
  ],
  "maxItems": 10
}' |
apify call thescrapelab/tumblr-blog-tag-search-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,thescrapelab/tumblr-blog-tag-search-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/cH22fIGYuhtTDgeuN/builds/moXJPygwLb2eenbGF/openapi.json
