# Tumblr Blog & Tag Scraper (`s-r/tumblr-blogs`) Actor

Read any Tumblr blog or tag page: post text, tags, posting date, media, and all four engagement counts (notes, likes, reblogs, replies). Tag pages work as a public search across the whole site. No API key.

- **URL**: https://apify.com/s-r/tumblr-blogs.md
- **Developed by:** [SR](https://apify.com/s-r) (community)
- **Categories:** Social media, News
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 actor run starteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Tumblr Blog & Tag Scraper

Read any Tumblr blog or tag page as structured data: the post text, its tags,
when it was posted, what media it carries, and all four engagement counts.

Give it blog names or tags. No API key, no login.

### Four counts, not one

Tumblr publishes four numbers per post and most tools report one of them:

| | example post |
|---|---|
| `note_count` | 11,952 |
| `like_count` | 9,106 |
| `reblog_count` | 2,037 |
| `reply_count` | 809 |

Notes is Tumblr's own rollup of everything that happened to a post. It is **not**
a reliable sum of the other three, and it is never computed here: the published
value is used.

The distinction that matters on Tumblr is between likes and **reblogs**. A like
means someone saw it. A reblog means someone put it in front of their own
followers, which is how anything travels on this site. A post with 9,000 likes
and 200 reblogs and one with 3,000 likes and 6,000 reblogs are very different
posts, and a single "engagement" number hides which is which.

### Tag pages are the search

Tumblr offers no keyword search to a signed-out visitor. What it does offer is
`/tagged/<tag>`, which is public and returns current posts carrying that tag
**from across the whole site**, not from one blog.

Pass `#photography` or a `/tagged/` link and that is what you get: one run of a
blog and a tag returned 30 posts from **11 different blogs**. Every row records
`found_on` and `source`, so blog posts and tag results stay distinguishable in
one table.

### The address that does not work, and what happens to it

A Tumblr blog has two addresses. `www.tumblr.com/staff` serves the posts.
`staff.tumblr.com` — the one in every link Tumblr itself writes — answers
**403**. So does the documented public API at `api.tumblr.com/v2`, without a key.

Paste either form. A subdomain link is **rewritten** to the path that works
rather than followed, so the link you copied from a post does the right thing
instead of failing.

### Post text is assembled, not read

A Tumblr post is a list of typed blocks: text, image, link, video. The words
have to be gathered from the text blocks, and a reblog's trail carries the
original post's text alongside the new comment. Both are included.

**A post that is only an image has no text**, and that is a fact about the post
rather than a failure to read it. `summary` — Tumblr's own one-line description
— is usually present even then, which is why both fields exist.

### What you get per post

- `text`, `summary`, `tags`
- `note_count`, `like_count`, `reblog_count`, `reply_count`
- `posted_at` as a UTC timestamp
- `media_type` and `media_url` for the first non-text block
- `is_reblog` and `parent_post_url` — reblogs are marked, with what they
  reblogged
- `blog_name`, `blog_title`, `blog_url`, `url`, `short_url`, `post_id`
- `found_on` and `source`

### Scale

A blog page carries about **20 posts**, a tag page about **10**. **Maximum
posts** is a ceiling across everything in the run, so a blog and a tag at 40
gives you 20 and 10 rather than a share of each.

Two targets and 30 posts took **fifteen seconds**. Fewer client identities are
served here than on most sites, so a page is sometimes requested more than once;
`requestsRetried` in the summary shows how often, and a rising figure across
scheduled runs is the early warning.

### What this does not cover

**No follower counts.** Tumblr does not publish them on these pages, so there is
no follower column rather than an empty one.

**No notes breakdown.** The note count is a total; who liked and who reblogged
is on a separate page per post and is not returned here.

**No dashboard, likes or following.** Those need an account and are not public.

**No blog search by name.** Tags are searchable, blogs are not: you name the
blogs you want.

### What people use this for

**Fandom and trend research.** Tag pages are where a topic is actually visible
on Tumblr, and the reblog counts show what is spreading rather than what is
merely being seen.

**Brand monitoring.** Watch a tag on a schedule and keep the rows. New posts
from blogs you have not seen before are the signal.

**Creator research.** Reblogs against likes across a blog's recent posts is a
better read on reach than a follower count, which Tumblr does not publish here
anyway.

**Content archiving.** Post text, tags, media URL and dates in one table, with
reblogs marked so originals can be separated from redistribution.

### Notes

Only public blogs are readable. A blog that is hidden, deleted or explicit-only
is reported by name rather than returned as an empty result.

Counts move, so a run is a snapshot; the posting date is on every row.

Media is identified rather than downloaded: `media_url` points at Tumblr's CDN.

Follower counts are not published on these pages and are therefore not returned.

# Actor input Schema

## `targets` (type: `array`):

Blog names (staff), blog links, or tags as #photography or a /tagged/ link. A link to a blog's own subdomain is rewritten to the path that works, because the subdomain itself refuses requests.

## `max_posts` (type: `integer`):

Stop after this many posts across all blogs and tags. A page carries about 20, a tag page about 10, so this is also the cost ceiling.

## `max_targets` (type: `integer`):

Stop after this many, counted after duplicates are removed.

## `attempts` (type: `integer`):

How often to retry a page that comes back without posts. Fewer identities are served here than on most sites, so six is the sensible default.

## `region` (type: `string`):

Two-letter country code the request should appear to come from.

## Actor input object example

```json
{
  "targets": [
    "https://www.tumblr.com/tagged/art"
  ],
  "max_posts": 50,
  "max_targets": 20,
  "attempts": 6,
  "region": "gb"
}
```

# Actor output Schema

## `posts` (type: `string`):

One row per post.

## `summary` (type: `string`):

Blogs and tags read, posts per target, how many were reblogs, and total notes and reblogs.

## `errors` (type: `string`):

Blogs or tags that could not be read.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "targets": [
        "staff",
        "#photography"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("s-r/tumblr-blogs").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "targets": [
        "staff",
        "#photography",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("s-r/tumblr-blogs").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "targets": [
    "staff",
    "#photography"
  ]
}' |
apify call s-r/tumblr-blogs --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,s-r/tumblr-blogs"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/IX26biclb0gxc229d/builds/qFdr3ya1ceWd9s9fW/openapi.json
