# Tumblr Scraper - Blog Posts, Tags, Notes & Media (`seemuapps/tumblr-scraper`) Actor

Scrape Tumblr blog posts and tagged posts with text, tags, note counts, reblog sources, image and video URLs plus blog info. No login needed.

- **URL**: https://apify.com/seemuapps/tumblr-scraper.md
- **Developed by:** [Seemu Scraping](https://apify.com/seemuapps) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.80 / 1,000 posts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Tumblr Scraper - Blog Posts, Tags, Notes & Media

Scrape posts from any public Tumblr blog, or pull the top and newest posts for any Tumblr tag. Every post comes with its full text, tags, note/like/reblog counts, the blog it was reblogged from, and direct image, video and audio URLs - plus the blog's title, description, avatar and total post count. No login and no Tumblr account needed.

Enter `nasa` and get the blog's posts newest-first. Enter the tag `photography`, pick **Top** or **Recent**, and get the most popular or the latest tagged posts - ready to export to JSON, CSV, Excel or Google Sheets.

### What you get

One row per post:

- **postId**, **url**, **shortUrl**, **postType** (`text`, `photo`, `video`, `audio`, `link`, `poll`), **createdAt**
- **summary**, **text** (the full post including any reblogged content) and **ownText** (only what this blog added)
- **tags**
- **noteCount**, **likeCount**, **reblogCount**, **replyCount**
- **isReblog**, **rebloggedFromBlog**, **rebloggedFromUrl**, **rebloggedRootBlog**, **rebloggedRootUrl** - where the post came from and who originally posted it
- **imageUrls**, **videoUrls**, **audioUrls**, **linkUrls** - full-size media links
- **isNsfw**
- Blog info on every row: **blogName**, **blogTitle**, **blogUrl**, **blogDescription**, **blogAvatarUrl**, **blogTotalPosts**, **blogUpdatedAt**, **blogIsAdult**
- **source** (`blog` or `tag`) and **sourceInput** - which blog or tag the row came from

### Use cases

- **Fandom and trend research** - see what's popular in a tag right now and which blogs drive it
- **Brand and creator monitoring** - archive a blog's posts with engagement counts for reporting
- **Influencer and artist discovery** - find active blogs posting in your niche, with their avatar and description
- **Content and media collection** - gather image and video URLs for a tag or blog for moodboards and datasets
- **Reblog and virality analysis** - trace posts back to their original author and compare note counts
- **Academic and social research** - build datasets of Tumblr posts on any topic

### How to use

1. Choose a **Mode**:
   - **Blog posts** - enter one or more **Blogs**: `nasa`, `nasa.tumblr.com`, `https://www.tumblr.com/nasa`, or a blog's custom domain.
   - **Tag search** - enter one or more **Tags** (with or without `#`) and pick **Tag sort**: **Top** for the most popular posts or **Recent** for the newest.
2. Set **Max posts per blog/tag** (default 100; `0` keeps going until there are no more posts).
3. In Blog posts mode, turn off **Include reblogs** to keep only a blog's original posts.
4. Run the actor - posts stream into the **Dataset** tab as they are found.

#### Fetching more on the next run

When you scrape a single blog or tag, the actor saves where it stopped. After the run finishes, open the **Key-value store** tab → copy the `NEXT_PAGE_ID` value → paste it into **Page ID** on your next run. If `NEXT_PAGE_ID` is `null` (or missing), you've fetched everything.

### Output format

```json
{
  "source": "blog",
  "sourceInput": "staff",
  "postId": "828009069026721792",
  "url": "https://staff.tumblr.com/post/828009069026721792/premium-just-got-better",
  "shortUrl": "https://tmblr.co/ZE5FbyjzhTn_8y00",
  "postType": "photo",
  "summary": "Premium just got better",
  "text": "Premium just got better\n\nYour feedback matters. Tumblr Premium now includes a few brand new, top-requested features...",
  "ownText": "Premium just got better\n\nYour feedback matters...",
  "tags": ["tumblr premium", "new features"],
  "noteCount": 3001,
  "likeCount": 2238,
  "reblogCount": 363,
  "replyCount": 400,
  "createdAt": "2026-09-17T13:16:22.000Z",
  "isReblog": false,
  "rebloggedFromBlog": null,
  "rebloggedFromUrl": null,
  "rebloggedRootBlog": null,
  "rebloggedRootUrl": null,
  "imageUrls": ["https://64.media.tumblr.com/8b352dd8.../s1280x1920/dbfb5aa1....pnj"],
  "videoUrls": [],
  "audioUrls": [],
  "linkUrls": [],
  "isNsfw": false,
  "blogName": "staff",
  "blogTitle": "Tumblr Staff",
  "blogUrl": "https://staff.tumblr.com/",
  "blogDescription": null,
  "blogAvatarUrl": "https://64.media.tumblr.com/dbc619ed.../s512x512u_c1/81e5a614....pnj",
  "blogTotalPosts": 2989,
  "blogUpdatedAt": "2026-09-17T13:16:22.000Z",
  "blogIsAdult": false
}
```

In Tag search mode `blogTotalPosts` is `null` - Tumblr only shares a blog's post count on its own page. Run Blog posts mode on any blog you find to get it.

### Pricing

You pay a small fee per post returned. No charge for blogs or tags that return nothing.

### Tips

- Tag search returns roughly 7-8 posts per page, so large **Recent** runs take a little longer than blog runs.
- Multi-word tags work as typed: `studio ghibli`, `artists on tumblr`.
- Use **Recent** on a schedule to monitor a tag; use **Top** for research.

### FAQ

**Do I need a Tumblr account?**
No. The actor only reads what Tumblr shows to logged-out visitors.

**What about private, password-protected or adult blogs and tags?**
Tumblr hides these from logged-out visitors. The actor logs a warning, skips them and carries on with the rest of your list.

**Are reblogs included?**
Yes by default, with the source and original blog attached. Turn off **Include reblogs** to keep only original posts.

# Actor input Schema

## `mode` (type: `string`):

Blog posts scrapes every post from the blogs you list. Tag search scrapes posts tagged with the tags you list.

## `blogs` (type: `array`):

Used in Blog posts mode. Blog names or URLs, one per line: 'nasa', 'nasa.tumblr.com', 'https://www.tumblr.com/nasa' or a custom domain.

## `tags` (type: `array`):

Used in Tag search mode. Tags to search, one per line, with or without '#', e.g. 'photography' or 'studio ghibli'.

## `tagSort` (type: `string`):

Used in Tag search mode. Top returns the most popular posts for the tag; Recent returns the newest posts first.

## `maxItems` (type: `integer`):

Maximum posts to return for each blog or tag. 0 = keep going until there are no more posts or the run times out.

## `includeReblogs` (type: `boolean`):

Keep reblogged posts in Blog posts mode. Turn off to return only the blog's original posts.

## `pageId` (type: `string`):

Paste NEXT\_PAGE\_ID from the previous run's Key-value store to fetch the next page. Works for runs with a single blog or tag.

## Actor input object example

```json
{
  "mode": "blog",
  "blogs": [
    "nasa"
  ],
  "tags": [
    "photography"
  ],
  "tagSort": "top",
  "maxItems": 50,
  "includeReblogs": true
}
```

# Actor output Schema

## `results` (type: `string`):

One post per record: source, sourceInput, postId, url, postType, summary, text, tags, noteCount, likeCount, reblogCount, createdAt, reblog source/root, imageUrls, videoUrls, audioUrls, linkUrls and blog info (name, title, description, avatar, total posts).

## `nextPageId` (type: `string`):

NEXT\_PAGE\_ID record in the default key-value store. Paste into Page ID on the next run to resume; null when the blog or tag is exhausted.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mode": "blog",
    "blogs": [
        "nasa"
    ],
    "tags": [
        "photography"
    ],
    "maxItems": 50
};

// Run the Actor and wait for it to finish
const run = await client.actor("seemuapps/tumblr-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "mode": "blog",
    "blogs": ["nasa"],
    "tags": ["photography"],
    "maxItems": 50,
}

# Run the Actor and wait for it to finish
run = client.actor("seemuapps/tumblr-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mode": "blog",
  "blogs": [
    "nasa"
  ],
  "tags": [
    "photography"
  ],
  "maxItems": 50
}' |
apify call seemuapps/tumblr-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,seemuapps/tumblr-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/GBBpd2tCgs3UNclmB/builds/vh6O9dC1kEZQz1Vq7/openapi.json
