Tumblr Scraper: Blogs, Tags, Search & Posts
Pricing
from $1.25 / 1,000 tumblr posts
Tumblr Scraper: Blogs, Tags, Search & Posts
Export public Tumblr posts and blog profiles from blogs, tags, search, and post URLs. Includes media metadata, reblog context, and optional repeat-run change monitoring. No Tumblr login or API key.
Pricing
from $1.25 / 1,000 tumblr posts
Rating
0.0
(0)
Developer
Inus Grobler
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
10 hours ago
Last modified
Categories
Share
Collect public Tumblr posts from blogs, individual post URLs, blog-specific tag pages, global tags, and search results. Get text, dates, tags, note counts, media URLs, links, and reblog context in an exportable dataset. Optional monitoring compares posts with the last complete run for the same source set.
No Tumblr account, Tumblr API key, or login cookies are required. Global tag and search discovery uses the public Tumblr site in a browser; blog archives and post pages use public HTML. Blogs that redirect to a custom domain are discovered through their public www.tumblr.com page when available. The scraper does not send direct requests to Tumblr's API.
Use cases
- Brand and topic monitoring: find public posts from Tumblr search and tags, then compare repeat runs for new or changed posts.
- Creator and blog research: export a public blog's recent posts, tags, media links, and available note counts.
- Content research: collect public posts from several blogs, tags, and searches in one dataset for analysis or reporting.
- Post lookup: extract one public post URL and its available text, media metadata, links, and reblog context.
Quick start in Apify Console
- Open the Input tab and enter a blog handle, public Tumblr URL,
#tag, orsearch:phrasein Tumblr sources. - Set Maximum post rows. Start with 10 posts to check that the source is public and available.
- Click Start. Open the Output tab to view or export the dataset. Check Run summary if a source returns fewer posts than expected.
The Console example requests up to 10 posts across the public Tumblr Staff blog, #photography, and search:independent artists, demonstrating all three source types. The Actor works with no Tumblr credentials. An Apify token is needed only if you choose to run it programmatically through the Apify API.
Input: choose Tumblr sources
Enter any mix of public Tumblr URLs, blog handles, #tags, and search:phrases in Tumblr sources. Duplicate sources are removed.
Supported examples:
https://staff.tumblr.com/— public blog and older archive pageshttps://staff.tumblr.com/tagged/tumblr%20premium— posts under one blog's taghttps://staff.tumblr.com/post/828009069026721792— one public posthttps://www.tumblr.com/tagged/photography— global tag discoveryhttps://www.tumblr.com/search/photography— global search discoverystaffor@staff— blog handle#photography— global tagsearch:independent artists— search phrase
Raise Maximum post rows for larger exports; the Actor adjusts blog pages and discovery scrolls to that limit. With multiple sources, it shares the result budget across sources so the first source does not consume the whole run. Monitoring mode defaults to an independent snapshot; choose compare for repeat-run changes. Enable Apify Proxy if direct requests are rate-limited.
{"startUrls": ["staff", "#photography", "search:independent artists"],"maxItems": 50}
Output dataset
The default dataset contains one row per unique post and a blog profile row for each explicitly selected blog or blog tag. Post rows include:
postId,postUrl,blogName,blogUrl,publishedAttitle,text, availablehtml,postType,tagsnoteCount,mediametadata, outboundlinks,isReblog,rootPostUrlsourceUrls,contentCompleteness,scrapedAtchangeTypeandnoteDeltawhen monitoring is enabled
{"recordType": "post","postId": "828009069026721792","postUrl": "https://staff.tumblr.com/post/828009069026721792","blogName": "staff","publishedAt": "2026-09-17T09:16:22.000Z","title": "Premium just got better","text": "Your feedback matters. Tumblr Premium now includes...","postType": "text","tags": ["tumblr premium", "new features"],"noteCount": 2948,"media": [{"type": "image", "url": "https://64.media.tumblr.com/..."}],"contentCompleteness": "full"}
RUN_SUMMARY in the run's key-value store reports counts, per-source status, and any access or depth limit. Overall status is bounded when an item, page, scroll, source share, or spending limit was reached; partial indicates a source or extraction gap. A source's publicMirrorUrl identifies fallback discovery through Tumblr's public page for a custom-domain blog. A result marked preview came from a public discovery card because the full post page was unavailable. Posts are saved as they are found, so a timed-out run can still leave partial results in the dataset. A browser scroll limit does not mean all historical global tag or search posts were collected.
Pricing
Pay $1.25 per 1,000 saved post rows ($0.00125 each), plus $0.00005 per Actor start per GB of memory. The default 2 GB setting charges two start events ($0.00010 total). Blog profile rows are free, and platform usage is included. The 10-post Console example costs $0.01260 if all 10 posts are saved, or $0.01255 if you select 1 GB for a blog-only run. Use Maximum cost per run to cap spending; Apify may stop the run when the cap is reached, and posts already saved remain in the dataset. Allow room above the expected post charges if you want the run to finish normally.
Compare repeat runs
Set Monitoring mode to compare. The Actor automatically keeps a separate baseline for each source set. The first successful run marks observed posts new. Later runs mark them new, changed, or unchanged; noteDelta reports engagement change separately. The saved baseline contains compact post hashes and note counts, not post bodies. It advances after complete or bounded runs, but not after source failures, parse gaps, or spending limits. The Actor does not create schedules.
Run through the Apify API
from apify_client import ApifyClientclient = ApifyClient("YOUR_APIFY_TOKEN")run = client.actor("thescrapelab/tumblr-blog-tag-search-scraper").call(run_input={"startUrls": ["staff"],"maxItems": 10,})for row in client.dataset(run["defaultDatasetId"]).iterate_items():if row["recordType"] == "post":print(row["postUrl"], row.get("noteCount"))
Public access limits
Tumblr blogs hidden from visitors without an account, login-only posts, and blocked pages cannot be collected. This Actor does not download media files or enumerate individual likes, reblogs, or replies. HTML themes and Tumblr's discovery UI can change; per-source status and content-completeness fields make those limits visible. Use modest limits for recurring monitoring, and respect the content owner's rights when reusing text or images.
Troubleshooting and support
- No posts or fewer than requested: check that the blog or post is public, then open Run summary for the source's status and stop reason. Tag and search pages may expose only a limited window of results.
- Rate-limited or blocked: enable Apify Proxy, lower the post limit, and allow time between repeated runs. Tumblr can still restrict access.
- Preview instead of full content:
contentCompleteness: "preview"means the public discovery card was available but the full post page was not. - Run marked ABORTED at its cost cap: saved rows remain available. Increase Maximum cost per run or request fewer posts for the next run.
- Monitoring did not advance: check Run summary for a source failure, extraction gap, or spending limit. The baseline advances only after a complete or bounded run.
For help with a specific run, share its Apify run ID and the source status shown in Run summary through the Actor's Issues tab. Do not include private credentials in an issue.