Tumblr Blog & Tag Scraper avatar

Tumblr Blog & Tag Scraper

Pricing

from $2.00 / 1,000 actor run starteds

Go to Apify Store
Tumblr Blog & Tag Scraper

Tumblr Blog & Tag Scraper

Read any Tumblr blog or tag page: post text, tags, posting date, media, and all four engagement counts (notes, likes, reblogs, replies). Tag pages work as a public search across the whole site. No API key.

Pricing

from $2.00 / 1,000 actor run starteds

Rating

0.0

(0)

Developer

SR

SR

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

Read any Tumblr blog or tag page as structured data: the post text, its tags, when it was posted, what media it carries, and all four engagement counts.

Give it blog names or tags. No API key, no login.

Four counts, not one

Tumblr publishes four numbers per post and most tools report one of them:

example post
note_count11,952
like_count9,106
reblog_count2,037
reply_count809

Notes is Tumblr's own rollup of everything that happened to a post. It is not a reliable sum of the other three, and it is never computed here: the published value is used.

The distinction that matters on Tumblr is between likes and reblogs. A like means someone saw it. A reblog means someone put it in front of their own followers, which is how anything travels on this site. A post with 9,000 likes and 200 reblogs and one with 3,000 likes and 6,000 reblogs are very different posts, and a single "engagement" number hides which is which.

Tumblr offers no keyword search to a signed-out visitor. What it does offer is /tagged/<tag>, which is public and returns current posts carrying that tag from across the whole site, not from one blog.

Pass #photography or a /tagged/ link and that is what you get: one run of a blog and a tag returned 30 posts from 11 different blogs. Every row records found_on and source, so blog posts and tag results stay distinguishable in one table.

The address that does not work, and what happens to it

A Tumblr blog has two addresses. www.tumblr.com/staff serves the posts. staff.tumblr.com — the one in every link Tumblr itself writes — answers 403. So does the documented public API at api.tumblr.com/v2, without a key.

Paste either form. A subdomain link is rewritten to the path that works rather than followed, so the link you copied from a post does the right thing instead of failing.

Post text is assembled, not read

A Tumblr post is a list of typed blocks: text, image, link, video. The words have to be gathered from the text blocks, and a reblog's trail carries the original post's text alongside the new comment. Both are included.

A post that is only an image has no text, and that is a fact about the post rather than a failure to read it. summary — Tumblr's own one-line description — is usually present even then, which is why both fields exist.

What you get per post

  • text, summary, tags
  • note_count, like_count, reblog_count, reply_count
  • posted_at as a UTC timestamp
  • media_type and media_url for the first non-text block
  • is_reblog and parent_post_url — reblogs are marked, with what they reblogged
  • blog_name, blog_title, blog_url, url, short_url, post_id
  • found_on and source

Scale

A blog page carries about 20 posts, a tag page about 10. Maximum posts is a ceiling across everything in the run, so a blog and a tag at 40 gives you 20 and 10 rather than a share of each.

Two targets and 30 posts took fifteen seconds. Fewer client identities are served here than on most sites, so a page is sometimes requested more than once; requestsRetried in the summary shows how often, and a rising figure across scheduled runs is the early warning.

What this does not cover

No follower counts. Tumblr does not publish them on these pages, so there is no follower column rather than an empty one.

No notes breakdown. The note count is a total; who liked and who reblogged is on a separate page per post and is not returned here.

No dashboard, likes or following. Those need an account and are not public.

No blog search by name. Tags are searchable, blogs are not: you name the blogs you want.

What people use this for

Fandom and trend research. Tag pages are where a topic is actually visible on Tumblr, and the reblog counts show what is spreading rather than what is merely being seen.

Brand monitoring. Watch a tag on a schedule and keep the rows. New posts from blogs you have not seen before are the signal.

Creator research. Reblogs against likes across a blog's recent posts is a better read on reach than a follower count, which Tumblr does not publish here anyway.

Content archiving. Post text, tags, media URL and dates in one table, with reblogs marked so originals can be separated from redistribution.

Notes

Only public blogs are readable. A blog that is hidden, deleted or explicit-only is reported by name rather than returned as an empty result.

Counts move, so a run is a snapshot; the posting date is on every row.

Media is identified rather than downloaded: media_url points at Tumblr's CDN.

Follower counts are not published on these pages and are therefore not returned.