Go to Apify Store
User picture

Data Mill

datamill

Structured data from sources most tools give up on. Built by someone who reads them.

ACTOR STATS

4 public Actors

8 total users

1 monthly user

75% runs succeeded

Data Mill

Messy pages in, clean JSON out.

I build scrapers for sources that are awkward on purpose: legacy encodings, undocumented internal APIs, search flows that only work in one language, sites that block anything running in a datacenter. Those are the ones worth doing, because they are the ones nobody else has bothered with.

  • 🧱 Reliability over cleverness — if a source blocks datacenter IPs, the Actor ships with the right proxy already switched on. You should not have to discover that yourself
  • Verified, not assumed — sold prices are confirmed sold, odds come from official endpoints, transcripts come from the API that still works
  • 🌍 English output, any source — Japanese, multilingual, whatever the origin. Clean English keys downstream
  • 🤖 Agent-ready — usable over API and MCP, so your agents can pull data on demand
  • 🛠️ Maintained — sites change. Open an issue and it usually gets fixed within a day
  • 📬 Need a source that doesn't exist yet? Open an issue on any Actor. Custom scrapers built on request

Actors

  • 🎬 YouTube Transcript Scraper — bulk transcripts from videos, playlists, channels and searches. Segments, plain text, SRT, VTT and RAG chunks
  • 💴 Mercari Sold Price Scraper — verified sold comps across Mercari, Yahoo Auctions and PayPay Flea Market
  • 🏨 Japan Hotel Scraper — Jalan rates, ratings and plans, including ryokan the global OTAs never list
  • 🏇 Japan Horse Racing Scraper — JRA race cards, live odds and results with full payouts

Public Actors