Wikidata Entity Search Scraper
Pricing
from $14.00 / 1,000 result items
Wikidata Entity Search Scraper
Search Wikidata's open knowledge graph of 100M+ entities (people, places, brands, books, films) by name. Returns Q-ID, label, description, aliases, all claims (P-properties), sitelinks to every Wikipedia language, structured facts and image. Filter by entity type, language and full-claims fetching.
Pricing
from $14.00 / 1,000 result items
Rating
0.0
(0)
Developer
ParseForge
Maintained by CommunityActor stats
0
Bookmarked
4
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share

๐ Wikidata Entity Search Scraper
๐ Search Wikidata's open knowledge graph of 100M+ entities by name.
The Wikidata Entity Search Scraper searches Wikidata's open knowledge graph of 100M+ entities by name. Output includes the canonical Q-ID, label, description, aliases, all claims (P-properties), sitelinks to every Wikipedia language edition, and structured facts.
Wikidata is the structured-data backbone of Wikipedia and one of the largest open knowledge graphs in the world. Filters run server-side, so a single run can resolve every entity matching a name, fetch full claim trees, or pull entities in non-English languages.
| ๐ฏ Target Audience | ๐ก Primary Use Cases |
|---|---|
| ML pipelines, knowledge-graph engineers, journalists, fact-checkers, content recommendation engines, search developers | Entity resolution, knowledge-graph augmentation, fact-checking, content enrichment, multilingual search, ML training datasets |
๐ What the Wikidata Entity Search Scraper does
Five filtering workflows in a single run:
- ๐ Free-text search. Match entity labels and aliases.
- ๐ Multilingual. Search in 20+ languages (en, es, fr, de, it, ja, zh, ko, ar, hi, pt, nl, ru).
- ๐ Item or property. Search Q-entities (items) or P-entities (properties).
- ๐ Full claims fetch. Optional: pull every statement, sitelink, and structured fact per entity.
- ๐ท๏ธ Image extraction. Auto-extracts the entity's primary image from claim P18.
๐ก Why it matters: clean, server-side filtering removes the parser-and-pagination work from your team and keeps your dataset fresh on every run.
๐ Data fields
Each record includes: aliasesText, claimCount, description, entityId, instanceOf, label, sitelinkCount, thumbnailUrl, wikidataUrl, wikipediaEnUrl. These field names come straight from the actor's dataset schema, so what you see here is what lands in your dataset.
๐ How to use
- ๐ Sign up. Create a free account with $5 credit (takes 2 minutes).
- ๐ Open the Actor. Go to the Wikidata Entity Search Scraper page on the Apify Store.
- ๐ฏ Set input. Pick your filters and
maxItems. - ๐ Run it. Click Start and let the Actor collect your data.
- ๐ฅ Download. Grab your results in the Dataset tab as CSV, Excel, JSON, or XML.
โฑ๏ธ Total time from signup to downloaded dataset: 3-5 minutes. No coding required.
๐ Recommended Actors
- ๐ Open Library Books - 30M+ books and editions
- ๐ Project Gutenberg Books - 75,000+ free public-domain books
- ๐จ Openverse Media - 800M+ openly licensed images and audio
- ๐ฐ Hacker News Search - Every HN story since 2007
- ๐ World Bank Indicators - Country economic indicators
๐ก Pro Tip: browse the complete ParseForge collection for more reference-data scrapers.
โ ๏ธ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by the Wikimedia Foundation, Wikidata, Wikipedia, or any contributing editor. All trademarks mentioned are the property of their respective owners. Only publicly available open data is collected.
๐ Need Help?
If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.
For faster answers, join our Discord. It's the best place to get support and suggest new actors.