Wikidata Entities Scraper avatar

Wikidata Entities Scraper

Pricing

from $0.84 / 1,000 results

Go to Apify Store
Wikidata Entities Scraper

Wikidata Entities Scraper

Full Wikidata entity data by QID: labels, descriptions and aliases in every published language, every claim (raw and simplified), and every cross-wiki sitelink. No API key. Detects silent item merges.

Pricing

from $0.84 / 1,000 results

Rating

0.0

(0)

Developer

Ibnu Adzim

Ibnu Adzim

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

7 days ago

Last modified

Share

Wikidata Entities Scraper (Labels, Claims & Sitelinks)

Full Wikidata entity data by QID: labels, descriptions and aliases in every published language, every claim the item carries (raw and simplified), and every cross-wiki sitelink.

No login. No API key. HTTP only — and deliberately narrow: there is no search mode. Wikidata's SPARQL endpoint and its label-search API are both robots.txt-disallowed (see below), so resolving ids you already have is the only in-policy operation. It pairs naturally with ../wikipedia-articles-scraper, which already returns the wikibase_item QID for an article but not its full entity data.

Record typeOne perCarries
ENTITYresolved idrequested/actual id, merge flag, labels/descriptions/aliases (all languages), sitelinks, claims (raw + simplified), revision info
ERRORfailed input_error code and an _errorDetail saying what to change

Input

Just entityIds — a bare id (Q42) or a full wikidata.org URL ending in one. preferredLanguage (default en) picks out labelPreferred / descriptionPreferred from the full multi-language maps every record already carries in full.

Things this endpoint will mislead you about

Each is measured, and each has a scenario in tests/smoke/wikidata-entities-scraper_traps.sh (8/8 passing).

Wikidata SPARQL — flagged in this repo's own research notes across many past sessions as "the strongest candidate left to build" — turns out to be closed, and always was. query.wikidata.org/robots.txt carries a blanket Disallow: /sparql for User-agent: *. www.wikidata.org/robots.txt separately disallows /w/ wholesale, which closes both the label-search API (wbsearchentities) and the newer Wikibase REST API. The only surface left Allow:ed is /wiki/Special:EntityData/<id>.<format> — one static file per entity, keyed on an id you already have. This Actor touches nothing else.

A merged item answers 200 and silently hands you a different entity. Wikidata periodically merges duplicate items. Q9270598 (merged) still resolves — but the response's entities object is keyed on Q13247166, the id it was merged INTO, with nothing at the top level marking a merge happened. wasMerged and entityId vs requestedId are how this Actor surfaces it — compare them yourself if you build on top of the raw JSON.

A missing label is not rare, and it is not a sign of a thin entity. Q42 (Douglas Adams, 75 label languages) and Q76 (Barack Obama, 112 label languages) — two of the most heavily cross-linked items on all of Wikidata — both lack a plain English label, despite both having an English Wikipedia sitelink. Once an item has enough interwiki links that everyone just reads its name off the Wikipedia sitelink, the formal labels.en field seems to go unmaintained. labelPreferred is reported honestly as null in that case rather than silently borrowed from the sitelink — sitelinkTitleEn is offered as a separate field for exactly this situation.

A nonexistent or malformed id is HTTP 400, not 404 — with an HTML (not JSON) error body naming the bad id. Reported as invalid_id, distinct from a transport failure.

Claim values are tagged with their own type, and unwrapping them one way loses information. An entity reference (P31, "instance of") simplifies cleanly to a bare QID string; a quantity carries a unit that would be lost by taking just the number; a time value carries a calendar model; a monolingual text carries its own language tag independent of the claim's property. claimsSimplified picks the right shape per claim rather than flattening everything the same way — claims still carries the full raw Wikibase structure alongside it.

Notes

No anti-bot layer was seen on four TLS profiles, so the proxy is off by default. One request per requested id — duplicate ids (after normalisation) are deduped before any request goes out.