Wikidata Entities Scraper
Pricing
from $0.84 / 1,000 results
Wikidata Entities Scraper
Full Wikidata entity data by QID: labels, descriptions and aliases in every published language, every claim (raw and simplified), and every cross-wiki sitelink. No API key. Detects silent item merges.
Pricing
from $0.84 / 1,000 results
Rating
0.0
(0)
Developer
Ibnu Adzim
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
7 days ago
Last modified
Categories
Share
Wikidata Entities Scraper (Labels, Claims & Sitelinks)
Full Wikidata entity data by QID: labels, descriptions and aliases in every published language, every claim the item carries (raw and simplified), and every cross-wiki sitelink.
No login. No API key. HTTP only — and deliberately narrow: there is no
search mode. Wikidata's SPARQL endpoint and its label-search API are both
robots.txt-disallowed (see below), so resolving ids you already have is
the only in-policy operation. It pairs naturally with
../wikipedia-articles-scraper,
which already returns the wikibase_item QID for an article but not its
full entity data.
| Record type | One per | Carries |
|---|---|---|
ENTITY | resolved id | requested/actual id, merge flag, labels/descriptions/aliases (all languages), sitelinks, claims (raw + simplified), revision info |
ERROR | failed input | _error code and an _errorDetail saying what to change |
Input
Just entityIds — a bare id (Q42) or a full wikidata.org URL ending in
one. preferredLanguage (default en) picks out labelPreferred /
descriptionPreferred from the full multi-language maps every record
already carries in full.
Things this endpoint will mislead you about
Each is measured, and each has a scenario in
tests/smoke/wikidata-entities-scraper_traps.sh (8/8 passing).
Wikidata SPARQL — flagged in this repo's own research notes across many
past sessions as "the strongest candidate left to build" — turns out to be
closed, and always was. query.wikidata.org/robots.txt carries a blanket
Disallow: /sparql for User-agent: *. www.wikidata.org/robots.txt
separately disallows /w/ wholesale, which closes both the label-search API
(wbsearchentities) and the newer Wikibase REST API. The only surface left
Allow:ed is /wiki/Special:EntityData/<id>.<format> — one static file per
entity, keyed on an id you already have. This Actor touches nothing else.
A merged item answers 200 and silently hands you a different entity.
Wikidata periodically merges duplicate items. Q9270598 (merged) still
resolves — but the response's entities object is keyed on Q13247166,
the id it was merged INTO, with nothing at the top level marking a merge
happened. wasMerged and entityId vs requestedId are how this Actor
surfaces it — compare them yourself if you build on top of the raw JSON.
A missing label is not rare, and it is not a sign of a thin entity.
Q42 (Douglas Adams, 75 label languages) and Q76 (Barack Obama, 112
label languages) — two of the most heavily cross-linked items on all of
Wikidata — both lack a plain English label, despite both having an
English Wikipedia sitelink. Once an item has enough interwiki links that
everyone just reads its name off the Wikipedia sitelink, the formal
labels.en field seems to go unmaintained. labelPreferred is reported
honestly as null in that case rather than silently borrowed from the
sitelink — sitelinkTitleEn is offered as a separate field for exactly
this situation.
A nonexistent or malformed id is HTTP 400, not 404 — with an HTML
(not JSON) error body naming the bad id. Reported as invalid_id, distinct
from a transport failure.
Claim values are tagged with their own type, and unwrapping them one way
loses information. An entity reference (P31, "instance of") simplifies
cleanly to a bare QID string; a quantity carries a unit that would be lost
by taking just the number; a time value carries a calendar model; a
monolingual text carries its own language tag independent of the claim's
property. claimsSimplified picks the right shape per claim rather than
flattening everything the same way — claims still carries the full raw
Wikibase structure alongside it.
Notes
No anti-bot layer was seen on four TLS profiles, so the proxy is off by default. One request per requested id — duplicate ids (after normalisation) are deduped before any request goes out.