# Wikidata Entities Scraper (`scrapyx/wikidata-entities-scraper`) Actor

Full Wikidata entity data by QID: labels, descriptions and aliases in every published language, every claim (raw and simplified), and every cross-wiki sitelink. No API key. Detects silent item merges.

- **URL**: https://apify.com/scrapyx/wikidata-entities-scraper.md
- **Developed by:** [Ibnu Adzim](https://apify.com/scrapyx) (community)
- **Categories:** Developer tools, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.84 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Wikidata Entities Scraper (Labels, Claims & Sitelinks)

Full **Wikidata** entity data by QID: labels, descriptions and aliases in
**every** published language, every claim the item carries (raw and
simplified), and every cross-wiki sitelink.

No login. No API key. HTTP only — and deliberately narrow: **there is no
search mode**. Wikidata's SPARQL endpoint and its label-search API are both
`robots.txt`-disallowed (see below), so resolving ids you already have is
the only in-policy operation. It pairs naturally with
[`reference/wikipedia-articles-scraper`](../wikipedia-articles-scraper),
which already returns the `wikibase_item` QID for an article but not its
full entity data.

| Record type | One per | Carries |
| --- | --- | --- |
| `ENTITY` | resolved id | requested/actual id, merge flag, labels/descriptions/aliases (all languages), sitelinks, claims (raw + simplified), revision info |
| `ERROR` | failed input | `_error` code and an `_errorDetail` saying what to change |

### Input

Just `entityIds` — a bare id (`Q42`) or a full `wikidata.org` URL ending in
one. `preferredLanguage` (default `en`) picks out `labelPreferred` /
`descriptionPreferred` from the full multi-language maps every record
already carries in full.

### Things this endpoint will mislead you about

Each is measured, and each has a scenario in
`tests/smoke/wikidata-entities-scraper_traps.sh` (8/8 passing).

**Wikidata SPARQL — flagged in this repo's own research notes across many
past sessions as "the strongest candidate left to build" — turns out to be
closed, and always was.** `query.wikidata.org/robots.txt` carries a blanket
`Disallow: /sparql` for `User-agent: *`. `www.wikidata.org/robots.txt`
separately disallows `/w/` wholesale, which closes both the label-search API
(`wbsearchentities`) and the newer Wikibase REST API. The only surface left
`Allow:`ed is `/wiki/Special:EntityData/<id>.<format>` — one static file per
entity, keyed on an id you already have. This Actor touches nothing else.

**A merged item answers 200 and silently hands you a different entity.**
Wikidata periodically merges duplicate items. `Q9270598` (merged) still
resolves — but the response's `entities` object is keyed on `Q13247166`,
the id it was merged INTO, with nothing at the top level marking a merge
happened. `wasMerged` and `entityId` vs `requestedId` are how this Actor
surfaces it — compare them yourself if you build on top of the raw JSON.

**A missing label is not rare, and it is not a sign of a thin entity.**
`Q42` (Douglas Adams, 75 label languages) and `Q76` (Barack Obama, 112
label languages) — two of the most heavily cross-linked items on all of
Wikidata — **both lack a plain English label**, despite both having an
English Wikipedia sitelink. Once an item has enough interwiki links that
everyone just reads its name off the Wikipedia sitelink, the formal
`labels.en` field seems to go unmaintained. `labelPreferred` is reported
honestly as `null` in that case rather than silently borrowed from the
sitelink — `sitelinkTitleEn` is offered as a **separate** field for exactly
this situation.

**A nonexistent or malformed id is HTTP 400, not 404** — with an HTML
(not JSON) error body naming the bad id. Reported as `invalid_id`, distinct
from a transport failure.

**Claim values are tagged with their own type, and unwrapping them one way
loses information.** An entity reference (`P31`, "instance of") simplifies
cleanly to a bare QID string; a quantity carries a unit that would be lost
by taking just the number; a time value carries a calendar model; a
monolingual text carries its own language tag independent of the claim's
property. `claimsSimplified` picks the right shape per claim rather than
flattening everything the same way — `claims` still carries the full raw
Wikibase structure alongside it.

### Notes

No anti-bot layer was seen on four TLS profiles, so the proxy is **off by
default**. One request per requested id — duplicate ids (after
normalisation) are deduped before any request goes out.

# Actor input Schema

## `entityIds` (type: `array`):

A Wikidata item id (e.g. 'Q42') or a full wikidata.org URL ending in one. This is the whole input — Wikidata's SPARQL endpoint and label-search API are both robots.txt-disallowed, so lookup by id is the only in-policy operation.

## `preferredLanguage` (type: `string`):

Picks out `labelPreferred`/`descriptionPreferred` from the full multi-language label/description maps, which every record carries in full regardless of this setting. A well-known entity can lack a label in your chosen language entirely — that is reported as null, not silently substituted.

## `maxConcurrency` (type: `integer`):

Parallel in-flight requests.

## `minRequestInterval` (type: `number`):

Politeness pacing for a public non-profit's infrastructure.

## `proxyConfiguration` (type: `object`):

Optional. No anti-bot layer was observed on four TLS profiles, so a proxy is OFF by default.

## Actor input object example

```json
{
  "entityIds": [
    "Q42"
  ],
  "preferredLanguage": "en",
  "maxConcurrency": 5,
  "minRequestInterval": 0.2,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `items` (type: `string`):

One row per scraped record. See the dataset's default view for field definitions.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "entityIds": [
        "Q42"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapyx/wikidata-entities-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "entityIds": ["Q42"] }

# Run the Actor and wait for it to finish
run = client.actor("scrapyx/wikidata-entities-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "entityIds": [
    "Q42"
  ]
}' |
apify call scrapyx/wikidata-entities-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapyx/wikidata-entities-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/dSPJP4Yhck9uv8S4x/builds/coMtc7r2YRAr32j1s/openapi.json
