# Wikidata Extractor (`cynix_dev/wikidata-extractor`) Actor

Extract structured knowledge from Wikidata — the free, open knowledge base. Search by term or fetch exact entities (Q-IDs) and get clean labels, descriptions, aliases, and a flattened claims map with human-readable property names. No API key required.

- **URL**: https://apify.com/cynix\_dev/wikidata-extractor.md
- **Developed by:** [Cynix Dev](https://apify.com/cynix_dev) (community)
- **Categories:** Automation, Developer tools, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.17 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Wikidata Extractor

Extract structured knowledge from **Wikidata** — the free, open knowledge base. Search by term or fetch exact entities (Q-IDs) and get clean labels, descriptions, aliases and a flattened claims map with human-readable property names. No API key required.

### What it does

Wikidata is the structured backbone behind Wikipedia — millions of entities with typed statements. This Actor searches it by term or fetches exact entities by Q-ID, and returns the useful parts flattened: label, description, aliases, sitelinks and a claims map where property IDs are resolved to readable names (so `P31` becomes "instance of").

It's the knowledge-enrichment primitive: resolve "Berlin" to structured facts, or pull a batch of entities to build a knowledge graph.

### Features

- **Search or fetch** — free-text term, or exact Q-IDs.
- **Flattened claims** — property IDs resolved to human-readable names.
- **Labels, descriptions, aliases** — in the language you choose.
- **Sitelinks** — links to Wikipedia and sister projects per wiki.
- **Keyless public API** — no API key, no proxy.

### What people use it for

- Knowledge-graph building — structured facts about entities at scale.
- Entity enrichment — attach Wikidata claims to your own records.
- Disambiguation — resolve a term to the right entity and Q-ID.
- NLP and QA — grounding answers in a structured source.
- Dataset linking — bridge internal IDs to Wikidata entities.

### Reading the claims map

Each claim is stored as a property ID + value. The Actor resolves the property ID to its name so the data is legible:

- `P31` → "instance of"
- `P569` → "date of birth"
- `P1082` → "population"

Values themselves may be entity IDs (`Q...`) or literals; for nested entities you'd resolve the Q-ID in a second call. This keeps one record per entity while preserving the structure.

#### Language

`language` controls labels, descriptions and aliases. Wikidata is multilingual; the same entity has labels in many languages, so pick the one your users read.

### Input

Provide `searchTerm` for a search, or `entityIds` (Q-IDs) for direct fetch — Q-IDs win when both are set. `language` selects label language.

| Field | Type | Default | What it does |
| --- | --- | --- | --- |
| `searchTerm` | string | — | Free-text search to resolve Wikidata entities, e.g. 'Berlin', 'Elon Musk', 'COVID-19'. Used when no entity IDs are supplied. |
| `entityIds` | array | `[]` | Exact Wikidata Q-IDs to fetch directly, e.g. \["Q64", "Q5"]. Wins over searchTerm if set. |
| `language` | string | `en` | Language code for labels/descriptions, e.g. en, de, fr. |
| `maxResults` | integer | `5` | Max entities when searching by term. Range 1–50. |
| `proxyConfiguration` | object | see below | Wikidata is a free, public API and does not require a proxy. Leave disabled. |

#### Input example

```json
{
  "searchTerm": "Berlin",
  "maxResults": 2,
  "language": "en",
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

### Output

One record per entity: id, label, description, aliases, a flattened claims map, and sitelinks.

Every dataset record contains: `entityId`, `label`, `description`, `aliases`, `claims`, `sitelinks`, `fetchedAt`.

#### Output example

A real record from a run of this Actor:

```json
{
  "entityId": "Q64",
  "label": "Berlin",
  "description": "federated state, capital and largest city of Germany",
  "aliases": "Berlin, Germany; DE-BE",
  "claims": "{\"highest point\":[\"Q19259618\"],\"topic's main Wikimedia portal\":[\"Q3248436\"],\"instance of\":[\"Q1901835\",\"Q200250\",\"Q1307779\",\"Q15974307\",\"Q42744322\",\"Q133442\",\"Q114401982\",\"Q51929311\",\"Q1221156\",\"Q257391\",\"Q707813\",\"Q67123 …",
  "sitelinks": "[\"itwikivoyage\",\"ukwikivoyage\",\"svwikivoyage\",\"ruwikivoyage\",\"rowikivoyage\",\"ptwikivoyage\",\"frwikivoyage\",\"hewikivoyage\",\"viwikivoyage\",\"zhwikivoyage\",\"dewikisource\",\"eowikiquote\",\"nnwikiquote\",\"enwikiquote\",\"itwikiquote …",
  "fetchedAt": "2026-08-21T13:10:47.284Z"
}
```

Export the dataset as JSON, CSV, Excel, XML or JSONL from the Console, or pull it programmatically through the Apify API and any of the official clients.

### How to use it

1. Click **Try for free** (or **Start** if you already have an Apify account).
2. Fill in the input fields described above — the defaults already produce a working run.
3. Press **Start** and watch the log; results stream into the dataset as they are found.
4. When the run finishes, open the **Output/Storage** tab and export as JSON, CSV or Excel.

Runs can be scheduled (hourly, daily, weekly) and wired into Slack, Google Sheets, Zapier, Make, webhooks or your own backend through Apify integrations. Everything the Console does is also available over the [Apify API](https://docs.apify.com/api/v2).

### Proxy configuration

This Actor accepts a standard Apify **proxy configuration** object. Residential proxy is the default because the target site rate-limits datacenter IP ranges; you can select a specific exit country or supply your own proxy URLs.

```json
{
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

### Pricing

This Actor is billed on Apify's **pay-per-event** model: a small charge when a run starts, plus a charge for each result written to the dataset. You only pay for records you actually receive — a run that finds nothing costs only the start event. Current rates are always shown on the **Pricing** tab of this page, and the run log prints your usage as it goes.

Free-plan credits from Apify cover a large amount of light usage, so you can evaluate the Actor before committing to anything.

### FAQ

#### Do I need an API key?

No. Wikidata's API is public and keyless.

#### Why are some claim values just Q-IDs?

Wikidata stores relationships as entity references. The Actor resolves property names for you, but resolving the *value* entity is a separate fetch — chain another call if you need the value's label.

#### How current is the data?

Wikidata is edited continuously by volunteers; popular entities are very fresh, obscure ones less so.

#### Is this affiliated with the Wikimedia Foundation?

No. It uses Wikidata's public API and is not affiliated with, endorsed by, or operated by the Wikimedia Foundation.

### Other Actors by cynix\_dev

| Actor | What it does |
| --- | --- |
| [arXiv Papers Extractor](https://apify.com/cynix_dev/arxiv-papers) | Search arXiv and extract papers as clean typed records: title, abstract, authors, categories, DOI, and direct PDF links. |
| [CoinGecko Markets — Crypto Data API](https://apify.com/cynix_dev/coingecko-markets) | Live cryptocurrency market data from CoinGecko as clean typed JSON: price, market cap, volume, 24h change, ATH/ATL, … |
| [FX Rates & History](https://apify.com/cynix_dev/fx-rates-history) | Latest and historical foreign exchange rates (ECB reference data) as clean, typed dataset records. |
| [USGS Earthquakes — GeoJSON Extractor](https://apify.com/cynix_dev/usgs-earthquakes) | Pull live and historical earthquakes from the USGS FDSN event service as clean typed JSON: magnitude, place, time, lat/lon/depth, … |
| [Launch Library 2 — Rocket Launch Tracker](https://apify.com/cynix_dev/launch-library-launches) | Upcoming, previous, and specific rocket launches from The Space Devs' Launch Library 2 API. |
| [Open Food Facts Extractor](https://apify.com/cynix_dev/open-food-facts) | Search and extract food-product data from Open Food Facts as clean typed JSON: name, brand, ingredients, allergens, nutrition … |

### Legal and responsible use

This Actor collects only publicly available information. You are responsible for how you use the data, including compliance with the target site's Terms of Service, robots directives, copyright, and data protection law such as GDPR and CCPA. Do not use it to gather personal data without a lawful basis.

### Support and feedback

Found a bug, hit a site change, or need an extra field? Open a ticket on the **Issues** tab of this Actor — issues are read and fixed. Feature requests and custom-scraper enquiries are welcome through the same channel.

# Actor input Schema

## `searchTerm` (type: `string`):

Free-text search to resolve Wikidata entities, e.g. 'Berlin', 'Elon Musk', 'COVID-19'. Used when no entity IDs are supplied.

## `entityIds` (type: `array`):

Exact Wikidata Q-IDs to fetch directly, e.g. \["Q64", "Q5"]. Wins over searchTerm if set.

## `language` (type: `string`):

Language code for labels/descriptions, e.g. en, de, fr.

## `maxResults` (type: `integer`):

Max entities when searching by term.

## `proxyConfiguration` (type: `object`):

Wikidata is a free, public API and does not require a proxy. Leave disabled.

## Actor input object example

```json
{
  "searchTerm": "Ada Lovelace",
  "entityIds": [],
  "language": "en",
  "maxResults": 5,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `dataset` (type: `string`):

Dataset containing all scraped records.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchTerm": "Ada Lovelace"
};

// Run the Actor and wait for it to finish
const run = await client.actor("cynix_dev/wikidata-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "searchTerm": "Ada Lovelace" }

# Run the Actor and wait for it to finish
run = client.actor("cynix_dev/wikidata-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchTerm": "Ada Lovelace"
}' |
apify call cynix_dev/wikidata-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,cynix_dev/wikidata-extractor"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/4NPz5Cdbh80Etyr8e/builds/8sGeOhDagqtvjDmXF/openapi.json
