Wikipedia Company Scraper — Founded, HQ, Revenue, Employees
Pricing
from $2.00 / 1,000 results
Wikipedia Company Scraper — Founded, HQ, Revenue, Employees
Extract structured company data from Wikipedia infoboxes — founded year, headquarters, revenue, employees, key people, industry, products, owner/parent, subsidiaries. Clean JSON from any Wikipedia article (English by default).
Pricing
from $2.00 / 1,000 results
Rating
0.0
(0)
Developer
Berkan Kaplan
Maintained by CommunityActor stats
0
Bookmarked
3
Total users
2
Monthly active users
17 days ago
Last modified
Categories
Share
Wikipedia Company Scraper 📚
foXLabs web & community series: GitHub trending · Hacker News · Community listening
🎉 Turn Wikipedia company articles into clean, structured data — no login, no API key, one row per company, with the summary, founded year, industry, headquarters, key people and links. Built for research, enrichment and knowledge-base building.
🔍 What is the Wikipedia Company Scraper — and when should you use it?
Give this actor company names and it returns matching companies from public Wikipedia company articles — as clean, deduplicated rows you can filter, export or feed to an AI agent. Every run reads the source live.
Use it when you need: a company list for outreach; a quick profile before a call; or a starting point for account research.
Use something else when: you need a registry record — Wikipedia is an encyclopedia, not an official registry.
🤖 Use with AI agents
Already on the Apify MCP server? Ask for this Actor by name: foxlabs/wikipedia-company-scraper.
Your agent can pay for its own runs. This Actor is pay-per-event with agentic payments, so an agent can discover it, run it and settle the bill over x402 (USDC on Base) or Skyfire — no Apify account or API token of its own. Billing is the same either way: per delivered record, never for errors.
Otherwise paste this into Claude, ChatGPT, Cursor or any MCP-enabled assistant:
I want to pull company records using the Apify Actor `foxlabs/wikipedia-company-scraper`.Input: `mode`, `companies`, `categories`, `categoryDepth` and more — see the Input table below. `maxResults` caps how many results are returned.Start with: {"mode":"lookup","companies":["Apple Inc.","Shopify","Siemens","Nestlé"],"maxResults":200}Ask me what to look up, run the Actor, then summarise the rows as a table.
The machine-readable API, MCP config and OpenAPI definition live at apify.com/foxlabs/wikipedia-company-scraper.md.
📋 Overview
Everything you need to turn public Wikipedia company articles into clean, structured data — in one actor, with no login, cookies or API key.
Why teams pick this actor:
- ✅ Whole source, one call — name or ID in, matching companies out.
- 🧹 No empty-promise columns — only fields this registry actually fills; degenerate columns are removed.
- 🔗 Stable identifiers — every row carries the source's own IDs, ready to join across runs and to other Fox Labs actors.
- 💰 Per-row pricing — a minimal price per delivered row, no subscription.
- 🤖 Agent-ready — MCP + x402 agentic payments.
✨ Features
- 🔍 Name or ID lookup — relevance-ranked name search or exact registry-ID lookup.
- 🏢 Full entity profile — status, legal form, formation date, address and the registry’s own contact fields.
- 🧹 Clean schema — deduplicated camelCase rows, ready for CSV/Excel/JSON.
🎬 Quick Start
curl -X POST "https://api.apify.com/v2/acts/foxlabs~wikipedia-company-scraper/runs?token=YOUR_TOKEN" \-H "Content-Type: application/json" \-d '{"mode":"lookup","companies":["Apple Inc.","Shopify","Siemens","Nestlé"],"maxResults":200}'
🚀 Getting Started (3 steps)
- Choose your targets — company names.
- Set the cap —
maxResultslimits how many results are returned. - Run and export — get a clean dataset as JSON, CSV or Excel.
📥 Input
{"mode":"lookup","companies":["Apple Inc.","Shopify","Siemens","Nestlé"],"maxResults":200}
| Field | Type | Description |
|---|---|---|
mode | string | Look up companies = you provide the names or Wikipedia URLs. Browse a category = the Actor walks a Wikipedia category (e.g. every company listed on the NYSE) and… |
companies | array | Lookup mode. Company names as they appear on Wikipedia ("Apple Inc.", "Nestlé") or full article URLs. Up to 50 are resolved per request, so long lists stay fast. |
categories | array | Category mode. Category names with or without the "Category:" prefix, e.g. "Companies listed on the New York Stock Exchange" or "Software companies of the United… |
categoryDepth | integer | Category mode. 0 = only the category itself (recommended). 1-3 also walks subcategories, which multiplies the result count fast. |
includeWikidata | boolean | Adds the joinable identifiers — Wikidata QID, LEI, ISIN, ticker, official website — plus industry, country, headquarters, inception, employees, parent,… |
includeInfobox | boolean | Also parses the article's infobox for figures Wikidata usually lacks — revenue, operating/net income, total assets, key people, products. Costs one extra page… |
pageviewMonths | integer | Adds monthly Wikipedia pageviews for each company — a public-attention signal you can trend over time. 0 = off. One extra request per company. |
maxResults | integer | Hard cap on dataset rows. Set 0 for unlimited (a large category can hold thousands of companies). |
📤 Output
One row per result, saved to the dataset. Every row carries scrapedAt. Lookups that cannot be completed are reported in the run log rather than silently dropped.
| Field | Description |
|---|---|
companyName | Company / entity name |
wikipediaUrl | Wikipedia Url |
summary | Summary |
wikidataQid | Wikidata Qid |
wikidataUrl | Wikidata Url |
officialWebsite | Official Website |
lei | Lei |
isin | Isin |
tickerSymbol | Ticker Symbol |
stockExchange | Stock Exchange |
legalForm | Legal form / entity type |
industry | Industry sector as published by the source — not a NACE or SIC code |
country | Country |
headquarters | Headquarters |
inception | Inception |
employees | Registered employee count |
subsidiaries | Subsidiaries |
founders | Founders |
pageId | Page Id |
language | Language |
license | License |
source | Source |
fetchedAt | Fetched At |
infobox | Infobox |
pageviewsMonthly | Pageviews Monthly |
pageviewsTotal | Pageviews Total |
pageviewsLatest | Pageviews Latest |
💼 Use cases
1. Company research — get a structured company overview. Input: company names. Output: summary + infobox facts. Use: a research brief.
2. Knowledge-base enrichment — attach Wikipedia facts to entities. Input: company names. Output: founded + industry + HQ. Use: an enriched KB.
3. Due diligence prep — get background on a company. Input: company names. Output: summary + key people. Use: a background sheet.
🔗 Integration
JavaScript / Node.js
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: 'YOUR_TOKEN' });const run = await client.actor('foxlabs/wikipedia-company-scraper').call({"mode":"lookup","companies":["Apple Inc.","Shopify","Siemens","Nestlé"],"maxResults":200});const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(items[0]);
Python
from apify_client import ApifyClientclient = ApifyClient('YOUR_TOKEN')run = client.actor('foxlabs/wikipedia-company-scraper').call(run_input={"mode":"lookup","companies":["Apple Inc.","Shopify","Siemens","Nestlé"],"maxResults":200})for item in client.dataset(run['defaultDatasetId']).iterate_items():print(item)
Automation (n8n / Zapier / Make): schedule or webhook → HTTP request to the actor API with your input → handle the JSON dataset → push to a sheet, CRM or dashboard.
📊 Pricing
Pay-per-event: per delivered record. Empty or failed lookups are never billed. View current pricing.
❓ FAQ
Do I need an account, login or API key? No. This reads public Wikipedia company articles.
What do I search by? Company names.
How current is the data? Every run queries the source live, so results are as fresh as the registry.
What fields are returned? Summary, founded year, industry, headquarters, key people and infobox links from the Wikipedia article.
Can I export to CSV / Excel / JSON? Yes — directly from the Apify dataset.
🐛 Troubleshooting
- Fewer rows than expected — raise
maxResults, or refine the input. - A name returns an unexpected entity — it matched a similar registered name; search the exact registry ID.
- No rows for a name — try the entity’s exact legal name or its registry ID.
⚖️ Is it legal to scrape this data?
This actor reads public Wikipedia company articles. Results can still contain personal data (e.g. a person’s name); personal data is protected by the GDPR and similar laws, so only process it with a legitimate basis. See Apify’s blog post on the legality of web scraping.
🤝 Support & contact
- 🌐 Website: data.foxlabs.com.tr
- 📧 Email: info@foxlabs.com.tr
- 🐛 Issues: open a ticket in the Actor’s Issues tab
- 🧰 More clean B2B data actors: Fox Labs on Apify
Changelog
0.2.12 — 2026-09-20 — README examples corrected against the real input schema
- The README's code examples did not match this Actor. They used
queriesandmaxResultsPerQuery— keys that do not exist in this Actor's input schema — with a placeholder value, and the input table listed those same phantom fields. Anyone who copied the AI-agent, cURL, JavaScript or Python example got a failing run. Every example now uses the real schema and matches the Console prefill:{"mode":"lookup","companies":["Apple Inc.","Shopify","Siemens","Nestlé"],"maxResults":200} - The input table is regenerated from
input_schema.json, so it lists the fields the Actor actually accepts. - Removed claims carried over from the same generator template where present: "formation / status monitoring", "a canonical registry record for KYB and due diligence", "every row carries
query", andindustrydescribed as a NACE code. - No code, output field or pricing change.
0.2 — 2026-09-07
- Dropped empty-promise columns. Removed
parentOrganization— public Wikipedia company articles does not carry them, so they were shipped as always-null columns. Only fields this source actually fills are now emitted. - Enabled AI-agent payments (x402) + rebuilt the README to the full standard (What-is / when, AI-agents + x402 agentic payments + MCP, Overview, Features, Use cases, Integration, FAQ, Troubleshooting, Support & contact).
0.0
- Initial release: data from public Wikipedia company articles by name or registry ID.