Wikipedia Company Scraper — Founded, HQ, Revenue, Employees
Pricing
from $2.00 / 1,000 results
Wikipedia Company Scraper — Founded, HQ, Revenue, Employees
Extract structured company data from Wikipedia infoboxes — founded year, headquarters, revenue, employees, key people, industry, products, owner/parent, subsidiaries. Clean JSON from any Wikipedia article (English by default).
Pricing
from $2.00 / 1,000 results
Rating
0.0
(0)
Developer
Berkan Kaplan
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
8 days ago
Last modified
Categories
Share
Wikipedia Company Scraper 📚
🎉 Turn Wikipedia company articles into clean, structured data — no login, no API key, one row per company, with the summary, founded year, industry, headquarters, key people and links. Built for research, enrichment and knowledge-base building.
🔍 What is the Wikipedia Company Scraper — and when should you use it?
Give this actor company names and it returns matching companies from public Wikipedia company articles — as clean, deduplicated rows you can filter, export or feed to an AI agent. Every run queries the source live, so the data is as fresh as the registry itself.
Use it when you need: a company company list for outreach; formation / status monitoring; or a canonical registry record for KYB and due diligence.
Use something else when: you need a registry record — Wikipedia is an encyclopedia, not an official registry.
🤖 Use with AI agents
Already on the Apify MCP server? Ask for this Actor by name: foxlabs/wikipedia-company-scraper.
Your agent can pay for its own runs. This Actor is pay-per-event with agentic payments, so an agent can discover it, run it and settle the bill over x402 (USDC on Base) or Skyfire — no Apify account or API token of its own. Billing is the same either way: per delivered record, never for errors.
Otherwise paste this into Claude, ChatGPT, Cursor or any MCP-enabled assistant:
I want to pull company company records using the Apify Actor `foxlabs/wikipedia-company-scraper`.Input: `queries` is a list of company names. `maxResultsPerQuery` caps rows per query.Start with: {"queries":["undefined"],"maxResultsPerQuery":50}Ask me what to look up, run the Actor, then summarise the rows as a table.
The machine-readable API, MCP config and OpenAPI definition live at apify.com/foxlabs/wikipedia-company-scraper.md.
📋 Overview
Everything you need to turn public Wikipedia company articles into clean, structured data — in one actor, with no login, cookies or API key.
Why teams pick this actor:
- ✅ Whole source, one call — name or ID in, matching companies out.
- 🧹 No empty-promise columns — only fields this registry actually fills; degenerate columns are removed.
- 🔗 Stable identifiers — every row carries the source's own IDs, ready to join across runs and to other Fox Labs actors.
- 💰 Pay only for results — per-row pricing, empty/failed lookups never billed.
- 🤖 Agent-ready — MCP + x402 agentic payments.
✨ Features
- 🔍 Name or ID lookup — relevance-ranked name search or exact registry-ID lookup.
- 🏢 Full entity profile — status, legal form, formation date, address and the registry’s own contact fields.
- 🧹 Clean schema — deduplicated camelCase rows, ready for CSV/Excel/JSON.
🎬 Quick Start
curl -X POST "https://api.apify.com/v2/acts/foxlabs~wikipedia-company-scraper/runs?token=YOUR_TOKEN" \-H "Content-Type: application/json" \-d '{"queries":["undefined"],"maxResultsPerQuery":50}'
🚀 Getting Started (3 steps)
- Choose your targets — company names.
- Set the cap —
maxResultsPerQuerylimits rows per query. - Run and export — get a clean dataset as JSON, CSV or Excel.
📥 Input
{"queries":["undefined"],"maxResultsPerQuery":50}
| Field | Type | Description |
|---|---|---|
queries | array | Company names. |
maxResultsPerQuery | integer | Caps rows per query. |
maxConcurrency | integer | How many queries to fetch at once. |
includeRaw | boolean | Attach the source’s untouched record under raw. |
📤 Output
One row per company, saved to the dataset. Every row also carries query, scrapedAt, and — when a lookup fails — an error explaining why (never silently dropped, never billed).
| Field | Description |
|---|---|
companyName | Company / entity name |
wikipediaUrl | Wikipedia Url |
summary | Summary |
wikidataQid | Wikidata Qid |
wikidataUrl | Wikidata Url |
officialWebsite | Official Website |
lei | Lei |
isin | Isin |
tickerSymbol | Ticker Symbol |
stockExchange | Stock Exchange |
legalForm | Legal form / entity type |
industry | Primary industry (NACE) |
country | Country |
headquarters | Headquarters |
inception | Inception |
employees | Registered employee count |
subsidiaries | Subsidiaries |
founders | Founders |
pageId | Page Id |
language | Language |
license | License |
source | Source |
fetchedAt | Fetched At |
infobox | Infobox |
pageviewsMonthly | Pageviews Monthly |
pageviewsTotal | Pageviews Total |
pageviewsLatest | Pageviews Latest |
💼 Use cases
1. Company research — get a structured company overview. Input: company names. Output: summary + infobox facts. Use: a research brief.
2. Knowledge-base enrichment — attach Wikipedia facts to entities. Input: company names. Output: founded + industry + HQ. Use: an enriched KB.
3. Due diligence prep — get background on a company. Input: company names. Output: summary + key people. Use: a background sheet.
🔗 Integration
JavaScript / Node.js
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: 'YOUR_TOKEN' });const run = await client.actor('foxlabs/wikipedia-company-scraper').call({"queries":["undefined"],"maxResultsPerQuery":50});const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(items[0]);
Python
from apify_client import ApifyClientclient = ApifyClient('YOUR_TOKEN')run = client.actor('foxlabs/wikipedia-company-scraper').call(run_input={"queries":["undefined"],"maxResultsPerQuery":50})for item in client.dataset(run['defaultDatasetId']).iterate_items():print(item)
Automation (n8n / Zapier / Make): schedule or webhook → HTTP request to the actor API with your queries → handle the JSON dataset → push to a sheet, CRM or dashboard.
📊 Pricing
Pay-per-event: per delivered record. Empty or failed lookups are never billed. View current pricing.
❓ FAQ
Do I need an account, login or API key? No. This reads public Wikipedia company articles.
What do I search by? Company names.
How current is the data? Every run queries the source live, so results are as fresh as the registry.
What fields are returned? Summary, founded year, industry, headquarters, key people and infobox links from the Wikipedia article.
Can I export to CSV / Excel / JSON? Yes — directly from the Apify dataset.
🐛 Troubleshooting
- Fewer rows than expected — raise
maxResultsPerQuery, or refine the name. - A name returns an unexpected entity — it matched a similar registered name; search the exact registry ID.
- No rows for a name — try the entity’s exact legal name or its registry ID.
⚖️ Is it legal to scrape this data?
This actor reads public Wikipedia company articles. Results can still contain personal data (e.g. a person’s name); personal data is protected by the GDPR and similar laws, so only process it with a legitimate basis. See Apify’s blog post on the legality of web scraping.
🤝 Support & contact
- 🌐 Website: data.foxlabs.com.tr
- 📧 Email: info@foxlabs.com.tr
- 🐛 Issues: open a ticket in the Actor’s Issues tab
- 🧰 More clean B2B data actors: Fox Labs on Apify
Changelog
0.2 — 2026-09-07
- Dropped empty-promise columns. Removed
parentOrganization— public Wikipedia company articles does not carry them, so they were shipped as always-null columns. Only fields this source actually fills are now emitted. - Enabled AI-agent payments (x402) + rebuilt the README to the full standard (What-is / when, AI-agents + x402 agentic payments + MCP, Overview, Features, Use cases, Integration, FAQ, Troubleshooting, Support & contact).
0.0
- Initial release: data from public Wikipedia company articles by name or registry ID.