Wikipedia Company Scraper — Founded, HQ, Revenue, Employees avatar

Wikipedia Company Scraper — Founded, HQ, Revenue, Employees

Pricing

from $2.00 / 1,000 results

Go to Apify Store
Wikipedia Company Scraper — Founded, HQ, Revenue, Employees

Wikipedia Company Scraper — Founded, HQ, Revenue, Employees

Extract structured company data from Wikipedia infoboxes — founded year, headquarters, revenue, employees, key people, industry, products, owner/parent, subsidiaries. Clean JSON from any Wikipedia article (English by default).

Pricing

from $2.00 / 1,000 results

Rating

0.0

(0)

Developer

Berkan Kaplan

Berkan Kaplan

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

2

Monthly active users

17 days ago

Last modified

Categories

Share

Wikipedia Company Scraper 📚

foXLabs web & community series: GitHub trending · Hacker News · Community listening

🎉 Turn Wikipedia company articles into clean, structured data — no login, no API key, one row per company, with the summary, founded year, industry, headquarters, key people and links. Built for research, enrichment and knowledge-base building.

🔍 What is the Wikipedia Company Scraper — and when should you use it?

Give this actor company names and it returns matching companies from public Wikipedia company articles — as clean, deduplicated rows you can filter, export or feed to an AI agent. Every run reads the source live.

Use it when you need: a company list for outreach; a quick profile before a call; or a starting point for account research.

Use something else when: you need a registry record — Wikipedia is an encyclopedia, not an official registry.

🤖 Use with AI agents

Already on the Apify MCP server? Ask for this Actor by name: foxlabs/wikipedia-company-scraper.

Your agent can pay for its own runs. This Actor is pay-per-event with agentic payments, so an agent can discover it, run it and settle the bill over x402 (USDC on Base) or Skyfire — no Apify account or API token of its own. Billing is the same either way: per delivered record, never for errors.

Otherwise paste this into Claude, ChatGPT, Cursor or any MCP-enabled assistant:

I want to pull company records using the Apify Actor `foxlabs/wikipedia-company-scraper`.
Input: `mode`, `companies`, `categories`, `categoryDepth` and more — see the Input table below. `maxResults` caps how many results are returned.
Start with: {"mode":"lookup","companies":["Apple Inc.","Shopify","Siemens","Nestlé"],"maxResults":200}
Ask me what to look up, run the Actor, then summarise the rows as a table.

The machine-readable API, MCP config and OpenAPI definition live at apify.com/foxlabs/wikipedia-company-scraper.md.

📋 Overview

Everything you need to turn public Wikipedia company articles into clean, structured data — in one actor, with no login, cookies or API key.

Why teams pick this actor:

  • ✅ Whole source, one call — name or ID in, matching companies out.
  • 🧹 No empty-promise columns — only fields this registry actually fills; degenerate columns are removed.
  • 🔗 Stable identifiers — every row carries the source's own IDs, ready to join across runs and to other Fox Labs actors.
  • 💰 Per-row pricing — a minimal price per delivered row, no subscription.
  • 🤖 Agent-ready — MCP + x402 agentic payments.

✨ Features

  • 🔍 Name or ID lookup — relevance-ranked name search or exact registry-ID lookup.
  • 🏢 Full entity profile — status, legal form, formation date, address and the registry’s own contact fields.
  • 🧹 Clean schema — deduplicated camelCase rows, ready for CSV/Excel/JSON.

🎬 Quick Start

curl -X POST "https://api.apify.com/v2/acts/foxlabs~wikipedia-company-scraper/runs?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"mode":"lookup","companies":["Apple Inc.","Shopify","Siemens","Nestlé"],"maxResults":200}'

🚀 Getting Started (3 steps)

  1. Choose your targets — company names.
  2. Set the cap — maxResults limits how many results are returned.
  3. Run and export — get a clean dataset as JSON, CSV or Excel.

📥 Input

{"mode":"lookup","companies":["Apple Inc.","Shopify","Siemens","Nestlé"],"maxResults":200}
FieldTypeDescription
modestringLook up companies = you provide the names or Wikipedia URLs. Browse a category = the Actor walks a Wikipedia category (e.g. every company listed on the NYSE) and…
companiesarrayLookup mode. Company names as they appear on Wikipedia ("Apple Inc.", "Nestlé") or full article URLs. Up to 50 are resolved per request, so long lists stay fast.
categoriesarrayCategory mode. Category names with or without the "Category:" prefix, e.g. "Companies listed on the New York Stock Exchange" or "Software companies of the United…
categoryDepthintegerCategory mode. 0 = only the category itself (recommended). 1-3 also walks subcategories, which multiplies the result count fast.
includeWikidatabooleanAdds the joinable identifiers — Wikidata QID, LEI, ISIN, ticker, official website — plus industry, country, headquarters, inception, employees, parent,…
includeInfoboxbooleanAlso parses the article's infobox for figures Wikidata usually lacks — revenue, operating/net income, total assets, key people, products. Costs one extra page…
pageviewMonthsintegerAdds monthly Wikipedia pageviews for each company — a public-attention signal you can trend over time. 0 = off. One extra request per company.
maxResultsintegerHard cap on dataset rows. Set 0 for unlimited (a large category can hold thousands of companies).

📤 Output

One row per result, saved to the dataset. Every row carries scrapedAt. Lookups that cannot be completed are reported in the run log rather than silently dropped.

FieldDescription
companyNameCompany / entity name
wikipediaUrlWikipedia Url
summarySummary
wikidataQidWikidata Qid
wikidataUrlWikidata Url
officialWebsiteOfficial Website
leiLei
isinIsin
tickerSymbolTicker Symbol
stockExchangeStock Exchange
legalFormLegal form / entity type
industryIndustry sector as published by the source — not a NACE or SIC code
countryCountry
headquartersHeadquarters
inceptionInception
employeesRegistered employee count
subsidiariesSubsidiaries
foundersFounders
pageIdPage Id
languageLanguage
licenseLicense
sourceSource
fetchedAtFetched At
infoboxInfobox
pageviewsMonthlyPageviews Monthly
pageviewsTotalPageviews Total
pageviewsLatestPageviews Latest

💼 Use cases

1. Company research — get a structured company overview. Input: company names. Output: summary + infobox facts. Use: a research brief.

2. Knowledge-base enrichment — attach Wikipedia facts to entities. Input: company names. Output: founded + industry + HQ. Use: an enriched KB.

3. Due diligence prep — get background on a company. Input: company names. Output: summary + key people. Use: a background sheet.

🔗 Integration

JavaScript / Node.js

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('foxlabs/wikipedia-company-scraper').call({"mode":"lookup","companies":["Apple Inc.","Shopify","Siemens","Nestlé"],"maxResults":200});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items[0]);

Python

from apify_client import ApifyClient
client = ApifyClient('YOUR_TOKEN')
run = client.actor('foxlabs/wikipedia-company-scraper').call(run_input={"mode":"lookup","companies":["Apple Inc.","Shopify","Siemens","Nestlé"],"maxResults":200})
for item in client.dataset(run['defaultDatasetId']).iterate_items():
print(item)

Automation (n8n / Zapier / Make): schedule or webhook → HTTP request to the actor API with your input → handle the JSON dataset → push to a sheet, CRM or dashboard.

📊 Pricing

Pay-per-event: per delivered record. Empty or failed lookups are never billed. View current pricing.

❓ FAQ

Do I need an account, login or API key? No. This reads public Wikipedia company articles.

What do I search by? Company names.

How current is the data? Every run queries the source live, so results are as fresh as the registry.

What fields are returned? Summary, founded year, industry, headquarters, key people and infobox links from the Wikipedia article.

Can I export to CSV / Excel / JSON? Yes — directly from the Apify dataset.

🐛 Troubleshooting

  • Fewer rows than expected — raise maxResults, or refine the input.
  • A name returns an unexpected entity — it matched a similar registered name; search the exact registry ID.
  • No rows for a name — try the entity’s exact legal name or its registry ID.

This actor reads public Wikipedia company articles. Results can still contain personal data (e.g. a person’s name); personal data is protected by the GDPR and similar laws, so only process it with a legitimate basis. See Apify’s blog post on the legality of web scraping.

🤝 Support & contact

Changelog

0.2.12 — 2026-09-20 — README examples corrected against the real input schema

  • The README's code examples did not match this Actor. They used queries and maxResultsPerQuery — keys that do not exist in this Actor's input schema — with a placeholder value, and the input table listed those same phantom fields. Anyone who copied the AI-agent, cURL, JavaScript or Python example got a failing run. Every example now uses the real schema and matches the Console prefill: {"mode":"lookup","companies":["Apple Inc.","Shopify","Siemens","Nestlé"],"maxResults":200}
  • The input table is regenerated from input_schema.json, so it lists the fields the Actor actually accepts.
  • Removed claims carried over from the same generator template where present: "formation / status monitoring", "a canonical registry record for KYB and due diligence", "every row carries query", and industry described as a NACE code.
  • No code, output field or pricing change.

0.2 — 2026-09-07

  • Dropped empty-promise columns. Removed parentOrganization — public Wikipedia company articles does not carry them, so they were shipped as always-null columns. Only fields this source actually fills are now emitted.
  • Enabled AI-agent payments (x402) + rebuilt the README to the full standard (What-is / when, AI-agents + x402 agentic payments + MCP, Overview, Features, Use cases, Integration, FAQ, Troubleshooting, Support & contact).

0.0

  • Initial release: data from public Wikipedia company articles by name or registry ID.