Wikipedia Company Scraper — Founded, HQ, Revenue, Employees avatar

Wikipedia Company Scraper — Founded, HQ, Revenue, Employees

Pricing

from $2.00 / 1,000 results

Go to Apify Store
Wikipedia Company Scraper — Founded, HQ, Revenue, Employees

Wikipedia Company Scraper — Founded, HQ, Revenue, Employees

Extract structured company data from Wikipedia infoboxes — founded year, headquarters, revenue, employees, key people, industry, products, owner/parent, subsidiaries. Clean JSON from any Wikipedia article (English by default).

Pricing

from $2.00 / 1,000 results

Rating

0.0

(0)

Developer

Berkan Kaplan

Berkan Kaplan

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

8 days ago

Last modified

Categories

Share

Wikipedia Company Scraper 📚

🎉 Turn Wikipedia company articles into clean, structured data — no login, no API key, one row per company, with the summary, founded year, industry, headquarters, key people and links. Built for research, enrichment and knowledge-base building.

🔍 What is the Wikipedia Company Scraper — and when should you use it?

Give this actor company names and it returns matching companies from public Wikipedia company articles — as clean, deduplicated rows you can filter, export or feed to an AI agent. Every run queries the source live, so the data is as fresh as the registry itself.

Use it when you need: a company company list for outreach; formation / status monitoring; or a canonical registry record for KYB and due diligence.

Use something else when: you need a registry record — Wikipedia is an encyclopedia, not an official registry.

🤖 Use with AI agents

Already on the Apify MCP server? Ask for this Actor by name: foxlabs/wikipedia-company-scraper.

Your agent can pay for its own runs. This Actor is pay-per-event with agentic payments, so an agent can discover it, run it and settle the bill over x402 (USDC on Base) or Skyfire — no Apify account or API token of its own. Billing is the same either way: per delivered record, never for errors.

Otherwise paste this into Claude, ChatGPT, Cursor or any MCP-enabled assistant:

I want to pull company company records using the Apify Actor `foxlabs/wikipedia-company-scraper`.
Input: `queries` is a list of company names. `maxResultsPerQuery` caps rows per query.
Start with: {"queries":["undefined"],"maxResultsPerQuery":50}
Ask me what to look up, run the Actor, then summarise the rows as a table.

The machine-readable API, MCP config and OpenAPI definition live at apify.com/foxlabs/wikipedia-company-scraper.md.

📋 Overview

Everything you need to turn public Wikipedia company articles into clean, structured data — in one actor, with no login, cookies or API key.

Why teams pick this actor:

  • Whole source, one call — name or ID in, matching companies out.
  • 🧹 No empty-promise columns — only fields this registry actually fills; degenerate columns are removed.
  • 🔗 Stable identifiers — every row carries the source's own IDs, ready to join across runs and to other Fox Labs actors.
  • 💰 Pay only for results — per-row pricing, empty/failed lookups never billed.
  • 🤖 Agent-ready — MCP + x402 agentic payments.

✨ Features

  • 🔍 Name or ID lookup — relevance-ranked name search or exact registry-ID lookup.
  • 🏢 Full entity profile — status, legal form, formation date, address and the registry’s own contact fields.
  • 🧹 Clean schema — deduplicated camelCase rows, ready for CSV/Excel/JSON.

🎬 Quick Start

curl -X POST "https://api.apify.com/v2/acts/foxlabs~wikipedia-company-scraper/runs?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"queries":["undefined"],"maxResultsPerQuery":50}'

🚀 Getting Started (3 steps)

  1. Choose your targets — company names.
  2. Set the capmaxResultsPerQuery limits rows per query.
  3. Run and export — get a clean dataset as JSON, CSV or Excel.

📥 Input

{"queries":["undefined"],"maxResultsPerQuery":50}
FieldTypeDescription
queriesarrayCompany names.
maxResultsPerQueryintegerCaps rows per query.
maxConcurrencyintegerHow many queries to fetch at once.
includeRawbooleanAttach the source’s untouched record under raw.

📤 Output

One row per company, saved to the dataset. Every row also carries query, scrapedAt, and — when a lookup fails — an error explaining why (never silently dropped, never billed).

FieldDescription
companyNameCompany / entity name
wikipediaUrlWikipedia Url
summarySummary
wikidataQidWikidata Qid
wikidataUrlWikidata Url
officialWebsiteOfficial Website
leiLei
isinIsin
tickerSymbolTicker Symbol
stockExchangeStock Exchange
legalFormLegal form / entity type
industryPrimary industry (NACE)
countryCountry
headquartersHeadquarters
inceptionInception
employeesRegistered employee count
subsidiariesSubsidiaries
foundersFounders
pageIdPage Id
languageLanguage
licenseLicense
sourceSource
fetchedAtFetched At
infoboxInfobox
pageviewsMonthlyPageviews Monthly
pageviewsTotalPageviews Total
pageviewsLatestPageviews Latest

💼 Use cases

1. Company research — get a structured company overview. Input: company names. Output: summary + infobox facts. Use: a research brief.

2. Knowledge-base enrichment — attach Wikipedia facts to entities. Input: company names. Output: founded + industry + HQ. Use: an enriched KB.

3. Due diligence prep — get background on a company. Input: company names. Output: summary + key people. Use: a background sheet.

🔗 Integration

JavaScript / Node.js

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('foxlabs/wikipedia-company-scraper').call({"queries":["undefined"],"maxResultsPerQuery":50});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items[0]);

Python

from apify_client import ApifyClient
client = ApifyClient('YOUR_TOKEN')
run = client.actor('foxlabs/wikipedia-company-scraper').call(run_input={"queries":["undefined"],"maxResultsPerQuery":50})
for item in client.dataset(run['defaultDatasetId']).iterate_items():
print(item)

Automation (n8n / Zapier / Make): schedule or webhook → HTTP request to the actor API with your queries → handle the JSON dataset → push to a sheet, CRM or dashboard.

📊 Pricing

Pay-per-event: per delivered record. Empty or failed lookups are never billed. View current pricing.

❓ FAQ

Do I need an account, login or API key? No. This reads public Wikipedia company articles.

What do I search by? Company names.

How current is the data? Every run queries the source live, so results are as fresh as the registry.

What fields are returned? Summary, founded year, industry, headquarters, key people and infobox links from the Wikipedia article.

Can I export to CSV / Excel / JSON? Yes — directly from the Apify dataset.

🐛 Troubleshooting

  • Fewer rows than expected — raise maxResultsPerQuery, or refine the name.
  • A name returns an unexpected entity — it matched a similar registered name; search the exact registry ID.
  • No rows for a name — try the entity’s exact legal name or its registry ID.

This actor reads public Wikipedia company articles. Results can still contain personal data (e.g. a person’s name); personal data is protected by the GDPR and similar laws, so only process it with a legitimate basis. See Apify’s blog post on the legality of web scraping.

🤝 Support & contact

Changelog

0.2 — 2026-09-07

  • Dropped empty-promise columns. Removed parentOrganization — public Wikipedia company articles does not carry them, so they were shipped as always-null columns. Only fields this source actually fills are now emitted.
  • Enabled AI-agent payments (x402) + rebuilt the README to the full standard (What-is / when, AI-agents + x402 agentic payments + MCP, Overview, Features, Use cases, Integration, FAQ, Troubleshooting, Support & contact).

0.0

  • Initial release: data from public Wikipedia company articles by name or registry ID.