# LinkedIn Company Scraper — Public Pages, No Login (`scrapersdelight/linkedin-company-scraper`) Actor

Bulk-scrape public logged-out LinkedIn company pages: industry, company size band, HQ and postal address, founded year, specialties, website, company type, followers and employees-on-LinkedIn. Companies, not people. No account and no cookie. $1.50 per 1,000 companies — the cheapest in the lane.

- **URL**: https://apify.com/scrapersdelight/linkedin-company-scraper.md
- **Developed by:** [Scrapers Delight](https://apify.com/scrapersdelight) (community)
- **Categories:** Lead generation, Business, Social media
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.50 / 1,000 per company returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## LinkedIn Company Scraper — Public Pages, No Login

Paste LinkedIn company URLs or slugs, get structured **company** records back:
**companyName, industry, companySize, employeesOnLinkedIn, headquarters, streetAddress,
addressLocality, addressRegion, postalCode, addressCountry, foundedYear, specialties,
website, companyType, followerCount, tagline, description, logoUrl** and the stable
numeric **companyId**.

Three things define this Actor, and they are the first thing on the page on purpose:

1. **It reads only the public, logged-out company page.** Exactly what a signed-out
   browser is served at `https://www.linkedin.com/company/<slug>/`. There is **no LinkedIn
   account, no member cookie, no `li_at`, no session token** anywhere in this Actor — not
   as a required field, not as an optional field, not as a documented "advanced" path.
   Nothing of yours can be rate-limited, restricted or banned, because nothing of yours is
   used.
2. **It returns companies, not people.** No employee names, no profiles, no staff rosters.
   `companySize` is the published band LinkedIn renders (`1,001-5,000 employees`). LinkedIn
   localises to the proxy exit IP, so the band and `industry` occasionally come back in
   another language (`11-50 працівників`) — filter on the numeric `companySizeMin` /
   `companySizeMax`, never on the band string. `employeesOnLinkedIn` is the count LinkedIn
   shows a signed-out visitor
   (*"View all 2,942 employees"*). Those are two numbers. They are not a list of humans.
3. **It does not search or discover LinkedIn.** You supply the companies. A multi-word
   company *name* is rejected by the input, because resolving a name means running LinkedIn
   search — the auth-gated part of the site, and the part that carries the real risk. Note
   that a **one-word** entry cannot be told apart from a slug and is fetched as one: typing
   `twilio` gets you whoever holds the `twilio` vanity slug, which may not be the company
   you meant. Paste the URL from the company's own LinkedIn page when you want certainty.

```json
{
  "companies": [
    "stripe",
    "https://www.linkedin.com/company/shopify/",
    "hubspot",
    "notionhq",
    "figma"
  ]
}
```

Click **Try for free** and hit **Start** — that block is literally the input this Actor
ships with. It returned **5 of 5 companies** and cost **$0.0075**.

**$0.0015 per company returned — $1.50 per 1,000.** A company that 404s, that fails every
proxy rung, or that duplicates a company already in your results is **never billed**.

***

### Read this before you buy rows

Five things that would otherwise turn into a refund request.

1. **LinkedIn's `robots.txt` disallows this path, and every other path.** Quoted verbatim
   below. It is not ambiguous and this page does not pretend otherwise — read it and make
   your own call before you run anything.
2. **`foundedYear` is present on about 56% of companies and `specialties` on about 82%** —
   and `foundedYear` collapses to 26% on a list of mega-caps (see the footnote below the
   table). These are optional fields a company fills in or does not. The full measured fill
   table is further down — it is sorted so the gaps are impossible to miss.
3. **`website` is whatever the company typed into LinkedIn's Website box, not their
   corporate domain.** Figma's currently reads `https://figma.bot/Config2026Recap` — a
   campaign link. Most companies put their homepage there; some do not. Same for
   `description`: it is the About text as written, and some companies use it for a
   campaign line rather than a company blurb.
4. **`employeesOnLinkedIn` and `companySize` disagree, and both are correct.**
   `companySize` is the self-declared band; `employeesOnLinkedIn` is how many LinkedIn
   members currently list that company. Anthropic reads `501-1,000 employees` with 5,700
   members. Use the band for firmographics and the count for reach.
5. **A retired vanity slug follows LinkedIn's redirect — check what it lands on.** When
   LinkedIn 301s a slug and then answers HTTP 999, this Actor re-fetches under the slug it
   was sent to and records the original in `redirectedFromSlug`. Sometimes that is a true
   rename (`/company/mailchimp` → `intuitmailchimp`, `/company/monzo-bank` → `monzo`).
   Sometimes the vanity slug now belongs to an unrelated company: `/company/gitlab` lands on
   `domifyio` ("Domify", 2-10 employees), `/company/basecamp` on `basecamppt` (a Portuguese
   agency), `/company/uber` on `ubercreativedigitalagency`. **Any row with a non-null
   `redirectedFromSlug` is billed and should be reviewed before you trust it** — filter on
   that field. Use the exact slug from the company's own LinkedIn URL to avoid this
   entirely. A slug that simply does not exist returns a logged 404 and no charge.

***

### robots.txt — the actual lines, unedited

Fetched from `https://www.linkedin.com/robots.txt` on **2026-08-15** (HTTP 200, 120,190
bytes). The file contains one `User-agent: *` group. It is the last group in the file, and
this is it in full:

```
User-agent: *
Disallow: /

## Notice: If you would like to crawl LinkedIn,
## please email whitelist-crawl@linkedin.com to apply
## for white listing.
```

**`Disallow: /` covers `/company/<slug>/`, the only path this Actor requests.** That
`User-agent: *` group contains no `Allow:` line at all, so nothing carves the path back out.
The file's only company-specific rules are `Disallow: /companyDir*`,
`Disallow: /pages-extensions/FollowCompany*` and a single `Allow: /companyDir*` — all of them
sit in *other, named* crawler groups, none of them names a path this Actor touches, and none
of them applies to a generic client, which falls into the catch-all group above.

So, stated plainly: **this Actor requests a path that LinkedIn's robots.txt disallows for
any crawler that is not on their whitelist.** It fetches only pages LinkedIn serves publicly
to a signed-out visitor, it does not log in, and it does not defeat any challenge — but it
is not robots-compliant, and LinkedIn's User Agreement prohibits automated collection
regardless of technical access. Whether that is acceptable is your decision, in your
jurisdiction, for your use case. Do not run this if the answer is no.

***

### What you get

One row per unique company. `scrapedAt` is a full UTC timestamp.

| Group | Fields | Example |
|---|---|---|
| **Identity** | `companyName`, `companyUrl`, `companySlug`, `companyId`, `inputSlug`, `redirectedFromSlug` | `Stripe` · `https://www.linkedin.com/company/stripe/` · `2135371` |
| **Positioning** | `tagline`, `description`, `industry` | `Help increase the GDP of the internet.` · `Technology, Information and Internet` |
| **Size** | `companySize`, `companySizeMin`, `companySizeMax`, `employeesOnLinkedIn` | `5,001-10,000 employees` · `5001` · `10000` · `17163` |
| **Location** | `headquarters`, `streetAddress`, `addressLocality`, `addressRegion`, `postalCode`, `addressCountry` | `South San Francisco, California` · `354 Oyster Point Blvd` · `94080` · `US` |
| **Firmographics** | `foundedYear`, `companyType`, `website`, `specialties`, `specialtiesCount` | `2010` · `Privately Held` · `https://stripe.com` |
| **Reach** | `followerCount`, `logoUrl` | `1,663,354` |

`companySizeMin` / `companySizeMax` are the band parsed into numbers so you can filter
without string-matching (`10,001+ employees` gives min `10001`, max `null`).

The dataset ships with a saved **table view** — *Companies* — so you do not have to
configure columns.

***

### Field fill — measured on 66 live companies

Every company this Actor has returned on the Apify platform: 66 distinct `companyId`s
across 55 runs, Apify **datacenter** proxies, 2026-08-15 → 2026-08-27. Sorted by fill, so
the sparse fields are impossible to miss.

| Field | Fill | Notes |
|---|---|---|
| `companyName` | **100%** | |
| `companyUrl` / `companySlug` | **100%** | |
| `companyId` | **100%** | the numeric `urn:li:organization:` id — stable across renames |
| `logoUrl` | **100%** | 200×200 CDN URL |
| `industry` | **100%** | LinkedIn's own taxonomy string |
| `companySize` / `companySizeMin` | **100%** | the published band |
| `employeesOnLinkedIn` | 98.5% | members listing this employer |
| `website` | 98.5% | see caveat 3 above |
| `companyType` | 98.5% | `Public Company`, `Privately Held`, `Nonprofit`, … |
| `followerCount` | 98.5% | |
| `description` | 97.0% | |
| `addressCountry` | 95.5% | |
| `headquarters` | 93.9% | city + region as LinkedIn prints it |
| `addressLocality` | 93.9% | |
| `specialties` | 81.8% | averaged 9.4 entries where present |
| `streetAddress` | 80.3% | |
| `postalCode` | 78.8% | |
| `addressRegion` | 77.3% | |
| `foundedYear` | 56.1% | optional field; many companies leave it blank |
| `tagline` | 53.0% | the one-line strapline under the company name |
| `companySizeMax` | 47.0% | `null` for `10,001+` bands by definition |

**Uniqueness: `companyId` is unique in every dataset — 27/27, 27/27, 18/18 on the largest
three runs.** Nothing was double-billed.

**Fill depends on company size — read this before you budget for the optional fields.**
The table above spans a *mix* of company sizes. The optional profile fields are the ones the
biggest companies most often leave blank, so an enterprise-only list fills far thinner than
that overall average. Measured on a 27-company mega-cap run (2026-08-19): `foundedYear`
**25.9%** (vs 56.1% overall) and `tagline` **14.8%** (vs 53.0% overall). `companySizeMax` is
`null` on **100%** of them, because `10,001+ employees` has no upper bound to extract.
The always-present fields — name, id, url, logo, industry, size band, type, followers,
employees-on-LinkedIn — are unaffected: **100%** on that same mega-cap run.

***

### How it works, and what it does not do

The signed-out company page server-renders everything above, in three places on the same
document:

- the **About us `<dl>`** — industry, size band, headquarters, type, founded, specialties,
  website;
- a **JSON-LD `Organization` node** — description, slogan, the full postal address, logo,
  employee count;
- the **top card** — name, tagline, follower count, and the employee count.

That is one HTTP GET per company on the happy path — up to four retries with a rotated
session, plus one capped residential pass, when LinkedIn serves the sign-in wall instead.
No browser, no Chromium, no JavaScript execution, no CAPTCHA solving, no challenge bypass —
and no login, which is the point.

`/company/<slug>/about/` is **not** used: for a signed-out visitor it 302s to
`/uas/login` (verified 2026-08-15 on `stripe`, `ibm` and `hubspot`). Everything this Actor
returns comes off the base page instead. If a page will not serve without signing in, the
Actor logs it and returns nothing for that company rather than reaching for a credential.

#### Transport, measured through Apify proxies (not a home IP)

| Rung | Sample | Result |
|---|---|---|
| Apify **DATACENTER**, rotating session | 60 slugs queued, 2026-08-20 (runs `6CQI5yHkcVYszpaob`, `ieITTQwtuuYNNW5CF`) | every company attempted inside the run's time budget returned a row; the runs stopped on the wall clock at 29 and 18 of 60 attempted, not on blocks |
| Apify **RESIDENTIAL**, one sticky session per company | bounded fallback only — never exercised on any billed run to date | — |

Residential is not the default: it exists only as a bounded fallback for the day LinkedIn
walls the shared datacenter pool, capped at `max(3, 10% of your list)` per run, and every
run logs how much of the cap it used.

***

### Pricing, and how it compares

**$0.0015 per company returned. $1.50 per 1,000.** Nothing else is charged — no run-start
fee, no per-dataset-item fee.

Checked live against the Apify Store on 2026-08-15:

| Actor | Price per company |
|---|---|
| **This Actor** | **$0.0015** |
| `datadoping/linkedin-company-scraper` | $0.00155 |
| `automation-lab/linkedin-company-scraper` | $0.00345 + $0.005 per run |
| `harvestapi/linkedin-company` (lane leader) | $0.004 + run-start |
| `data-slayer/linkedin-company-scraper` | $0.004 + run-start |
| `unseenuser/LinkedIn-Company-Scraper` | $0.004 |
| `scraper-engine/linkedin-company-about-scraper` | $0.00499 + run-start |
| `scrapeverse/linkedin-company-profile-scraper-pay-per-event` | $0.006 + run-start |

**An honest note on "no cookies":** most of the company-detail Actors in this lane also run
without a member cookie — it is table stakes here, not a unique feature, and you should
discount anyone selling it as one (including this page). What is *not* table stakes is the
rest of it: a published per-field fill table measured on 66 live companies, every redirect
disclosed in `redirectedFromSlug`, de-duplication on the numeric company id **before**
anything is billed, and the robots.txt above quoted in full rather than summarised.

Cookie-based tooling (browser extensions, PhantomBuster-style session hijacking) is a
different category and does carry account risk. This Actor is not in that category.

***

### Input

| Field | What it does |
|---|---|
| `companies` | One company per line: a slug (`stripe`), a full company or showcase URL, a country-subdomain URL (`uk.linkedin.com/company/…`), or the bare `linkedin.com/company/x` form. |
| `startUrls` | Bulk path: paste many URLs, upload a `.txt`/`.csv`, or link a Google Sheet. Entries must be full LinkedIn company URLs. Merged with `companies` and de-duplicated. |
| `proxyConfiguration` | Leave on the default (Apify datacenter + capped residential fallback). |
| `requestConcurrency` | Default 5. Measured on the platform at this setting: 12 companies in 8.9 s (run `Fd9jLQe6qMfI9gfpZ`), 5 companies in 6.9 s. |
| `requestDelayMs` | Default 400 ms per worker. |

Both input fields are merged and de-duplicated **before** any request is made, so pasting
`stripe` and `https://www.linkedin.com/company/stripe/` in the same run costs you one row,
not two.

### Output

Every company that returns is one dataset item, pushed and charged in the same call, so a
charge cap can never leave you with unpaid rows.

If you start the Actor with **no input at all**, it scrapes the documented 5-company sample
above so you can see the output shape, and says so in the run status — you are billed for
those 5 rows and nothing else. If a run returns **nothing** (every company 404'd or every
proxy rung was walled), it ends with a status message asking you to re-run, and **nothing is
billed**.

### Limits

- Public company pages only. Anything a signed-out visitor cannot see is out of scope, permanently.
- No employees, no people, no posts, no jobs, no follower lists.
- No search, no discovery, no enumeration — you bring the companies.
- Optional profile fields (`specialties`, `foundedYear`, `streetAddress`, `tagline`) are
  blank when the company left them blank. See the fill table.
- LinkedIn's robots.txt disallows this path. See the section above; it is your call.

# Actor input Schema

## `companies` (type: `array`):

One company per line. Accepts a bare slug (`stripe`), a full company-page URL (`https://www.linkedin.com/company/shopify/`), a showcase-page URL, a country-subdomain URL (`https://uk.linkedin.com/company/…`) or the bare `linkedin.com/company/x` form. A company **name** is rejected on purpose — resolving a name means searching LinkedIn, which this Actor never does. You are billed per company **returned**, so this list is also your cost ceiling: 1,000 companies = $1.50. Duplicate lines and two aliases of the same company collapse to a single billed row.

## `startUrls` (type: `array`):

The bulk path: paste many company-page URLs at once, upload a .txt/.csv of them, or link a Google Sheet. Entries must be full LinkedIn company or showcase URLs (`https://www.linkedin.com/company/<slug>`) — a bare slug inside a file is ignored, because a spreadsheet column of words cannot be told apart from data. Merged with the field above and de-duplicated. Reading the list is free; you are billed only for companies returned.

## `proxyConfiguration` (type: `object`):

Leave this on the default. The Actor runs on standard Apify DATACENTER proxies (measured 47 of 50 varied company slugs returned usable rows, 2026-08-15) and, only if LinkedIn walls the datacenter pool, retries the walled companies through RESIDENTIAL — hard-capped at max(3, 10% of your list) per run. Residential measured no better than datacenter on this surface, so switching this to Residential yourself usually just costs more.

## `requestConcurrency` (type: `integer`):

How many company pages to fetch at once. The default of 5 measured 10 companies in 3.6 seconds. Raising it shortens the run but gives LinkedIn more reason to serve the sign-in wall, and a walled page returns nothing — so it costs you rows, not money.

## `requestDelayMs` (type: `integer`):

Pacing per worker, applied after each company. 0 is allowed and is faster; the default 400 ms is the pacing the measured success rate was taken at.

## Actor input object example

```json
{
  "companies": [
    "stripe",
    "https://www.linkedin.com/company/shopify/",
    "hubspot"
  ],
  "startUrls": [],
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "requestConcurrency": 5,
  "requestDelayMs": 400
}
```

# Actor output Schema

## `companies` (type: `string`):

Company name, industry, size band, HQ and postal address, founded year, specialties, website, type, followers and employees-on-LinkedIn.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "companies": [
        "stripe",
        "https://www.linkedin.com/company/shopify/",
        "hubspot",
        "notionhq",
        "figma"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapersdelight/linkedin-company-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "companies": [
        "stripe",
        "https://www.linkedin.com/company/shopify/",
        "hubspot",
        "notionhq",
        "figma",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("scrapersdelight/linkedin-company-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "companies": [
    "stripe",
    "https://www.linkedin.com/company/shopify/",
    "hubspot",
    "notionhq",
    "figma"
  ]
}' |
apify call scrapersdelight/linkedin-company-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapersdelight/linkedin-company-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/zMT1MEKCu8YkAspYS/builds/r73Bw7g8mi1fTfCSd/openapi.json
