# Startups List Scraper (`crawlerbros/startups-list-scraper`) Actor

Scrape Startups List (startups-list.com) - a public directory of thousands of early-stage startups across 75+ cities worldwide. Browse by city and industry category to get name, tagline, description, website, logo, and category tags.

- **URL**: https://apify.com/crawlerbros/startups-list-scraper.md
- **Developed by:** [Crawler Bros](https://apify.com/crawlerbros) (community)
- **Categories:** Lead generation, Jobs, Other
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Startups List Scraper

Scrape [Startups List](https://www.startups-list.com) — a public, community-maintained directory of thousands of early-stage startups across 75+ cities worldwide (Austin, San Francisco, London, Berlin, Singapore, Tokyo, and more), plus two curated topic lists (Transportation, Writing Apps). Browse by city and, optionally, filter by industry category. HTTP-only, no login, no API key, no proxy required.

### What this actor does

- **Browse startups by city** — pick from 75+ cities across North America, Europe, Asia, Latin America, Africa, and Oceania
- **Filter by industry category** — 40+ curated categories (SaaS, FinTech, AI, Big Data, Health Care, E-Commerce, etc.), plus a free-text override for any of the platform's 700+ raw category slugs
- **Full startup profile per record** — name, tagline, description, website, logo, social links, and every industry tag
- **Sponsored listings flagged** — `isPromoted` marks paid placements so you can filter them out
- **Empty fields are omitted** — every record only contains fields that were actually extracted

### Output per startup

- `id` — Startups List internal ID
- `name` — startup name
- `tagline` — short one-line pitch (bolded lead-in on the site)
- `description` — longer description
- `website` — the startup's own website URL
- `categories[]` — industry tags (e.g. `SaaS`, `Big Data`, `Risk Management`)
- `imageUrl` — logo/thumbnail CDN URL
- `twitterUrl`, `facebookUrl`, `linkedinUrl`, `crunchbaseUrl`, `angelListUrl` — social/profile links, when present
- `city`, `citySlug` — the browsed city, human-readable and slug form
- `sourceUrl` — the Startups List listing page the record was scraped from
- `recordType: "startup"`, `scrapedAt`

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| `citySlug` | select | – (required, no default) | City (or topic list) to browse — required |
| `category` | select | `""` (all) | Curated common category filter |
| `customCategorySlug` | string | – | Raw category slug override, for categories outside the curated list |
| `maxItems` | int | `50` | Hard cap on emitted records (1–10000) |

#### Example: startups in Berlin

```json
{
  "citySlug": "berlin",
  "maxItems": 100
}
```

#### Example: SaaS startups in Austin

```json
{
  "citySlug": "austin",
  "category": "SaaS",
  "maxItems": 200
}
```

#### Example: a category outside the curated list, in London

```json
{
  "citySlug": "london",
  "customCategorySlug": "virtual_reality",
  "maxItems": 50
}
```

### Use cases

- **Market research** — map the startup landscape of a specific city or industry
- **Lead generation** — build outreach lists of early-stage startups by category
- **Investor sourcing** — scan multiple cities for startups in a target vertical
- **Competitive intelligence** — track which startups are active in a given space
- **Academic / journalism research** — analyze startup ecosystem density by geography

### FAQ

**What's Startups List?** A long-running, community-curated directory of local startup ecosystems, organized by city (startups-list.com), similar in spirit to F6S's company directory but with a simpler, fully public browsing surface that requires no account.

**Do I need an API key or login?** No — every listing page is publicly browsable HTML.

**Why is `category` a curated dropdown instead of the full list?** The site exposes 700+ raw category slugs, most highly specific and only present on a handful of city pages. The dropdown covers the ~40 most common, broadly-applicable categories; use `customCategorySlug` to target any other slug directly.

**What if a category doesn't exist for the chosen city?** The actor returns 0 records and sets a clear status message — this is expected behavior, not an error, since category availability varies by city.

**How many startups are in each city?** It varies widely — from a few dozen in smaller markets to 8,000+ in hubs like NYC and London. Use `maxItems` to cap large cities.

**Are all startups still active?** The directory includes both recently-added and long-standing entries; some older listings may reference companies that have since shut down or been acquired. All data reflects what is currently published on startups-list.com.

**What are the two "topic list" entries in `citySlug`?** `transport` and `writers` are curated cross-city lists (Transportation startups, Writing-app startups) served through the same URL pattern as city pages.

**Does the output ever include sponsored/ad placements?** No. Some city pages inject a small number of sponsored/paid-promotion cards alongside the real directory entries; the actor detects and excludes these unconditionally so every record you get back is a genuine entry from the curated startup directory.

# Actor input Schema

## `citySlug` (type: `string`):

Which city (or curated topic list) to browse startups from.

## `category` (type: `string`):

Only emit startups tagged with this industry category. Leave blank for all categories. For a category outside this curated list, use `customCategorySlug` instead.

## `customCategorySlug` (type: `string`):

Override `category` with any raw Startups List category slug not in the curated list above (e.g. `virtual_reality`, `dating`). Takes priority over `category` when set.

## `maxItems` (type: `integer`):

Hard cap on the number of startup records emitted. The largest cities (e.g. NYC, London) list 8,000+ startups on a single page, so this can go well above 3,000 if you want the full list.

## Actor input object example

```json
{
  "citySlug": "austin",
  "category": "",
  "maxItems": 50
}
```

# Actor output Schema

## `startups` (type: `string`):

Dataset containing all scraped startup records.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "citySlug": "austin",
    "category": "",
    "maxItems": 50
};

// Run the Actor and wait for it to finish
const run = await client.actor("crawlerbros/startups-list-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "citySlug": "austin",
    "category": "",
    "maxItems": 50,
}

# Run the Actor and wait for it to finish
run = client.actor("crawlerbros/startups-list-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "citySlug": "austin",
  "category": "",
  "maxItems": 50
}' |
apify call crawlerbros/startups-list-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=crawlerbros/startups-list-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ZboGkuMc0VcRNaA6j/builds/3xyehLtkxGTFdbauV/openapi.json
