# DirectIndustry Scraper - Industrial Products & Specs (`crawloop/directindustry-scraper`) Actor

Scrape DirectIndustry industrial products for OEMs and buyers. Export title, model, manufacturer, technical characteristics, specs, images and PDF catalogs. Keyword, listing, stand or product URL — DirectIndustry scraper / API alternative.

- **URL**: https://apify.com/crawloop/directindustry-scraper.md
- **Developed by:** [Andrej Kiva](https://apify.com/crawloop) (community)
- **Categories:** E-commerce, Lead generation, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.40 / 1,000 product records

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## DirectIndustry Scraper — Industrial Products & Specs

> **Disclaimer:** Unofficial integration for publicly accessible sources. Trademarks belong to their respective owners. Provided for informational use only; users must comply with applicable platform terms and laws.

> **Crawloop B2B industrial data** — product catalogs (DirectIndustry) plus company directories (Europages / WLW).

| DirectIndustry (catalog) | Europages (EU directory) | WLW (DACH directory) |
| :--- | :--- | :--- |
| **DirectIndustry Scraper** ◄── you are here | [Europages Scraper](https://apify.com/crawloop/europages-scraper) | [WLW Scraper](https://apify.com/crawloop/wlw-scraper) |
| Products, tech specs, PDF catalogs, manufacturers | EU companies, VAT, contacts | DE / AT / CH suppliers |

**DirectIndustry scraper** for Apify — a practical **DirectIndustry API alternative** that turns industrial product listings into structured JSON. Extract **product title, model, manufacturer, technical characteristics, numeric specifications, images, linked PDF catalogs**, and **company websites** from keyword listings, category pages, manufacturer stands, or direct product URLs.

Built for **OEM competitive intelligence**, **supplier discovery by specs**, **BOM / sourcing research**, and **catalog monitoring**. Run from the Console, **Python**, **Node.js**, or **MCP**. Fast HTTP crawl via `curl_cffi` — parses VirtualExpo product payloads (no headless browser).

### Use cases

| Use case | What you get |
| :--- | :--- |
| **Product shortlists by type** | Keyword → industrial-manufacturer listing → product rows |
| **Tech-spec comparison** | Characteristics (Technology, Medium, ATEX…) plus min/max specs |
| **Manufacturer catalog pull** | All products on a stand URL with optional full PDP enrichment |
| **PDF catalog harvest** | Linked datasheet / brochure catalog titles and viewer URLs |
| **Website enrichment** | External manufacturer website + off-platform product link |
| **Category deep-dive** | `/cat/` pages expand into child product-type listings |

### When to use this Actor

- You need **DirectIndustry product data** as dataset rows (not just company contacts)
- You have **keywords**, **listing URLs**, **manufacturer stands**, or **product PDPs**
- You want **technical characteristics and PDF catalogs** alongside titles
- You prefer a **browser-free** crawl on Apify

### When not to use this Actor

- **Guaranteed live prices / stock** — most listings are RFQ / price-on-request
- **Sending RFQs** through the portal contact form — this Actor is read-only extraction
- **Company-directory firmographics (VAT, phone)** — use Europages or WLW instead
- **Authenticated MySpace-only fields** — public pages only

### Key features

- **Keyword search** — resolves via DirectIndustry kwref sitemaps to listing URLs
- **Listing & category URLs** — paginated `industrial-manufacturer` pages; `/cat/` expands to children
- **Manufacturer stands** — crawl all product cards on a company stand
- **Product detail enrichment** — `fetchDetails` parses `__preloadData__` for full specs
- **VirtualExpo portal switch** — `portal` input prepares MedicalExpo / AeroExpo / … hosts
- **Streaming results** — dataset rows appear while the run is in progress
- **Deduped push** — unique by `portal` + `productId` within a run
- **Pagination controls** — `maxPages` + `maxItems` for predictable run size
- **Lightweight & resilient** — Chrome TLS fingerprinting, proxy session rotation on WAF challenges

### Input parameters

| Parameter | Type | Default | Description |
| :--- | :--- | :--- | :--- |
| `searchKeywords` | Array | `["float level sensor"]` | Keywords → kwref listing URLs. |
| `startUrls` | Array | `[]` | Mixed product / manufacturer / listing / category URLs. |
| `listingUrls` | Array | `[]` | industrial-manufacturer and/or `/cat/` URLs. |
| `productUrls` | Array | `[]` | Direct product detail URLs. |
| `manufacturerUrls` | Array | `[]` | Manufacturer stand URLs. |
| `portal` | String | `"directindustry"` | VirtualExpo host (`directindustry`, `medicalexpo`, …). |
| `fetchDetails` | Boolean | `true` | Open PDPs for specs, description, images, catalogs. |
| `maxItems` | Integer | `50` | Max dataset rows (`0` = unlimited within `maxPages`). |
| `maxPages` | Integer | `3` | Max listing pages per list URL. |
| `concurrency` | Integer | `3` | Parallel PDP workers (1–15). |
| `proxyConfiguration` | Object | residential | Apify Proxy settings (residential recommended). |

#### Example — keyword product crawl

```json
{
  "searchKeywords": ["float level sensor"],
  "fetchDetails": true,
  "maxItems": 50,
  "maxPages": 3,
  "concurrency": 3,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": ["RESIDENTIAL"]
  }
}
```

#### Example — manufacturer stand + product URLs

```json
{
  "manufacturerUrls": [
    { "url": "https://www.directindustry.com/prod/flygt-113401.html" }
  ],
  "productUrls": [
    { "url": "https://www.directindustry.com/prod/flygt/product-113401-1101505.html" }
  ],
  "fetchDetails": true,
  "maxItems": 100,
  "concurrency": 3
}
```

### Output

Each dataset item is one industrial product.

| Field | Description |
| :--- | :--- |
| `title` / `model` | Product label and model designation |
| `companyName` / `companyId` | Manufacturer stand name and id |
| `url` / `companyUrl` | Product PDP and manufacturer stand URLs |
| `companyWebsite` | External manufacturer website when published |
| `features` | Characteristic rows (Technology, Medium, ATEX, …) |
| `specifications` | Numeric / range specs (`min` / `max` / `raw`) |
| `images` | Product image URLs |
| `catalogs` | Linked PDF catalog title, URL, pages, language |
| `description` | Full product description from the detail page |
| `category` / `breadcrumbs` | Listing category and navigation path |
| `enriched` | `true` when fields come from a product detail page |

Example (illustrative):

```json
{
  "recordType": "product",
  "portal": "directindustry",
  "productId": "1101505",
  "companyId": "113401",
  "title": "Float level sensor",
  "model": "ENM 10",
  "companyName": "FLYGT",
  "groupCompanyName": "Xylem",
  "url": "https://www.directindustry.com/prod/flygt/product-113401-1101505.html",
  "companyUrl": "https://www.directindustry.com/prod/flygt-113401.html",
  "companyWebsite": "https://www.xylem.com/en-us/brands/flygt",
  "description": "PRODUCT FEATURES …",
  "category": "Float level sensor",
  "features": [
    { "name": "Technology", "value": "float" },
    { "name": "Other characteristics", "value": "ATEX, IECEx" }
  ],
  "specifications": [
    {
      "name": "Process temperature",
      "min": "Min.: 0 °C (32 °F)",
      "max": "Max.: 60 °C (140 °F)"
    }
  ],
  "images": ["https://img.directindustry.com/images_di/photo-g/113401-14661051.jpg"],
  "catalogs": [
    {
      "title": "Flygt C-pumps 3068–3800",
      "url": "https://pdf.directindustry.com/pdf/flygt/flygt-c-pumps-3068-3800/113401-865179.html",
      "pages": 8
    }
  ],
  "enriched": true,
  "scrapedAt": "2026-08-05T12:00:00Z"
}
```

### Integration examples

#### Node.js

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('crawloop/directindustry-scraper').call({
  searchKeywords: ['float level sensor'],
  fetchDetails: true,
  maxItems: 50,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.slice(0, 5));
```

#### Python

```python
from apify_client import ApifyClient

client = ApifyClient(token)
run = client.actor("crawloop/directindustry-scraper").call(
    run_input={
        "searchKeywords": ["float level sensor"],
        "fetchDetails": True,
        "maxItems": 50,
    }
)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item.get("title"), item.get("model"), item.get("companyName"))
```

#### cURL

```bash
curl "https://api.apify.com/v2/acts/crawloop~directindustry-scraper/runs?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"searchKeywords":["float level sensor"],"fetchDetails":true,"maxItems":50}'
```

### MCP and AI assistants

Use this Actor from AI tools via [Apify MCP](https://docs.apify.com/platform/integrations/mcp). Connect your Apify account, then call `crawloop/directindustry-scraper`.

Example prompts:

- "Run DirectIndustry Scraper for keyword float level sensor, max 30, return title, model, companyName, features"
- "Scrape products from a DirectIndustry manufacturer stand URL and summarize ATEX-related specs"
- "Chain DirectIndustry Scraper then Europages Scraper to map product catalogs to EU company contacts"

### Suite next step

For Europe-wide **company contacts / VAT / firmographics**, run [Europages Scraper](https://apify.com/crawloop/europages-scraper). For DACH-only suppliers, use [WLW Scraper](https://apify.com/crawloop/wlw-scraper).

### Related Actors

| Actor | Use for |
| :--- | :--- |
| **DirectIndustry Scraper** ◄── you are here | Industrial product catalog, specs, PDF catalogs |
| [Europages Scraper](https://apify.com/crawloop/europages-scraper) | Europe-wide B2B directory, multi-locale, VAT & contacts |
| [WLW Scraper](https://apify.com/crawloop/wlw-scraper) | DACH (DE / AT / CH) B2B suppliers from Wer liefert was |

### FAQ

**Is this a DirectIndustry API?**\
No official public product API is required. This Actor is a **DirectIndustry scraper / API alternative** that returns structured dataset rows you can call from Python, Node.js, cURL, or MCP.

**Does keyword search need exact listing URLs?**\
No — keywords are matched against DirectIndustry kwref sitemaps (e.g. `pump` → pump listing). Prefer exact listing URLs when you already have them.

**Can I scrape MedicalExpo with the same Actor?**\
Set `portal` to `medicalexpo` (and sister portals). Page patterns are shared across VirtualExpo; DirectIndustry is the primary tested host.

**Why are some rows missing phone/email?**\
This Actor targets **product catalog** fields. Public RFQ flows do not expose manufacturer phones on product pages the way Europages company profiles do.

# Actor input Schema

## `searchKeywords` (type: `array`):

Product-type keywords resolved via DirectIndustry kwref sitemaps to industrial-manufacturer listing pages (e.g. "float level sensor", "pump", "pressure transmitter").

## `startUrls` (type: `array`):

Mix of product PDPs, manufacturer stands, industrial-manufacturer listings, and/or category pages.

## `listingUrls` (type: `array`):

industrial-manufacturer listing pages and/or /cat/ category pages (categories expand to child listings).

## `productUrls` (type: `array`):

Direct product detail URLs (/prod/{brand}/product-{companyId}-{productId}.html).

## `manufacturerUrls` (type: `array`):

Manufacturer stand URLs (/prod/{brand}-{companyId}.html) — expands to product cards on the stand.

## `portal` (type: `string`):

Which VirtualExpo marketplace host to use. DirectIndustry is the primary catalog; sister portals share the same page patterns.

## `fetchDetails` (type: `boolean`):

When true, open each product PDP and parse window.**preloadData** for full specs, description, images and PDF catalogs. When false, emit listing-card fields only.

## `maxItems` (type: `integer`):

Hard cap on dataset rows (0 = unlimited within maxPages).

## `maxPages` (type: `integer`):

Max pagination pages per industrial-manufacturer listing URL.

## `concurrency` (type: `integer`):

Parallel PDP workers (1–15). Keep low if Cloudflare challenges appear.

## `proxyConfiguration` (type: `object`):

Apify Proxy settings. Residential recommended.

## Actor input object example

```json
{
  "searchKeywords": [
    "float level sensor"
  ],
  "startUrls": [],
  "listingUrls": [],
  "productUrls": [],
  "manufacturerUrls": [],
  "portal": "directindustry",
  "fetchDetails": true,
  "maxItems": 50,
  "maxPages": 3,
  "concurrency": 3,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `results` (type: `string`):

Default dataset items (product records).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("crawloop/directindustry-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("crawloop/directindustry-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call crawloop/directindustry-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,crawloop/directindustry-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/z5AuLLUmpD3ZNBurP/builds/nwWmvcwMhA74xgqUj/openapi.json
