# E-commerce Store Intelligence: Leads & Media Readiness (`toolsheder/shop-intel`) Actor

One row per online shop: platform, how well its catalogue is photographed and measured, static accessibility signals, product structured data, and contact details. Built to turn a list of domains into a list of prospects.

- **URL**: https://apify.com/toolsheder/shop-intel.md
- **Developed by:** [Kostas Skutulas](https://apify.com/toolsheder) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.00 / 1,000 analysed stores

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## E-commerce Store Intelligence: Leads, Media Readiness & Accessibility Signals

This actor takes a list of online store domains (Shopify, WooCommerce or another
platform) and returns one flat row per shop: the platform it runs on, how many
photos its products have and whether they state their dimensions, which
accessibility checks its HTML fails, whether its products have structured data,
and who to contact.

### Introduction

It is built for anyone who sells to online stores: accessibility remediation,
SEO, product photography, 3D and AR. It shows which shops have the problem you
fix before you spend a call finding out. The result is a list of online store
leads, one row per shop, where each column is a reason to contact the shop or
to skip it.

Different sellers work from different columns of the same run. Accessibility
agencies use `a11yChecksFailed`, SEO agencies use `schemaProductJsonLd`, and
anyone selling product media uses `mediaReadiness`. Everyone uses `emails`.

It analyses the shops you give it. It does not search for new ones.

### What it finds

- **Platform**: Shopify, WooCommerce, PrestaShop, Magento, OpenCart,
  BigCommerce, Wix, Squarespace and others, and whether the shop already runs a
  3D viewer such as model-viewer or Sketchfab (`has3dViewer`).
- **Product photos and dimensions**: the product count where the platform
  publishes one, the average number of photos per product, the share of
  products with three or more photos, the share that state their dimensions,
  and a 0 to 100 `mediaReadiness` score built from these.
- **Accessibility**: eight signals read from the HTML (page language, image
  alt text, form labels, a single h1, skip link, pinch zoom, generic link texts
  and page title) and `a11yChecksFailed`, the number that fail.
- **Product structured data**: whether the product page has Product JSON-LD
  (`schemaProductJsonLd`), and whether it gives a price, an image and a brand
  or product code (GTIN or MPN).
- **Contacts**: emails, with an MX check on the first one, phone numbers in
  international format and links to social profiles.
- **Status**: `ok`, `unreachable`, `blocked` or `not-a-shop`, plus `warnings`
  that say what could not be read and why.

### Tutorial

#### 1. Bring a list of domains

Anything with a host in it works: `ledinis.lt`, `https://www.shop.lt/kontaktai`,
`shop.lt/`. Everything after the host is dropped, because the host is what you
will join the results back onto.

#### 2. Run it

The defaults are set for a first pass: 120 products sampled per shop and a hard
limit of 25 requests per shop. Fifty shops take about a minute.

#### 3. Sort by the column you sell against

Use `mediaReadiness` if you sell product photography, 3D or AR,
`a11yChecksFailed` for accessibility work and `schemaProductJsonLd` for SEO.
Filter on `has3dViewer` to drop the shops that have already bought what you
sell.

#### 4. Check `status` before you trust a blank

A blank cell can mean two different things. Shops with the status
`unreachable` or `blocked` were never measured. An `ok` shop with an empty
`mediaReadiness` is one whose catalogue could not be read. The `warnings`
column says which.

### Measured on 50 Lithuanian shops

One run on real domains with the default settings. These are measured numbers
to plan a job with:

| | |
|---|---|
| Wall clock | about 57 s for 50 shops at concurrency 8 |
| Reached | 48 analysed, 1 unreachable, 1 blocked |
| Requests | 431 total, 8.6 per shop, 25 at the most (the cap) |
| Time per shop | 6.6 s median, 22.7 s at the 95th percentile |
| Contacts found | an email on all 48, every one with live MX; a phone on 45, every number valid |
| Media readiness scored | 45 of 48 |
| Accessibility | 2.4 of 8 signals failing on average; 11 shops failing 4 or more |
| Product structured data | 23 of 48 |
| Already running a 3D viewer | 0 of 48 |

Readiness across the 45 shops that could be scored: ten under 25, nine between
25 and 49, eleven between 50 and 74, fifteen at 75 or above. In the best shop
every sampled product had three or more photos and stated its dimensions. The
worst published no sizes at all.

Phone numbers are parsed with libphonenumber, using the country named by the
domain or by the page's language, and a number has to be written as a phone
number. An earlier version matched any eight digits after an 8, and a quarter
of what it returned were prices and product codes.

Three of the 48 could not be scored: no store API, no sitemap listing
products, and no product page among the links on the home page and its first
categories. Shops without a JSON catalogue are read from product pages, found
through the sitemap or, if there is none, through the shop's own links. A page
counts as a product only if it says so: structured data, Open Graph, or one
machine-readable price next to a basket button.

On nine of the shops read this way the photo count could not be read reliably
(one image in the structured data, and a gallery named by upload time). Their
image columns are left empty instead of showing 1, and their score is based on
dimensions.

Dimensions are looked for only in the text a shopper reads. In raw HTML, a
page's own stylesheet (`max-width: 900px; height: 1em`) counted as a stated
size, and an earlier version of this table was wrong because of it.

### Pricing

$5 per 1,000 shops analysed: $0.005 for each shop that was reached and read.
Unreachable and blocked domains, and sites that turn out not to be shops, are
still recorded free of charge, so the dead entries in an old list cost nothing. The fifty shops above come to $0.24 at
most. Apify adds its usual start fee of $0.00005 a run, and nothing else:
compute is included in the price.

### Input and output

#### Input

```json
{
  "domains": ["ledinis.lt", "https://www.hovden.lt", "sofaforma.lt"],
  "maxProductsPerStore": 120,
  "maxRequestsPerStore": 25,
  "concurrency": 8
}
```

#### Output

```json
{
  "domain": "sofaforma.lt",
  "finalUrl": "https://sofaforma.lt/",
  "status": "ok",
  "platform": "woocommerce",
  "has3dViewer": false,
  "productCount": 843,
  "productsSampled": 120,
  "estimated": true,
  "avgImagesPerProduct": 8.5,
  "pctWith3PlusImages": 96,
  "pctWithDimensions": 76,
  "mediaReadiness": 91,
  "schemaProductJsonLd": true,
  "a11yChecksFailed": 4,
  "emails": ["shop@sofaforma.lt"],
  "emailHasMx": true,
  "requestsMade": 4,
  "elapsedMs": 9639
}
```

`productCount` is exact only where the platform publishes a total, which in
practice means WooCommerce. Elsewhere it is `null`. `estimated` is `true`
whenever the percentages come from a sample. `mediaReadiness` gives a weight of
one half to the share of products with three or more photos, three tenths to
the share that state dimensions, and one fifth to the average photo count.

### Actor recommendations

**`a11yChecksFailed` does not measure compliance.** The European Accessibility
Act has applied to e-commerce since June 2025 and enforcement has begun, which
is why this column is worth money. Most of WCAG cannot be checked without a
rendered page, a keyboard and a person: contrast needs computed styles, focus
order needs the Tab key, and only a person can judge whether alternative text
describes its image. The column counts eight facts readable from the source,
which is where failures cluster. A shop failing five of them is not compliant,
and one failing none still needs a full audit.

**An empty `mediaReadiness` means the catalogue could not be read.** A 0 is a
score for a catalogue that was read. In the measured run, three of the 48
shops had no score. Sorting blanks as zeros would put the wrong shops at the
bottom of your list.

**Raise `maxRequestsPerStore` only for shops with no store API.** WooCommerce
and Shopify publish JSON catalogues, so those shops are done in four or five
requests whatever the limit. The limit only matters for other shops, where the
figures come from opening product pages one at a time.

**Do not raise `concurrency` to finish a big list sooner.** Requests within one
shop are sequential by design, with a pause between them. `concurrency` sets
how many different small shops are visited at the same time, and those shops
did not ask to be visited.

**Use `has3dViewer` to filter shops out.** A shop already running
model-viewer, Sketchfab or a paid viewer has bought what a 3D vendor sells.
Filtering those shops out is usually worth more than any ranking of the rest.

### FAQ and support

**Does it find shops as well as analyse them?**
No. It takes the list you give it. Store finders and lead databases already do
that job well; this actor fills in what their lists leave blank.

**Why does a shop show `blocked`?**
It answered with 401, 403, 407 or 429, or with a 503 from Cloudflare, usually because a firewall turns away
automated visitors. The row records the refusal, so a blocked shop is not
reported as unreachable.

**Does it respect robots.txt?**
Yes, including a group that names `shop-intel` specifically. robots.txt is
fetched outside the request budget, which only limits the pages taken from the
site.

**Why is `productCount` empty on a Shopify shop?**
Shopify's open catalogue endpoint lists products but publishes no total, so
there is no count to report. The sampled percentages are still valid.

### Related tools

- [Website Lead Qualifier](https://apify.com/toolsheder/website-lead-qualifier),
  for any business website or a Google Maps list: what to pitch, emails, tech
  stack.
- [Lead Qualifier Chrome extension](https://itneeds.lt/lead-qualifier/), the
  same kind of check for the one site open in your browser.
- [Photo to 3D](https://apify.com/toolsheder/photo-to-3d), product photos to
  web-ready 3D models at real size.
- Made by [ITneeds](https://itneeds.lt/).

# Actor input Schema

## `domains` (type: `array`):

One shop per line. Domains or full URLs both work; anything after the host is ignored, because the host is what you will join the results back onto. This actor analyses the shops you give it and does not search for new ones.

## `maxProductsPerStore` (type: `integer`):

How many products to read before the media figures are settled. Shops on WooCommerce and Shopify publish a JSON catalogue, so a higher number costs one more request per 100 products on WooCommerce and per 50 on Shopify; everywhere else the sample is twenty product pages regardless.

## `maxRequestsPerStore` (type: `integer`):

The hard limit on how many pages are fetched from one shop. It keeps the load on small shops low and does not make runs faster: requests within a store are always sequential, with a pause between them. Raising it gives deeper catalogue sampling on shops with no store API.

## `concurrency` (type: `integer`):

How many different shops to visit in parallel. Within each shop, requests are still made one at a time, with a pause between them.

## Actor input object example

```json
{
  "domains": [
    "ledinis.lt",
    "https://www.hovden.lt",
    "sofaforma.lt"
  ],
  "maxProductsPerStore": 120,
  "maxRequestsPerStore": 25,
  "concurrency": 8
}
```

# Actor output Schema

## `stores` (type: `string`):

One row per shop: platform, media readiness, accessibility signals, product structured data and contacts.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "ledinis.lt",
        "https://www.hovden.lt",
        "sofaforma.lt"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("toolsheder/shop-intel").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "domains": [
        "ledinis.lt",
        "https://www.hovden.lt",
        "sofaforma.lt",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("toolsheder/shop-intel").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "ledinis.lt",
    "https://www.hovden.lt",
    "sofaforma.lt"
  ]
}' |
apify call toolsheder/shop-intel --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,toolsheder/shop-intel"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/do5PSWYBk5hUrQwLm/builds/thfYazkDKIcWWee1D/openapi.json
