# FDA Recalls API & CPSC Recalls Scraper - NHTSA Car Recalls (`snow_leo_data/product-recalls-scraper`) Actor

127,501 recall records from five US feeds in one 53-field schema. openFDA skip stops at 25,000; this pages by cursor and reads the whole archive. Medical device recalls, drug recalls, food recalls, CPSC consumer product recalls and NHTSA vehicle recalls, with hazard labels.

- **URL**: https://apify.com/snow\_leo\_data/product-recalls-scraper.md
- **Developed by:** [Snow Leo Data](https://apify.com/snow_leo_data) (community)
- **Categories:** News, E-commerce
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.75 / 1,000 recalls

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## US Product Recalls: FDA + CPSC + NHTSA in one dataset

Five official United States recall feeds, collected in a single run and returned
as one table with one schema: **FDA food enforcement**, **FDA drug enforcement**,
**FDA medical device enforcement**, **CPSC consumer product recalls** and
**NHTSA vehicle, tyre, equipment and child seat campaigns**.

Every number on this page comes from a live request you can repeat yourself. The
five feeds held **127,501 records** when this Actor was last measured on
2026-09-11:

| Feed | Agency API | Records |
|---|---|---|
| Food | `api.fda.gov/food/enforcement.json` | 29,386 |
| Drug | `api.fda.gov/drug/enforcement.json` | 17,938 |
| Medical device | `api.fda.gov/device/enforcement.json` | 39,885 |
| Consumer products | `saferproducts.gov/RestWebServices/Recall` | 10,002 |
| Vehicles and equipment | `datahub.transportation.gov` dataset `6axg-epim` | 30,290 |

No proxy, no browser, no API key: all five endpoints answer a plain HTTPS
request with JSON, and this Actor uses nothing but the Python standard library
to read them.

#### What problem does this Actor actually solve?

Product safety data in the United States is split across three agencies that
never agreed on a format. The FDA calls a recall an *enforcement report* and
grades it Class I, II or III. The CPSC publishes a press release with a list of
products, hazards, injuries, remedies and photographs. The NHTSA publishes a
campaign with a defect summary, a consequence summary and the number of vehicles
potentially affected. The field names overlap in none of the three.

If you are a compliance team, a marketplace, an insurer or an importer, you do
not care which agency publishes the row. You care whether anything you sell,
carry, insure or drive has been recalled. That means one table, one date field,
one company field, one severity signal, and a way to ask "what changed since
yesterday" without paying for the whole archive again.

That is what this Actor does: it normalises all five feeds into a single
**53-field** schema, tags every row with derived hazard and category labels,
scores severity with fixed rules, and remembers what it already delivered so a
daily schedule only bills you for what is new.

#### How does it get past the openFDA 25,000 record limit?

This is the measurable thing this Actor does that a straightforward openFDA
client cannot.

openFDA paginates with `skip`, and `skip` is hard-capped:

```
GET https://api.fda.gov/device/enforcement.json?limit=1&skip=25001
HTTP 400  {"error":{"code":"BAD_REQUEST","message":"Skip value must 25000 or less."}}
```

The device enforcement feed alone holds 39,885 records. An actor that pages with
`skip` therefore cannot reach roughly **14,885 of them** — about 37% of that
feed — no matter how long it runs.

openFDA also returns a cursor in the `Link` response header:

```
Link: <https://api.fda.gov/device/enforcement.json?limit=1000&sort=report_date%3Aasc
       &skip=0&search_after=0%3D1340150400000%3B1%3D0f251e78...>; rel="next"
```

This Actor follows that header instead of counting `skip`. Measured on
2026-09-11: 26 consecutive pages of 1,000 records, **26,000 unique records**
collected in 121 seconds, with a further `rel="next"` cursor still waiting — 1,000
records past the point where `skip` refuses to answer at all. The live test
`tests/test_live.py::test_cursor_walks_past_the_cap` reproduces exactly that walk
on every run and fails if the cursor ever stops working.

The other two feeds have no comparable ceiling, and the Actor proves it rather
than assuming it: the CPSC test sums three date windows and checks the total
against one unbounded request (10,002 either way), and the NHTSA test compares
the claimed total against a live `count(1)` from the dataset.

#### What comes back in each row?

53 fields, the same names whichever agency the row came from. Empty means the
agency does not publish that value for that row — nothing is inferred or filled
in by a language model.

**Identity.** `recall_id`, `source`, `agency`, `product_type`, `recall_number`,
`event_id`.

**What was recalled and by whom.** `title`, `product_description`, `brands`,
`models`, `product_upcs`, `product_quantity`, `units_affected` (parsed to a real
integer, so `About 179,739` becomes `179739`), `company`, `company_city`,
`company_state`, `company_country`, `manufacturer_countries`, `retailers`,
`importers`, `distributors`.

**Why, and how dangerous.** `reason`, `hazards`, `hazard_tags`, `consequence`,
`injuries`, `injury_reported`, `classification`, `severity_score`,
`severity_label`, `category`, `do_not_drive`, `fire_risk_when_parked`,
`component`.

**What owners should do, and where it stands.** `status`, `remedy`,
`consumer_contact`, `voluntary_mandated`, `initial_firm_notification`,
`distribution_pattern`, `distribution_states`. Only the FDA publishes a lifecycle
status of its own (Ongoing, Completed, Terminated); CPSC rows are labelled
`Announced` and NHTSA rows `Open` by this Actor so that one column means
something in every row.

**Dates, all normalised to `YYYY-MM-DD`.** `recall_date`, `report_date`,
`recall_initiation_date`, `center_classification_date`, `termination_date`,
`last_published_date`.

**Links and bookkeeping.** `url`, `api_url`, `images`, `change_type`,
`scraped_at`, `source_record`.

Numeric fields are numbers or `null`, never an empty string: an empty string in a
float column breaks Excel, BigQuery and pandas alike. Reversed ranges, HTML
entities and non-breaking spaces are normalised before anything is compared, so
the same recall never arrives twice under two spellings.

#### How is the severity score calculated?

By fixed, published rules over agency fields, not by a model. Every point is
traceable, which means you can filter on it and defend the filter:

- FDA classification: Class I adds 60, Class II adds 35, Class III adds 15.
- Rows with no FDA class start from a base by agency: CPSC 40, NHTSA 35.
- NHTSA `do_not_drive` adds 30; `fire_risk_when_parked` adds 15.
- An injury actually reported adds 15; any mention of death adds 25.
- An open recall (`status` beginning "Ongoing") adds 10.
- Units affected: one million or more adds 10, 100,000 adds 6, 10,000 adds 3.
- A contamination or undeclared allergen tag adds 5.

The total is clamped to 0-100 and labelled in `severity_label`: 75 and up is
`Critical`, 50 is `High`, 25 is `Medium`, below that `Low`. Set
**Minimum severity score** in the input to have anything below your threshold
dropped before it reaches the dataset, and before you are billed for it.

#### How do the hazard and category tags work?

`hazard_tags` holds every tag that matches the hazard, defect and consequence
text: choking, strangulation, suffocation, entrapment, fire, burn, electrical,
laceration, chemical, fall, poisoning, drowning, impact, crash, allergen,
contamination, mislabeling, sterility, battery. `category` holds one of food,
drug, medical device, vehicle, children's products, toys, electronics,
appliances, furniture, clothing, sports and recreation, tools, cosmetics,
household or other.

Matching is by whole word, never substring. That sounds pedantic until you watch
a substring rule tag "leadership training" as a **lead** chemical hazard. Food,
drug, medical device and vehicle categories are not guessed at all: they come
from the feed the row arrived in.

Both tag sets are derived by this Actor, and the input form says so. The agencies
publish neither.

#### How does the incremental mode save money?

Switch on **Only what changed since the last run**. The Actor keeps a compact
fingerprint of every recall it has delivered in a named key-value store that
survives between runs, and on the next run it returns only rows that are new or
whose meaningful content changed. Each row carries `change_type` of `NEW`,
`UPDATED` or `UNCHANGED`.

The fingerprint deliberately ignores cosmetic edits. It covers status,
classification, termination date, remedy and the first 500 characters of the
product description, because those are the parts that mean "this recall moved".
Agencies fiddle with punctuation constantly; without that restraint every row
would look updated every single day and you would pay for the whole archive
again each morning.

Order of operations matters more than it looks: rows are pushed to the dataset
**first** and only then marked as delivered. A run that dies halfway therefore
repeats a few rows on the next run rather than losing them forever. The
lifecycle test kills a run mid-push and asserts that memory never runs ahead of
delivery.

#### How much does a run cost and how is spending controlled?

Pricing is pay per result, so the only thing that costs money is a row that
actually reaches your dataset. Four separate brakes exist and all of them run
before billing:

1. **Maximum rows in total** stops the whole run.
2. **Maximum rows per agency feed** stops one large feed from swallowing the
   entire limit. Without it, a five-agency run comes back looking like a
   single-agency scrape.
3. **Every filter** — keywords, company, class, status, category, hazard,
   severity, state, injuries, do-not-drive, date window — removes rows before
   they are pushed. The run report counts exactly how many each one removed.
4. **The platform spending limit** you set on the run is read at start and used
   as a hard row budget. The platform stops charging when your limit is reached
   but does not stop the run; this Actor stops it.

With no date window, no keywords and no limit at all, a run is treated as a trial
and stops at 200 rows. Pressing Start by accident cannot bill you for 127,501
records.

#### What does the run report contain?

Every run writes a `REPORT` record to the key-value store: rows fetched and
pushed per feed, how many cursor pages were walked, whether the run went past the
openFDA skip ceiling, how many rows each filter removed, how many empty or
duplicate rows were dropped before billing, and the new/updated/unchanged counts
from the incremental mode. If a host asked us to slow down, that is in there too.

#### How is this different from the single-agency recall scrapers?

There are competent Actors for individual feeds. The difference is arithmetic:

- **Five feeds in one row format.** The nearest FDA-only competitor covers the
  three openFDA feeds. This Actor adds CPSC and NHTSA — 40,292 more records —
  without a second run, a second schema or a second subscription.
- **No 25,000 ceiling.** Measured above.
- **53 fields.** Including images, UPC codes, retailers, importers, distributors,
  manufacturing countries, distribution states, do-not-drive and park-outside
  flags, and the untouched agency record on request.
- **Incremental monitoring with change types**, so a daily schedule costs a few
  rows rather than the whole archive.
- **No proxy and no browser.** Government APIs answer plain requests; adding a
  headless browser to this job only adds failure modes and cost.

#### What does this Actor not do?

Stated plainly, because finding out later is worse.

- **It is United States federal data only.** It does not cover the EU Safety
  Gate (RAPEX), Health Canada, UK OPSS, ACCC Australia, NITE Japan, KATS Korea,
  Product Safety New Zealand, Singapore or INMETRO Brazil. An aggregator that
  offers those exists and uses a proxy to reach them.
- **Two US feeds are missing for a measurable reason.** USDA FSIS meat and
  poultry recalls answer `HTTP 403 Access Denied` to a plain request, and the EU
  Safety Gate REST endpoint answers `HTTP 460 This action is not authorized`.
  Both were probed on 2026-09-11 and both are re-probed by the live test suite,
  which fails if either ever opens up — so this paragraph cannot quietly go
  stale.
- **No VIN lookup.** NHTSA rows are campaign-level. Asking "is this exact vehicle
  affected" needs the per-VIN endpoint, which is a different product.
- **No translation.** Rows come back in the language the agency published them
  in, which for all five feeds is English.
- **No official web page for FDA rows.** openFDA publishes none, so `url` is
  empty for FDA rows and `api_url` holds the exact API query that returns that
  single record instead. CPSC and NHTSA rows do carry a real `url`.
- **Categories, hazard tags and the severity score are ours, not the
  agencies'.** They are rule-based and documented above so you can audit them.

#### Who uses recall data on a schedule?

- **Marketplaces and retailers** delisting recalled stock before a regulator or a
  journalist finds it still for sale.
- **Compliance and quality teams** watching their own suppliers, contract
  manufacturers and competitors across all three agencies at once.
- **Importers and customs brokers** checking manufacturing countries and brand
  names against incoming shipments.
- **Insurers and claims analysts** pricing exposure with `units_affected`,
  `severity_score` and injury reports.
- **Fleet operators** watching `do_not_drive` and `fire_risk_when_parked`
  campaigns for the makes they run.
- **Newsrooms and researchers** who need the archive once and the delta daily.

#### How do I run it on a schedule?

Set a date window of the last week or two, switch on **Only what changed since
the last run**, and schedule it daily. The first run fills the memory, and every
run after it returns a handful of rows: new recalls plus recalls whose status,
classification or remedy moved. Add **Minimum severity score** at 50 or 75 if you
only want to be woken up for the serious ones, and **Any of these words** with
your brands if you are watching a specific catalogue.

### FAQ

#### Do I need an API key or a proxy?

No. All five endpoints are public government APIs that answer unauthenticated
requests. This Actor sends no credentials and uses no proxy, which is also why it
has no per-request cost floor to pass on to you.

#### How current is the data?

Each feed carries its own publication cadence and this Actor never caches. The
openFDA feeds expose a `last_updated` date in their metadata — 2026-09-02 for the
enforcement feeds and 2026-09-10 for device recalls when this page was written.
CPSC had recalls dated 2026-09-10 on that day and the NHTSA dataset had campaigns
received 2026-09-09.

#### How far back does the archive go?

CPSC goes back to a Tappan oven recall dated 1973-06-08. NHTSA campaign data goes
back to 1966 per the dataset description. The openFDA enforcement feeds start in
the early 2000s. Leave the date window empty and you get all of it, subject to
your row limits.

#### Can I get only food recalls, or only vehicles?

Yes — untick the other feeds under **Agency feeds**. Each feed is a separate
government API and you are only charged for rows that reach your dataset, so
narrowing the feeds narrows the bill.

#### Why are some rows missing a company name?

Because the agency did not publish one. CPSC in particular often names only the
retailer or importer; in that case this Actor falls back to the importer and then
the retailer rather than leaving the field blank, and the original lists stay
available in `importers` and `retailers`.

#### What is `source_record` for?

Switch on **Attach the original agency record** and every row carries the raw
JSON exactly as the agency returned it. Use it when you need a field this Actor
does not normalise, or when you want to prove what the agency said on the day you
collected it.

#### Does the compact mode change the price?

No. Price is per row, not per byte. Compact mode exists because 17 fields are
easier for an agent or a diffing job to handle than 53, not to save money.

#### What happens if a feed is down?

The run continues with the feeds that answer, records the failure in the run
report under `failures`, and never silently reports a partial collection as a
complete one. Rate limiting is handled per host with a shared cooldown: a
`429 Too Many Requests` from any agency slows every thread touching that host,
not just the one that met it, and the cooldown shows up in the run report under
`throttle`. openFDA throttles unauthenticated clients, so a very large archive
pull may pause and resume rather than fail.

#### Can I use this data commercially?

The underlying records are United States federal government publications. Read
each agency's terms — openFDA publishes its own at `open.fda.gov/terms` — and
note that openFDA explicitly says its data is not validated for medical
decisions. This Actor passes the data through unchanged and adds its derived
fields on top; the disclaimers of the source apply to it.

# Actor input Schema

## `sources` (type: `array`):

Pick one or more. Every feed is a separate government API - competitors sell them as separate Actors.

## `dateFrom` (type: `string`):

YYYY-MM-DD. Pushed into the API query itself, so older records are never downloaded. Leave empty for the whole archive.

## `dateTo` (type: `string`):

YYYY-MM-DD. A reversed window is swapped silently.

## `maxItems` (type: `integer`):

Hard stop for the whole run. 0 means no limit. With no date window, no keywords and no limit the run stops at 200 rows, so an accidental Start never bills for the whole archive.

## `maxItemsPerSource` (type: `integer`):

Quota for each feed. Without it the largest feed eats the whole limit and a five-agency run looks like a single-agency scrape.

## `keywords` (type: `array`):

Matched on whole words across title, description, reason, hazard, company, brand, model, retailer and UPC. `lead` will not match `leadership`.

## `excludeKeywords` (type: `array`):

Same whole-word matching, in reverse.

## `company` (type: `string`):

Recalling firm, manufacturer or retailer, matched as a substring.

## `classification` (type: `array`):

Class I is the most serious: reasonable probability of serious injury or death. FDA feeds only - CPSC and NHTSA do not publish a class.

## `status` (type: `array`):

FDA publishes Ongoing, Completed or Terminated. CPSC rows carry Announced, NHTSA rows carry Open.

## `categories` (type: `array`):

Derived by this Actor from the product text, not published by the agencies. Food, drug, medical device and vehicle come straight from the feed the row arrived in.

## `hazardTags` (type: `array`):

Derived by this Actor from hazard, defect and consequence text with whole-word rules. A row keeps every tag that matches.

## `minSeverity` (type: `integer`):

0-100, computed by fixed rules from the agency fields: FDA class, NHTSA do-not-drive and park-outside flags, injuries and deaths mentioned, status and units affected. 75 and above is labelled Critical.

## `state` (type: `string`):

Two-letter codes, comma separated. Matches the recalling firm's state or any state named in the FDA distribution pattern.

## `onlyWithInjuries` (type: `boolean`):

CPSC injury reports plus any row whose text mentions deaths. `None reported.` counts as no injuries.

## `onlyDoNotDrive` (type: `boolean`):

The rare campaigns where NHTSA tells owners to stop driving the vehicle.

## `onlyNew` (type: `boolean`):

Remembers what it already delivered in a named key-value store and skips it next time. Each row is marked NEW or UPDATED in `change_type`.

## `emitUnchanged` (type: `boolean`):

Only meaningful together with the previous option. Off by default: you should not pay twice for the same row.

## `compactOutput` (type: `boolean`):

17 key fields instead of 53. Useful for AI agents and for cheap diffing.

## `excludeEmptyFields` (type: `boolean`):

Leaves out keys the agency did not publish for that row instead of returning them empty.

## `includeSourceRecord` (type: `boolean`):

Adds `source_record` with the untouched JSON exactly as the agency returned it, so nothing this Actor normalises is lost.

## Actor input object example

```json
{
  "sources": [
    "fda_food",
    "fda_drug",
    "fda_device",
    "cpsc",
    "nhtsa"
  ],
  "dateFrom": "",
  "dateTo": "",
  "maxItems": 0,
  "maxItemsPerSource": 0,
  "keywords": [],
  "excludeKeywords": [],
  "company": "",
  "classification": [],
  "status": [],
  "categories": [],
  "hazardTags": [],
  "minSeverity": 0,
  "state": "",
  "onlyWithInjuries": false,
  "onlyDoNotDrive": false,
  "onlyNew": false,
  "emitUnchanged": false,
  "compactOutput": false,
  "excludeEmptyFields": false,
  "includeSourceRecord": false
}
```

# Actor output Schema

## `results` (type: `string`):

All collected rows

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "sources": [
        "fda_food",
        "fda_drug",
        "fda_device",
        "cpsc",
        "nhtsa"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("snow_leo_data/product-recalls-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "sources": [
        "fda_food",
        "fda_drug",
        "fda_device",
        "cpsc",
        "nhtsa",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("snow_leo_data/product-recalls-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "sources": [
    "fda_food",
    "fda_drug",
    "fda_device",
    "cpsc",
    "nhtsa"
  ]
}' |
apify call snow_leo_data/product-recalls-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,snow_leo_data/product-recalls-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/I37a4zhxpCYgMJUkp/builds/HRyEbmD1qat7fnQJU/openapi.json
