# EPA ECHO Facility Compliance & Enforcement Scraper (`haketa/epa-echo-scraper`) Actor

Scrape US EPA ECHO facility data: compliance status (air, water, hazardous waste), violations, inspections, penalties ($), major-polluter flag, %minority, coordinates and program IDs. For ESG, due diligence and compliance research. Not affiliated with the EPA.

- **URL**: https://apify.com/haketa/epa-echo-scraper.md
- **Developed by:** [Haketa](https://apify.com/haketa) (community)
- **Categories:** Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.75 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## EPA ECHO Facility Compliance & Enforcement Scraper

> **Search and extract US EPA environmental-compliance data for any facility: compliance status (air, water, hazardous waste, drinking water), violations, inspections, formal actions, penalties ($), major-polluter flag, environmental-justice %minority, coordinates and program IDs.** Search by company name and/or location (state, city, ZIP, NAICS/SIC). Clean JSON/CSV/Excel in seconds — built for ESG, due diligence, risk and compliance research.

[![EPA ECHO](https://img.shields.io/badge/EPA-ECHO%20Compliance-006747)]()
[![Violations & Penalties](https://img.shields.io/badge/Violations%20%2B%20Penalties-c8102e)]()
[![ESG & Due Diligence](https://img.shields.io/badge/ESG%20%2F%20Due%20Diligence-8250df)]()
[![Export](https://img.shields.io/badge/Export-JSON%20%2F%20CSV%20%2F%20Excel-fb8500)]()

***

### What This Actor Does

This Actor searches the US EPA's **Enforcement & Compliance History Online (ECHO)** database and returns rich, structured records for regulated facilities. For each facility it captures:

- **Identity & location** — facility name, full address, county, EPA region, coordinates, FRS Registry ID
- **Industry** — SIC and NAICS codes
- **Compliance status** — overall plus per-statute: **air (Clean Air Act)**, **water (Clean Water Act)**, **hazardous waste (RCRA)** and **drinking water (SDWA)**
- **Violations** — significant non-complier flag, quarters in non-compliance, 3-year compliance history
- **Enforcement** — inspection count and last inspection date, formal action count and date
- **Penalties** — total penalties (USD), penalty count, last penalty date and amount
- **Risk & ESG** — major-polluter flag, surrounding **% minority** (environmental-justice indicator)
- **Program IDs** — air, water (NPDES), hazardous waste (RCRA), toxics (TRI) and greenhouse-gas (GHG) identifiers
- **Detailed report** — a direct link to the facility's full ECHO report

***

### Why Use This

- **See who is in violation.** Filter to facilities with current violations or significant non-compliance — the exact list a risk, ESG or compliance team needs.
- **Quantify enforcement.** Penalties in dollars, inspection counts and formal actions — hard numbers, not prose.
- **Screen supply chains & counterparties.** Check any company's US facilities for environmental risk before you onboard, invest or lend.
- **Official & free.** Reads EPA's own public database — no key, no anti-bot, no browser.

***

### Quick Start

#### Run it in the console (no code)

1. Add a **facility name** (e.g. `exxon mobil`) and/or a **location** filter (state, city, ZIP, NAICS/SIC).
2. Optionally tick **Only facilities with violations** or **Only major facilities**.
3. Set **Max facilities**, click **Start**, and export as **JSON, CSV, Excel or HTML**.

#### Find environmental violators in a state (Python)

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")

run_input = {
    "states": ["TX", "CA"],
    "onlyWithViolations": True,
    "maxItems": 1000,
}

run = client.actor("YOUR_USERNAME/epa-echo-scraper").call(run_input=run_input)

for f in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(f["name"], "·", f["complianceStatus"], "·", f["totalPenaltiesUsd"], "USD")
```

#### Screen one company's facilities (Python)

```python
run = client.actor("YOUR_USERNAME/epa-echo-scraper").call(run_input={
    "searchTerms": ["exxon mobil"], "states": ["TX"], "maxItems": 500,
})
for f in client.dataset(run["defaultDatasetId"]).iterate_items():
    if f.get("significantViolator") or (f.get("totalPenaltiesUsd") or 0) > 0:
        print(f["name"], f["city"], "→", f["totalPenaltiesUsd"], "USD", f["detailedReportUrl"])
```

***

### Input Parameters

| Field | Type | Description |
|---|---|---|
| `searchTerms` | array | Facility/company name keywords. Combined with the location filters. |
| `states` | array | Two-letter US state codes (e.g. `TX`, `CA`). |
| `cities` | array | City names (optional). |
| `zipCodes` | array | ZIP codes (optional). |
| `naicsCodes` | array | NAICS industry codes (optional). |
| `sicCodes` | array | SIC industry codes (optional). |
| `onlyMajor` | boolean | Only facilities flagged as major. |
| `onlyWithViolations` | boolean | Only facilities with a current violation / significant non-compliance / recent NC quarters. |
| `maxItems` | integer | Max facilities across all searches. `0` = no limit. |
| `proxyConfiguration` | object | Apify Proxy. Datacenter is enough (public API). |

Provide at least one criterion — a facility name and/or a location filter.

***

### Output

Each facility is one record:

```json
{
  "registryId": "110000599273",
  "name": "BAYTOWN OLEFINS PLANT",
  "street": "5000 BAYWAY DR", "city": "BAYTOWN", "state": "TX", "zip": "77520",
  "county": "HARRIS", "epaRegion": "06",
  "latitude": 29.74, "longitude": -95.01,
  "naicsCodes": ["325110"], "sicCodes": ["2869"],
  "majorFacility": true,
  "complianceStatus": "Significant Violation",
  "airComplianceStatus": "Significant Violation",
  "waterComplianceStatus": "No Violation Identified",
  "significantViolator": true,
  "quartersInNonCompliance": 4,
  "inspectionCount": 12, "lastInspectionDate": "09/18/2025",
  "formalActionCount": 2,
  "totalPenaltiesUsd": 2204913, "penaltyCount": 3, "lastPenaltyAmountUsd": 1500000,
  "percentMinority": 62.4,
  "airIds": ["TX0001"], "toxicsReleaseIds": ["77520XXNPL"],
  "detailedReportUrl": "https://echo.epa.gov/detailed-facility-report?fid=110000599273"
}
```

**About coverage:** identity, address and coordinates are present for essentially every facility. Compliance status, penalties and program IDs are populated by EPA for regulated/assessed facilities (industrial and major sites) — small retail sites (e.g. some gas stations) legitimately have little or no compliance history. Use `onlyWithViolations` or `onlyMajor` to focus on facilities with rich enforcement data.

***

### Use Cases

#### 1. ESG & sustainability screening

Score companies and portfolios on environmental compliance, violations and penalties across their US facilities.

#### 2. Due diligence & risk

Check a target, supplier or counterparty for environmental violations and enforcement history before onboarding, investing or lending.

#### 3. Compliance & legal research

Build lists of significant non-compliers, penalty actions and inspection histories by state, industry or region.

#### 4. Journalism & research

Map polluters, penalties and environmental-justice indicators (% minority) by geography and sector.

***

### Tips

- **`onlyWithViolations`** is the fastest way to a list of active environmental violators.
- **`totalPenaltiesUsd`** and **`penaltyCount`** quantify enforcement in dollars.
- **`significantViolator`** flags the EPA's "significant non-complier" designation.
- **`naicsCodes` / `sicCodes`** let you focus on an industry (e.g. refineries, chemicals, power).
- **`detailedReportUrl`** links straight to the full ECHO report for any facility.
- **Schedule it** with Apify Schedules to monitor a state or company over time.

***

### Frequently Asked Questions

**Do I need an account or key?**
No. The ECHO database is public and free — no login, key or anti-bot.

**What is a Registry ID?**
It is the facility's EPA FRS (Facility Registry Service) identifier — a stable ID used to join across EPA program systems.

**Why do some facilities have empty compliance fields?**
EPA populates compliance, penalty and program data for regulated and assessed facilities. Smaller sites with no regulated activity legitimately have little history. Filter with `onlyWithViolations` or `onlyMajor` for the data-rich subset.

**Which environmental programs are covered?**
Clean Air Act (air), Clean Water Act (water/NPDES), RCRA (hazardous waste), SDWA (drinking water), plus TRI (toxics) and GHG identifiers.

**What export formats are supported?**
JSON, CSV, Excel, HTML, or via API — plus Google Sheets, webhooks, Make and Zapier.

***

### Legal & Responsible Use

This Actor is an independent tool and is **not affiliated with, endorsed by, or sponsored by the US Environmental Protection Agency (EPA)**. All trademarks are the property of their respective owners. It reads only public compliance records. Use the data responsibly and in line with applicable terms and laws.

# Actor input Schema

## `searchTerms` (type: `array`):

Facility/company name keywords (e.g. "chevron", "3M"). Combined with the location filters below. Leave empty to pull all facilities matching the location filters.

## `states` (type: `array`):

Two-letter US state codes to filter by (e.g. TX, CA, NY).

## `cities` (type: `array`):

City names to filter by (optional).

## `zipCodes` (type: `array`):

ZIP codes to filter by (optional).

## `naicsCodes` (type: `array`):

NAICS industry codes to filter by (optional).

## `sicCodes` (type: `array`):

SIC industry codes to filter by (optional).

## `onlyMajor` (type: `boolean`):

Return only facilities flagged as major (larger regulated sites).

## `onlyWithViolations` (type: `boolean`):

Return only facilities that currently have a violation, are significant non-compliers, or have recent non-compliance quarters.

## `maxItems` (type: `integer`):

Maximum facilities across all searches. 0 = no limit.

## `proxyConfiguration` (type: `object`):

Apify Proxy. The API is public — datacenter is enough and enabled by default.

## Actor input object example

```json
{
  "searchTerms": [
    "exxon mobil"
  ],
  "states": [
    "TX"
  ],
  "onlyMajor": false,
  "onlyWithViolations": false,
  "maxItems": 200,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `registryId` (type: `string`):

EPA FRS Registry ID

## `name` (type: `string`):

Facility name

## `street` (type: `string`):

Street address

## `city` (type: `string`):

City

## `state` (type: `string`):

State

## `zip` (type: `string`):

ZIP code

## `county` (type: `string`):

County

## `epaRegion` (type: `string`):

EPA region

## `latitude` (type: `string`):

Latitude

## `longitude` (type: `string`):

Longitude

## `sicCodes` (type: `string`):

SIC codes

## `naicsCodes` (type: `string`):

NAICS codes

## `federalAgency` (type: `string`):

Federal agency (if federal)

## `inIndianCountry` (type: `string`):

In Indian country

## `majorFacility` (type: `string`):

Major-polluter flag

## `complianceStatus` (type: `string`):

Overall compliance

## `airComplianceStatus` (type: `string`):

Clean Air Act status

## `waterComplianceStatus` (type: `string`):

Clean Water Act status

## `hazWasteComplianceStatus` (type: `string`):

RCRA status

## `drinkingWaterComplianceStatus` (type: `string`):

SDWA status

## `significantViolator` (type: `string`):

Significant non-complier

## `quartersInNonCompliance` (type: `string`):

Quarters in non-compliance

## `threeYearComplianceHistory` (type: `string`):

12-quarter compliance string

## `inspectionCount` (type: `string`):

Inspection count

## `lastInspectionDate` (type: `string`):

Last inspection date

## `formalActionCount` (type: `string`):

Formal enforcement actions

## `lastFormalActionDate` (type: `string`):

Last formal action date

## `totalPenaltiesUsd` (type: `string`):

Total penalties (USD)

## `penaltyCount` (type: `string`):

Number of penalties

## `lastPenaltyDate` (type: `string`):

Last penalty date

## `lastPenaltyAmountUsd` (type: `string`):

Last penalty amount (USD)

## `percentMinority` (type: `string`):

Surrounding %minority (EJ)

## `airIds` (type: `string`):

Clean Air Act program IDs

## `waterIds` (type: `string`):

NPDES program IDs

## `hazWasteIds` (type: `string`):

RCRA program IDs

## `toxicsReleaseIds` (type: `string`):

TRI program IDs

## `greenhouseGasIds` (type: `string`):

Greenhouse-gas program IDs

## `detailedReportUrl` (type: `string`):

ECHO detailed facility report

## `scrapedAt` (type: `string`):

ISO timestamp

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchTerms": [
        "exxon mobil"
    ],
    "states": [
        "TX"
    ],
    "maxItems": 200,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("haketa/epa-echo-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchTerms": ["exxon mobil"],
    "states": ["TX"],
    "maxItems": 200,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("haketa/epa-echo-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchTerms": [
    "exxon mobil"
  ],
  "states": [
    "TX"
  ],
  "maxItems": 200,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call haketa/epa-echo-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,haketa/epa-echo-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/XHDxvsmuBCxDZIhIh/builds/b5Oj1BhWzThGcp4M3/openapi.json
