# Federal Register Scraper (`aurenic/federal-register-scraper`) Actor

Extract every US Federal Register document since 1994 — rules, proposed rules, notices, presidential documents. Auto-partitions past the 2,000-doc window. Keyless, no browser, no proxy.

- **URL**: https://apify.com/aurenic/federal-register-scraper.md
- **Developed by:** [Aurenic](https://apify.com/aurenic) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.30 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Federal Register Scraper

Extract every US Federal Register document since 1994 — rules, proposed rules, notices, presidential documents. Auto-partitions past the 2,000-doc window. Keyless, no browser, no proxy.

### What does Federal Register Scraper do?

The Federal Register is the daily journal of the US federal government — every rule, proposed rule, notice, and presidential document published since 1994 lives here. This actor wraps the **official keyless Federal Register API** in four modes:

- **Search documents** — filter by full-text term, agency, document type, significance flag, and publication date range. **Auto-partitions by date** so queries returning more than the API's 2,000-document window still complete fully.
- **Single document** — fetch one document by its Federal Register number.
- **List agencies** — every federal agency with its slug, description, and parent-child structure.
- **List document types** — the canonical value list (RULE, PRORULE, NOTICE, PRESDOCU).

Every document record includes number, title, type, abstract, agencies, publication date, **comment deadline and days remaining**, significance flag, docket IDs, RINs, CFR references, page range, and direct HTML/PDF/raw-text URLs.

### Output fields

#### Document

| Field | Description |
|---|---|
| documentNumber | Federal Register citation number |
| title | Document title |
| type / subtype | RULE, PRORULE, NOTICE, PRESDOCU |
| abstract | Official abstract |
| publicationDate | Publication date |
| agencies | Array of `{ id, name, url, parentId, slug }` |
| agencyNames | Flat list of issuing agency names |
| action | Stated action (e.g. "Final rule") |
| effectiveOn | Effective date |
| commentsCloseOn | Public comment deadline |
| commentWindowOpen | Boolean — true while the deadline is today or later |
| commentDaysRemaining | Days until deadline (negative after) |
| significance | EO 12866 significant-regulatory-action flag |
| docketIds | Rulemaking dockets |
| regulationIdNumbers | RINs — track a rule across its lifecycle |
| cfrReferences | Array of `{ title, part, chapter }` |
| citation | Full Federal Register citation |
| startPage / endPage / pageLength | Page range |
| volume | FR volume number |
| htmlUrl / pdfUrl / rawTextUrl | Canonical document URLs |
| executiveOrderNumber | EO number, when applicable |
| presidentialDocumentNumber | Presidential doc number, when applicable |
| topics | Topic tags |

#### Agency

| Field | Description |
|---|---|
| id / name / shortName / slug | Identity |
| url / recentArticlesUrl | Links |
| parentId | Parent agency (null for top-level) |
| description | Agency description |

#### Document type

`{ value, label, description }` — the canonical enum.

### Who is it for?

- **Regulatory affairs and compliance teams** tracking rules that affect their industry
- **Law firms and lobbyists** monitoring proposed rules and comment windows
- **Policy researchers** building longitudinal rulemaking datasets
- **Data journalists** investigating agency activity
- **Algo trading and quant funds** parsing regulatory signals
- **Government contractors** tracking procurement and compliance obligations

### Pricing

**$1.40 per 1,000 results.** No subscription.

| Results | Cost |
|---|---|
| 100 | $0.14 |
| 1,000 | $1.40 |
| 10,000 | $14.00 |

### How to use it

1. Pick a **Mode**.
2. For search: enter a **Search Term** and/or set **Agency**, **Document Type**, and **Date Range**.
3. For document: enter **Document Number**.
4. Set **Max Items** (default 2000).
5. Click **Start**.

### Output example

```json
{
  "recordType": "document",
  "documentNumber": "2026-18046",
  "title": "Energy Conservation Program: Test Procedures for Consumer Refrigerators",
  "type": "Proposed Rule",
  "subtype": "",
  "abstract": "The U.S. Department of Energy proposes to amend its test procedures for consumer refrigerators...",
  "publicationDate": "2026-09-18",
  "agencies": [
    { "id": 197, "name": "Energy Department", "url": "https://www.federalregister.gov/agencies/energy-department", "parentId": null, "slug": "energy-department" }
  ],
  "agencyNames": ["Energy Department"],
  "action": "Notice of proposed rulemaking and public meeting",
  "dates": "Comments due on or before October 20, 2026.",
  "effectiveOn": "",
  "commentsCloseOn": "2026-10-20",
  "commentWindowOpen": true,
  "commentDaysRemaining": 24,
  "significance": true,
  "docketIds": ["EERE-2024-BT-TP-0005"],
  "regulationIdNumbers": ["1904-AD95"],
  "cfrReferences": [{ "title": 10, "part": 430, "chapter": null }],
  "citation": "91 FR 45120",
  "startPage": 45120,
  "endPage": 45156,
  "pageLength": 36,
  "volume": 91,
  "htmlUrl": "https://www.federalregister.gov/documents/2026/09/18/2026-18046/...",
  "pdfUrl": "https://www.govinfo.gov/content/pkg/FR-2026-09-18/pdf/2026-18046.pdf",
  "rawTextUrl": "https://www.federalregister.gov/documents/full_text/text/2026/09/18/2026-18046.txt",
  "executiveOrderNumber": "",
  "presidentialDocumentNumber": "",
  "topics": [],
  "scrapedAt": "2026-09-26T12:00:00.000Z"
}
```

### Technical details

- **Official Federal Register API** — `https://www.federalregister.gov/api/v1`. Public, keyless, unauthenticated.
- **No documented rate limit.** The actor spaces requests ≥250ms apart and backs off on 429.
- **2,000-document window** — the API returns HTTP 400 when `page × per_page` exceeds 2,000. **The actor auto-partitions by publication date and recurses until each slice fits under the ceiling.** A query matching 10,000+ documents completes fully.
- **404 means "no documents matched"** — the API's documented convention. The actor treats 404 as an empty result set, not an error.
- **`per_page` maxes at 100.** The actor uses 100 by default.
- **Sparse fields** — different document types include different fields. Every optional field is emitted as `null` or `[]`, never omitted, so downstream parsers don't break on sparse records.
- **No browser, no proxy** — pure REST JSON.

### Known limits

- **Historical coverage starts at 1994-01-01.** Earlier documents are not in the API.
- **`raw_text_url` is a pointer, not the text.** The actor returns the URL; fetch it downstream if you need the full body.
- **Some documents have no comment period.** `commentsCloseOn` is null for final rules and most notices.
- **Agency slugs must be exact.** Use agencies mode to list valid values — `epa` won't work, `environmental-protection-agency` will.
- **Very broad queries with no date range** auto-partition the full 1994→today span and can take several minutes.

### FAQ

**Do I need an API key?** No. The Federal Register API is fully keyless.

**Do I need a proxy?** No. Datacenter IPs work.

**How do I track a rule across its lifecycle?** Use `regulationIdNumbers` (RINs) — they persist from the proposed rule through the final rule.

**Why does a query return 0 results?** Either the query matched nothing (the API returns 404), or your agency slug or document type is invalid. Verify with agencies mode.

**How far back does the data go?** 1994-01-01.

**How do I export data?** After a run, go to Storage → Export as JSON, CSV, Excel.

### Support

Open an issue on the Actor's page for bugs or feature requests.

# Actor input Schema

## `mode` (type: `string`):

What to fetch.

## `query` (type: `string`):

Full-text search term (e.g. 'clean energy', 'emissions').

## `agency` (type: `string`):

Agency slug (e.g. environmental-protection-agency, securities-and-exchange-commission, department-of-energy). Use agencies mode to list all.

## `documentType` (type: `string`):

Document type.

## `documentNumber` (type: `string`):

Federal Register document number (e.g. 2026-18046). Used in document mode.

## `dateFrom` (type: `string`):

Earliest publication date (YYYY-MM-DD). Earliest possible is 1994-01-01.

## `dateTo` (type: `string`):

Latest publication date (YYYY-MM-DD).

## `significant` (type: `string`):

Filter by 'significant regulatory action' flag (EO 12866).

## `includeFullText` (type: `boolean`):

Populate `fullText` with the raw\_text\_url. Full text is not fetched (use the URL downstream).

## `maxItems` (type: `integer`):

Hard cap on records per run.

## `requestDelayMs` (type: `integer`):

Delay between requests. No documented rate limit; 250ms is polite.

## Actor input object example

```json
{
  "mode": "search",
  "query": "clean energy",
  "agency": "",
  "documentType": "",
  "documentNumber": "",
  "dateFrom": "2026-01-01",
  "dateTo": "2026-09-26",
  "significant": "",
  "includeFullText": false,
  "maxItems": 2000,
  "requestDelayMs": 250
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("aurenic/federal-register-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("aurenic/federal-register-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call aurenic/federal-register-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,aurenic/federal-register-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/WhQtNHaO3hmHwsbQj/builds/CEf4qr8HiAJTTfF8M/openapi.json
