# Federal Register Scraper - Rules and Comment Dates (`s-r/federalregister-scraper`) Actor

Search US Federal Register documents: rules, proposed rules, notices and presidential documents. Returns agencies, effective dates, public comment deadlines with days remaining, docket ids and CFR references.

- **URL**: https://apify.com/s-r/federalregister-scraper.md
- **Developed by:** [SR](https://apify.com/s-r) (community)
- **Categories:** Business
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 run start fees

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Federal Register Scraper

Search the **US Federal Register**: rules, proposed rules, notices and
presidential documents. Agencies, effective dates, **public comment deadlines
with days remaining**, docket ids and the sections of federal regulation
affected.

Reads the Register's own API. No key, no login.

### The comment deadline is the field with consequences

A proposed rule is open for public comment until a date, and after that date the
opportunity to influence it is gone. That is the one piece of this data people
actually act on, and it is one of the fields the API leaves out unless you ask
for it by name.

Every row carries `comments_close_on`, plus two derived fields:

- **`comment_days_remaining`** — days until the deadline, **negative once it has
  passed**. A closed deadline being visible matters as much as an open one.
- **`comments_open`** — whether the window is still open today.

Set `comments_open_only` and you get exactly the documents you can still respond
to. A run on that filter returned **191 proposed rules currently open**, with
deadlines from 13 to 45 days out.

### Ask for your fields or you will not get them

The API returns a small default set and **silently omits** `comments_close_on`,
`effective_on`, `docket_ids`, `regulation_id_numbers`, `significant`,
`cfr_references` and more. Nothing in the response indicates anything was left
out, so a scraper that does not name its fields concludes those values do not
exist for these documents.

This Actor always requests the full set.

### Fields

- **Identity**: `document_number`, `title`, `type`, `citation`
- **Content**: `abstract`, `action`, `topics`, `dates_text`
- **Who**: `agencies`, `president`
- **When**: `publication_date`, `effective_on`, `comments_close_on`,
  `comment_days_remaining`, `comments_open`
- **Regulatory**: `significant`, `docket_ids`, `regulation_id_numbers`,
  `cfr_references`
- **Where to read it**: `html_url`, `pdf_url`, **`raw_text_url`**
- **Extent**: `page_length`, `start_page`, `end_page`

`cfr_references` is returned readable — `40 CFR 60` rather than a nested object —
because that is the form a compliance team recognises.

`raw_text_url` is the plain-text version, which is the one to feed to anything
doing full-text analysis rather than the PDF.

`page_length` is a rough proxy for how substantial a document is: a
three-page notice and a three-hundred-page rule are different animals.

### A note on the reported count

The Register **caps its reported count at 10,000**. Four different queries — a
term search, a type filter, a date filter and an agency filter — each came back
with exactly 10,000, which is a display ceiling rather than a measurement.

Narrower queries return real numbers: `artificial intelligence` reports 1,554,
and open proposed rules 191. So the count is trustworthy below the ceiling and a
floor at it. The run summary flags which case you are in rather than letting you
treat 10,000 as a total.

### Input reference

| Field | Type | Default |
|---|---|---|
| `search` | free-text term | — |
| `doc_type` | rule, proposed\_rule, notice, presidential\_document | any |
| `agency` | agency slug, e.g. `environmental-protection-agency` | — |
| `comments_open_only` | only documents still accepting comment | `false` |
| `date_from`, `date_to` | YYYY-MM-DD | — |
| `order` | newest, oldest, relevance | newest |
| `limit` | 1-5000 | 100 |
| `retries` | 1-6 | 3 |

Agency names are slugified for you, so `Environmental Protection Agency` and
`environmental-protection-agency` both work.

### Why this Actor connects directly

across every `.gov` host tested:

| Route | Result |
|---|---|
| Direct | **HTTP 200** |
| Through a residential proxy | `CONNECT tunnel failed, response 491` |

That 491 is our proxy refusing to tunnel to the host, not the government
refusing us. The Register publishes this API for public use and the Actor takes
it directly, paced politely.

### Typical uses

- **Regulatory monitoring.** Filter by agency and `comments_open_only`, run
  daily, and you have every rule your industry can still respond to, with the
  deadline attached.
- **Compliance calendars.** `effective_on` is when a rule bites.
  `cfr_references` says which parts of the code change.
- **Policy research.** Term search across decades, with `raw_text_url` for the
  documents you want to read in full.
- **Competitive and lobbying intelligence.** `docket_ids` link Register
  documents to the dockets where comments are filed.
- **Executive action tracking.** `doc_type: presidential_document` with
  `president` returns executive orders and proclamations.

### Notes

`significant` is the Register's own flag and is frequently null rather than
false; absent is not the same as "not significant" and is returned as-is.

An empty result comes back as a 404 from the API, which this Actor treats as
zero documents rather than a failure. A 400 means the search conditions were
rejected and is not retried, since retrying a bad condition cannot help.

Coverage is the US Federal Register only. State registers are published
separately and are not in this data.

### Finding an agency slug

Agency slugs are the agency name lower-cased with hyphens:
`environmental-protection-agency`, `food-and-drug-administration`,
`securities-and-exchange-commission`, `federal-communications-commission`. The
Actor slugifies whatever you type, so the plain name works too.

A document is frequently issued by **several** agencies at once, and `agencies`
returns all of them. Filtering by one agency returns documents where it is any
of the issuers, not only the lead, which is usually what you want and
occasionally a surprise.

### What this Actor does not do

**No document full text.** `raw_text_url` gives you the address of the plain
text, and fetching it is a separate job with a different cost profile: some
rules run to hundreds of pages.

**No public comments.** The Register publishes the documents and the deadlines;
the comments themselves live on Regulations.gov, which is a different API.

**No historical amendments.** A rule's row describes the rule as published. How
it was later amended lives in the eCFR, not here.

# Actor input Schema

## `search` (type: `string`):

Free-text search across documents, for example artificial intelligence or PFAS.

## `doc_type` (type: `string`):

Restrict to one type. Proposed rules are the ones with comment deadlines.

## `agency` (type: `string`):

Agency slug, for example environmental-protection-agency or food-and-drug-administration.

## `comments_open_only` (type: `boolean`):

Return only documents whose public comment period is still open today.

## `date_from` (type: `string`):

Only documents published on or after this date, as YYYY-MM-DD.

## `date_to` (type: `string`):

Only documents published on or before this date, as YYYY-MM-DD.

## `order` (type: `string`):

How results are sorted.

## `limit` (type: `integer`):

How many documents to return.

## `retries` (type: `integer`):

Retries with backoff before a request is reported as an error.

## Actor input object example

```json
{
  "search": "artificial intelligence",
  "doc_type": "",
  "agency": "environmental-protection-agency",
  "comments_open_only": false,
  "order": "newest",
  "limit": 100,
  "retries": 3
}
```

# Actor output Schema

## `documents` (type: `string`):

One row per Federal Register document.

## `summary` (type: `string`):

Counts, document type breakdown and how many are open for comment.

## `errors` (type: `string`):

Failures with a code and a redacted message.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "limit": 100,
    "retries": 3
};

// Run the Actor and wait for it to finish
const run = await client.actor("s-r/federalregister-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "limit": 100,
    "retries": 3,
}

# Run the Actor and wait for it to finish
run = client.actor("s-r/federalregister-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "limit": 100,
  "retries": 3
}' |
apify call s-r/federalregister-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,s-r/federalregister-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/WfdjZjFQsgVv1ME2U/builds/KsAejaFJhgkJcVUfX/openapi.json
