Federal Register Scraper - Rules and Comment Dates
Pricing
from $1.00 / 1,000 run start fees
Federal Register Scraper - Rules and Comment Dates
Search US Federal Register documents: rules, proposed rules, notices and presidential documents. Returns agencies, effective dates, public comment deadlines with days remaining, docket ids and CFR references.
Federal Register Scraper
Search the US Federal Register: rules, proposed rules, notices and presidential documents. Agencies, effective dates, public comment deadlines with days remaining, docket ids and the sections of federal regulation affected.
Reads the Register's own API. No key, no login.
The comment deadline is the field with consequences
A proposed rule is open for public comment until a date, and after that date the opportunity to influence it is gone. That is the one piece of this data people actually act on, and it is one of the fields the API leaves out unless you ask for it by name.
Every row carries comments_close_on, plus two derived fields:
comment_days_remaining— days until the deadline, negative once it has passed. A closed deadline being visible matters as much as an open one.comments_open— whether the window is still open today.
Set comments_open_only and you get exactly the documents you can still respond
to. A run on that filter returned 191 proposed rules currently open, with
deadlines from 13 to 45 days out.
Ask for your fields or you will not get them
The API returns a small default set and silently omits comments_close_on,
effective_on, docket_ids, regulation_id_numbers, significant,
cfr_references and more. Nothing in the response indicates anything was left
out, so a scraper that does not name its fields concludes those values do not
exist for these documents.
This Actor always requests the full set.
Fields
- Identity:
document_number,title,type,citation - Content:
abstract,action,topics,dates_text - Who:
agencies,president - When:
publication_date,effective_on,comments_close_on,comment_days_remaining,comments_open - Regulatory:
significant,docket_ids,regulation_id_numbers,cfr_references - Where to read it:
html_url,pdf_url,raw_text_url - Extent:
page_length,start_page,end_page
cfr_references is returned readable — 40 CFR 60 rather than a nested object —
because that is the form a compliance team recognises.
raw_text_url is the plain-text version, which is the one to feed to anything
doing full-text analysis rather than the PDF.
page_length is a rough proxy for how substantial a document is: a
three-page notice and a three-hundred-page rule are different animals.
A note on the reported count
The Register caps its reported count at 10,000. Four different queries — a term search, a type filter, a date filter and an agency filter — each came back with exactly 10,000, which is a display ceiling rather than a measurement.
Narrower queries return real numbers: artificial intelligence reports 1,554,
and open proposed rules 191. So the count is trustworthy below the ceiling and a
floor at it. The run summary flags which case you are in rather than letting you
treat 10,000 as a total.
Input reference
| Field | Type | Default |
|---|---|---|
search | free-text term | — |
doc_type | rule, proposed_rule, notice, presidential_document | any |
agency | agency slug, e.g. environmental-protection-agency | — |
comments_open_only | only documents still accepting comment | false |
date_from, date_to | YYYY-MM-DD | — |
order | newest, oldest, relevance | newest |
limit | 1-5000 | 100 |
retries | 1-6 | 3 |
Agency names are slugified for you, so Environmental Protection Agency and
environmental-protection-agency both work.
Why this Actor connects directly
across every .gov host tested:
| Route | Result |
|---|---|
| Direct | HTTP 200 |
| Through a residential proxy | CONNECT tunnel failed, response 491 |
That 491 is our proxy refusing to tunnel to the host, not the government refusing us. The Register publishes this API for public use and the Actor takes it directly, paced politely.
Typical uses
- Regulatory monitoring. Filter by agency and
comments_open_only, run daily, and you have every rule your industry can still respond to, with the deadline attached. - Compliance calendars.
effective_onis when a rule bites.cfr_referencessays which parts of the code change. - Policy research. Term search across decades, with
raw_text_urlfor the documents you want to read in full. - Competitive and lobbying intelligence.
docket_idslink Register documents to the dockets where comments are filed. - Executive action tracking.
doc_type: presidential_documentwithpresidentreturns executive orders and proclamations.
Notes
significant is the Register's own flag and is frequently null rather than
false; absent is not the same as "not significant" and is returned as-is.
An empty result comes back as a 404 from the API, which this Actor treats as zero documents rather than a failure. A 400 means the search conditions were rejected and is not retried, since retrying a bad condition cannot help.
Coverage is the US Federal Register only. State registers are published separately and are not in this data.
Finding an agency slug
Agency slugs are the agency name lower-cased with hyphens:
environmental-protection-agency, food-and-drug-administration,
securities-and-exchange-commission, federal-communications-commission. The
Actor slugifies whatever you type, so the plain name works too.
A document is frequently issued by several agencies at once, and agencies
returns all of them. Filtering by one agency returns documents where it is any
of the issuers, not only the lead, which is usually what you want and
occasionally a surprise.
What this Actor does not do
No document full text. raw_text_url gives you the address of the plain
text, and fetching it is a separate job with a different cost profile: some
rules run to hundreds of pages.
No public comments. The Register publishes the documents and the deadlines; the comments themselves live on Regulations.gov, which is a different API.
No historical amendments. A rule's row describes the rule as published. How it was later amended lives in the eCFR, not here.