Federal Register Documents Scraper avatar

Federal Register Documents Scraper

Pricing

from $2.00 / 1,000 results

Go to Apify Store
Federal Register Documents Scraper

Federal Register Documents Scraper

Scrape US Federal Register rules, proposed rules, notices and presidential documents: title, abstract, agencies, dockets, CFR references, comment deadlines and full-text links. Keyless public API, no login.

Pricing

from $2.00 / 1,000 results

Rating

0.0

(0)

Developer

Farhan Febrian Nauval

Farhan Febrian Nauval

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Categories

Share

Federal Register Documents Scraper — US Rules, Notices & Presidential Documents

Scrape the US Federal Register: final and proposed rules, notices and presidential documents, with titles, abstracts, agencies, docket and RIN numbers, CFR references, comment deadlines, page citations and links to the full text in HTML, XML, plain text and PDF.

Keyless public government API. No login, no browser, and no proxy needed — there is no anti-bot layer on this source.

Input

Every field is optional, but you must give at least one of a search term, an agency, a document type or a date bound.

FieldTypeDefaultDescription
searchTermsarray—One full-text search per entry
agenciesarray—Agency slugs, e.g. securities-and-exchange-commission
documentTypesarray—RULE, PRORULE, NOTICE, PRESDOCU
publishedFrom / publishedTostring—YYYY-MM-DD bounds
orderstringnewestnewest, oldest or relevance
maxItemsPerSearchinteger500Capped by the API at 10,000
pageSizeinteger1000Up to 2000

Searching by filters alone is a first-class use: leave searchTerms empty and ask for everything one agency published in a date range.

Output

{
"_input": "climate disclosure",
"_source": "S1-federalregister-api",
"_scrapedAt": "2026-09-09T12:04:18Z",
"documentNumber": "2026-18317",
"title": "Updated Definition of \"Waters of the United States\"",
"type": "Proposed Rule", "subtype": null,
"abstract": "…", "action": "Proposed rule.",
"publicationDate": "2026-09-09",
"signingDate": null, "effectiveOn": null,
"commentsCloseOn": "2026-10-09",
"datesText": "Comments must be received on or before October 9, 2026.",
"agencyNames": ["Defense Department", "Engineers Corps", "Environmental Protection Agency"],
"agencySlugs": ["defense-department", "engineers-corps", "environmental-protection-agency"],
"presidentName": "Donald Trump", "presidentId": "donald-trump",
"docketId": "EPA-HQ-OW-2025-0117",
"docketIds": ["…"], "regulationIdNumbers": ["2040-AGngs"],
"cfrReferences": [{ "title": 33, "part": "328", "chapter": null, "citation_url": null }],
"topics": ["Water pollution control"], "significant": true,
"citation": "91 FR 57284", "volume": 91,
"startPage": 57284, "endPage": 57310, "pageLength": 27,
"htmlUrl": "…", "pdfUrl": "…", "jsonUrl": "…",
"rawTextUrl": "…", "fullTextXmlUrl": "…", "bodyHtmlUrl": "…",
"publicInspectionPdfUrl": "…",
"commentsCount": 0,
"regulationsDotGovDocumentId": "EPA-HQ-OW-2025-0117-0001",
"correctionOf": "https://www.federalregister.gov/api/v1/documents/2026-16628",
"correctionOfDocumentNumber": "2026-16628",
"corrections": [], "imageCount": 3
}

Three things worth knowing

Past the last page, the API silently serves page 1 again. With per_page=20 the response reports total_pages: 50, and pages 51, 52, 100, 200 and 500 all returned the identical first page — HTTP 200, no error, no warning. A crawler that pages until it sees an empty response will re-emit the same documents forever and count them as new. This actor stops at total_pages and also de-duplicates by document number. (At per_page=2000 the same overrun returns an honest HTTP 400 instead, so the behaviour depends on page size.)

A count of exactly 10,000 means "at least 10,000". The API clamps it: unrelated broad searches all report exactly that number, while narrow ones report real totals — 2,005 for a single month, 8,560 for presidential documents. Paging is capped at the same 10,000. The actor logs at least N rather than N when it sees the ceiling, and tells you to narrow the search; a date range is the most effective way to do that.

correctionOf is a URL, not a document number. Correction C1-2026-16628 carries …/api/v1/documents/2026-16628. The bare number is provided as correctionOfDocumentNumber so corrections can be joined to what they correct without every consumer parsing the URL.

Errors

_errorMeaning
no_resultsThe search ran and matched no documents
unexpected_shape200 without the results list — the API changed
blockedEvery TLS profile was refused
network_errorThe ladder never reached the server — DNS, timeout, reset

If every search fails, the run itself fails rather than reporting success over an empty dataset.