Federal Register Scraper avatar

Federal Register Scraper

Pricing

from $1.53 / 1,000 federal register document extracteds

Go to Apify Store
Federal Register Scraper

Federal Register Scraper

Search official Federal Register rules, proposed rules, notices, and presidential documents with agency, docket, date, CFR, RIN, abstract, and source-link metadata.

Pricing

from $1.53 / 1,000 federal register document extracteds

Rating

0.0

(0)

Developer

Stas Persiianenko

Stas Persiianenko

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Categories

Share

Search and export official Federal Register rules, proposed rules, notices, and presidential documents as clean dataset records. Filter documents by keyword, agency, type, publication date, comment deadline, or exact document number, then collect abstracts, docket references, RINs, CFR citations, important dates, and official source links.

The Actor uses the anonymous FederalRegister.gov API. It adds validated inputs, pagination, detail enrichment, deduplication, Apify scheduling, exports, webhooks, API access, and MCP access without requiring a browser, proxy, or source account.

What can this Federal Register scraper do?

  • Search rules, proposed rules, notices, and presidential documents by topic.
  • Limit results to one or more FederalRegister.gov agency slugs.
  • Filter publication dates and public-comment closing dates.
  • Retrieve exact Federal Register document numbers.
  • Sort search results by newest, oldest, or relevance.
  • Enrich each result with docket IDs, Regulation Identifier Numbers, CFR references, topics, actions, dates, and official text links.
  • Save an exact maximum number of unique documents.
  • Export JSON, CSV, Excel, XML, RSS, or JSONL from the default dataset.
  • Run on a schedule for regulatory monitoring workflows.

Who is it for?

Regulatory affairs teams can collect new rules and proposed rules by agency or subject.

Legal and compliance researchers can create source-backed datasets with citations, dockets, effective dates, and comment deadlines.

Policy analysts and journalists can search abstracts and official metadata without maintaining Federal Register API pagination code.

Government-relations teams can schedule recurring searches and compare stable document numbers between runs.

Developers and data teams can feed normalized records into databases, dashboards, spreadsheets, LLM workflows, or alerting systems.

Why use this Actor instead of calling the Federal Register API directly?

FederalRegister.gov already provides an excellent public API. This Actor packages that source into a reusable product with an input form, strict validation, bounded retries, pagination, optional detail requests, normalized output, usage-based charging, datasets, schedules, webhooks, and integrations.

Every result keeps official identifiers and URLs. The Actor does not scrape visual page markup or hide the data source.

What Federal Register data is extracted?

FieldDescription
title, abstract, actionOfficial title, abstract, and action statement
documentNumberStable Federal Register document number
documentType, subtypeRule, proposed rule, notice, or presidential document classification
publicationDateDate published in the Federal Register
effectiveDateEffective date when supplied
commentsCloseDatePublic-comment deadline when supplied
datesText, signingDateOfficial dates narrative and signing date
agencies, agencyNamesStructured issuing-agency metadata
topicsFederalRegister.gov indexed topics
docketIds, docketsDocket references and structured docket metadata
regulationIdsRegulation Identifier Numbers (RINs)
cfrReferences, citationCFR references and Federal Register citation
executiveOrderNumberExecutive order number when applicable
htmlUrl, pdfUrlOfficial HTML and PDF links
jsonUrl, rawTextUrl, fullTextXmlUrlOfficial machine-readable source links
regulationsGovUrl, commentUrlRelated Regulations.gov and comment links when available
sourceUrl, scrapedAtAPI provenance and collection timestamp

Fields are omitted when the official source does not provide them. Arrays remain empty when no agency, topic, docket, RIN, or CFR reference is present.

How to search Federal Register rules and notices

  1. Open the Actor input page.
  2. Enter a topic such as artificial intelligence.
  3. Select the document types you need.
  4. Add agency slugs or date bounds if required.
  5. Keep Fetch full metadata enabled for docket, RIN, CFR, topic, and official-text details.
  6. Set a small maxItems value for your first run.
  7. Start the run and open the default dataset.

Example input:

{
"searchTerm": "artificial intelligence",
"documentTypes": ["RULE", "PRORULE", "NOTICE"],
"sortBy": "newest",
"includeDetails": true,
"maxItems": 25
}

Input parameters

FieldTypeDefaultDescription
searchTermstringTopic or words searched in Federal Register documents.
agencySlugsstring[][]FederalRegister.gov slugs such as environmental-protection-agency.
documentTypesstring[]rules, proposed rules, noticesAny of RULE, PRORULE, NOTICE, or PRESDOCU.
publicationDateFromdateEarliest publication date in YYYY-MM-DD format.
publicationDateTodateLatest publication date in YYYY-MM-DD format.
commentDateFromdateEarliest public-comment closing date.
commentDateTodateLatest public-comment closing date.
documentNumbersstring[][]Exact source document numbers. Other supplied filters still apply.
sortBystringnewestnewest, oldest, or relevance.
includeDetailsbooleantrueRequest full document metadata for each search result.
maxItemsinteger100Maximum unique saved records, from 1 to 10,000.

When documentNumbers is supplied, the Actor retrieves those documents directly instead of running a search. All agency, type, topic, and date filters still apply to direct results.

Output example

A real result has this shape (abbreviated):

{
"title": "Request for Information on the Development of an Artificial Intelligence Action Plan",
"documentNumber": "2025-02305",
"documentType": "Notice",
"publicationDate": "2025-02-06",
"agencies": [
{
"name": "Executive Office of the President",
"slug": "executive-office-of-the-president"
}
],
"agencyNames": ["Executive Office of the President"],
"topics": [],
"docketIds": [],
"regulationIds": [],
"cfrReferences": [],
"htmlUrl": "https://www.federalregister.gov/documents/2025/02/06/2025-02305/request-for-information-on-the-development-of-an-artificial-intelligence-action-plan",
"jsonUrl": "https://www.federalregister.gov/api/v1/documents/2025-02305.json",
"sourceUrl": "https://www.federalregister.gov/api/v1/documents/2025-02305.json",
"scrapedAt": "2026-08-28T20:00:00.000Z"
}

The exact record content is controlled by the official source and can change when FederalRegister.gov updates metadata.

How much does it cost to export Federal Register documents?

Pricing uses one small start event plus one item event for every accepted dataset record. Duplicate, rejected, empty, or failed records are not charged as items.

The exact tier applying to your account is shown by Apify before a run. At the current Bronze configuration, the start fee is $0.005 and each accepted document is $0.002544. That makes 25 documents about $0.0686, 100 documents about $0.2594, and 1,000 documents about $2.549 in Actor charges. Higher usage tiers have lower per-document prices.

Apify platform usage is included in pay-per-event pricing. The upstream Federal Register API does not require a paid key.

Agency and document-type workflows

Use agency slugs exactly as FederalRegister.gov identifies them. For example:

{
"agencySlugs": ["environmental-protection-agency"],
"documentTypes": ["RULE", "PRORULE"],
"publicationDateFrom": "2025-01-01",
"publicationDateTo": "2025-12-31",
"maxItems": 100
}

For open-comment research, combine proposed rules or notices with commentDateFrom and commentDateTo. Some documents do not have a comment deadline; those records are correctly excluded when a comment-date filter is active.

Exact document lookup

Use stable document numbers when another system already knows the records to retrieve:

{
"documentNumbers": ["2025-02305"],
"includeDetails": true,
"maxItems": 1
}

Exact lookup is useful for enrichment, verification, and repeatable source citations. A mismatching topic, agency, type, or date filter can intentionally produce an empty dataset.

Scheduling and regulatory monitoring

Save a tested input as an Apify Task and schedule it daily, weekly, or monthly. Use documentNumber as the stable key when comparing datasets across runs. Send a webhook after successful runs to trigger a database load, spreadsheet update, Slack workflow, or internal alert.

The Actor exports snapshots. It does not itself calculate changes, continuously watch the source, or send legal alerts. Your downstream workflow should compare datasets and decide which changes matter.

Use from the Apify API with cURL

curl -X POST \
"https://api.apify.com/v2/acts/automation-lab~federal-register-rules-notices/runs?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"searchTerm":"clean energy","documentTypes":["RULE","PRORULE"],"maxItems":25}'

Poll the returned run ID or configure a webhook before production use. Keep your Apify token in an environment variable or secret manager.

Use from JavaScript

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('automation-lab/federal-register-rules-notices').call({
searchTerm: 'artificial intelligence',
documentTypes: ['RULE', 'PRORULE', 'NOTICE'],
includeDetails: true,
maxItems: 50,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);

Use from Python

import os
from apify_client import ApifyClient
client = ApifyClient(os.environ['APIFY_TOKEN'])
run = client.actor('automation-lab/federal-register-rules-notices').call(run_input={
'agencySlugs': ['environmental-protection-agency'],
'documentTypes': ['RULE', 'PRORULE'],
'maxItems': 100,
})
items = client.dataset(run['defaultDatasetId']).list_items().items
print(items)

Use with MCP and AI assistants

Add the Actor to Claude Code:

claude mcp add --transport http apify \
"https://mcp.apify.com?tools=automation-lab/federal-register-rules-notices"

For Claude Desktop, Cursor, or VS Code, use:

{
"mcpServers": {
"apify": {
"url": "https://mcp.apify.com?tools=automation-lab/federal-register-rules-notices"
}
}
}

Example prompts:

  • “Search Federal Register proposed rules about artificial intelligence and summarize agency, docket, and comment deadlines.”
  • “Export 2025 EPA final rules and return document numbers, effective dates, CFR references, and official URLs.”
  • “Retrieve Federal Register document 2025-02305 and cite its official HTML and PDF sources.”

Reliability, retries, and source limits

The Actor requests the official JSON API directly. HTTP 429 and 5xx responses are retried up to three times with bounded exponential backoff. Permanent API errors fail the run instead of silently producing incomplete output.

Search pages contain up to 100 records. Detail enrichment runs at bounded concurrency. maxItems stops accepted output at the requested count, up to 10,000.

FederalRegister.gov controls availability, coverage, indexing, field definitions, and freshness. A successful query with no matching documents produces an empty dataset. The Actor does not bypass source restrictions or enrich records from private systems.

Disable includeDetails only when faster list-level metadata is sufficient. Dockets, RINs, CFR references, topics, action text, and machine-readable links may require detail enrichment.

Federal Register documents are official public U.S. government records. Use the data responsibly, retain source URLs for auditability, and verify consequential conclusions against the official document text.

This Actor is a retrieval and normalization tool, not legal advice. A missing or delayed source record does not prove that no regulatory action exists. Respect Apify terms, FederalRegister.gov policies, applicable laws, and your organization's retention and compliance requirements.

Troubleshooting

The dataset is empty. Remove filters one at a time. Confirm agency slugs, date order, document types, and exact document numbers. Comment-date filters intentionally reject documents without a matching closing date.

The run reports an invalid date. Use complete dates in YYYY-MM-DD format and ensure each “from” date is not later than its “to” date.

The run fails with 429 or 5xx. The Actor retries transient failures. If the source remains unavailable, wait before rerunning rather than starting repeated parallel jobs.

Some enriched fields are empty. The official source does not provide every field for every document. Keep includeDetails enabled for the richest available metadata.

FAQ

Does the Actor require a FederalRegister.gov API key?

No. It uses public anonymous FederalRegister.gov API endpoints.

Does it download full document text?

It exports official HTML, PDF, raw-text, and XML links when available. It does not download and embed entire PDF or XML bodies in each dataset row.

Can it monitor executive orders?

Yes. Include PRESDOCU in documentTypes, optionally add a keyword, and schedule the input. Presidential-document records can include executiveOrderNumber when the source provides it.

Can I filter by topic and agency together?

Yes. Search, agency, type, publication-date, and comment-date filters can be combined. Every accepted row must satisfy the supplied filters.

How do I avoid duplicate records?

The Actor deduplicates by documentNumber within each run. For recurring monitoring, compare that stable field across run datasets.

Is this an official Federal Register product?

No. This is an independent Apify Actor that retrieves public official data from FederalRegister.gov and preserves links back to the source.