# SEC Filings Scraper - EDGAR Full Text Search, XBRL, 10-K, 8-K (`snow_leo_data/sec-edgar-scraper`) Actor

EDGAR full text search stops at 10,000 hits per query; splitting the period returned 27,224 documents, 2.72x more. SEC EDGAR filings API for company filings since 1993, 8-K item codes, insider Form 4 and XBRL financial statements API. 56 fields, no API key. Official SEC filings, nothing second-hand.

- **URL**: https://apify.com/snow\_leo\_data/sec-edgar-scraper.md
- **Developed by:** [Snow Leo Data](https://apify.com/snow_leo_data) (community)
- **Categories:** News, Integrations
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$2.50 / 1,000 filings

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## SEC EDGAR Filings Search

Filings, company histories and financial numbers straight from the **official SEC
EDGAR APIs** — `efts.sec.gov` for full-text search, `data.sec.gov` for submissions
and XBRL. No API key, no proxy, no headless browser, nothing scraped out of HTML
that can silently change shape overnight.

**$2.50/1K rows.** Press **Start** with nothing filled in and you get a sample of
the newest 8-K current reports, so you can see the shape of a row before you
decide anything.

### The one number that matters

EDGAR full-text search refuses to page past **10,000 documents** for any single
query. Worse, once you are over that line it stops telling you the truth about
the size of your own result set: the response comes back with
`{"value": 10000, "relation": "gte"}` — *at least* ten thousand, and it will not
say how many more. Ask for page 101 and you get an error, not data:

```
search_phase_execution_exception: Result window is too large,
from + size must be less than or equal to: [10000] but was [10100]
```

This Actor treats that as a problem to solve rather than a limit to live with.
It asks the index how big the period is, and while the answer is "at least ten
thousand" it cuts the period in half and asks again, until every window is one
the API will actually serve in full.

Measured on 2026-09-12, searching the exact phrase `"revenue"` in 8-K current
reports filed between 2026-01-01 and 2026-09-10:

| | Documents reachable |
|---|---|
| One query, the way the API is normally used | 10,000, with no idea how many were missed |
| This Actor, same query, period split automatically | **27,224** |

Four windows, seven counting queries, about six seconds of work. That is 2.72x
past the cap, and the run report tells you exactly how the period was cut, so
the number is auditable rather than a marketing claim. Reproduce it yourself
with `python3 tests/test_live.py`.

#### What can I actually collect with it?

Three modes, chosen with one dropdown.

**`fullTextSearch`** looks *inside* the documents. Every word of every filing
since 2001 is indexed by SEC, including exhibits, press releases and the
footnotes nobody reads. Search for `"material weakness"`, `"going concern"`,
`"ransomware"`, the name of a supplier, a drug, a competitor or a law firm, and
you get back every filing that says it, with the form type, the filer, the
filing date, the 8-K item numbers and a direct link to the document.

**`companyFilings`** walks the complete submission history of the companies you
name, by ticker or by CIK. That is every form an issuer has ever filed, back to
1993 — annual reports, quarterly reports, current reports, insider Forms 3, 4
and 5, proxy statements, registration statements, correspondence with the staff.
Each row carries the company profile too: SIC industry code and its description,
EIN, state of incorporation, fiscal year end, filer category, business address
and every former name the company traded under.

**`xbrlFacts`** returns the numbers companies tagged in their own financial
statements. Either one concept across every period a company has reported it
(give tickers), or one concept across every company that reported it in a
period (give a period such as `CY2025Q1` and no tickers). That second form is
the cheapest market-wide financial snapshot there is: one request returned 1,934
companies' quarterly revenue when this was written.

#### Which fields do I get back?

**56 fields** in total across the three modes, all of them straight from SEC and
none of them inferred. The ones people ask about first:

- identity — `accession_number`, `document_file`, `cik`, `company`, `ticker`,
  `tickers`, `exchanges`, `other_filers`, `all_ciks`
- the filing — `form`, `root_form`, `form_meaning`, `filed_at`, `period_ending`,
  `accepted_at`, `act`, `file_number`, `film_number`, `size_bytes`
- 8-K events — `items` and `items_explained`
- the filer — `sic`, `sic_description`, `ein`, `entity_type`,
  `state_of_incorporation`, `fiscal_year_end`, `filer_category`, `phone`,
  `former_names`, `business_street`, `business_city`, `business_state`,
  `business_zip`
- machine-readable data — `is_xbrl`, `is_inline_xbrl`, and in XBRL mode
  `taxonomy`, `tag`, `label`, `value`, `unit`, `period_start`, `period_end`,
  `fiscal_year`, `fiscal_period`, `frame`, `location`
- the document itself — `document_url`, `filing_index_url`, and optionally
  `document_text` with `document_text_chars`
- run bookkeeping — `record_type`, `relevance_score`, `change_type`

An empty field means SEC did not publish that value for that record. Nothing is
guessed, modelled or filled in from elsewhere. A numeric field that has no value
is `null`, never an empty string, so the dataset loads into pandas, BigQuery or
Excel without a type error on the first row.

#### Why are the 8-K item codes worth anything?

Because a current report is meaningless until you know which item it was filed
under, and EDGAR gives you only the number. A filing tagged `1.05` is a company
telling the market it has had a material cybersecurity incident. A filing tagged
`4.02` is a company saying its previously issued financial statements can no
longer be relied on. A filing tagged `5.02` is a director or officer leaving.

Every row carries both. `items` holds the raw codes so you can filter on them,
and `items_explained` pairs each code with its official title, so a row reads as
`{"code": "1.05", "title": "Material Cybersecurity Incidents"}` without anyone
having to keep a lookup table. 32 official 8-K item numbers are covered, and a
code that is not in the list still comes back with its number and an honest
`null` title rather than a made-up one.

The same idea applies to form types: `form_meaning` says in plain English what
the form is for, and 28 form types are named — so `SC 13G` reads as beneficial
ownership above five percent with passive intent, and `NT 10-K` reads as a
notification of a late annual report. An amendment keeps its own `form` (`8-K/A`)
while `root_form` and `form_meaning` come from the parent, so you can group
amendments with what they amend without string-slicing form names yourself.

#### How do I run this on a schedule without paying twice?

Switch on **Only what changed since the last run**. The Actor keeps a compact
memory of what it has already handed you in a named key-value store that
survives between runs, and on the next run it delivers only what is new or
amended. Each row is tagged `NEW`, `UPDATED` or `UNCHANGED`, and unchanged rows
are not returned at all unless you ask for them.

This is what makes a daily schedule affordable. A search that matches thirty
thousand filings costs you thirty thousand rows once; after that you pay for the
few dozen that appeared overnight. The memory holds 70,000 keys and, when it
fills, drops the oldest first — recent filings are the ones a monitor cares
about.

The fingerprint behind `UPDATED` deliberately ignores fields that churn. A
relevance score that shifts because the index was rebuilt is not a change to the
filing, and treating it as one would mean charging you for the whole result set
every night. Form, dates, item codes, document size and, for XBRL, the value
itself are what count as a change.

Two documents belonging to the same filing are remembered separately. A single
8-K can carry nine exhibits; keyed on the accession number alone, eight of them
would vanish and you would never know they existed.

#### What happens when the run breaks in the middle?

You keep everything that was delivered and nothing is silently lost. The order
of operations is fixed and tested: rows are pushed to the dataset **first**, and
only then marked as delivered in memory. If the container is moved, the run
times out or the network drops, the next run picks up exactly the rows that
never made it out, and no row is marked delivered that a buyer never received.

That order is not obvious and it is easy to get backwards, so the test suite
breaks a run on purpose halfway through and fails if memory ever runs ahead of
delivery.

#### Will I be charged for rows I did not want?

Filtering happens before anything is written to the dataset, so anything you
filtered out never reaches your bill. `keywords`, `excludeKeywords`,
`excludeForms`, `items`, `sicCodes`, `states`, `minValue` and `maxValue` all
apply at that point, and the run report lists how many rows each one removed —
so an empty result reads as "your filter was strict", not as "the source is
broken".

Two more guards sit in the same place. Rows with no accession number and no CIK
are dropped rather than billed: EDGAR does occasionally return a hit you cannot
build a link from, and a row you cannot follow is worth nothing. And your Apify
spending limit is enforced by this Actor itself — the platform stops *charging*
at your limit but does not stop the *run*, which means an unguarded Actor keeps
burning compute on rows nobody is paying for.

#### How do I stop one huge company eating the whole run?

Use **Rows per company**. Ask for twenty tickers with a limit of a thousand rows
and, without a quota, the first large issuer hands you a thousand filings on its
own and the other nineteen never get a turn. Left at zero, the run limit is
shared out evenly between the companies you named; set it explicitly to take, say,
the fifty most recent filings from each.

By default company mode returns the most recent 1,000 filings per company,
because that is what the submissions feed serves in one request. Switch on
**Include the archive** to walk the older pages back to 1993 as well — one extra
request per company, and for a long-lived issuer over a thousand extra rows.

#### Do I need a proxy, a key or an account with SEC?

No. All three endpoints are public and free, and this Actor uses only the Python
standard library to reach them. What SEC does ask for is that automated callers
identify themselves in the `User-Agent` header — its Fair Access policy is
explicit about it, and a request without one is refused with HTTP 403. A working
identifier is already set, and the **Your contact email** field lets you put your
own address there instead, which is what SEC prefers.

The same policy caps callers at ten requests a second. This Actor holds itself
to eight, shared across every thread, because exceeding the limit gets the
address blocked for ten minutes and that costs you a whole run to save a few
seconds.

#### Can I get the text of the filing itself?

Yes — switch on **Include the document text**. Each row then carries
`document_text`, the document with its markup stripped, capped at 200000
characters, plus `document_text_chars` so you know whether the cap bit. It is off
by default because it costs one extra request per row and makes rows far larger;
for most monitoring the link in `document_url` is enough, and for language work
the text is the whole point.

#### What does this Actor not do?

Named honestly, because finding out after you have paid is worse than reading it
here.

- **It does not parse Form 4 transaction tables.** You get the insider filing,
  its date, the issuer and the link, but not the parsed rows of shares, price
  and transaction code. Actors dedicated to insider trading do that.
- **It does not parse 13F holdings tables.** Same story: the quarterly
  institutional report comes back as a filing, not as a list of positions.
- **It does not normalise financial statements.** XBRL facts come back exactly
  as the company tagged them, concept by concept. It does not assemble them into
  a balance sheet or an income statement, and it does not reconcile different
  companies' tagging choices.
- **It writes no AI summaries.** Every value in every row came from SEC. Nothing
  in the output is generated, scored or interpreted by a model, which is a
  limitation if you wanted a summary and a feature if you are feeding a
  compliance process.
- **Full-text search starts in 2001.** SEC does not index document bodies before
  then. Company filing history reaches back to 1993, because that feed does.
- **It has no proxy configuration.** None is needed for these endpoints, and
  offering a setting that does nothing would be worse than leaving it out.

#### How much does a run cost?

$2.50/1K rows, charged per row written to the dataset. A daily monitor in
incremental mode typically writes tens of rows, not thousands, once the first
run has filled its memory. Counting queries used to split a period are not
billed — only rows you receive are.

With no search phrase and no limit of your own, a run is treated as a trial and
stops at 300 rows, so a first press of Start cannot produce a surprise.

### FAQ

#### Which is faster, searching by phrase or by company?

By company. `companyFilings` is one request per company and returns up to a
thousand filings from it, while a full-text search over a wide period has to
count the period first and then page through it a hundred documents at a time.
If you know exactly whose filings you want, name the tickers.

#### Can I search for a phrase within one company only?

Yes. Put the phrase in **Search phrase** and the company's CIK in **CIK
numbers**, and SEC restricts the search itself rather than the Actor filtering
afterwards — which means you are not billed for documents from other filers.
**Company name contains** does the same thing by name when you do not have the
CIK to hand.

#### Why does my search return fewer rows than the report says exist?

Because the report tells you how many documents the index holds for your period,
which is deliberately not the same as how many you asked for. It is there so you
can see what a wider limit would buy you before you pay for it. Raise **Maximum
rows**, or narrow the period.

#### What if a single day holds more than the cap by itself?

Then the API genuinely cannot return the rest of that day, and the run says so
instead of pretending otherwise: the day appears in the report under
`capped_days` and a warning is written to the log. Splitting a period stops at
one day because that is the smallest window EDGAR will filter on. In practice
this only happens on very common words with no form filter.

#### Do amendments show up separately?

Yes, as their own rows, which is what you want — an amended 8-K is news in its
own right. `form` holds `8-K/A`, `root_form` holds `8-K`, and a form filter of
`8-K` matches both. In incremental mode an amendment to something you already
received arrives tagged `UPDATED`.

#### Which XBRL tag should I ask for?

`Revenues`, `Assets`, `Liabilities`, `NetIncomeLoss`, `StockholdersEquity` and
`CashAndCashEquivalentsAtCarryingValue` cover most questions. Concepts live in
the `us-gaap` taxonomy for US filers, `ifrs-full` for foreign private issuers,
and `dei` for entity facts such as the share count. Not every company tags every
concept, so a company missing from a market-wide frame usually tagged something
more specific rather than reporting nothing.

#### What is the difference between CY2025Q1 and CY2025Q1I?

`CY2025Q1` is a period — something that happened over the quarter, like revenue.
`CY2025Q1I` is an instant — something measured at a point in time, like the
balance of accounts payable on the last day. Asking for a flow with an instant
frame, or the reverse, returns nothing; that is SEC's convention, not a fault in
the run.

#### Can I feed the rows straight to an AI agent?

Switch on **Compact rows** and **Drop empty fields**. The first keeps only the
identifying fields and the links, the second leaves out keys with no value
instead of returning them as null. Together they cut a row to roughly a fifth of
its full size, which matters when every field costs context.

#### Is the data official?

Yes, and only that. Every row comes from `efts.sec.gov` or `data.sec.gov`, both
operated by the U.S. Securities and Exchange Commission and both free to the
public. This Actor adds structure, splitting, filtering and memory between runs;
it adds no data of its own, and every value can be checked against the link in
`filing_index_url`.

# Actor input Schema

## `mode` (type: `string`):

Full-text search looks inside the documents themselves (2001 to today). Company filings walks the complete submission history of a ticker or CIK (1993 to today). XBRL facts returns the numbers companies tagged in their financial statements.

## `query` (type: `string`):

Words to find inside the filing text. Wrap in double quotes for an exact phrase: `"material weakness"`. Without quotes any of the words counts. Leave empty to take every filing of the chosen form types.

## `companyName` (type: `string`):

Narrow a full-text search to filers whose registered name contains this, for example `Apple`. Matched by SEC itself, so spelling follows the EDGAR register.

## `forms` (type: `array`):

EDGAR form names such as `8-K`, `10-K`, `10-Q`, `4`, `13F-HR`, `SC 13D`, `DEF 14A`, `D`. Amendments are included with their parent form. Empty means every form — in company mode that is the whole filing history, not just current reports.

## `tickers` (type: `array`):

Stock tickers such as `AAPL`, `MSFT`. Resolved against the official SEC ticker list. Used by the company filings and XBRL modes; in full-text search use CIKs instead.

## `ciks` (type: `array`):

SEC central index keys, with or without leading zeros: `320193` and `0000320193` both work. Use these for filers that have no ticker, such as funds and individual insiders.

## `filedWithinDays` (type: `integer`):

Used when no explicit start date is given. The default keeps a scheduled run cheap instead of quietly pulling twenty years of archive.

## `filedFrom` (type: `string`):

Earliest filing date, `YYYY-MM-DD`. Full-text search has no documents before 2001-01-01, so anything earlier is raised to that date.

## `filedTo` (type: `string`):

Latest filing date, `YYYY-MM-DD`. Empty means today.

## `locationCode` (type: `string`):

Two-letter state or country code of the filer's principal office, for example `CA` or `NY`. Full-text search only.

## `includeDocumentText` (type: `boolean`):

Fetch each filing document and return its text with the markup stripped, capped at 200,000 characters per document. Off by default: it costs one extra request per row and makes rows much larger.

## `includeHistory` (type: `boolean`):

Company filings mode returns the most recent 1,000 filings by default. Switch this on to walk the archive pages back to 1993 as well — one extra request per company and, for a large issuer, over a thousand extra rows.

## `perCompanyLimit` (type: `integer`):

Cap on rows taken from each ticker or CIK. Without it the first large issuer eats the whole run limit and a twenty-ticker request looks like a one-company export. Left at zero the limit is shared out evenly.

## `taxonomy` (type: `string`):

`us-gaap`, `ifrs-full`, `dei` or `srt`. XBRL facts mode only.

## `tag` (type: `string`):

The tagged concept, for example `Revenues`, `Assets`, `NetIncomeLoss`, `AccountsPayableCurrent`. XBRL facts mode only.

## `unit` (type: `string`):

Unit of measure, usually `USD`. Use `shares` for share counts. XBRL facts mode only.

## `period` (type: `string`):

Calendar period such as `CY2025Q1` for a quarter, `CY2025Q1I` for a balance-sheet instant, `CY2024` for a year. Leave empty and give tickers instead to get one company across all periods; fill it in with no tickers to get every company that reported that period in one request.

## `keywords` (type: `array`):

Keep only rows whose company name, form, description or XBRL label contains at least one of these. Applied before you are charged.

## `excludeKeywords` (type: `array`):

Drop rows matching any of these. Applied before you are charged.

## `excludeForms` (type: `array`):

Form names to throw away, for example `4` to drop routine insider paperwork from a company's history.

## `items` (type: `array`):

Keep only current reports carrying one of these item numbers, for example `1.05` for a material cybersecurity incident or `5.02` for a change of directors or officers. Every code that comes back is also returned with its official title.

## `sicCodes` (type: `array`):

Keep only filers in these industries, for example `7372` for prepackaged software.

## `states` (type: `array`):

Two-letter codes matched against the filer's business address and state of incorporation, for example `CA`, `NY`, `DE`.

## `minValue` (type: `integer`):

XBRL facts mode: drop facts below this number.

## `maxValue` (type: `integer`):

XBRL facts mode: drop facts above this number.

## `xbrlOnly` (type: `boolean`):

Company filings mode: keep only submissions that carry machine-readable XBRL financial data.

## `maxItems` (type: `integer`):

Hard stop for this run. Zero means no limit of your own; the run still stops at your spending limit. With no search phrase and no limit the run is treated as a trial and stops at 300 rows.

## `onlyNew` (type: `boolean`):

Remembers what it already handed you in a named storage that survives between runs, and returns only new and amended rows. This is what makes a daily schedule cheap: you pay for the dozens that appeared, not the thousands that did not move.

## `emitUnchanged` (type: `boolean`):

Incremental mode normally hides rows that have not changed. Switch this on to get them too, each tagged `NEW`, `UPDATED` or `UNCHANGED`.

## `compactOutput` (type: `boolean`):

Keep only the identifying fields and the links. Useful when the rows are fed to an AI agent and every field costs context.

## `excludeEmptyFields` (type: `boolean`):

Leave out keys that have no value instead of returning them as null.

## `contactEmail` (type: `string`):

SEC asks every automated caller to identify itself in the User-Agent header. A working default is already set; put your own address here to identify the traffic as yours, which is what the SEC Fair Access policy prefers.

## Actor input object example

```json
{
  "mode": "fullTextSearch",
  "query": "\"material weakness\"",
  "forms": [
    "8-K"
  ],
  "tickers": [
    "AAPL"
  ],
  "filedWithinDays": 30,
  "includeDocumentText": false,
  "includeHistory": false,
  "perCompanyLimit": 0,
  "taxonomy": "us-gaap",
  "tag": "Revenues",
  "unit": "USD",
  "xbrlOnly": false,
  "maxItems": 0,
  "onlyNew": false,
  "emitUnchanged": false,
  "compactOutput": false,
  "excludeEmptyFields": false
}
```

# Actor output Schema

## `results` (type: `string`):

All collected rows

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mode": "fullTextSearch",
    "query": "\"material weakness\"",
    "forms": [
        "8-K"
    ],
    "tickers": [
        "AAPL"
    ],
    "filedWithinDays": 30
};

// Run the Actor and wait for it to finish
const run = await client.actor("snow_leo_data/sec-edgar-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "mode": "fullTextSearch",
    "query": "\"material weakness\"",
    "forms": ["8-K"],
    "tickers": ["AAPL"],
    "filedWithinDays": 30,
}

# Run the Actor and wait for it to finish
run = client.actor("snow_leo_data/sec-edgar-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mode": "fullTextSearch",
  "query": "\\"material weakness\\"",
  "forms": [
    "8-K"
  ],
  "tickers": [
    "AAPL"
  ],
  "filedWithinDays": 30
}' |
apify call snow_leo_data/sec-edgar-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,snow_leo_data/sec-edgar-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/D1hOYX911cB70gUWX/builds/md1wbLJzqzbj2wFgd/openapi.json
