SEC Filings Scraper - EDGAR Full Text Search, XBRL, 10-K, 8-K avatar

SEC Filings Scraper - EDGAR Full Text Search, XBRL, 10-K, 8-K

Pricing

$2.50 / 1,000 filings

Go to Apify Store
SEC Filings Scraper - EDGAR Full Text Search, XBRL, 10-K, 8-K

SEC Filings Scraper - EDGAR Full Text Search, XBRL, 10-K, 8-K

EDGAR full text search stops at 10,000 hits per query; splitting the period returned 27,224 documents, 2.72x more. SEC EDGAR filings API for company filings since 1993, 8-K item codes, insider Form 4 and XBRL financial statements API. 56 fields, no API key. Official SEC filings, nothing second-hand.

Pricing

$2.50 / 1,000 filings

Rating

0.0

(0)

Developer

Snow Leo Data

Snow Leo Data

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Share

SEC EDGAR Filings Search

Filings, company histories and financial numbers straight from the official SEC EDGAR APIs — efts.sec.gov for full-text search, data.sec.gov for submissions and XBRL. No API key, no proxy, no headless browser, nothing scraped out of HTML that can silently change shape overnight.

$2.50/1K rows. Press Start with nothing filled in and you get a sample of the newest 8-K current reports, so you can see the shape of a row before you decide anything.

The one number that matters

EDGAR full-text search refuses to page past 10,000 documents for any single query. Worse, once you are over that line it stops telling you the truth about the size of your own result set: the response comes back with {"value": 10000, "relation": "gte"} — at least ten thousand, and it will not say how many more. Ask for page 101 and you get an error, not data:

search_phase_execution_exception: Result window is too large,
from + size must be less than or equal to: [10000] but was [10100]

This Actor treats that as a problem to solve rather than a limit to live with. It asks the index how big the period is, and while the answer is "at least ten thousand" it cuts the period in half and asks again, until every window is one the API will actually serve in full.

Measured on 2026-09-12, searching the exact phrase "revenue" in 8-K current reports filed between 2026-01-01 and 2026-09-10:

Documents reachable
One query, the way the API is normally used10,000, with no idea how many were missed
This Actor, same query, period split automatically27,224

Four windows, seven counting queries, about six seconds of work. That is 2.72x past the cap, and the run report tells you exactly how the period was cut, so the number is auditable rather than a marketing claim. Reproduce it yourself with python3 tests/test_live.py.

What can I actually collect with it?

Three modes, chosen with one dropdown.

fullTextSearch looks inside the documents. Every word of every filing since 2001 is indexed by SEC, including exhibits, press releases and the footnotes nobody reads. Search for "material weakness", "going concern", "ransomware", the name of a supplier, a drug, a competitor or a law firm, and you get back every filing that says it, with the form type, the filer, the filing date, the 8-K item numbers and a direct link to the document.

companyFilings walks the complete submission history of the companies you name, by ticker or by CIK. That is every form an issuer has ever filed, back to 1993 — annual reports, quarterly reports, current reports, insider Forms 3, 4 and 5, proxy statements, registration statements, correspondence with the staff. Each row carries the company profile too: SIC industry code and its description, EIN, state of incorporation, fiscal year end, filer category, business address and every former name the company traded under.

xbrlFacts returns the numbers companies tagged in their own financial statements. Either one concept across every period a company has reported it (give tickers), or one concept across every company that reported it in a period (give a period such as CY2025Q1 and no tickers). That second form is the cheapest market-wide financial snapshot there is: one request returned 1,934 companies' quarterly revenue when this was written.

Which fields do I get back?

56 fields in total across the three modes, all of them straight from SEC and none of them inferred. The ones people ask about first:

  • identity — accession_number, document_file, cik, company, ticker, tickers, exchanges, other_filers, all_ciks
  • the filing — form, root_form, form_meaning, filed_at, period_ending, accepted_at, act, file_number, film_number, size_bytes
  • 8-K events — items and items_explained
  • the filer — sic, sic_description, ein, entity_type, state_of_incorporation, fiscal_year_end, filer_category, phone, former_names, business_street, business_city, business_state, business_zip
  • machine-readable data — is_xbrl, is_inline_xbrl, and in XBRL mode taxonomy, tag, label, value, unit, period_start, period_end, fiscal_year, fiscal_period, frame, location
  • the document itself — document_url, filing_index_url, and optionally document_text with document_text_chars
  • run bookkeeping — record_type, relevance_score, change_type

An empty field means SEC did not publish that value for that record. Nothing is guessed, modelled or filled in from elsewhere. A numeric field that has no value is null, never an empty string, so the dataset loads into pandas, BigQuery or Excel without a type error on the first row.

Why are the 8-K item codes worth anything?

Because a current report is meaningless until you know which item it was filed under, and EDGAR gives you only the number. A filing tagged 1.05 is a company telling the market it has had a material cybersecurity incident. A filing tagged 4.02 is a company saying its previously issued financial statements can no longer be relied on. A filing tagged 5.02 is a director or officer leaving.

Every row carries both. items holds the raw codes so you can filter on them, and items_explained pairs each code with its official title, so a row reads as {"code": "1.05", "title": "Material Cybersecurity Incidents"} without anyone having to keep a lookup table. 32 official 8-K item numbers are covered, and a code that is not in the list still comes back with its number and an honest null title rather than a made-up one.

The same idea applies to form types: form_meaning says in plain English what the form is for, and 28 form types are named — so SC 13G reads as beneficial ownership above five percent with passive intent, and NT 10-K reads as a notification of a late annual report. An amendment keeps its own form (8-K/A) while root_form and form_meaning come from the parent, so you can group amendments with what they amend without string-slicing form names yourself.

How do I run this on a schedule without paying twice?

Switch on Only what changed since the last run. The Actor keeps a compact memory of what it has already handed you in a named key-value store that survives between runs, and on the next run it delivers only what is new or amended. Each row is tagged NEW, UPDATED or UNCHANGED, and unchanged rows are not returned at all unless you ask for them.

This is what makes a daily schedule affordable. A search that matches thirty thousand filings costs you thirty thousand rows once; after that you pay for the few dozen that appeared overnight. The memory holds 70,000 keys and, when it fills, drops the oldest first — recent filings are the ones a monitor cares about.

The fingerprint behind UPDATED deliberately ignores fields that churn. A relevance score that shifts because the index was rebuilt is not a change to the filing, and treating it as one would mean charging you for the whole result set every night. Form, dates, item codes, document size and, for XBRL, the value itself are what count as a change.

Two documents belonging to the same filing are remembered separately. A single 8-K can carry nine exhibits; keyed on the accession number alone, eight of them would vanish and you would never know they existed.

What happens when the run breaks in the middle?

You keep everything that was delivered and nothing is silently lost. The order of operations is fixed and tested: rows are pushed to the dataset first, and only then marked as delivered in memory. If the container is moved, the run times out or the network drops, the next run picks up exactly the rows that never made it out, and no row is marked delivered that a buyer never received.

That order is not obvious and it is easy to get backwards, so the test suite breaks a run on purpose halfway through and fails if memory ever runs ahead of delivery.

Will I be charged for rows I did not want?

Filtering happens before anything is written to the dataset, so anything you filtered out never reaches your bill. keywords, excludeKeywords, excludeForms, items, sicCodes, states, minValue and maxValue all apply at that point, and the run report lists how many rows each one removed — so an empty result reads as "your filter was strict", not as "the source is broken".

Two more guards sit in the same place. Rows with no accession number and no CIK are dropped rather than billed: EDGAR does occasionally return a hit you cannot build a link from, and a row you cannot follow is worth nothing. And your Apify spending limit is enforced by this Actor itself — the platform stops charging at your limit but does not stop the run, which means an unguarded Actor keeps burning compute on rows nobody is paying for.

How do I stop one huge company eating the whole run?

Use Rows per company. Ask for twenty tickers with a limit of a thousand rows and, without a quota, the first large issuer hands you a thousand filings on its own and the other nineteen never get a turn. Left at zero, the run limit is shared out evenly between the companies you named; set it explicitly to take, say, the fifty most recent filings from each.

By default company mode returns the most recent 1,000 filings per company, because that is what the submissions feed serves in one request. Switch on Include the archive to walk the older pages back to 1993 as well — one extra request per company, and for a long-lived issuer over a thousand extra rows.

Do I need a proxy, a key or an account with SEC?

No. All three endpoints are public and free, and this Actor uses only the Python standard library to reach them. What SEC does ask for is that automated callers identify themselves in the User-Agent header — its Fair Access policy is explicit about it, and a request without one is refused with HTTP 403. A working identifier is already set, and the Your contact email field lets you put your own address there instead, which is what SEC prefers.

The same policy caps callers at ten requests a second. This Actor holds itself to eight, shared across every thread, because exceeding the limit gets the address blocked for ten minutes and that costs you a whole run to save a few seconds.

Can I get the text of the filing itself?

Yes — switch on Include the document text. Each row then carries document_text, the document with its markup stripped, capped at 200000 characters, plus document_text_chars so you know whether the cap bit. It is off by default because it costs one extra request per row and makes rows far larger; for most monitoring the link in document_url is enough, and for language work the text is the whole point.

What does this Actor not do?

Named honestly, because finding out after you have paid is worse than reading it here.

  • It does not parse Form 4 transaction tables. You get the insider filing, its date, the issuer and the link, but not the parsed rows of shares, price and transaction code. Actors dedicated to insider trading do that.
  • It does not parse 13F holdings tables. Same story: the quarterly institutional report comes back as a filing, not as a list of positions.
  • It does not normalise financial statements. XBRL facts come back exactly as the company tagged them, concept by concept. It does not assemble them into a balance sheet or an income statement, and it does not reconcile different companies' tagging choices.
  • It writes no AI summaries. Every value in every row came from SEC. Nothing in the output is generated, scored or interpreted by a model, which is a limitation if you wanted a summary and a feature if you are feeding a compliance process.
  • Full-text search starts in 2001. SEC does not index document bodies before then. Company filing history reaches back to 1993, because that feed does.
  • It has no proxy configuration. None is needed for these endpoints, and offering a setting that does nothing would be worse than leaving it out.

How much does a run cost?

$2.50/1K rows, charged per row written to the dataset. A daily monitor in incremental mode typically writes tens of rows, not thousands, once the first run has filled its memory. Counting queries used to split a period are not billed — only rows you receive are.

With no search phrase and no limit of your own, a run is treated as a trial and stops at 300 rows, so a first press of Start cannot produce a surprise.

FAQ

Which is faster, searching by phrase or by company?

By company. companyFilings is one request per company and returns up to a thousand filings from it, while a full-text search over a wide period has to count the period first and then page through it a hundred documents at a time. If you know exactly whose filings you want, name the tickers.

Can I search for a phrase within one company only?

Yes. Put the phrase in Search phrase and the company's CIK in CIK numbers, and SEC restricts the search itself rather than the Actor filtering afterwards — which means you are not billed for documents from other filers. Company name contains does the same thing by name when you do not have the CIK to hand.

Why does my search return fewer rows than the report says exist?

Because the report tells you how many documents the index holds for your period, which is deliberately not the same as how many you asked for. It is there so you can see what a wider limit would buy you before you pay for it. Raise Maximum rows, or narrow the period.

What if a single day holds more than the cap by itself?

Then the API genuinely cannot return the rest of that day, and the run says so instead of pretending otherwise: the day appears in the report under capped_days and a warning is written to the log. Splitting a period stops at one day because that is the smallest window EDGAR will filter on. In practice this only happens on very common words with no form filter.

Do amendments show up separately?

Yes, as their own rows, which is what you want — an amended 8-K is news in its own right. form holds 8-K/A, root_form holds 8-K, and a form filter of 8-K matches both. In incremental mode an amendment to something you already received arrives tagged UPDATED.

Which XBRL tag should I ask for?

Revenues, Assets, Liabilities, NetIncomeLoss, StockholdersEquity and CashAndCashEquivalentsAtCarryingValue cover most questions. Concepts live in the us-gaap taxonomy for US filers, ifrs-full for foreign private issuers, and dei for entity facts such as the share count. Not every company tags every concept, so a company missing from a market-wide frame usually tagged something more specific rather than reporting nothing.

What is the difference between CY2025Q1 and CY2025Q1I?

CY2025Q1 is a period — something that happened over the quarter, like revenue. CY2025Q1I is an instant — something measured at a point in time, like the balance of accounts payable on the last day. Asking for a flow with an instant frame, or the reverse, returns nothing; that is SEC's convention, not a fault in the run.

Can I feed the rows straight to an AI agent?

Switch on Compact rows and Drop empty fields. The first keeps only the identifying fields and the links, the second leaves out keys with no value instead of returning them as null. Together they cut a row to roughly a fifth of its full size, which matters when every field costs context.

Is the data official?

Yes, and only that. Every row comes from efts.sec.gov or data.sec.gov, both operated by the U.S. Securities and Exchange Commission and both free to the public. This Actor adds structure, splitting, filtering and memory between runs; it adds no data of its own, and every value can be checked against the link in filing_index_url.