SEC Filings Scraper - EDGAR Full Text Search, XBRL, 10-K, 8-K
Pricing
$2.50 / 1,000 filings
SEC Filings Scraper - EDGAR Full Text Search, XBRL, 10-K, 8-K
EDGAR full text search stops at 10,000 hits per query; splitting the period returned 27,224 documents, 2.72x more. SEC EDGAR filings API for company filings since 1993, 8-K item codes, insider Form 4 and XBRL financial statements API. 56 fields, no API key. Official SEC filings, nothing second-hand.
Pricing
$2.50 / 1,000 filings
Rating
0.0
(0)
Developer
Snow Leo Data
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
SEC EDGAR Filings Search
Filings, company histories and financial numbers straight from the official SEC
EDGAR APIs — efts.sec.gov for full-text search, data.sec.gov for submissions
and XBRL. No API key, no proxy, no headless browser, nothing scraped out of HTML
that can silently change shape overnight.
$2.50/1K rows. Press Start with nothing filled in and you get a sample of the newest 8-K current reports, so you can see the shape of a row before you decide anything.
The one number that matters
EDGAR full-text search refuses to page past 10,000 documents for any single
query. Worse, once you are over that line it stops telling you the truth about
the size of your own result set: the response comes back with
{"value": 10000, "relation": "gte"} — at least ten thousand, and it will not
say how many more. Ask for page 101 and you get an error, not data:
search_phase_execution_exception: Result window is too large,from + size must be less than or equal to: [10000] but was [10100]
This Actor treats that as a problem to solve rather than a limit to live with. It asks the index how big the period is, and while the answer is "at least ten thousand" it cuts the period in half and asks again, until every window is one the API will actually serve in full.
Measured on 2026-09-12, searching the exact phrase "revenue" in 8-K current
reports filed between 2026-01-01 and 2026-09-10:
| Documents reachable | |
|---|---|
| One query, the way the API is normally used | 10,000, with no idea how many were missed |
| This Actor, same query, period split automatically | 27,224 |
Four windows, seven counting queries, about six seconds of work. That is 2.72x
past the cap, and the run report tells you exactly how the period was cut, so
the number is auditable rather than a marketing claim. Reproduce it yourself
with python3 tests/test_live.py.
What can I actually collect with it?
Three modes, chosen with one dropdown.
fullTextSearch looks inside the documents. Every word of every filing
since 2001 is indexed by SEC, including exhibits, press releases and the
footnotes nobody reads. Search for "material weakness", "going concern",
"ransomware", the name of a supplier, a drug, a competitor or a law firm, and
you get back every filing that says it, with the form type, the filer, the
filing date, the 8-K item numbers and a direct link to the document.
companyFilings walks the complete submission history of the companies you
name, by ticker or by CIK. That is every form an issuer has ever filed, back to
1993 — annual reports, quarterly reports, current reports, insider Forms 3, 4
and 5, proxy statements, registration statements, correspondence with the staff.
Each row carries the company profile too: SIC industry code and its description,
EIN, state of incorporation, fiscal year end, filer category, business address
and every former name the company traded under.
xbrlFacts returns the numbers companies tagged in their own financial
statements. Either one concept across every period a company has reported it
(give tickers), or one concept across every company that reported it in a
period (give a period such as CY2025Q1 and no tickers). That second form is
the cheapest market-wide financial snapshot there is: one request returned 1,934
companies' quarterly revenue when this was written.
Which fields do I get back?
56 fields in total across the three modes, all of them straight from SEC and none of them inferred. The ones people ask about first:
- identity —
accession_number,document_file,cik,company,ticker,tickers,exchanges,other_filers,all_ciks - the filing —
form,root_form,form_meaning,filed_at,period_ending,accepted_at,act,file_number,film_number,size_bytes - 8-K events —
itemsanditems_explained - the filer —
sic,sic_description,ein,entity_type,state_of_incorporation,fiscal_year_end,filer_category,phone,former_names,business_street,business_city,business_state,business_zip - machine-readable data —
is_xbrl,is_inline_xbrl, and in XBRL modetaxonomy,tag,label,value,unit,period_start,period_end,fiscal_year,fiscal_period,frame,location - the document itself —
document_url,filing_index_url, and optionallydocument_textwithdocument_text_chars - run bookkeeping —
record_type,relevance_score,change_type
An empty field means SEC did not publish that value for that record. Nothing is
guessed, modelled or filled in from elsewhere. A numeric field that has no value
is null, never an empty string, so the dataset loads into pandas, BigQuery or
Excel without a type error on the first row.
Why are the 8-K item codes worth anything?
Because a current report is meaningless until you know which item it was filed
under, and EDGAR gives you only the number. A filing tagged 1.05 is a company
telling the market it has had a material cybersecurity incident. A filing tagged
4.02 is a company saying its previously issued financial statements can no
longer be relied on. A filing tagged 5.02 is a director or officer leaving.
Every row carries both. items holds the raw codes so you can filter on them,
and items_explained pairs each code with its official title, so a row reads as
{"code": "1.05", "title": "Material Cybersecurity Incidents"} without anyone
having to keep a lookup table. 32 official 8-K item numbers are covered, and a
code that is not in the list still comes back with its number and an honest
null title rather than a made-up one.
The same idea applies to form types: form_meaning says in plain English what
the form is for, and 28 form types are named — so SC 13G reads as beneficial
ownership above five percent with passive intent, and NT 10-K reads as a
notification of a late annual report. An amendment keeps its own form (8-K/A)
while root_form and form_meaning come from the parent, so you can group
amendments with what they amend without string-slicing form names yourself.
How do I run this on a schedule without paying twice?
Switch on Only what changed since the last run. The Actor keeps a compact
memory of what it has already handed you in a named key-value store that
survives between runs, and on the next run it delivers only what is new or
amended. Each row is tagged NEW, UPDATED or UNCHANGED, and unchanged rows
are not returned at all unless you ask for them.
This is what makes a daily schedule affordable. A search that matches thirty thousand filings costs you thirty thousand rows once; after that you pay for the few dozen that appeared overnight. The memory holds 70,000 keys and, when it fills, drops the oldest first — recent filings are the ones a monitor cares about.
The fingerprint behind UPDATED deliberately ignores fields that churn. A
relevance score that shifts because the index was rebuilt is not a change to the
filing, and treating it as one would mean charging you for the whole result set
every night. Form, dates, item codes, document size and, for XBRL, the value
itself are what count as a change.
Two documents belonging to the same filing are remembered separately. A single 8-K can carry nine exhibits; keyed on the accession number alone, eight of them would vanish and you would never know they existed.
What happens when the run breaks in the middle?
You keep everything that was delivered and nothing is silently lost. The order of operations is fixed and tested: rows are pushed to the dataset first, and only then marked as delivered in memory. If the container is moved, the run times out or the network drops, the next run picks up exactly the rows that never made it out, and no row is marked delivered that a buyer never received.
That order is not obvious and it is easy to get backwards, so the test suite breaks a run on purpose halfway through and fails if memory ever runs ahead of delivery.
Will I be charged for rows I did not want?
Filtering happens before anything is written to the dataset, so anything you
filtered out never reaches your bill. keywords, excludeKeywords,
excludeForms, items, sicCodes, states, minValue and maxValue all
apply at that point, and the run report lists how many rows each one removed —
so an empty result reads as "your filter was strict", not as "the source is
broken".
Two more guards sit in the same place. Rows with no accession number and no CIK are dropped rather than billed: EDGAR does occasionally return a hit you cannot build a link from, and a row you cannot follow is worth nothing. And your Apify spending limit is enforced by this Actor itself — the platform stops charging at your limit but does not stop the run, which means an unguarded Actor keeps burning compute on rows nobody is paying for.
How do I stop one huge company eating the whole run?
Use Rows per company. Ask for twenty tickers with a limit of a thousand rows and, without a quota, the first large issuer hands you a thousand filings on its own and the other nineteen never get a turn. Left at zero, the run limit is shared out evenly between the companies you named; set it explicitly to take, say, the fifty most recent filings from each.
By default company mode returns the most recent 1,000 filings per company, because that is what the submissions feed serves in one request. Switch on Include the archive to walk the older pages back to 1993 as well — one extra request per company, and for a long-lived issuer over a thousand extra rows.
Do I need a proxy, a key or an account with SEC?
No. All three endpoints are public and free, and this Actor uses only the Python
standard library to reach them. What SEC does ask for is that automated callers
identify themselves in the User-Agent header — its Fair Access policy is
explicit about it, and a request without one is refused with HTTP 403. A working
identifier is already set, and the Your contact email field lets you put your
own address there instead, which is what SEC prefers.
The same policy caps callers at ten requests a second. This Actor holds itself to eight, shared across every thread, because exceeding the limit gets the address blocked for ten minutes and that costs you a whole run to save a few seconds.
Can I get the text of the filing itself?
Yes — switch on Include the document text. Each row then carries
document_text, the document with its markup stripped, capped at 200000
characters, plus document_text_chars so you know whether the cap bit. It is off
by default because it costs one extra request per row and makes rows far larger;
for most monitoring the link in document_url is enough, and for language work
the text is the whole point.
What does this Actor not do?
Named honestly, because finding out after you have paid is worse than reading it here.
- It does not parse Form 4 transaction tables. You get the insider filing, its date, the issuer and the link, but not the parsed rows of shares, price and transaction code. Actors dedicated to insider trading do that.
- It does not parse 13F holdings tables. Same story: the quarterly institutional report comes back as a filing, not as a list of positions.
- It does not normalise financial statements. XBRL facts come back exactly as the company tagged them, concept by concept. It does not assemble them into a balance sheet or an income statement, and it does not reconcile different companies' tagging choices.
- It writes no AI summaries. Every value in every row came from SEC. Nothing in the output is generated, scored or interpreted by a model, which is a limitation if you wanted a summary and a feature if you are feeding a compliance process.
- Full-text search starts in 2001. SEC does not index document bodies before then. Company filing history reaches back to 1993, because that feed does.
- It has no proxy configuration. None is needed for these endpoints, and offering a setting that does nothing would be worse than leaving it out.
How much does a run cost?
$2.50/1K rows, charged per row written to the dataset. A daily monitor in incremental mode typically writes tens of rows, not thousands, once the first run has filled its memory. Counting queries used to split a period are not billed — only rows you receive are.
With no search phrase and no limit of your own, a run is treated as a trial and stops at 300 rows, so a first press of Start cannot produce a surprise.
FAQ
Which is faster, searching by phrase or by company?
By company. companyFilings is one request per company and returns up to a
thousand filings from it, while a full-text search over a wide period has to
count the period first and then page through it a hundred documents at a time.
If you know exactly whose filings you want, name the tickers.
Can I search for a phrase within one company only?
Yes. Put the phrase in Search phrase and the company's CIK in CIK numbers, and SEC restricts the search itself rather than the Actor filtering afterwards — which means you are not billed for documents from other filers. Company name contains does the same thing by name when you do not have the CIK to hand.
Why does my search return fewer rows than the report says exist?
Because the report tells you how many documents the index holds for your period, which is deliberately not the same as how many you asked for. It is there so you can see what a wider limit would buy you before you pay for it. Raise Maximum rows, or narrow the period.
What if a single day holds more than the cap by itself?
Then the API genuinely cannot return the rest of that day, and the run says so
instead of pretending otherwise: the day appears in the report under
capped_days and a warning is written to the log. Splitting a period stops at
one day because that is the smallest window EDGAR will filter on. In practice
this only happens on very common words with no form filter.
Do amendments show up separately?
Yes, as their own rows, which is what you want — an amended 8-K is news in its
own right. form holds 8-K/A, root_form holds 8-K, and a form filter of
8-K matches both. In incremental mode an amendment to something you already
received arrives tagged UPDATED.
Which XBRL tag should I ask for?
Revenues, Assets, Liabilities, NetIncomeLoss, StockholdersEquity and
CashAndCashEquivalentsAtCarryingValue cover most questions. Concepts live in
the us-gaap taxonomy for US filers, ifrs-full for foreign private issuers,
and dei for entity facts such as the share count. Not every company tags every
concept, so a company missing from a market-wide frame usually tagged something
more specific rather than reporting nothing.
What is the difference between CY2025Q1 and CY2025Q1I?
CY2025Q1 is a period — something that happened over the quarter, like revenue.
CY2025Q1I is an instant — something measured at a point in time, like the
balance of accounts payable on the last day. Asking for a flow with an instant
frame, or the reverse, returns nothing; that is SEC's convention, not a fault in
the run.
Can I feed the rows straight to an AI agent?
Switch on Compact rows and Drop empty fields. The first keeps only the identifying fields and the links, the second leaves out keys with no value instead of returning them as null. Together they cut a row to roughly a fifth of its full size, which matters when every field costs context.
Is the data official?
Yes, and only that. Every row comes from efts.sec.gov or data.sec.gov, both
operated by the U.S. Securities and Exchange Commission and both free to the
public. This Actor adds structure, splitting, filtering and memory between runs;
it adds no data of its own, and every value can be checked against the link in
filing_index_url.