NY Business Entity Status Delta avatar

NY Business Entity Status Delta

Pricing

$10.00 / 1,000 entity status changeds

Go to Apify Store
NY Business Entity Status Delta

NY Business Entity Status Delta

Watches named New York corporations/LLCs in the official NY DOS entity registry and alerts only on a genuine status change: dissolved, suspended, reinstated, or discontinued. Built for due-diligence teams, lenders, title companies, and compliance/KYB screening. Free when nothing changes.

Pricing

$10.00 / 1,000 entity status changeds

Rating

0.0

(0)

Developer

Radu Furtuna

Radu Furtuna

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

9 days ago

Last modified

Share

Durable monitor for the official NY Department of State "Corporations and Other Entities: All Filings – Entity Status History" registry, published as open data on data.ny.gov. Watch specific corporations, LLCs, and other registered entities by their DOS ID number and get notified only when their registry status genuinely changes — dissolved, suspended, reinstated, or discontinued. No API key needed, no login, no captcha.

Source

https://data.ny.gov/resource/3gg2-jgnp.json — the Socrata Open Data API for "Corporations and Other Entities: All Filings – Entity Status History" (dataset id 3gg2-jgnp). Confirmed live 13.09.2026: 20,929,658 rows, 1,467 new rows in the last 7 days — an actively updated filing history, not a static snapshot. corpid_num is the dataset's unique, permanent DOS ID number for one entity — not its name, which is never unique and can be reused after dissolution. The status field observed live (13.09.2026) takes these values on status-bearing filings: Active (17,423,251 rows), Inactive (2,861,262), Suspended (12,072), Discontinued (3,945) — real transitions, not a static roster.

How it works

  1. Each watch names exactly one entity by corpidNum (its NY DOS ID Number — look it up at apps.dos.ny.gov/publicInquiry if you only have a business name). One watch = one point query for that entity's most recent status-bearing filing = one HTTP request per run.
  2. This registry is a full filing history, not a one-row-per-entity table — a busy entity can have hundreds of filing rows spanning decades, and not every filing carries a status (address/agent changes don't). Each check server-side queries for the single most recent row that DOES carry a status ($where=corpid_num=N AND status IS NOT NULL, ordered by filing date then filing number) — never a full download of the entity's history, and never the 20.9M-row dataset.
  3. The first check of a new watch establishes a baseline (no charge). Every later check compares the current status against the durable record of what it was last time.
  4. Billing is tied only to the status field — not to the filing date or filing number of whichever document carried it. A routine re-filing that re-confirms the same status (e.g. a biennial statement confirming Active) never bills; only a genuine status transition does. A status that changes and later reverts (e.g. suspended, then reinstated, then suspended again) bills every genuine transition, never silently deduplicated against an earlier occurrence of the same status value.
  5. Genuinely new (first found) or status-changed entities are pushed to the dataset and billed once each (entity-status-changed); a check that finds nothing new costs nothing beyond the fixed platform run cost.

Input

{
"monitorId": "my-entity-watch",
"watches": [
{ "watchId": "borrower-abc-llc", "corpidNum": "1130455" }
],
"notifyOn": "new_alerts",
"webhookUrl": "https://example.com/webhook"
}

Add more entities later under the same monitorId — each watch keeps its own independent history. A watchId is permanently bound to the corpidNum it first saw; pointing the same watchId at a different DOS ID number later fails the run instead of silently mixing histories.

socrataAppToken is optional — data.ny.gov does not require a key for this dataset, but a free Socrata app token from your own account raises the anonymous request-rate ceiling if you run many watches across many monitors.

Output row (per change)

watchId, corpidNum, changeType ("new"|"status_changed"), status, previousStatus, filmNum, dateFiled, modCertCode, monitorId, runId, discoveredAt, eventId, billed

Billing

Pay-per-event: entity-status-changed — charged only for a watch's first found status (baseline is free) or a genuine status transition since the previous check. Failed/blocked checks are never charged.

Important — read before relying on this for any lending/leasing/title/compliance decision

This registry is a publication of DOS filing status, not a certification of good standing, not a lien search, and not a complete legal/financial picture of the entity. It does not cover UCC filings, litigation, tax liens, or beneficial ownership. The public dataset reflects what NY DOS has published, which can lag real-world events. This actor is an informational monitor of CHANGES to that public publication — it is NOT a current "good standing" certificate, NOT a lien/litigation search, and NOT legal or financial advice. Always confirm directly at apps.dos.ny.gov/publicInquiry (or order an official Certificate of Status from NY DOS) before acting on any single entry — especially for lending, leasing, title insurance, or M&A due-diligence decisions.

Delivery guarantee: at-most-once (we would rather lose an alert than bill you twice)

Each computed change is delivered to the dataset and charged at most once, for as long as the monitor's claim log exists (see the boundary below). Before any irreversible step (writing the row, charging the event) the run takes an atomic claim on that exact change, using the only atomic primitive the Apify platform offers: a request queue's unique-key insert. Exactly one run can win that claim for a given change. The claim log is never consumed, deleted or rotated by this actor; it is a permanent record of what was already attempted, and coverage.claimJournalSize reports its size each run so you can watch it grow (the platform's counter is eventually consistent, so treat it as a lagging estimate, not an exact count).

Where that guarantee ends — the honest boundary. The claim log lives in a named request queue (<prefix>-<monitorId>-claims) in your own account. The at-most-once guarantee holds as long as that queue keeps existing. If you — or any process holding your account credentials — delete, rename or re-create it from the Console or API, the log starts empty and previously delivered changes can be delivered and charged again. That is the unavoidable boundary of any durable storage, not a loophole in the protocol. For the same reason, the actor's storage prefix and internal claim namespace are frozen after release: changing either would create a fresh, empty log with exactly the same effect.

The response the platform returns for each claim is interpreted strictly: only a real boolean false grants the right to write and charge, only a real boolean true denies it, and anything else — a missing field, null, 0, an empty string, a changed SDK response shape — aborts the run's delivery for that item with claim_protocol_error before any row or charge. An answer we do not fully understand is never read as "you may charge".

One thing we deliberately do not claim: the monitor's lease makes overlapping runs a fail-closed exception rather than a fact of life, but between the moment a run verifies it still holds the lease and the moment the dataset write or charge actually lands there is an unavoidable time gap (the platform offers no fencing token for datasets or billing). So "a run that lost the lease can never write another row" would be an overstatement. What actually protects your money is the claim above: the key is already taken, so even a ghost run cannot charge for the same change twice.

The honest consequence, stated plainly: if a run dies after taking the claim but before finishing, that one change is lost. It is recorded as dataset_unknown or charge_unknown and it is not re-delivered on the next run — the next run moves on to the entity's next status transition. We deliberately chose possible loss of one alert over the possibility of charging you twice for the same event. This is at-most-once delivery, not exactly-once; any actor that claims exactly-once over a store without compare-and-swap is overstating what the platform can do.

Practically this only happens if the Apify run is killed mid-delivery (platform abort, timeout, migration). Every such case is visible: the run's coverage and run_summary report it, and run_summary.eventsBilled plus Apify's own billing ledger remain the source of truth for what you actually paid for.

Honest limits

  • The durable dataset is a delivery-attempt log, not a guaranteed mirror of the default dataset. Each row is written to the durable dataset first, then mirrored to the run's default dataset before billing proceeds for that row. If the durable write succeeds but the default-dataset mirror write fails (e.g. transient Apify storage error), the item is marked dataset_unknown, billing for it is permanently blocked (fail-closed — we never charge for a row we can't confirm was delivered), and the run is not retried into re-creating that exact row. The durable dataset can therefore end up with a small number of orphan rows that were never mirrored and never billed. The default dataset is the canonical log of rows successfully written to this run's output (see its run_summary row) — but a default-dataset row does not by itself prove the row was billed: the row is written before Actor.charge() runs, so if charging then fails or comes back charge_unknown, the row is present but not confirmably paid. run_summary.eventsBilled and Apify's own billing ledger are the source of truth for confirmed payment, not the presence of a row in either dataset.
  • One watch = one entity, one request. There is no bulk/roster mode — to track a portfolio, add one watch per entity (up to 30 per run). This keeps the network cost fixed and predictable regardless of the registry's total size (20.9M+ filing rows), and keeps each entity's history independently auditable.
  • A "not found" result on the very first check of a watch is reported honestly, not as an error — coverage.watches[].matched: false. This can mean the corpidNum was mistyped, or — more rarely — the entity exists but has never had a status-bearing filing. The registry does not delete filing rows (confirmed live: filings from 1893 remain in the dataset), so a typo'd corpidNum will simply never match; it costs nothing and is safe to correct and retry under the same watchId.
  • A corpidNum that WAS matched on a previous check but is NOT found on a later one is treated as source_access_limited for that watch this run — no baseline/history update, no billing. Since the registry doesn't delete filing rows, this should never happen from a genuine data change; it's the honest fallback if the source ever answers unexpectedly.
  • Billing tracks only the status field, deliberately excluding the filing date and filing number of whichever document carried it (a routine re-filing that re-confirms the same status is not a change). Those fields are still delivered in every row for context.
  • We don't invent data: if the API ever returns something other than a bare JSON array, two most-recent rows with an identical filing date AND filing number (violating the dataset's own uniqueness contract for a single filed document), a row missing corpid_num/film_num/date_filed/status, or a row whose corpid_num doesn't exactly match the one requested (even after canonicalizing leading zeros/whitespace the same way input is canonicalized), the run reports it honestly (source_access_limited) instead of guessing which record it actually found.

Author: OmniCoder (https://t.me/OmniCoder)