# NY Business Entity Status Delta (`titan_coder/ny-business-entity-status-delta`) Actor

Watches named New York corporations/LLCs in the official NY DOS entity registry and alerts only on a genuine status change: dissolved, suspended, reinstated, or discontinued. Built for due-diligence teams, lenders, title companies, and compliance/KYB screening. Free when nothing changes.

- **URL**: https://apify.com/titan\_coder/ny-business-entity-status-delta.md
- **Developed by:** [Radu Furtuna](https://apify.com/titan_coder) (community)
- **Categories:** Business, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$10.00 / 1,000 entity status changeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## NY Business Entity Status Delta

Durable monitor for the official **NY Department of State "Corporations and Other Entities: All Filings –
Entity Status History"** registry, published as open data on `data.ny.gov`. Watch specific corporations,
LLCs, and other registered entities by their DOS ID number and get notified only when their registry
status genuinely changes — dissolved, suspended, reinstated, or discontinued. No API key needed, no login,
no captcha.

### Source

`https://data.ny.gov/resource/3gg2-jgnp.json` — the Socrata Open Data API for "Corporations and Other
Entities: All Filings – Entity Status History" (dataset id `3gg2-jgnp`). Confirmed live 13.09.2026:
**20,929,658 rows**, 1,467 new rows in the last 7 days — an actively updated filing history, not a static
snapshot. `corpid_num` is the dataset's unique, permanent DOS ID number for one entity — not its name,
which is never unique and can be reused after dissolution. The `status` field observed live (13.09.2026)
takes these values on status-bearing filings: `Active` (17,423,251 rows), `Inactive` (2,861,262),
`Suspended` (12,072), `Discontinued` (3,945) — real transitions, not a static roster.

### How it works

1. Each `watch` names exactly **one** entity by `corpidNum` (its NY DOS ID Number — look it up at
   [apps.dos.ny.gov/publicInquiry](https://apps.dos.ny.gov/publicInquiry) if you only have a business
   name). One watch = one point query for that entity's most recent status-bearing filing = one HTTP
   request per run.
2. This registry is a **full filing history**, not a one-row-per-entity table — a busy entity can have
   hundreds of filing rows spanning decades, and not every filing carries a status (address/agent changes
   don't). Each check server-side queries for the single most recent row that DOES carry a status
   (`$where=corpid_num=N AND status IS NOT NULL`, ordered by filing date then filing number) — never a
   full download of the entity's history, and never the 20.9M-row dataset.
3. The first check of a new watch establishes a baseline (no charge). Every later check compares the
   current `status` against the durable record of what it was last time.
4. Billing is tied **only to the `status` field** — not to the filing date or filing number of whichever
   document carried it. A routine re-filing that re-confirms the same status (e.g. a biennial statement
   confirming `Active`) never bills; only a genuine status transition does. A status that changes and
   later reverts (e.g. suspended, then reinstated, then suspended again) bills every genuine transition,
   never silently deduplicated against an earlier occurrence of the same status value.
5. Genuinely new (first found) or status-changed entities are pushed to the dataset and billed once each
   (`entity-status-changed`); a check that finds nothing new costs nothing beyond the fixed platform run
   cost.

### Input

```json
{
  "monitorId": "my-entity-watch",
  "watches": [
    { "watchId": "borrower-abc-llc", "corpidNum": "1130455" }
  ],
  "notifyOn": "new_alerts",
  "webhookUrl": "https://example.com/webhook"
}
```

Add more entities later under the same `monitorId` — each watch keeps its own independent history. A
`watchId` is permanently bound to the `corpidNum` it first saw; pointing the same `watchId` at a different
DOS ID number later fails the run instead of silently mixing histories.

`socrataAppToken` is optional — `data.ny.gov` does not require a key for this dataset, but a free Socrata
app token from your own account raises the anonymous request-rate ceiling if you run many watches across
many monitors.

### Output row (per change)

`watchId, corpidNum, changeType ("new"|"status_changed"), status, previousStatus, filmNum, dateFiled,
modCertCode, monitorId, runId, discoveredAt, eventId, billed`

### Billing

Pay-per-event: `entity-status-changed` — charged only for a watch's first found status (baseline is free)
or a genuine `status` transition since the previous check. Failed/blocked checks are never charged.

### Important — read before relying on this for any lending/leasing/title/compliance decision

**This registry is a publication of DOS filing status, not a certification of good standing, not a lien
search, and not a complete legal/financial picture of the entity.** It does not cover UCC filings,
litigation, tax liens, or beneficial ownership. The public dataset reflects what NY DOS has published,
which can lag real-world events. **This actor is an informational monitor of CHANGES to that public
publication — it is NOT a current "good standing" certificate, NOT a lien/litigation search, and NOT
legal or financial advice.** Always confirm directly at
[apps.dos.ny.gov/publicInquiry](https://apps.dos.ny.gov/publicInquiry) (or order an official Certificate
of Status from NY DOS) before acting on any single entry — especially for lending, leasing, title
insurance, or M\&A due-diligence decisions.

### Delivery guarantee: at-most-once (we would rather lose an alert than bill you twice)

Each computed change is delivered to the dataset and charged **at most once**, for as long as the
monitor's claim log exists (see the boundary below). Before any irreversible step (writing the row,
charging the event) the run takes an **atomic claim** on that exact change, using the only atomic
primitive the Apify platform offers: a request queue's unique-key insert. Exactly one run can win that
claim for a given change. The claim log is never consumed, deleted or rotated by this actor; it is a
permanent record of what was already attempted, and `coverage.claimJournalSize` reports its size each run
so you can watch it grow (the platform's counter is eventually consistent, so treat it as a lagging
estimate, not an exact count).

**Where that guarantee ends — the honest boundary.** The claim log lives in a *named* request queue
(`<prefix>-<monitorId>-claims`) in your own account. The at-most-once guarantee holds as long as that
queue keeps existing. If you — or any process holding your account credentials — delete, rename or
re-create it from the Console or API, the log starts empty and previously delivered changes can be
delivered and charged again. That is the unavoidable boundary of *any* durable storage, not a loophole in
the protocol. For the same reason, the actor's storage prefix and internal claim namespace are frozen
after release: changing either would create a fresh, empty log with exactly the same effect.

The response the platform returns for each claim is interpreted **strictly**: only a real boolean `false`
grants the right to write and charge, only a real boolean `true` denies it, and anything else — a missing
field, `null`, `0`, an empty string, a changed SDK response shape — aborts the run's delivery for that
item with `claim_protocol_error` **before** any row or charge. An answer we do not fully understand is
never read as "you may charge".

One thing we deliberately do **not** claim: the monitor's lease makes overlapping runs a fail-closed
exception rather than a fact of life, but between the moment a run verifies it still holds the lease and
the moment the dataset write or charge actually lands there is an unavoidable time gap (the platform
offers no fencing token for datasets or billing). So "a run that lost the lease can never write another
row" would be an overstatement. What actually protects your money is the claim above: the key is already
taken, so even a ghost run cannot charge for the same change twice.

The honest consequence, stated plainly: **if a run dies after taking the claim but before finishing, that
one change is lost**. It is recorded as `dataset_unknown` or `charge_unknown` and it is **not**
re-delivered on the next run — the next run moves on to the entity's next status transition. We
deliberately chose possible loss of one alert over the possibility of charging you twice for the same
event. This is *at-most-once* delivery, not *exactly-once*; any actor that claims exactly-once over a
store without compare-and-swap is overstating what the platform can do.

Practically this only happens if the Apify run is killed mid-delivery (platform abort, timeout, migration).
Every such case is visible: the run's `coverage` and `run_summary` report it, and `run_summary.eventsBilled`
plus Apify's own billing ledger remain the source of truth for what you actually paid for.

### Honest limits

- **The durable dataset is a delivery-attempt log, not a guaranteed mirror of the default dataset.**
  Each row is written to the durable dataset first, then mirrored to the run's default dataset before
  billing proceeds for that row. If the durable write succeeds but the default-dataset mirror write
  fails (e.g. transient Apify storage error), the item is marked `dataset_unknown`, billing for it is
  permanently blocked (fail-closed — we never charge for a row we can't confirm was delivered), and the
  run is not retried into re-creating that exact row. The durable dataset can therefore end up with a
  small number of orphan rows that were never mirrored and never billed. The **default dataset is the
  canonical log of rows successfully written to this run's output** (see its `run_summary` row) — but a
  default-dataset row does not by itself prove the row was billed: the row is written before
  `Actor.charge()` runs, so if charging then fails or comes back `charge_unknown`, the row is present but
  not confirmably paid. **`run_summary.eventsBilled` and Apify's own billing ledger are the source of
  truth for confirmed payment**, not the presence of a row in either dataset.
- **One watch = one entity, one request.** There is no bulk/roster mode — to track a portfolio, add one
  watch per entity (up to 30 per run). This keeps the network cost fixed and predictable regardless of the
  registry's total size (20.9M+ filing rows), and keeps each entity's history independently auditable.
- **A "not found" result on the very first check of a watch is reported honestly, not as an error** —
  `coverage.watches[].matched: false`. This can mean the `corpidNum` was mistyped, or — more rarely — the
  entity exists but has never had a status-bearing filing. The registry does not delete filing rows
  (confirmed live: filings from 1893 remain in the dataset), so a typo'd `corpidNum` will simply never
  match; it costs nothing and is safe to correct and retry under the same `watchId`.
- **A `corpidNum` that WAS matched on a previous check but is NOT found on a later one** is treated as
  `source_access_limited` for that watch this run — no baseline/history update, no billing. Since the
  registry doesn't delete filing rows, this should never happen from a genuine data change; it's the
  honest fallback if the source ever answers unexpectedly.
- **Billing tracks only the `status` field**, deliberately excluding the filing date and filing number of
  whichever document carried it (a routine re-filing that re-confirms the same status is not a change).
  Those fields are still delivered in every row for context.
- We don't invent data: if the API ever returns something other than a bare JSON array, two most-recent
  rows with an identical filing date AND filing number (violating the dataset's own uniqueness contract
  for a single filed document), a row missing `corpid_num`/`film_num`/`date_filed`/`status`, or a row whose
  `corpid_num` doesn't exactly match the one requested (even after canonicalizing leading zeros/whitespace
  the same way input is canonicalized), the run reports it honestly (`source_access_limited`) instead of
  guessing which record it actually found.

Author: OmniCoder (https://t.me/OmniCoder)

# Actor input Schema

## `monitorId` (type: `string`):

Name of this monitor's durable history (a-z, 0-9, dash; up to 40 chars).

## `watches` (type: `array`):

1-30 objects: {"watchId": "borrower-abc-llc", "corpidNum": "1130455"}. corpidNum is the entity's NY DOS ID number (the corpid\_num field in the official registry — look it up at apps.dos.ny.gov/publicInquiry if you only have a business name). One watch = one entity. New entities can be added later under the same monitorId.

## `socrataAppToken` (type: `string`):

Optional. data.ny.gov does not require a key for this dataset, but a free Socrata app token (from your own data.ny.gov account) raises the anonymous request-rate ceiling if you run many watches across many monitors. Leave empty for normal use — a handful of watches per run works fine without one.

## `notifyOn` (type: `string`):

new\_alerts — post the webhook only when new/changed paid status changes were delivered; always — post it every run; never — do not call webhookUrl at all.

## `webhookUrl` (type: `string`):

Optional. Receives a digest of delivered (paid) status changes as JSON. HTTPS only.

## Actor input object example

```json
{
  "monitorId": "my-entity-watch",
  "watches": [
    {
      "watchId": "example-entity",
      "corpidNum": "1130455"
    }
  ],
  "notifyOn": "new_alerts"
}
```

# Actor output Schema

## `results` (type: `string`):

Every row this run produced. Key fields: watchId, corpidNum, changeType (new|status\_changed), status, previousStatus, filmNum, dateFiled, modCertCode. Informational only — not a good-standing certificate or legal advice.

## `coverage` (type: `string`):

What this run actually covered and what it charged for: per-watch status/reason/matched, records delivered and billed, requested/attempted/succeeded/failed watch counts. Enough to reconcile every charge against every row.

## `digest` (type: `string`):

A short human-readable summary of what this run found, written every run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "monitorId": "my-entity-watch",
    "watches": [
        {
            "watchId": "example-entity",
            "corpidNum": "1130455"
        }
    ],
    "notifyOn": "new_alerts"
};

// Run the Actor and wait for it to finish
const run = await client.actor("titan_coder/ny-business-entity-status-delta").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "monitorId": "my-entity-watch",
    "watches": [{
            "watchId": "example-entity",
            "corpidNum": "1130455",
        }],
    "notifyOn": "new_alerts",
}

# Run the Actor and wait for it to finish
run = client.actor("titan_coder/ny-business-entity-status-delta").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "monitorId": "my-entity-watch",
  "watches": [
    {
      "watchId": "example-entity",
      "corpidNum": "1130455"
    }
  ],
  "notifyOn": "new_alerts"
}' |
apify call titan_coder/ny-business-entity-status-delta --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,titan_coder/ny-business-entity-status-delta"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/8oxXHKoKmMDHZ0dRG/builds/tELFb2RJTwVH3rG5N/openapi.json
