# CMS CLIA Laboratory Registry (`nexgenwatch/cms-clia-laboratory-registry`) Actor

One clia\_lab\_record per CLIA-certified laboratory: CLIA number, facility name, full address, certification date, termination status, and lab classification, read from the official CMS dataset API (10-second crawl-delay honoured).

- **URL**: https://apify.com/nexgenwatch/cms-clia-laboratory-registry.md
- **Developed by:** [NexGen Watch](https://apify.com/nexgenwatch) (community)
- **Categories:** Business
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $13.40 / 1,000 registry records

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## 🔔 CMS CLIA Laboratory Registry

One clia\_lab\_record per CLIA-certified laboratory: CLIA number, facility name, full address, certification date, termination status, and lab classification, read from the official CMS dataset API (10-second crawl-delay honoured).

One structured `registry_record` per record from the CMS Provider of Services File - Clinical Laboratories dataset (data.cms.gov data-api, ~681,000 rows). Public, logged-out, robots-honoured.

Output is one `clia_lab_record` row per result; billing is pay-per-event, the value event being one registry record (a $0.02 start fee per run, then $0.02 per registry record).

No login, no API key and no CAPTCHA solving are involved: the source is read logged-out with an identified contact User-Agent.

### 📊 Sample Output

[![CMS CLIA Laboratory Registry sample output — a table of real registry record rows (clia\_number, facility\_name, street, city) from run TJPoFURgLt9xT8bam on build 0.1.5](https://api.apify.com/v2/key-value-stores/8gLgXMBveEI1tTz1z/records/cms-clia-laboratory-registry-sample)](https://apify.com/nexgenwatch/cms-clia-laboratory-registry?fpr=2ayu9b)

Real rows from run `TJPoFURgLt9xT8bam` on build 0.1.5 (2026-09-18), the same input as the Quick start below — every value is as the source published it (emails masked, long text shortened):

| clia\_number | facility\_name | street | city | state | certification\_date |
|---|---|---|---|---|---|
| 01D0026356 | SHELBY BAPTIST MEDICAL CENTER | 1000 FIRST STREET NORTH | ALABASTER | AL | 1993-02-23 |
| 01D0026428 | WOODLAND COMMUNITY HOSPITAL LAB | 1910 CHEROKEE AVE SW | CULLMAN | AL | 1992-12-24 |
| 01D0026438 | CULLMAN REGIONAL MED CTR/RESP CARE DEP | 1912 AL HIGHWAY 157 | CULLMAN | AL | 1992-12-24 |
| 01D0026498 | ST VINCENT'S ST CLAIR HOSPITAL LABORATORY | 7063 VETERANS HWY ATTN WENDY JENKINS | PELL CITY | AL | 2006-12-13 |
| 01D0026531 | BAPTIST HEALTH CITIZENS | 604 STONE AVENUE | TALLADEGA | AL | 2009-06-24 |
| 01D0026562 | UAB HIGHLANDS LABORATORY | 1201 11TH AVENUE SOUTH | BIRMINGHAM | AL | 1992-12-28 |
| 01D0026563 | ASCENSION ST VINCENT'S BIRMINGHAM | 810 ST VINCENTS DRIVE | BIRMINGHAM | AL | 1994-03-02 |
| 01D0026581 | AMER CAST IRON PIPE CO HLTH SERV LAB | 1501 31ST AVENUE NORTH | BIRMINGHAM | AL | 1993-01-08 |

The run finished with the status message: `CAPPED: delivered 50 of 51 unique rows (maxRecords=50). billable=50`

### ✅ What you get

Each row is flat JSON with these fields (from the dataset schema and the sample run; a field the source does not publish for a given row is `null`):

- `record_type` (string) — e.g. `clia_lab_record`
- `clia_number` (string/null) — e.g. `01D0026356`
- `facility_name` (string/null) — e.g. `SHELBY BAPTIST MEDICAL CENTER`
- `street` (string/null) — e.g. `1000 FIRST STREET NORTH`
- `city` (string/null) — e.g. `ALABASTER`
- `state` (string/null) — e.g. `AL`
- `zip` (string/null) — e.g. `35007`
- `certification_date` (string/null) — e.g. `1993-02-23`
- `termination_code` (string/null) — e.g. `00`
- `lab_classification_code` (string/null) — e.g. `00`
- `provider_category_code` (string/null) — e.g. `22`
- `source` (string/null) — e.g. `cms-clia-laboratory-registry`
- `release_date` (string/null) — null in every sample row
- `source_url` (string/null) — e.g. `https://data.cms.gov/data-api/v1/dataset/d3eb38ac-d8e9-40d3-b7b7-6205d3d1dc16/da`
- `observed_at` (string/null) — e.g. `2026-09-18T05:50:42Z`
- `terminal` (string/integer/null) — null in every sample row
- `reported_total` (string/integer/null) — null in every sample row
- `raw_seen` (string/integer/null) — null in every sample row
- `duplicates` (string/integer/null) — null in every sample row
- `unique_found` (string/integer/null) — null in every sample row
- `delivered` (string/integer/null) — null in every sample row
- `scope` (string/integer/null) — null in every sample row
- `note` (string/integer/null) — null in every sample row

**What you get**

One `registry_record` per entity, deduplicated on the source-native id. Fields the source does not publish for a row are honest `null`. A `run_receipt` row carries the source URL, any reported total, and the delivered/duplicate counts.

### ⚙️ Sample inputs

**1. Quick start — the Store example (this is what the sample above came from)**

```json
{
  "maxRecords": 50
}
```

The sample run charged exactly: 1 × $0.02 apify-actor-start + 50 × $0.02 registry\_record = **$1.02** on the Free tier — every delivered row was billed.

**2. A smaller, narrowed run**

```json
{
  "maxRecords": 5
}
```

Caps the run at 5 rows — about $0.12 on the Free tier ($0.02 start + 5 × $0.02).

**3. A full-size run**

```json
{
  "maxRecords": 2000
}
```

Up to 2000 rows (the schema default for `maxRecords`) — about $40.02 on the Free tier ($0.02 start + 2000 × $0.02) if the source has that many.

### 🧾 JSON sample record

One real record from run `TJPoFURgLt9xT8bam`, exactly as it lands in the dataset (emails masked, long text shortened):

```json
{
  "record_type": "clia_lab_record",
  "clia_number": "01D0026356",
  "facility_name": "SHELBY BAPTIST MEDICAL CENTER",
  "street": "1000 FIRST STREET NORTH",
  "city": "ALABASTER",
  "state": "AL",
  "zip": "35007",
  "certification_date": "1993-02-23",
  "termination_code": "00",
  "lab_classification_code": "00",
  "provider_category_code": "22",
  "source": "cms-clia-laboratory-registry",
  "release_date": null,
  "source_url": "https://data.cms.gov/data-api/v1/dataset/d3eb38ac-d8e9-40d3-b7b7-6205d3d1dc16/data",
  "observed_at": "2026-09-18T05:50:42Z"
}
```

### 🔧 How it works

**Transport.** Plain HTTPS from the Apify platform, no proxy. robots.txt is read first and a disallowed path is never fetched. Every request carries an identified contact User-Agent. Pacing: `CRAWL_DELAY`=10.0.

**Charging.** Each registry record is charged at the moment it is pushed (`registry_record`); a row that fails to charge is not delivered, so the dataset count always equals the charged count.

**How it behaves**

The dataset is ~681,000 rows. Runs page size/offset and stream each page; the crawl-delay is honoured. Bound the run with maxRecords. `maxRecords` caps the count. Duplicates, outages, and schema drift charge nothing.

**What is not done.** No login, no cookie or CAPTCHA bypass, no private or personal-account data, no browser automation.

### 💰 Pricing example

| Event | Free | Bronze | Silver | Gold |
|---|---|---|---|---|
| Actor Start (`apify-actor-start`) | $0.02 | $0.02 | $0.02 | $0.02 |
| Registry Record (`registry_record`) | $0.02 | $0.02 | $0.02 | $0.01 |

Worked at the live Free-tier price:

- 8 registry records: $0.02 start + 8 × $0.02 = **$0.18**
- 25 registry records: $0.02 start + 25 × $0.02 = **$0.52**
- 2000 registry records: $0.02 start + 2000 × $0.02 = **$40.02**

A run that delivers zero rows charges the $0.02 start fee only. A BLOCKED run (source refused) fails loud and charges no value event. The start fee is charged once per GB of run memory; the default run memory is 4096 MB.

Yield on the sample run: `CAPPED: delivered 50 of 51 unique rows (maxRecords=50). billable=50`. `maxRecords` is a hard ceiling on what is delivered and billed, never a target.

### ⚖️ Legal & ToS

This actor reads public data only. It collects only what the source publishes to any visitor, keeps to the source's robots rules (checked on every run), identifies itself with a contact User-Agent, and does not access accounts, private data or anything behind authentication. Use the output in line with the source's terms and your local law; the intended use is B2B research and monitoring.

### ❓ FAQ

**Q: Do I need an API key or a login?**\
A: No. the input schema has no key field and the actor carries no secrets.

**Q: Why did my run return 0 rows?**\
A: Read the run's status message. GENUINE\_EMPTY means the source was read and had nothing in scope for your input; BLOCKED means the source refused and the run failed without billing a value event — retry later or narrow the input. A zero-row run bills the start fee only.

**Q: How many rows can one run return?**\
A: Up to `maxRecords` (default 2000). Raise the cap for a bigger run; you pay per delivered row.

**Q: How fresh is the data?**\
A: Every run reads the source live at run time; nothing is cached between runs. Put it on a schedule for a continuous feed.

**Q: What formats can I export?**\
A: The dataset downloads as JSON, CSV, Excel, XML or RSS from the run's Dataset tab or the Apify API, and any run can push to a webhook or integration.

**Q: How is this different from the other NexGen Watch actors actors?**\
A: Same output shape and billing model; this one covers CMS CLIA Laboratory Registry. The siblings under **Related Actors** cover the other sources or slices — run several on one schedule for a combined feed.

**Q: Are there rate limits?**\
A: The actor paces itself against the source (`CRAWL_DELAY`=10.0) and honours its robots rules; there is no per-buyer limit beyond your Apify plan's concurrency.

### 🆘 Troubleshooting

- **Run FAILED with BLOCKED** → the source refused the request or changed its page shape → nothing was billed beyond the start fee; retry after a while, and if it persists open an Issue with the run id.
- **Status says CAPPED** → your cap (`maxRecords`) was reached → raise it for a bigger run.
- **Input validation error on start** → a field is outside the schema's allowed values → start from the Quick start block and change one field at a time.
- **Run TIMED-OUT** → a very wide request on a slow day → raise the run timeout in Run options or narrow the input; what was delivered before the timeout is still in the dataset.

### 🔗 Related Actors

- [Brazil PNCP Awarded Contracts](https://apify.com/nexgenwatch/brazil-pncp-awarded-contracts?fpr=2ayu9b) — One structured registryrecord per record from the official public source, logged-out. Public, robots-honoured, keyless
- [Brazil PNCP Price-Registration Minutes](https://apify.com/nexgenwatch/brazil-pncp-price-registration-minutes?fpr=2ayu9b) — One structured registryrecord per record from the official public source, logged-out. Public, robots-honoured, keyless
- [Brazil PNCP Tender Notices](https://apify.com/nexgenwatch/brazil-pncp-tender-notices?fpr=2ayu9b) — One structured registryrecord per record from the official public source, logged-out. Public, robots-honoured, keyless
- [CMS Dialysis Facilities](https://apify.com/nexgenwatch/cms-dialysis-facilities?fpr=2ayu9b) — A clean, structured roster of Medicare-certified dialysis facilities in the United States,
- [CMS Hospice Facilities](https://apify.com/nexgenwatch/cms-hospice-facilities?fpr=2ayu9b) — A clean, structured roster of Medicare-certified hospice facilities in the United States,
- [Czech ARES Company Registry](https://apify.com/nexgenwatch/czech-ares-company-registry?fpr=2ayu9b) — One structured registryrecord per record from the keyless ares.gov.cz REST economic-subjects API, logged-out. Public, logged-out, robots-honoured
- 🏢 **About NexGenData** — NexGen Watch is NexGenData's fleet of 256 public monitoring and lookup actors built on official sources, pay-per-result. Browse the catalog at [apify.com/nexgenwatch](https://apify.com/nexgenwatch?fpr=2ayu9b).

### ⭐ Found this useful?

If this actor saved you a manual check, a quick **[review on the Apify Store](https://apify.com/nexgenwatch/cms-clia-laboratory-registry?fpr=2ayu9b)** helps other teams find it. Feature request or a source that changed? Open it from the **Issues** tab — every one is read.

# Actor input Schema

## `maxRecords` (type: `integer`):

Upper bound on records delivered (buyer cap).

## `userAgent` (type: `string`):

Optional identified contact User-Agent.

## Actor input object example

```json
{
  "maxRecords": 50
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxRecords": 50
};

// Run the Actor and wait for it to finish
const run = await client.actor("nexgenwatch/cms-clia-laboratory-registry").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "maxRecords": 50 }

# Run the Actor and wait for it to finish
run = client.actor("nexgenwatch/cms-clia-laboratory-registry").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxRecords": 50
}' |
apify call nexgenwatch/cms-clia-laboratory-registry --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,nexgenwatch/cms-clia-laboratory-registry"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/bPthE5mmwbM918jjO/builds/MDZMWYUS1GgEe7LVI/openapi.json
