# RAG Incremental Reindexer - Only Changed Pages & Chunks (`dvorzhik/rag-incremental-reindexer`) Actor

Tells you which of your RAG sources actually changed since the last run, and returns embedding-ready chunks only for the delta. Saves embedding costs by skipping unchanged pages.

- **URL**: https://apify.com/dvorzhik/rag-incremental-reindexer.md
- **Developed by:** [Aleksandr Dvorzhitckii](https://apify.com/dvorzhik) (community)
- **Categories:** AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## RAG Incremental Reindexer - Only Changed Pages & Chunks

**Which of the pages in my RAG index changed since the last crawl?**

That is the only question this Actor answers, and it answers it without
re-embedding your whole corpus. Point it at the pages that feed your vector
database. On the first run it builds a baseline. On every run after that it
returns **only the pages whose content actually changed**, with
**embedding-ready chunks** for that delta.

On a 5,000-page documentation site, only 20-80 pages change on a typical day. You
embed those 20-80 instead of all 5,000.

### What problem this solves

Every "crawl to Markdown" Actor hands you the whole corpus on every run. Build a
RAG pipeline on that and one of two things happens:

- you re-embed everything on every run and pay for it, or
- you write your own diffing layer, then spend weeks discovering that raw HTML
  "changes" on every single request because of view counters, timestamps, ad
  slots and A/B markers.

This Actor is that diffing layer, done properly. The difference is not marginal:
it is the difference between paying to embed 5,000 pages a day and paying to
embed 60.

### How the change detection works

The comparison is **not** a raw HTML diff. Raw HTML is useless for this job.

1. The page is fetched over HTTP (no browser) and the article container is extracted.
2. Boilerplate is stripped: navigation, footers, cookie banners, scripts, ads.
3. Visible text is extracted with block-level structure preserved.
4. **Volatile fragments are masked**: dates, timestamps, relative times ("5 minutes
   ago") and view/like/comment counters become placeholders.
5. Whitespace is normalized, and the result is SHA-256 hashed.
6. The hash is compared against the stored baseline for that URL.

So when a news page changes its timestamp from 09:14 to 18:47 and its counter from
`1 234 views` to `9 871 views`, the fingerprint stays **identical** - and you are not
charged for a change that did not happen.

Two output fields make the detection auditable: `changePercent` (how much of the text
really differs) and `maskedFragments` (how much noise was filtered out).

### Persistent state

The baseline lives in a **named Key-Value Store** (`stateStoreName`), which is retained
indefinitely and can be shared across runs, schedules and even Actor versions.

- One store holds the state for one corpus. Use a different name per project.
- State is written as shard records (500 URLs each), because a Key-Value Store allows
  only ~200 operations per second.
- Use `resetBaseline: true` for a deliberate full re-index, or to rebuild a corrupted
  baseline.

### Output

By default the dataset contains **only changed pages**. Each item has the URL, the
change type (`new`, `content`, `metadata`, `removed`, `unchanged`, `error`),
`changePercent`, the extracted text and - if enabled - embedding-ready `chunks` with
character offsets.

`chunks` are only produced for changed pages, and are chunked on paragraph or sentence
boundaries, with configurable size and overlap.

Error rows are also emitted (`success: false` with an `errorType`), and are never billed.

### Pricing - what a re-index actually costs

The Actor is **pay per event**. Three events exist:

| Event | Charged when | Price |
| --- | --- | --- |
| `page-checked` | each URL fetched and compared against the baseline | $0.50 per 1,000 |
| `change-detected` | each page that genuinely changed | $1.00 per 1,000 |
| `chunk-exported` | each embedding-ready chunk emitted for a changed page | $0.05 per 1,000 |

Plus standard Apify platform usage (compute units). The Actor runs on **256 MB of
memory** - measured across live runs it uses about 0.07 GB - so compute cost is a
small fraction of the total. Checking 1,000 pages takes roughly 0.11 compute unit,
which is about **$0.02** on the Free plan.

Those figures come from measured runs, not estimates: pages are fetched in **1.55 s
on average** (Crawlee request statistics), and a run carries about 5.5 s of container
start-up. Since memory is fixed at 256 MB, platform cost grows linearly with the page
count and is roughly **$0.00002 per page** - about 4% of the event price.

#### Worked example: 1,000 pages, 20 changed, 3 chunks each

| Charge | Amount |
| --- | --- |
| `page-checked` 1,000 | $0.5000 |
| `change-detected` 20 | $0.0200 |
| `chunk-exported` 60 | $0.0030 |
| Platform usage | ~$0.0208 |
| **Total** | **~$0.54** |

#### Worked example: nightly re-index of 5,000 pages

The point of this Actor is the second run and every run after it.

| Scenario | Event charges | Platform | Total |
| --- | --- | --- | --- |
| Full re-embed of 5,000 pages | $8.25 | ~$0.10 | ~$8.35 |
| **Incremental** re-index (2% changed) | $2.62 | ~$0.10 | ~$2.72 |

You cut roughly **68%** of the cost - and you also cut 98% of the embedding
spend that your vector database charges you, which is usually the larger bill.

Costs above are computed from the live pricing constants in
`_cli/lib/pricing.mjs` and verified against real run statistics by
`npm run cost:audit`. You can cap the spend per run in the Apify Console; the Actor
stops cleanly when the budget is reached instead of failing.

### Try it in 2 minutes

1. Open the **Try it** panel. It is pre-filled with two documentation pages and a
   fresh baseline store name.
2. Press **Start**. Every page is reported as `new` and you get chunks back - this is
   the cold start, and it is what a full re-index looks like.
3. Press **Start** again. Nothing has changed, so the dataset is **empty** and you are
   charged only for the checks - not for changes, and not for chunks.
4. Change `webhookUrl` to your own endpoint, or schedule the Actor, and you have a
   working incremental pipeline.

The empty second run is the whole product in one step: you paid for 2 page checks
and nothing else.

Measured on this Actor, over two back-to-back runs of the same two URLs:

| | Pages checked | Changed | Chunks | Compute | Events charged |
| --- | --- | --- | --- | --- | --- |
| First run (cold) | 2 | 2 | 3 | 0.001922 CU | 2 + 2 + 3 |
| **Second run** | 2 | **0** | **0** | **0.000452 CU** | **2 + 0 + 0** |

The second run cost about a quarter of the compute and nothing at all for changes or
chunks - because nothing had changed.

### Tutorial: a nightly incremental RAG re-index

**Step 1 - seed the baseline.** Run once with your full URL list and
`resetBaseline: true`. This indexes everything and stores a fingerprint per URL in
the named store `stateStoreName`. Export the dataset and embed it as usual.

**Step 2 - schedule the incremental runs.** Create a schedule that runs the Actor
daily with the same URLs and the same `stateStoreName`. Leave `resetBaseline: false`.

**Step 3 - consume the delta.** The dataset now contains only changed pages. For each
item, embed `chunks` and upsert them into your vector database keyed by `url`. Pages
that were removed from your source list appear with `changeType: "removed"` - use that
to delete the corresponding vectors.

**Step 4 - automate the trigger.** Set `webhookUrl` to a small endpoint of yours. The
Actor POSTs a summary containing the changed-page count and the dataset ID when the run
finishes, so your pipeline can start re-embedding without polling.

#### Connecting it to your vector database

The output is deliberately plain - no vendor coupling:

- LangChain: read the dataset, build `Document(page_content=chunk.text,
  metadata={"source": item.url, "chunk_index": chunk.index})`, then
  `vectorstore.add_documents(...)`.
- LlamaIndex: same idea with `TextNode(text=..., metadata=...)` and
  `index.insert_nodes(...)`.
- Raw Qdrant / Pinecone / pgvector: upsert `chunk.text` with `url` plus `charStart` /
  `charEnd` as payload, using `contentHash` as a cheap idempotency key.

Because `charStart` and `charEnd` are included, you can also map a retrieved chunk
back to the exact passage on the live page.

### Example: nightly incremental re-index

1. Schedule the Actor daily with your documentation URLs and `stateStoreName`.
2. Set `outputMode: "changedOnly"` so the dataset is exactly the work queue.
3. In your pipeline, read the dataset, embed `chunks`, and upsert into your vector DB.
4. Optionally set `webhookUrl` to be notified when the run finishes.

### Running it from the API

```bash
curl -X POST "https://api.apify.com/v2/acts/dvorzhik~rag-incremental-reindexer/runs?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "startUrls": [{ "url": "https://docs.apify.com/platform/actors" }],
    "stateStoreName": "my-docs-baseline",
    "emitChunks": true,
    "outputMode": "changedOnly",
    "webhookUrl": "https://example.com/reindex-hook"
  }'
```

The complete input reference is the **Input** tab; every field is documented there.

### Limitations - read this before using it

- **No browser is used.** Pages whose text is rendered by JavaScript will come back
  empty. The Actor reports those explicitly as `EMPTY_CONTENT` errors rather than
  pretending they changed. If your site is a SPA, you need a browser-based Actor.
- **Masking is pattern-based.** Versions a, b and c of this feature cover dates,
  counters, relative times, user-supplied regexes and CSS selectors to ignore. Banner
  rotation and numeric normalization beyond counters are not handled yet.
- **Not a crawler.** It checks the URLs you give it. Feed it a sitemap-derived list
  or a fixed URL set. Crawling link graphs is a different job.
- **`text` is truncated in state**, not in output: the baseline keeps up to 200,000
  characters per page for the diff percentage calculation. Pages larger than that still
  compare correctly by hash.

### Frequently asked questions

#### How do I know which pages need re-embedding without crawling the whole site?

Give this Actor the URL list that backs your index and reuse the same
`stateStoreName` on every run. It returns the changed subset, so the dataset *is* your
re-embedding queue. You do not need a crawler, a sitemap diff, or custom hash
bookkeeping.

#### Why not just compare HTML hashes myself?

Because HTML hash comparison reports a change on nearly every page on every run.
Analytics IDs, cache-busting query strings, CSRF tokens, relative timestamps and view
counters all vary per request. This Actor compares *masked visible text* instead, and
tells you how much noise it removed via `maskedFragments`.

#### Does it work with my vector database?

There is no integration to configure. You get JSON records with `url`, text `chunks`
and character offsets, so it drops into LangChain, LlamaIndex, Qdrant, Pinecone,
Weaviate, pgvector or a plain script.

#### What happens on the very first run?

Everything is new, so everything is returned. That is the correct cold-start
behaviour - you need a full index before you can have a delta. Set
`resetBaseline: true` if you want to force this deliberately later.

#### Can I control the cost?

Yes. `maxPages` caps the number of pages checked, and the Apify Console lets you set a
maximum charge per run; the Actor stops gracefully when it is reached. Pages that fail
to load are never billed.

#### Does it work with JavaScript-heavy sites?

No - and it says so instead of silently returning nothing. It uses a plain HTTP client,
so text rendered client-side comes back as an explicit `EMPTY_CONTENT` error. Use a
browser-based Actor for SPAs.

### Integrations and scheduling

- **Schedules** - run it nightly, hourly or weekly against the same baseline.
- **Webhooks** - get a POST when the delta is ready, so downstream embedding starts
  immediately.
- **API** - every run is a REST call; see the example above.
- **Apify MCP / AI agents** - the Actor exposes a typed output schema, so an agent can
  call it as a tool and read `changeType`, `changePercent` and `chunks` directly.
- **Make / n8n / Zapier** - trigger the run and forward the dataset to your ingestion
  step.

### Support

Found a page that reports as changed every run, or one that changed but was not
detected? Open an issue on the **Issues** tab of this Actor and include the URL. Masking
rules are pattern-based, so a concrete counterexample is usually enough to fix it.

When reporting a problem, the run log already contains the numbers needed to diagnose
it: `changePercent`, `maskedFragments` and the decided `changeType` for every page.

### Input highlights

- `startUrls` - pages to check.
- `stateStoreName` - the named store holding your baseline.
- `maskVolatile` - dates, counters, extra regexes, CSS selectors to strip.
- `changeThresholdPercent` - ignore differences smaller than this (0 = report anything).
- `emitChunks`, `chunkSize`, `chunkOverlap` - chunk output for your vector DB.
- `emitCleanMarkdown` - also return the extracted article as Markdown.
- `outputMode` - `changedOnly` (default) or `all`.

# Actor input Schema

## `startUrls` (type: `array`):

One or more pages to compare against your stored baseline. Only server-rendered / static HTML is supported - JavaScript-only pages cannot be read without a browser.

## `maxPages` (type: `integer`):

Safety limit. The Actor stops after checking this many pages, which prevents accidental runaway costs on large URL lists.

## `crawlerType` (type: `string`):

Cheerio is a plain HTTP client: cheapest and fastest, but returns nothing for sites that render text with JavaScript. jsdom can execute some scripts in-process but is noticeably slower and less reliable. Sites built as SPAs will need a browser-based Actor instead of this one.

## `proxyConfiguration` (type: `object`):

Optional. Enable if the sites you check block requests from datacenter IPs.

## `stateStoreName` (type: `string`):

Name of the named Key-Value Store that persists your baseline between runs. This store is retained indefinitely, so scheduled runs always have something to compare against. Use a different store name for each separate project or corpus.

## `resetBaseline` (type: `boolean`):

When enabled, the Actor discards the stored baseline and treats every page as changed. Use it on the very first run if you want a full corpus, or after a major content migration.

## `compareMode` (type: `string`):

Comparing visible text instead of raw HTML is what makes the detection useful: raw HTML 'changes' on every page view because of ads, counters and nonces. Use 'text-and-metadata' if a title or description change should also trigger re-indexing.

## `maskVolatile` (type: `object`):

Patterns that are replaced with placeholders before hashing, so they never count as a real change.

## `changeThresholdPercent` (type: `number`):

A page is only reported as changed when at least this percentage of the visible text differs. Set to 0 to report any difference. Raising it (for example to 2) filters out trivial wording tweaks that are not worth re-embedding.

## `emitChunks` (type: `boolean`):

Include a list of text chunks (with character offsets) for every changed page, ready to be embedded and upserted. Charges one 'chunk-exported' event per chunk.

## `chunkSize` (type: `integer`):

Approximate number of characters per chunk. Chunking tries to end on a paragraph or sentence boundary rather than cutting mid-word. Set the value your embedding model expects (around 2000 characters is comfortable for most models).

## `chunkOverlap` (type: `integer`):

How many characters each chunk repeats from the previous one. A little overlap helps retrieval quality by keeping sentences that straddle a boundary intact.

## `emitCleanMarkdown` (type: `boolean`):

Also return the extracted article body as Markdown for changed pages. Useful when you rebuild your index from Markdown rather than from raw text.

## `outputMode` (type: `string`):

By default the dataset contains only changed pages, which keeps the output small and directly usable as a re-indexing queue. Switch to 'all' when you want a record of every page checked.

## `webhookUrl` (type: `string`):

Optional. When set, the Actor sends a small POST request here after the run finishes with the number of changed pages and the dataset ID, so your pipeline can trigger re-embedding automatically.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://docs.apify.com/platform/actors"
    },
    {
      "url": "https://docs.apify.com/academy/concepts"
    }
  ],
  "maxPages": 500,
  "crawlerType": "cheerio",
  "proxyConfiguration": {
    "useApifyProxy": false
  },
  "stateStoreName": "rag-reindexer-baseline",
  "resetBaseline": false,
  "compareMode": "text",
  "maskVolatile": {
    "dates": true,
    "dateFormats": [
      "iso",
      "slash",
      "dots",
      "long"
    ],
    "viewCounters": true
  },
  "changeThresholdPercent": 0,
  "emitChunks": true,
  "chunkSize": 2000,
  "chunkOverlap": 200,
  "emitCleanMarkdown": false,
  "outputMode": "changedOnly",
  "webhookUrl": ""
}
```

# Actor output Schema

## `changedPages` (type: `string`):

One item per changed page with its chunks. This is the re-indexing work queue.

## `changedPagesView` (type: `string`):

Table of changed pages with change type and percentage.

## `errors` (type: `string`):

Pages that failed to fetch or parse. Errors are never billed and are retried on the next run.

## `runSummary` (type: `string`):

Counters for this run: pages checked, changed, unchanged and failed.

## `baseline` (type: `string`):

Content fingerprints written by this run. The next run compares against these.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://docs.apify.com/platform/actors"
        },
        {
            "url": "https://docs.apify.com/academy/concepts"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("dvorzhik/rag-incremental-reindexer").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [
        { "url": "https://docs.apify.com/platform/actors" },
        { "url": "https://docs.apify.com/academy/concepts" },
    ] }

# Run the Actor and wait for it to finish
run = client.actor("dvorzhik/rag-incremental-reindexer").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://docs.apify.com/platform/actors"
    },
    {
      "url": "https://docs.apify.com/academy/concepts"
    }
  ]
}' |
apify call dvorzhik/rag-incremental-reindexer --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,dvorzhik/rag-incremental-reindexer"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/BwChBCN9iS7aVriCj/builds/3fichzI8CO9ipTzCN/openapi.json
