RAG Incremental Reindexer - Only Changed Pages & Chunks avatar

RAG Incremental Reindexer - Only Changed Pages & Chunks

Pricing

Pay per event

Go to Apify Store
RAG Incremental Reindexer - Only Changed Pages & Chunks

RAG Incremental Reindexer - Only Changed Pages & Chunks

Tells you which of your RAG sources actually changed since the last run, and returns embedding-ready chunks only for the delta. Saves embedding costs by skipping unchanged pages.

Pricing

Pay per event

Rating

0.0

(0)

Developer

Aleksandr Dvorzhitckii

Aleksandr Dvorzhitckii

Maintained by Community

Actor stats

1

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

Which of the pages in my RAG index changed since the last crawl?

That is the only question this Actor answers, and it answers it without re-embedding your whole corpus. Point it at the pages that feed your vector database. On the first run it builds a baseline. On every run after that it returns only the pages whose content actually changed, with embedding-ready chunks for that delta.

On a 5,000-page documentation site, only 20-80 pages change on a typical day. You embed those 20-80 instead of all 5,000.

What problem this solves

Every "crawl to Markdown" Actor hands you the whole corpus on every run. Build a RAG pipeline on that and one of two things happens:

  • you re-embed everything on every run and pay for it, or
  • you write your own diffing layer, then spend weeks discovering that raw HTML "changes" on every single request because of view counters, timestamps, ad slots and A/B markers.

This Actor is that diffing layer, done properly. The difference is not marginal: it is the difference between paying to embed 5,000 pages a day and paying to embed 60.

How the change detection works

The comparison is not a raw HTML diff. Raw HTML is useless for this job.

  1. The page is fetched over HTTP (no browser) and the article container is extracted.
  2. Boilerplate is stripped: navigation, footers, cookie banners, scripts, ads.
  3. Visible text is extracted with block-level structure preserved.
  4. Volatile fragments are masked: dates, timestamps, relative times ("5 minutes ago") and view/like/comment counters become placeholders.
  5. Whitespace is normalized, and the result is SHA-256 hashed.
  6. The hash is compared against the stored baseline for that URL.

So when a news page changes its timestamp from 09:14 to 18:47 and its counter from 1 234 views to 9 871 views, the fingerprint stays identical - and you are not charged for a change that did not happen.

Two output fields make the detection auditable: changePercent (how much of the text really differs) and maskedFragments (how much noise was filtered out).

Persistent state

The baseline lives in a named Key-Value Store (stateStoreName), which is retained indefinitely and can be shared across runs, schedules and even Actor versions.

  • One store holds the state for one corpus. Use a different name per project.
  • State is written as shard records (500 URLs each), because a Key-Value Store allows only ~200 operations per second.
  • Use resetBaseline: true for a deliberate full re-index, or to rebuild a corrupted baseline.

Output

By default the dataset contains only changed pages. Each item has the URL, the change type (new, content, metadata, removed, unchanged, error), changePercent, the extracted text and - if enabled - embedding-ready chunks with character offsets.

chunks are only produced for changed pages, and are chunked on paragraph or sentence boundaries, with configurable size and overlap.

Error rows are also emitted (success: false with an errorType), and are never billed.

Pricing - what a re-index actually costs

The Actor is pay per event. Three events exist:

EventCharged whenPrice
page-checkedeach URL fetched and compared against the baseline$0.50 per 1,000
change-detectedeach page that genuinely changed$1.00 per 1,000
chunk-exportedeach embedding-ready chunk emitted for a changed page$0.05 per 1,000

Plus standard Apify platform usage (compute units). The Actor runs on 256 MB of memory - measured across live runs it uses about 0.07 GB - so compute cost is a small fraction of the total. Checking 1,000 pages takes roughly 0.11 compute unit, which is about $0.02 on the Free plan.

Those figures come from measured runs, not estimates: pages are fetched in 1.55 s on average (Crawlee request statistics), and a run carries about 5.5 s of container start-up. Since memory is fixed at 256 MB, platform cost grows linearly with the page count and is roughly $0.00002 per page - about 4% of the event price.

Worked example: 1,000 pages, 20 changed, 3 chunks each

ChargeAmount
page-checked 1,000$0.5000
change-detected 20$0.0200
chunk-exported 60$0.0030
Platform usage~$0.0208
Total~$0.54

Worked example: nightly re-index of 5,000 pages

The point of this Actor is the second run and every run after it.

ScenarioEvent chargesPlatformTotal
Full re-embed of 5,000 pages$8.25~$0.10~$8.35
Incremental re-index (2% changed)$2.62~$0.10~$2.72

You cut roughly 68% of the cost - and you also cut 98% of the embedding spend that your vector database charges you, which is usually the larger bill.

Costs above are computed from the live pricing constants in _cli/lib/pricing.mjs and verified against real run statistics by npm run cost:audit. You can cap the spend per run in the Apify Console; the Actor stops cleanly when the budget is reached instead of failing.

Try it in 2 minutes

  1. Open the Try it panel. It is pre-filled with two documentation pages and a fresh baseline store name.
  2. Press Start. Every page is reported as new and you get chunks back - this is the cold start, and it is what a full re-index looks like.
  3. Press Start again. Nothing has changed, so the dataset is empty and you are charged only for the checks - not for changes, and not for chunks.
  4. Change webhookUrl to your own endpoint, or schedule the Actor, and you have a working incremental pipeline.

The empty second run is the whole product in one step: you paid for 2 page checks and nothing else.

Measured on this Actor, over two back-to-back runs of the same two URLs:

Pages checkedChangedChunksComputeEvents charged
First run (cold)2230.001922 CU2 + 2 + 3
Second run2000.000452 CU2 + 0 + 0

The second run cost about a quarter of the compute and nothing at all for changes or chunks - because nothing had changed.

Tutorial: a nightly incremental RAG re-index

Step 1 - seed the baseline. Run once with your full URL list and resetBaseline: true. This indexes everything and stores a fingerprint per URL in the named store stateStoreName. Export the dataset and embed it as usual.

Step 2 - schedule the incremental runs. Create a schedule that runs the Actor daily with the same URLs and the same stateStoreName. Leave resetBaseline: false.

Step 3 - consume the delta. The dataset now contains only changed pages. For each item, embed chunks and upsert them into your vector database keyed by url. Pages that were removed from your source list appear with changeType: "removed" - use that to delete the corresponding vectors.

Step 4 - automate the trigger. Set webhookUrl to a small endpoint of yours. The Actor POSTs a summary containing the changed-page count and the dataset ID when the run finishes, so your pipeline can start re-embedding without polling.

Connecting it to your vector database

The output is deliberately plain - no vendor coupling:

  • LangChain: read the dataset, build
    Document(page_content=chunk.text, metadata={"source": item.url, "chunk_index": chunk.index})
    , then vectorstore.add_documents(...).
  • LlamaIndex: same idea with TextNode(text=..., metadata=...) and index.insert_nodes(...).
  • Raw Qdrant / Pinecone / pgvector: upsert chunk.text with url plus charStart / charEnd as payload, using contentHash as a cheap idempotency key.

Because charStart and charEnd are included, you can also map a retrieved chunk back to the exact passage on the live page.

Example: nightly incremental re-index

  1. Schedule the Actor daily with your documentation URLs and stateStoreName.
  2. Set outputMode: "changedOnly" so the dataset is exactly the work queue.
  3. In your pipeline, read the dataset, embed chunks, and upsert into your vector DB.
  4. Optionally set webhookUrl to be notified when the run finishes.

Running it from the API

curl -X POST "https://api.apify.com/v2/acts/dvorzhik~rag-incremental-reindexer/runs?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"startUrls": [{ "url": "https://docs.apify.com/platform/actors" }],
"stateStoreName": "my-docs-baseline",
"emitChunks": true,
"outputMode": "changedOnly",
"webhookUrl": "https://example.com/reindex-hook"
}'

The complete input reference is the Input tab; every field is documented there.

Limitations - read this before using it

  • No browser is used. Pages whose text is rendered by JavaScript will come back empty. The Actor reports those explicitly as EMPTY_CONTENT errors rather than pretending they changed. If your site is a SPA, you need a browser-based Actor.
  • Masking is pattern-based. Versions a, b and c of this feature cover dates, counters, relative times, user-supplied regexes and CSS selectors to ignore. Banner rotation and numeric normalization beyond counters are not handled yet.
  • Not a crawler. It checks the URLs you give it. Feed it a sitemap-derived list or a fixed URL set. Crawling link graphs is a different job.
  • text is truncated in state, not in output: the baseline keeps up to 200,000 characters per page for the diff percentage calculation. Pages larger than that still compare correctly by hash.

Frequently asked questions

How do I know which pages need re-embedding without crawling the whole site?

Give this Actor the URL list that backs your index and reuse the same stateStoreName on every run. It returns the changed subset, so the dataset is your re-embedding queue. You do not need a crawler, a sitemap diff, or custom hash bookkeeping.

Why not just compare HTML hashes myself?

Because HTML hash comparison reports a change on nearly every page on every run. Analytics IDs, cache-busting query strings, CSRF tokens, relative timestamps and view counters all vary per request. This Actor compares masked visible text instead, and tells you how much noise it removed via maskedFragments.

Does it work with my vector database?

There is no integration to configure. You get JSON records with url, text chunks and character offsets, so it drops into LangChain, LlamaIndex, Qdrant, Pinecone, Weaviate, pgvector or a plain script.

What happens on the very first run?

Everything is new, so everything is returned. That is the correct cold-start behaviour - you need a full index before you can have a delta. Set resetBaseline: true if you want to force this deliberately later.

Can I control the cost?

Yes. maxPages caps the number of pages checked, and the Apify Console lets you set a maximum charge per run; the Actor stops gracefully when it is reached. Pages that fail to load are never billed.

Does it work with JavaScript-heavy sites?

No - and it says so instead of silently returning nothing. It uses a plain HTTP client, so text rendered client-side comes back as an explicit EMPTY_CONTENT error. Use a browser-based Actor for SPAs.

Integrations and scheduling

  • Schedules - run it nightly, hourly or weekly against the same baseline.
  • Webhooks - get a POST when the delta is ready, so downstream embedding starts immediately.
  • API - every run is a REST call; see the example above.
  • Apify MCP / AI agents - the Actor exposes a typed output schema, so an agent can call it as a tool and read changeType, changePercent and chunks directly.
  • Make / n8n / Zapier - trigger the run and forward the dataset to your ingestion step.

Support

Found a page that reports as changed every run, or one that changed but was not detected? Open an issue on the Issues tab of this Actor and include the URL. Masking rules are pattern-based, so a concrete counterexample is usually enough to fix it.

When reporting a problem, the run log already contains the numbers needed to diagnose it: changePercent, maskedFragments and the decided changeType for every page.

Input highlights

  • startUrls - pages to check.
  • stateStoreName - the named store holding your baseline.
  • maskVolatile - dates, counters, extra regexes, CSS selectors to strip.
  • changeThresholdPercent - ignore differences smaller than this (0 = report anything).
  • emitChunks, chunkSize, chunkOverlap - chunk output for your vector DB.
  • emitCleanMarkdown - also return the extracted article as Markdown.
  • outputMode - changedOnly (default) or all.