Knowledge Delta | RAG Knowledge Base Change Detection avatar

Knowledge Delta | RAG Knowledge Base Change Detection

Pricing

$0.10 / complete sync plan

Go to Apify Store
Knowledge Delta | RAG Knowledge Base Change Detection

Knowledge Delta | RAG Knowledge Base Change Detection

Compare crawler exports and saved snapshots. Generate stable chunk upserts, metadata updates and guarded deletes for RAG pipelines. Process up to 1,000 documents per plan; no model API required.

Pricing

$0.10 / complete sync plan

Rating

0.0

(0)

Developer

Liam King

Liam King

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

10 days ago

Last modified

Share

Knowledge Delta — crawler exports to AI knowledge-base updates

Turn a new content export and the last applied snapshot into a deterministic update plan. Knowledge Delta tells your pipeline which chunks need new embeddings, which only need metadata updates, and which old chunk IDs to remove.

Designed for: developers and automation agencies maintaining retrieval-augmented generation (RAG) knowledge bases from documentation or help-center exports.

Bring your existing crawler and destination integration. This tool does not crawl, embed, summarize, or write to your vector database. It prepares a portable plan so those steps remain under your control.

Is this the missing step in your RAG pipeline?

Use Knowledge Delta when you already have exported Markdown/text and need a repeatable way to keep an AI knowledge base current:

Your workflowWhat the plan provides
A help-center policy changesUpdated paragraph chunks and the IDs of superseded chunks
Documentation is exported again without changesZero upserts or deletes, so your integration can skip embedding work
An agency maintains several client knowledge basesSeparate namespaces and explicit snapshots for each client
A title or URL changes but content does notMetadata updates marked as not requiring a new embedding

Try the bundled sample first: select demo and inspect the complete plan. The sample has no custom plan charge. Real comparisons are US$0.10 each, including unchanged results, for up to 1,000 documents within the stated size limits.

If your current indexing framework already handles incremental updates reliably, you may not need another tool. Knowledge Delta is for teams that want this comparison step as an Apify job and already have a destination integration.

Quick start on Apify

  1. Click Try for free or open the Actor in Apify Console.
  2. Leave Source mode set to Free sample, then click Start. The sample changes one paragraph: expect one upsert, one delete and five unchanged chunks.
  3. For your own content, choose Paste documents or select an Existing Apify dataset, and set a stable namespace.
  4. Read the complete plan in Output, then apply its actions in your existing knowledge-base pipeline.
  5. Save the resulting snapshot only after applying every action. Supply that snapshot on the next comparison.

Example: changed returns policy

A help center changes its returns window from 14 days to 30 days. Knowledge Delta emits the updated paragraph and the superseded chunk ID, while leaving the five unchanged sample chunks alone. Your integration can embed the updated text and remove the old chunk without rebuilding the whole sample knowledge base.

This is a developer tool for an existing ingestion pipeline. It does not include a crawler, embedding service or vector-database connector.

Input formats

Supply an array, a { "documents": [...] } or { "items": [...] } wrapper in the preview, or JSONL with one document per line. In the Actor, paste an array into Content documents or choose Existing Apify dataset.

[
{
"url": "https://example.com/help/returns",
"title": "Returns policy",
"markdown": "# Returns\n\nUnopened items can be returned within 30 days."
}
]

Each document needs:

FieldMeaning
id or urlStable identity. A supplied string id takes precedence. Without an ID, a canonical HTTP(S) URL is used; its fragment is removed.
markdown or textNonempty exported content. Nonempty Markdown takes precedence. HTML-only rows are rejected.
titleOptional title; metadata.title is accepted as a fallback.

Other fields are ignored. IDs must be unique within the namespace. Fetch-error rows are rejected rather than silently treated as missing pages. Only title and URL are retained as metadata; custom metadata and document ordering are not synchronized.

Apify usage

  1. Free sample: keep mode demo. This uses bundled sample data and produces a nonempty plan without a custom event charge. Platform usage can still apply according to the pricing configuration.
  2. First real import: choose documents or dataset, set a stable namespace, and omit previous snapshot inputs.
  3. Apply the plan to your destination using the procedure below.
  4. Repeat: pass the previous SNAPSHOT object, or the previous run's key-value store ID as previousSnapshotStoreId, only after that plan was successfully applied.

For dataset mode, use an immutable dataset from a finished crawl. The storage pickers grant read-only access to the selected dataset and previous snapshot store. The Actor retains limited permissions; no account-wide token is needed. A dataset changing during the read is rejected. Row-count limits are enforced without truncation.

{
"mode": "documents",
"namespace": "help-center-production",
"chunkSize": 1600,
"documents": [
{ "id": "returns", "title": "Returns", "text": "Unopened items can be returned within 30 days." }
]
}

For repeated jobs, add previousSnapshot with the original snapshot object. For cloud workflows, use previousSnapshotStoreId instead; it is the key-value store ID, not the Actor/run/dataset ID. Use only one of these two snapshot inputs.

Output

The default dataset contains one complete plan per run, not one row per chunk. A ready plan includes:

  • summary: source counts, upsert/delete counts, metadata changes and new embeddings required.
  • actions: ordered upsert operations followed by delete operations.
  • documents: a per-document review list.
  • snapshot: a compact manifest for the next comparison; it stores hashes and metadata, not original text.
  • planId: deterministic identity for this baseline-to-next-snapshot transition.

The default key-value store also contains PLAN, ACTIONS, SNAPSHOT and SUMMARY for convenient retrieval. The full dataset record remains a recovery path if a later artifact write fails.

An upsert includes id, namespace, documentId, sourceId, url, title, contentHash, text and requiresEmbedding. If requiresEmbedding is false, update metadata while retaining the existing vector. A delete contains the old chunk ID and either replaced_chunk or missing_document as its reason.

Applying plans correctly

  1. Serialize updates for each namespace. Keep the last successfully applied snapshot in your own workflow state.
  2. Generate a plan against that snapshot.
  3. Embed and upsert entries with requiresEmbedding: true using the supplied IDs.
  4. For metadata-only entries, reuse the vector and replace metadata. If your database cannot patch metadata separately, read the existing vector and upsert it with the new metadata.
  5. After all upserts succeed, apply the delete instructions only in this namespace.
  6. Only after every operation succeeds, commit the new snapshot to your workflow state.

If any step fails, keep the old snapshot and retry. IDs are deterministic, so repeating successful upserts and deletions is safe when the destination supports idempotent operations. Do not treat a successful Actor run as confirmation that your destination updated.

This tool generates instructions; it never automatically advances a shared cloud baseline or changes a customer's database. See INTEGRATION.md for an integration contract and example consumer.

Deletion protection

Missing documents are retained by default. Replacing text within a supplied document still emits removal instructions for the superseded chunks.

To remove absent documents, set both completeSource and removeMissing to true. A guard blocks the entire plan if more than maxDeletionRatio of the prior document count would disappear (default 0.2 = 20%). This guard concerns missing documents, not ordinary edited chunks. Empty sources are separately blocked unless allowEmptySource is true.

Completeness is your assertion. The tool cannot know whether your crawler missed a page. Never assert completeness for a filtered, paginated, failed or partial export. No external deletion is performed by this Actor.

Limits and tradeoffs

  • At most 1,000 documents per namespace, 200,000 characters per document and 2 million total source text characters.
  • At most 10,000 chunks in a snapshot. Input files are limited to 12 MB; cloud plan records to 8 MB after JSON serialization.
  • Paragraph-based splitting, with long paragraphs split at nearby spaces or bounded character offsets. The size limit is in characters, not model tokens. There is no overlap or semantic chunking.
  • Chunk IDs are content-addressed within each document. Adding a paragraph usually preserves other IDs. Editing a long split paragraph can change several chunks. Paragraph order alone is not represented as a destination update.
  • Only line endings and outer whitespace are normalized. HTML extraction, boilerplate removal, crawl quality and access permissions belong to the upstream crawler.
  • Keep namespace and chunk size fixed with a snapshot. Changing them requires an intentional rebuild in a separate destination namespace.
  • No embeddings, database credentials, automatic notifications, model access, or direct Pinecone/Qdrant integrations are included. Any embedding service and existing crawler remain separate costs.
  • Plan storage contains source text and URLs. Use source material you have permission to process, and protect exported plans and Apify storage appropriately.

There is no claim of guaranteed dollar savings: embedding prices can be very low, and this product's fee may exceed embedding savings. Its proposed value is simpler, repeatable synchronization.

Pricing

US$0.10 for one complete real sync plan, including an unchanged comparison, within the above limits. One event named sync-plan; no extra start or dataset-row event. The sample is free of this custom charge. Invalid inputs and budget-blocked runs do not produce a plan charge.

Always check the actual Apify pricing tab before running. Platform usage during the run is included in the event price; post-run storage and downloads follow Apify's standard rules. The source includes guards to stop before reading a dataset when the run budget cannot cover one plan.

Try the working preview

On Windows, double-click START-APP.cmd. Alternatively, with Node.js 22 or newer:

npm run app

Open http://127.0.0.1:4317 and select Load sample comparison. One returns-policy paragraph changes from 14 to 30 days. The result contains one upsert, one deletion and five unchanged chunks. This is synthetic sample content, not a customer case study.

The local browser preview uses no external requests for your data, no telemetry and no API keys. Files remain in browser memory until you close/reload the page. Downloads go to your usual browser download folder.

If an embedded browser blocks downloads, use View or copy an export if downloads are blocked to copy the JSON, or open the preview in your normal browser. The CLI also writes files directly.

Development

The preview, CLI and engine tests use only Node.js. Install the pinned cloud runtime when you want to run the Actor locally:

npm ci --ignore-scripts
npm test
npm run demo
npm run smoke
npm run benchmark

The CLI accepts an input JSON file and a new output directory:

npm run plan -- examples/first-import.json results/my-first-plan

It refuses to overwrite an existing output directory. Deployment uses the included Dockerfile, lockfile, input schema, dataset view and output schema.

Support

Before your first real comparison

  • Can I connect it directly to Pinecone or Qdrant? It outputs portable instructions; your integration calls your database. There is no built-in database connector.
  • Will it delete content if a crawl is incomplete? Missing documents are retained by default. Removing them requires an explicit complete-source assertion and passes a deletion-ratio guard.
  • Which snapshot should I reuse? Only the snapshot whose actions your destination has successfully applied. A successful comparison is not a successful database update.
  • Does it need an AI API key? No. Any embeddings are generated by your own downstream pipeline.
  • What if I need a different input format? Open an Issue with the field names and a small synthetic example. Do not upload private source content.

Trying it with an existing pipeline? In the Issues tab, tell us your source format, destination and the step that was difficult. Feedback on real integration blockers helps prioritize improvements.

Report a reproducible problem through the Actor's Issues tab. Include sanitized input, the expected result and run ID. Never post tokens, confidential content, credentials or private customer data in an issue.

Worked integration walkthrough

Knowledge Delta compares exported text. Bring an existing crawler, embedding service and database integration. It does not connect those services for you.

1. Inspect the sample

Open https://apify.com/futurefortune/knowledge-delta and select the bundled sample. Its synthetic returns policy changes from 14 to 30 days. Inspect Output: one upsert, one deletion and five unchanged chunks. The sample has no custom plan charge; check the live pricing tab for storage and usage details.

2. Make a small first import

Choose Paste documents and use a new test namespace such as my-help-center-test. Paste:

[
{
"id": "returns",
"title": "Returns policy",
"text": "# Returns\n\nUnopened items can be returned within 14 days.\n\nContact support for assistance."
}
]

Leave missing-document removal off. Real comparisons cost US$0.10 each, including unchanged comparisons. Three real runs below cost US$0.30 in Actor events; crawler, embedding, database and post-run storage costs are separate where applicable. You can instead inspect the bundled local examples without buying cloud comparisons.

Require a successful run and status: ready, demo: false. Download the full plan from Output. Do not mistake this for a completed database update.

3. Apply and acknowledge the first plan

In your existing workflow, process actions:

ActionDestination operation
upsert, requiresEmbedding: trueEmbed text, then write the vector using the exact namespace and id; retain text, title and URL as metadata.
upsert, requiresEmbedding: falseReplace metadata while retaining the existing vector.
deleteRemove that exact ID only in the action's namespace, after all upserts succeed.

After every operation succeeds, save the plan's snapshot as your committed baseline. If anything fails, keep the old baseline and retry idempotently. Serialize all work for a namespace. The adapter interface in INTEGRATION.md shows the lock, baseline check and commit order.

4. Change one paragraph

Run again with the same namespace, document ID and chunk size. Change 14 days to 30 days and supply the committed snapshot as previousSnapshot. Alternatively select the previous applied run's key-value store as previousSnapshotStoreId; use only one snapshot input.

For this short three-paragraph example, expect one updated chunk and one superseded chunk deletion. Apply upserts, then deletes, then commit the new snapshot. Never commit a snapshot merely because the Actor succeeded.

5. Check an unchanged repeat

Run the same content against the newly committed snapshot. Expect zero upserts and zero deletions. This is still a paid real comparison. Your downstream workflow can skip embedding and database writes when no actions are needed.

6. Connect a real completed crawl

Select Existing Apify dataset from a finished, immutable crawl. Each row must have stable id or url and nonempty markdown or text. Keep one namespace per knowledge base. The picker requests read-only access. Stay within 1,000 documents, 200,000 characters per document and 2 million total source characters, plus the other limits in the listing.

Keep missing documents by default. Only enable missing-document deletion for a verified complete export; a partial crawl must not be treated as a complete inventory.

When this is not a fit

If your framework already handles incremental indexing reliably, another comparison step may be unnecessary. This tool does not offer semantic chunking, embeddings, direct vector-database connectors or guaranteed cost savings. Its purpose is a deterministic update plan inside an existing Apify workflow.

Need help with a source format? Open an Issue on the listing with a small synthetic example and field names. Do not post private documents or credentials.