# Knowledge Delta | RAG Knowledge Base Change Detection (`futurefortune/knowledge-delta`) Actor

Compare crawler exports and saved snapshots. Generate stable chunk upserts, metadata updates and guarded deletes for RAG pipelines. Process up to 1,000 documents per plan; no model API required.

- **URL**: https://apify.com/futurefortune/knowledge-delta.md
- **Developed by:** [Liam King](https://apify.com/futurefortune) (community)
- **Categories:** AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.10 / complete sync plan

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Knowledge Delta — crawler exports to AI knowledge-base updates

Turn a new content export and the last applied snapshot into a deterministic update plan. Knowledge Delta tells your pipeline which chunks need new embeddings, which only need metadata updates, and which old chunk IDs to remove.

**Designed for:** developers and automation agencies maintaining retrieval-augmented generation (RAG) knowledge bases from documentation or help-center exports.

**Bring your existing crawler and destination integration.** This tool does not crawl, embed, summarize, or write to your vector database. It prepares a portable plan so those steps remain under your control.

### Is this the missing step in your RAG pipeline?

Use Knowledge Delta when you already have exported Markdown/text and need a repeatable way to keep an AI knowledge base current:

| Your workflow | What the plan provides |
| --- | --- |
| A help-center policy changes | Updated paragraph chunks and the IDs of superseded chunks |
| Documentation is exported again without changes | Zero upserts or deletes, so your integration can skip embedding work |
| An agency maintains several client knowledge bases | Separate namespaces and explicit snapshots for each client |
| A title or URL changes but content does not | Metadata updates marked as not requiring a new embedding |

**Try the bundled sample first:** select `demo` and inspect the complete plan. The sample has no custom plan charge. Real comparisons are **US$0.10 each**, including unchanged results, for up to 1,000 documents within the stated size limits.

If your current indexing framework already handles incremental updates reliably, you may not need another tool. Knowledge Delta is for teams that want this comparison step as an Apify job and already have a destination integration.

### Quick start on Apify

1. Click **Try for free** or open the Actor in Apify Console.
2. Leave **Source mode** set to **Free sample**, then click **Start**. The sample changes one paragraph: expect one upsert, one delete and five unchanged chunks.
3. For your own content, choose **Paste documents** or select an **Existing Apify dataset**, and set a stable namespace.
4. Read the complete plan in **Output**, then apply its actions in your existing knowledge-base pipeline.
5. Save the resulting snapshot only after applying every action. Supply that snapshot on the next comparison.

#### Example: changed returns policy

A help center changes its returns window from 14 days to 30 days. Knowledge Delta emits the updated paragraph and the superseded chunk ID, while leaving the five unchanged sample chunks alone. Your integration can embed the updated text and remove the old chunk without rebuilding the whole sample knowledge base.

This is a developer tool for an existing ingestion pipeline. It does not include a crawler, embedding service or vector-database connector.

### Input formats

Supply an array, a `{ "documents": [...] }` or `{ "items": [...] }` wrapper in the preview, or JSONL with one document per line. In the Actor, paste an array into **Content documents** or choose **Existing Apify dataset**.

```json
[
  {
    "url": "https://example.com/help/returns",
    "title": "Returns policy",
    "markdown": "# Returns\n\nUnopened items can be returned within 30 days."
  }
]
```

Each document needs:

| Field | Meaning |
| --- | --- |
| `id` or `url` | Stable identity. A supplied string `id` takes precedence. Without an ID, a canonical HTTP(S) URL is used; its fragment is removed. |
| `markdown` or `text` | Nonempty exported content. Nonempty Markdown takes precedence. HTML-only rows are rejected. |
| `title` | Optional title; `metadata.title` is accepted as a fallback. |

Other fields are ignored. IDs must be unique within the namespace. Fetch-error rows are rejected rather than silently treated as missing pages. Only title and URL are retained as metadata; custom metadata and document ordering are not synchronized.

### Apify usage

1. **Free sample:** keep mode `demo`. This uses bundled sample data and produces a nonempty plan without a custom event charge. Platform usage can still apply according to the pricing configuration.
2. **First real import:** choose `documents` or `dataset`, set a stable namespace, and omit previous snapshot inputs.
3. Apply the plan to your destination using the procedure below.
4. **Repeat:** pass the previous `SNAPSHOT` object, or the previous run's key-value store ID as `previousSnapshotStoreId`, only after that plan was successfully applied.

For dataset mode, use an immutable dataset from a finished crawl. The storage pickers grant read-only access to the selected dataset and previous snapshot store. The Actor retains limited permissions; no account-wide token is needed. A dataset changing during the read is rejected. Row-count limits are enforced without truncation.

```json
{
  "mode": "documents",
  "namespace": "help-center-production",
  "chunkSize": 1600,
  "documents": [
    { "id": "returns", "title": "Returns", "text": "Unopened items can be returned within 30 days." }
  ]
}
```

For repeated jobs, add `previousSnapshot` with the original snapshot object. For cloud workflows, use `previousSnapshotStoreId` instead; it is the key-value store ID, not the Actor/run/dataset ID. Use only one of these two snapshot inputs.

### Output

The default dataset contains **one complete plan per run**, not one row per chunk. A ready plan includes:

- `summary`: source counts, upsert/delete counts, metadata changes and new embeddings required.
- `actions`: ordered `upsert` operations followed by `delete` operations.
- `documents`: a per-document review list.
- `snapshot`: a compact manifest for the next comparison; it stores hashes and metadata, not original text.
- `planId`: deterministic identity for this baseline-to-next-snapshot transition.

The default key-value store also contains `PLAN`, `ACTIONS`, `SNAPSHOT` and `SUMMARY` for convenient retrieval. The full dataset record remains a recovery path if a later artifact write fails.

An upsert includes `id`, `namespace`, `documentId`, `sourceId`, `url`, `title`, `contentHash`, `text` and `requiresEmbedding`. If `requiresEmbedding` is false, update metadata while retaining the existing vector. A delete contains the old chunk ID and either `replaced_chunk` or `missing_document` as its reason.

### Applying plans correctly

1. Serialize updates for each namespace. Keep the last successfully applied snapshot in your own workflow state.
2. Generate a plan against that snapshot.
3. Embed and upsert entries with `requiresEmbedding: true` using the supplied IDs.
4. For metadata-only entries, reuse the vector and replace metadata. If your database cannot patch metadata separately, read the existing vector and upsert it with the new metadata.
5. After all upserts succeed, apply the delete instructions **only in this namespace**.
6. Only after every operation succeeds, commit the new snapshot to your workflow state.

If any step fails, keep the old snapshot and retry. IDs are deterministic, so repeating successful upserts and deletions is safe when the destination supports idempotent operations. Do not treat a successful Actor run as confirmation that your destination updated.

This tool generates instructions; it never automatically advances a shared cloud baseline or changes a customer's database. See `INTEGRATION.md` for an integration contract and example consumer.

### Deletion protection

Missing documents are retained by default. Replacing text within a supplied document still emits removal instructions for the superseded chunks.

To remove absent documents, set both `completeSource` and `removeMissing` to true. A guard blocks the entire plan if more than `maxDeletionRatio` of the prior document count would disappear (default 0.2 = 20%). This guard concerns missing documents, not ordinary edited chunks. Empty sources are separately blocked unless `allowEmptySource` is true.

Completeness is your assertion. The tool cannot know whether your crawler missed a page. Never assert completeness for a filtered, paginated, failed or partial export. No external deletion is performed by this Actor.

### Limits and tradeoffs

- At most 1,000 documents per namespace, 200,000 characters per document and 2 million total source text characters.
- At most 10,000 chunks in a snapshot. Input files are limited to 12 MB; cloud plan records to 8 MB after JSON serialization.
- Paragraph-based splitting, with long paragraphs split at nearby spaces or bounded character offsets. The size limit is in characters, not model tokens. There is no overlap or semantic chunking.
- Chunk IDs are content-addressed within each document. Adding a paragraph usually preserves other IDs. Editing a long split paragraph can change several chunks. Paragraph order alone is not represented as a destination update.
- Only line endings and outer whitespace are normalized. HTML extraction, boilerplate removal, crawl quality and access permissions belong to the upstream crawler.
- Keep namespace and chunk size fixed with a snapshot. Changing them requires an intentional rebuild in a separate destination namespace.
- No embeddings, database credentials, automatic notifications, model access, or direct Pinecone/Qdrant integrations are included. Any embedding service and existing crawler remain separate costs.
- Plan storage contains source text and URLs. Use source material you have permission to process, and protect exported plans and Apify storage appropriately.

There is no claim of guaranteed dollar savings: embedding prices can be very low, and this product's fee may exceed embedding savings. Its proposed value is simpler, repeatable synchronization.

### Pricing

**US$0.10 for one complete real sync plan**, including an unchanged comparison, within the above limits. One event named `sync-plan`; no extra start or dataset-row event. The sample is free of this custom charge. Invalid inputs and budget-blocked runs do not produce a plan charge.

Always check the actual Apify pricing tab before running. Platform usage during the run is included in the event price; post-run storage and downloads follow Apify's standard rules. The source includes guards to stop before reading a dataset when the run budget cannot cover one plan.

### Try the working preview

On Windows, double-click `START-APP.cmd`. Alternatively, with Node.js 22 or newer:

```sh
npm run app
```

Open http://127.0.0.1:4317 and select **Load sample comparison**. One returns-policy paragraph changes from 14 to 30 days. The result contains one upsert, one deletion and five unchanged chunks. This is synthetic sample content, not a customer case study.

The local browser preview uses no external requests for your data, no telemetry and no API keys. Files remain in browser memory until you close/reload the page. Downloads go to your usual browser download folder.

If an embedded browser blocks downloads, use **View or copy an export if downloads are blocked** to copy the JSON, or open the preview in your normal browser. The CLI also writes files directly.

### Development

The preview, CLI and engine tests use only Node.js. Install the pinned cloud runtime when you want to run the Actor locally:

```sh
npm ci --ignore-scripts
npm test
npm run demo
npm run smoke
npm run benchmark
```

The CLI accepts an input JSON file and a new output directory:

```sh
npm run plan -- examples/first-import.json results/my-first-plan
```

It refuses to overwrite an existing output directory. Deployment uses the included Dockerfile, lockfile, input schema, dataset view and output schema.

### Support

#### Before your first real comparison

- **Can I connect it directly to Pinecone or Qdrant?** It outputs portable instructions; your integration calls your database. There is no built-in database connector.
- **Will it delete content if a crawl is incomplete?** Missing documents are retained by default. Removing them requires an explicit complete-source assertion and passes a deletion-ratio guard.
- **Which snapshot should I reuse?** Only the snapshot whose actions your destination has successfully applied. A successful comparison is not a successful database update.
- **Does it need an AI API key?** No. Any embeddings are generated by your own downstream pipeline.
- **What if I need a different input format?** Open an Issue with the field names and a small synthetic example. Do not upload private source content.

Trying it with an existing pipeline? In the Issues tab, tell us your source format, destination and the step that was difficult. Feedback on real integration blockers helps prioritize improvements.

Report a reproducible problem through the Actor's Issues tab. Include sanitized input, the expected result and run ID. Never post tokens, confidential content, credentials or private customer data in an issue.

### Worked integration walkthrough

Knowledge Delta compares exported text. Bring an existing crawler, embedding service and database integration. It does not connect those services for you.

#### 1. Inspect the sample

Open https://apify.com/futurefortune/knowledge-delta and select the bundled sample. Its synthetic returns policy changes from 14 to 30 days. Inspect Output: one upsert, one deletion and five unchanged chunks. The sample has no custom plan charge; check the live pricing tab for storage and usage details.

#### 2. Make a small first import

Choose Paste documents and use a new test namespace such as `my-help-center-test`. Paste:

```json
[
  {
    "id": "returns",
    "title": "Returns policy",
    "text": "# Returns\n\nUnopened items can be returned within 14 days.\n\nContact support for assistance."
  }
]
```

Leave missing-document removal off. Real comparisons cost US$0.10 each, including unchanged comparisons. Three real runs below cost US$0.30 in Actor events; crawler, embedding, database and post-run storage costs are separate where applicable. You can instead inspect the bundled local examples without buying cloud comparisons.

Require a successful run and `status: ready`, `demo: false`. Download the full plan from Output. Do not mistake this for a completed database update.

#### 3. Apply and acknowledge the first plan

In your existing workflow, process `actions`:

| Action | Destination operation |
| --- | --- |
| `upsert`, `requiresEmbedding: true` | Embed `text`, then write the vector using the exact `namespace` and `id`; retain text, title and URL as metadata. |
| `upsert`, `requiresEmbedding: false` | Replace metadata while retaining the existing vector. |
| `delete` | Remove that exact ID only in the action's namespace, after all upserts succeed. |

After every operation succeeds, save the plan's `snapshot` as your committed baseline. If anything fails, keep the old baseline and retry idempotently. Serialize all work for a namespace. The adapter interface in INTEGRATION.md shows the lock, baseline check and commit order.

#### 4. Change one paragraph

Run again with the same namespace, document ID and chunk size. Change `14 days` to `30 days` and supply the committed snapshot as `previousSnapshot`. Alternatively select the previous applied run's key-value store as `previousSnapshotStoreId`; use only one snapshot input.

For this short three-paragraph example, expect one updated chunk and one superseded chunk deletion. Apply upserts, then deletes, then commit the new snapshot. Never commit a snapshot merely because the Actor succeeded.

#### 5. Check an unchanged repeat

Run the same content against the newly committed snapshot. Expect zero upserts and zero deletions. This is still a paid real comparison. Your downstream workflow can skip embedding and database writes when no actions are needed.

#### 6. Connect a real completed crawl

Select Existing Apify dataset from a finished, immutable crawl. Each row must have stable `id` or `url` and nonempty `markdown` or `text`. Keep one namespace per knowledge base. The picker requests read-only access. Stay within 1,000 documents, 200,000 characters per document and 2 million total source characters, plus the other limits in the listing.

Keep missing documents by default. Only enable missing-document deletion for a verified complete export; a partial crawl must not be treated as a complete inventory.

#### When this is not a fit

If your framework already handles incremental indexing reliably, another comparison step may be unnecessary. This tool does not offer semantic chunking, embeddings, direct vector-database connectors or guaranteed cost savings. Its purpose is a deterministic update plan inside an existing Apify workflow.

Need help with a source format? Open an Issue on the listing with a small synthetic example and field names. Do not post private documents or credentials.

# Actor input Schema

## `mode` (type: `string`):

Demo is a free synthetic comparison. Documents accepts pasted exports; Dataset reads an existing finished crawler dataset in your account.

## `documents` (type: `array`):

JSON rows with id or url and markdown or text. Optional title. Maximum 1,000 documents and 2 million total text characters. Used in Documents mode.

## `sourceDatasetId` (type: `string`):

17-character ID of an immutable dataset from a finished crawl. Used only in Dataset mode.

## `namespace` (type: `string`):

Stable name identifying one destination knowledge base. Keep unchanged on repeat runs; isolate different sources.

## `previousSnapshot` (type: `object`):

Paste a SNAPSHOT produced by this Actor only after its actions were successfully applied. Omit for a first import. Use either this field or a snapshot store ID.

## `previousSnapshotStoreId` (type: `string`):

17-character key-value store ID from the last fully applied plan; reads its SNAPSHOT record. Do not use the latest run unless the destination applied it.

## `chunkSize` (type: `integer`):

Paragraph-based chunk limit, in characters rather than tokens. Keep unchanged when reusing a snapshot.

## `completeSource` (type: `boolean`):

Assert only when the export contains every document in this namespace. Never assert for limited, failed or partial crawls.

## `removeMissing` (type: `boolean`):

Emit delete instructions for absent documents only with a complete source. Default keeps missing documents. Replaced chunks inside present documents are always removed.

## `maxDeletionRatio` (type: `number`):

Block the entire plan if missing-document deletions exceed this fraction of the prior document count. 0.2 means 20%. Raise only after verifying completeness.

## `allowEmptySource` (type: `boolean`):

Permit an empty documents array. Removing the whole knowledge base also requires complete source, remove missing and a deletion fraction of 1.

## Actor input object example

```json
{
  "mode": "demo",
  "namespace": "my-knowledge-base",
  "chunkSize": 1600,
  "completeSource": false,
  "removeMissing": false,
  "maxDeletionRatio": 0.2,
  "allowEmptySource": false
}
```

# Actor output Schema

## `plan` (type: `string`):

No description

## `actions` (type: `string`):

No description

## `snapshot` (type: `string`):

No description

## `summary` (type: `string`):

No description

## `dataset` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("futurefortune/knowledge-delta").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("futurefortune/knowledge-delta").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call futurefortune/knowledge-delta --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,futurefortune/knowledge-delta"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/itZhtrYY0MZJZRdB0/builds/E7p9SrVNpKxNoDZWN/openapi.json
