# Documentation RAG Update Packager (`autonome_bots/documentation-rag-update-packager`) Actor

Turn supplied Markdown documentation into version-specific chunks for RAG. Preserve code blocks, source citations and hashes; export JSONL updates for new or changed documents. $0.50 per completed package, up to 25 documents. No crawling or embeddings.

- **URL**: https://apify.com/autonome\_bots/documentation-rag-update-packager.md
- **Developed by:** [Autonome Bots](https://apify.com/autonome_bots) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.50 / completed documentation package

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Documentation RAG Update Packager

Prepare documentation for a retrieval pipeline without rewriting your chunking and update scripts. Supply Markdown, keep each documentation version separate, and receive source-cited chunks plus a file containing only new or changed document chunks.

Use it after your crawler or export step. **Source URLs are citations; this Actor never visits them.** It does not crawl websites, generate embeddings, call an LLM, or update your vector database.

### Price

**$0.50 per completed package**, including Apify platform usage. One package can contain up to 25 documents and 2 MiB of combined Markdown within the limits below. There is no per-chunk or start charge. A successfully completed empty or unchanged package is still one package.

The completion event is sent once, after the output files and dataset have been saved and read back. Rejected input does not trigger that event. Set the run maximum cost to at least $0.50. Use the default 256 MB memory, 60-second timeout and restart-on-error off.

### Quick start

1. Paste your documents into **Markdown documents**. Each needs an ID, source URL, version and Markdown text.
2. Confirm that you have rights to process the supplied material. Do not include credentials or sensitive personal information.
3. Run the Actor. Open **Version-specific chunks** to inspect the results or export the dataset as JSON/CSV. The other output tabs provide JSONL files, the manifest, change statuses and summary.
4. For the next update, paste the exact prior `manifest.json` into **Optional previous manifest** and provide the new complete document text. Keep the same chunk-size setting.

Example input, using synthetic documentation:

```json
{
  "rightsConfirmed": true,
  "maxChunkBytes": 4096,
  "documents": [
    {
      "id": "api-guide",
      "sourceUrl": "https://docs.example.com/api",
      "version": "v1",
      "markdown": "# API\n\n## Requests\n\nSend a SKU and quantity.\n"
    }
  ]
}
```

Start with a small representative batch. Recognized code blocks and tables stay intact; an individual block larger than the chosen chunk size rejects rather than being silently split.

### Results

Each chunk contains `documentId`, `documentContentHash`, `chunkId`, `contentHash`, `sourceId`, `sourceUrl`, `version`, `index`, `headingPath`, `bytes` and the original `markdown`. Version-specific IDs prevent v1 and v2 documentation from sharing an identity. Hashes identify content; they do not verify that a source is authentic.

For example, supplying one v1 guide and a separate v2 guide creates separate document identities. On a subsequent run, an unchanged v1 guide remains in `chunks.jsonl` but contributes no rows to `upserts.jsonl`. A changed v2 guide contributes all of its current chunks to the update file.

`rightsConfirmed: true` is your assertion that you may process the supplied material, not an independent rights check. Instructions, scripts, links, HTML and formulas within documents remain untrusted text. The Actor never executes or follows them. Downstream models, spreadsheets and renderers must maintain their own content safeguards; packaging is not prompt-injection sanitization.

### Input contract

Only these root fields are accepted: `rightsConfirmed`, `documents`, optional `maxChunkBytes`, optional `previousManifest`.

Each document requires exactly `id`, `sourceUrl`, `version`, `markdown`, all strings. IDs and versions are nonempty, at most 128 UTF-8 bytes, with no control characters or surrounding whitespace. The same human-facing ID can be used for separate versions. Source URLs must be absolute HTTP(S), at most 2,048 UTF-8 bytes, without credentials, whitespace, controls or backslashes. URLs are canonically normalized **only for identity comparison**; the supplied citation text is retained exactly. A supplied real fragment is retained, and no heading anchor is invented.

Duplicate canonical `(sourceUrl, version)` pairs reject the entire package. Versions never share a document identity or chunk identity. Empty document text and an empty supplied snapshot are valid; neither asserts deletion of a source.

| Limit | Enforced bound |
| --- | --- |
| Documents per input / previous manifest | 25 |
| Markdown per document | 512 KiB |
| Combined Markdown | 2 MiB, measured as UTF-8 |
| Lines per document / combined | 10,000 / 25,000, checked before line/atom allocation |
| Encoded CLI JSON input file | 8 MiB |
| Chunk Markdown | Default 4,096 bytes; selectable integer 128–32,768 |
| Heading metadata | 256 UTF-8 bytes per level, up to six levels |
| Chunks per package | 4,096 |
| Combined serialized output artifacts | 16 MiB |

Chunk size is a **byte budget, not a tokenizer or embedding limit**. Metadata is additional to each chunk's Markdown byte count and included in the total output budget. A large indivisible block rejects with `ATOMIC_BLOCK_TOO_LARGE`; it is never cut apart or truncated. Invalid Unicode, unsupported control characters, and lone carriage returns reject rather than silently changing bytes.

### Exact text and stable identities

Concatenating a document's chunk `markdown` fields in `index` order reproduces its supplied text exactly, including CRLF, whitespace and a missing final newline. The parser recognizes top-level fenced code blocks (backticks or tildes), indented code blocks, pipe tables, ATX headings and one-line Setext headings. Recognized code blocks/tables are indivisible. Ordinary long prose may split between Unicode code points. Heading boundaries begin a new chunk so the heading path describes its text.

This is a deliberately bounded Markdown parser, not a full CommonMark implementation. Nested blockquote/list fences reject. Unsupported constructs such as multiline Setext headings are preserved as text but are not guaranteed equivalent structural metadata. Unclosed fences, invalid backtick fence info strings and oversized blocks reject with a fixed diagnostic.

`documentId` is SHA-256 of the canonical source URL and version and remains stable across content edits. `documentContentHash` is SHA-256 of the complete original Markdown. Each chunk has a `contentHash` of its own exact Markdown; its `chunkId` binds the document identity, index and chunk content hash. IDs and artifact ordering are deterministic and unaffected by input document order. These hashes are consistency identifiers, not source-authenticity or copyright attestations.

### Files and incremental updates

| Artifact | Contents |
| --- | --- |
| `chunks.jsonl` | Every current supplied chunk, including source citation, version, heading path and hashes |
| `upserts.jsonl` | All chunks for only new or changed supplied documents |
| `manifest.json` | Current supplied documents, identities, content hashes, counts and chunking settings |
| `changes.json` | `new`, `changed`, `unchanged`, or `missing-unconfirmed` for each compared document identity |
| `summary.json` | Counts, including retained unchanged documents and candidate upsert chunks |

Pass the exact prior `manifest.json` object as `previousManifest` to compare packages. `changed` covers source-byte changes and supplied ID/citation-text changes. Different `maxChunkBytes` rejects with `MANIFEST_CHUNKING_MISMATCH`; omit the previous manifest to explicitly generate a full repack. The previous manifest is caller-supplied comparison evidence, not a live-source observation or trusted signature.

**No automatic deletion or retirement is implemented.** `missing-unconfirmed` means only that a previously listed source/version was not supplied this time. It cannot establish that a page was deleted, that the crawl was complete, or that an index row should be removed. The new manifest contains the current supplied batch, so preserve previous records separately when processing partial batches; it is not a complete inventory by assertion.

The upsert file is a candidate input for an external retrieval pipeline, not a database mutation. For a changed document, old chunk IDs may remain in a downstream index; a shorter or empty replacement can leave obsolete chunks. Before activating an update, the consumer must validate the complete supplied document revision and use `(documentId, documentContentHash)` to select its active revision, or use a separately authorized replacement process. Blindly merging the JSONL into an existing index does not perform stale-chunk cleanup. Empty changed documents have no chunk rows; their explicit manifest/change entry must still be handled by that consumer.

### Support and troubleshooting

Open an issue on this Actor with the run ID, fixed error code and a small synthetic example. Remove private document content, credentials and personal information. If a run reports an uncertain billing outcome, inspect its event details before starting another run; do not resurrect it to retry a charge.

An Apify run status alone is not proof of completed delivery. Check the package summary and application `OUTPUT.status`. A stopped or rejected run can leave partial artifacts; use a package only when its application result is `Completed`.

### Local development

The pure core uses Node.js 24 built-ins. Run `npm test` and `npm run lint` from this package. The CLI accepts an input JSON file and requires a new output directory:

```sh
node src/cli.js --input fixtures/example.json --output-dir example-output
```

With the pinned SDK installed, `npm run test:integration` exercises the adapter against isolated local storage and synthetic platform metadata. Local tests and publisher runs do not establish customer demand or independently verify customer billing/downloads.

# Actor input Schema

## `rightsConfirmed` (type: `boolean`):

The prefilled synthetic example is authorized. If you replace it, confirm you have permission to process your own documents. Do not include credentials, secrets or sensitive personal information.

## `documents` (type: `array`):

Up to 25 objects with id, sourceUrl, version, markdown. Supply text directly; URLs are citations and are never fetched.

## `maxChunkBytes` (type: `integer`):

UTF-8 bytes, not tokens. A larger single code block/table rejects rather than being cut apart.

## `previousManifest` (type: `object`):

The exact prior manifest.json object. Missing sources are flagged missing-unconfirmed and never become deletion instructions.

## Actor input object example

````json
{
  "rightsConfirmed": true,
  "documents": [
    {
      "id": "example-api",
      "sourceUrl": "https://docs.example.com/api",
      "version": "v1",
      "markdown": "# Sample API\n\n## Request\n\n```json\n{\"sku\": \"DEMO-001\"}\n```\n"
    }
  ],
  "maxChunkBytes": 4096
}
````

# Actor output Schema

## `chunkTable` (type: `string`):

Inspect the completed package in a table; JSONL files remain available as separate outputs.

## `chunks` (type: `string`):

No description

## `upserts` (type: `string`):

No description

## `manifest` (type: `string`):

No description

## `changes` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

````javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "rightsConfirmed": true,
    "documents": [
        {
            "id": "example-api",
            "sourceUrl": "https://docs.example.com/api",
            "version": "v1",
            "markdown": "# Sample API\n\n## Request\n\n```json\n{\"sku\": \"DEMO-001\"}\n```\n"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("autonome_bots/documentation-rag-update-packager").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

````

## Python example

````python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "rightsConfirmed": True,
    "documents": [{
            "id": "example-api",
            "sourceUrl": "https://docs.example.com/api",
            "version": "v1",
            "markdown": """# Sample API

## Request

```json
{\"sku\": \"DEMO-001\"}
````

""",
}],
}

# Run the Actor and wait for it to finish

run = client.actor("autonome\_bots/documentation-rag-update-packager").call(run\_input=run\_input)

# Fetch and print Actor results from the run's dataset (if there are any)

print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default\_dataset\_id}")
for item in client.dataset(run.default\_dataset\_id).iterate\_items():
print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

````

## CLI example

```bash
echo '{
  "rightsConfirmed": true,
  "documents": [
    {
      "id": "example-api",
      "sourceUrl": "https://docs.example.com/api",
      "version": "v1",
      "markdown": "# Sample API\\n\\n## Request\\n\\n```json\\n{\\"sku\\": \\"DEMO-001\\"}\\n```\\n"
    }
  ]
}' |
apify call autonome_bots/documentation-rag-update-packager --silent --output-dataset

````

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,autonome_bots/documentation-rag-update-packager"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/LZJpwA6eKSBKd0eYB/builds/eXe4Ag6Zbajgald2e/openapi.json
