Documentation RAG Update Packager
Pricing
$0.50 / completed documentation package
Documentation RAG Update Packager
Turn supplied Markdown documentation into version-specific chunks for RAG. Preserve code blocks, source citations and hashes; export JSONL updates for new or changed documents. $0.50 per completed package, up to 25 documents. No crawling or embeddings.
Pricing
$0.50 / completed documentation package
Rating
0.0
(0)
Developer
Autonome Bots
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
6 days ago
Last modified
Categories
Share
Prepare documentation for a retrieval pipeline without rewriting your chunking and update scripts. Supply Markdown, keep each documentation version separate, and receive source-cited chunks plus a file containing only new or changed document chunks.
Use it after your crawler or export step. Source URLs are citations; this Actor never visits them. It does not crawl websites, generate embeddings, call an LLM, or update your vector database.
Price
$0.50 per completed package, including Apify platform usage. One package can contain up to 25 documents and 2 MiB of combined Markdown within the limits below. There is no per-chunk or start charge. A successfully completed empty or unchanged package is still one package.
The completion event is sent once, after the output files and dataset have been saved and read back. Rejected input does not trigger that event. Set the run maximum cost to at least $0.50. Use the default 256 MB memory, 60-second timeout and restart-on-error off.
Quick start
- Paste your documents into Markdown documents. Each needs an ID, source URL, version and Markdown text.
- Confirm that you have rights to process the supplied material. Do not include credentials or sensitive personal information.
- Run the Actor. Open Version-specific chunks to inspect the results or export the dataset as JSON/CSV. The other output tabs provide JSONL files, the manifest, change statuses and summary.
- For the next update, paste the exact prior
manifest.jsoninto Optional previous manifest and provide the new complete document text. Keep the same chunk-size setting.
Example input, using synthetic documentation:
{"rightsConfirmed": true,"maxChunkBytes": 4096,"documents": [{"id": "api-guide","sourceUrl": "https://docs.example.com/api","version": "v1","markdown": "# API\n\n## Requests\n\nSend a SKU and quantity.\n"}]}
Start with a small representative batch. Recognized code blocks and tables stay intact; an individual block larger than the chosen chunk size rejects rather than being silently split.
Results
Each chunk contains documentId, documentContentHash, chunkId, contentHash, sourceId, sourceUrl, version, index, headingPath, bytes and the original markdown. Version-specific IDs prevent v1 and v2 documentation from sharing an identity. Hashes identify content; they do not verify that a source is authentic.
For example, supplying one v1 guide and a separate v2 guide creates separate document identities. On a subsequent run, an unchanged v1 guide remains in chunks.jsonl but contributes no rows to upserts.jsonl. A changed v2 guide contributes all of its current chunks to the update file.
rightsConfirmed: true is your assertion that you may process the supplied material, not an independent rights check. Instructions, scripts, links, HTML and formulas within documents remain untrusted text. The Actor never executes or follows them. Downstream models, spreadsheets and renderers must maintain their own content safeguards; packaging is not prompt-injection sanitization.
Input contract
Only these root fields are accepted: rightsConfirmed, documents, optional maxChunkBytes, optional previousManifest.
Each document requires exactly id, sourceUrl, version, markdown, all strings. IDs and versions are nonempty, at most 128 UTF-8 bytes, with no control characters or surrounding whitespace. The same human-facing ID can be used for separate versions. Source URLs must be absolute HTTP(S), at most 2,048 UTF-8 bytes, without credentials, whitespace, controls or backslashes. URLs are canonically normalized only for identity comparison; the supplied citation text is retained exactly. A supplied real fragment is retained, and no heading anchor is invented.
Duplicate canonical (sourceUrl, version) pairs reject the entire package. Versions never share a document identity or chunk identity. Empty document text and an empty supplied snapshot are valid; neither asserts deletion of a source.
| Limit | Enforced bound |
|---|---|
| Documents per input / previous manifest | 25 |
| Markdown per document | 512 KiB |
| Combined Markdown | 2 MiB, measured as UTF-8 |
| Lines per document / combined | 10,000 / 25,000, checked before line/atom allocation |
| Encoded CLI JSON input file | 8 MiB |
| Chunk Markdown | Default 4,096 bytes; selectable integer 128–32,768 |
| Heading metadata | 256 UTF-8 bytes per level, up to six levels |
| Chunks per package | 4,096 |
| Combined serialized output artifacts | 16 MiB |
Chunk size is a byte budget, not a tokenizer or embedding limit. Metadata is additional to each chunk's Markdown byte count and included in the total output budget. A large indivisible block rejects with ATOMIC_BLOCK_TOO_LARGE; it is never cut apart or truncated. Invalid Unicode, unsupported control characters, and lone carriage returns reject rather than silently changing bytes.
Exact text and stable identities
Concatenating a document's chunk markdown fields in index order reproduces its supplied text exactly, including CRLF, whitespace and a missing final newline. The parser recognizes top-level fenced code blocks (backticks or tildes), indented code blocks, pipe tables, ATX headings and one-line Setext headings. Recognized code blocks/tables are indivisible. Ordinary long prose may split between Unicode code points. Heading boundaries begin a new chunk so the heading path describes its text.
This is a deliberately bounded Markdown parser, not a full CommonMark implementation. Nested blockquote/list fences reject. Unsupported constructs such as multiline Setext headings are preserved as text but are not guaranteed equivalent structural metadata. Unclosed fences, invalid backtick fence info strings and oversized blocks reject with a fixed diagnostic.
documentId is SHA-256 of the canonical source URL and version and remains stable across content edits. documentContentHash is SHA-256 of the complete original Markdown. Each chunk has a contentHash of its own exact Markdown; its chunkId binds the document identity, index and chunk content hash. IDs and artifact ordering are deterministic and unaffected by input document order. These hashes are consistency identifiers, not source-authenticity or copyright attestations.
Files and incremental updates
| Artifact | Contents |
|---|---|
chunks.jsonl | Every current supplied chunk, including source citation, version, heading path and hashes |
upserts.jsonl | All chunks for only new or changed supplied documents |
manifest.json | Current supplied documents, identities, content hashes, counts and chunking settings |
changes.json | new, changed, unchanged, or missing-unconfirmed for each compared document identity |
summary.json | Counts, including retained unchanged documents and candidate upsert chunks |
Pass the exact prior manifest.json object as previousManifest to compare packages. changed covers source-byte changes and supplied ID/citation-text changes. Different maxChunkBytes rejects with MANIFEST_CHUNKING_MISMATCH; omit the previous manifest to explicitly generate a full repack. The previous manifest is caller-supplied comparison evidence, not a live-source observation or trusted signature.
No automatic deletion or retirement is implemented. missing-unconfirmed means only that a previously listed source/version was not supplied this time. It cannot establish that a page was deleted, that the crawl was complete, or that an index row should be removed. The new manifest contains the current supplied batch, so preserve previous records separately when processing partial batches; it is not a complete inventory by assertion.
The upsert file is a candidate input for an external retrieval pipeline, not a database mutation. For a changed document, old chunk IDs may remain in a downstream index; a shorter or empty replacement can leave obsolete chunks. Before activating an update, the consumer must validate the complete supplied document revision and use (documentId, documentContentHash) to select its active revision, or use a separately authorized replacement process. Blindly merging the JSONL into an existing index does not perform stale-chunk cleanup. Empty changed documents have no chunk rows; their explicit manifest/change entry must still be handled by that consumer.
Support and troubleshooting
Open an issue on this Actor with the run ID, fixed error code and a small synthetic example. Remove private document content, credentials and personal information. If a run reports an uncertain billing outcome, inspect its event details before starting another run; do not resurrect it to retry a charge.
An Apify run status alone is not proof of completed delivery. Check the package summary and application OUTPUT.status. A stopped or rejected run can leave partial artifacts; use a package only when its application result is Completed.
Local development
The pure core uses Node.js 24 built-ins. Run npm test and npm run lint from this package. The CLI accepts an input JSON file and requires a new output directory:
node src/cli.js --input fixtures/example.json --output-dir example-output
With the pinned SDK installed, npm run test:integration exercises the adapter against isolated local storage and synthetic platform metadata. Local tests and publisher runs do not establish customer demand or independently verify customer billing/downloads.