Crawl Quality Gate: Check Exports Before RAG Refresh
Pricing
$100.00 / 1,000 completed audits
Crawl Quality Gate: Check Exports Before RAG Refresh
Check crawler exports for missing required pages, coverage loss, duplicate URLs, source errors and empty text before updating an AI knowledge base. JSON and readable report; no model API required.
Pricing
$100.00 / 1,000 completed audits
Rating
0.0
(0)
Developer
Tom Hester
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
Crawl Quality Gate
Check a crawler export before refreshing an AI knowledge base. Get a clear
gatePassed result, missing-page evidence and a readable HTML report.
A crawler can finish with fewer pages than expected. If your next step replaces the entire knowledge base, that partial export can remove useful content. This Actor checks the export against your required URLs and previously accepted manifest, and flags duplicate URLs, recorded HTTP/source errors and short text.
Try the example
Run Free synthetic demo. It contains one healthy help page, one failed page
and one absent required page. The expected result is BLOCK, with two required
pages lacking valid content. This is a synthetic example, not a customer result.
Use your own crawl
- Wait for your crawler to finish successfully. Check its failed-request count.
- Choose Existing Apify dataset and select the finished, immutable dataset, or choose Paste documents and supply an array.
- Set Source finished successfully and Failed upstream requests to the actual source outcome. These values are your assertions; this Actor does not fetch or verify the upstream run status.
- Add required page URLs, a previous accepted manifest, or both. Optionally restrict allowed origins and choose minimum text and document counts.
- Read the default dataset's single report, or
REPORT/REPORT.htmlin the default key-value store.
Each row needs url and nonempty markdown or text. Nonempty Markdown takes
precedence. Website Content Crawler exports use this shape. The Actor also checks
optional statusCode, crawl.statusCode, crawl.httpStatusCode, error,
crawl.error, errorMessages, and crawl.loadedUrl fields. Missing HTTP metadata
cannot establish that the original request succeeded.
{"mode": "documents","namespace": "client-help-center","sourceFinished": true,"expectedUrls": ["https://example.com/help/returns"],"documents": [{"url": "https://example.com/help/returns","text": "Unopened products can be returned within thirty days of delivery. Contact support for a return label and retain your order number."}]}
Connect the result to a workflow
For n8n, Make, or your own code, place this check between crawling and updating
the destination. Continue only if both the Actor run succeeded and the
report has gatePassed === true. An Actor run can succeed while its report says
BLOCK: that means the audit worked and found a problem.
On BLOCK, keep your existing destination and previous manifest. On PASS,
review warnings and apply your intended destination update. Save the returned
MANIFEST only after that update completes successfully. Pass it back as
previousManifest next time, using the same namespace.
This Actor never authorizes deletions, writes to your vector database, advances a shared baseline, sends notifications, generates embeddings or calls an LLM. Integrating the gate is required: merely running an audit does not stop another independent workflow.
n8n template: stop an incomplete crawl before a chatbot refresh
Get the free template from the n8n library and use Use for free → Copy template to clipboard. The template is free; real completed audits cost $0.10 each, including BLOCK results. The workflow below is the same tested example, also provided here for direct import.
Import the workflow JSON below into n8n. It reads a finished source run, checks its export, and routes PASS and BLOCK separately. Both branches initially end without changing your knowledge base.
- In Configure source, set your completed
sourceRunId, a stablenamespace, and the requiredexpectedUrls. Review the crawler's failed request count and explicitly setupstreamFailedRequests; null stops the workflow. The count is your assertion, while the source run's successful completion is checked through the API. - Create an n8n Header Auth credential: header
Authorization, valueBearer YOUR_APIFY_TOKEN. Select it in all three HTTP Request nodes. Keep the token in credentials, never in the workflow JSON or a URL. Use a token with access to your source run/dataset and the audit run and its outputs. - Run manually. A complete export should pass the configured checks. Add a required URL absent from that same export and run again to demonstrate BLOCK. These are two real audits at $0.10 each, including the BLOCK result. The separate free synthetic Store demo has no custom audit charge.
- Connect only PASS - connect your refresh here to your existing refresh.
The output provides
sourceDatasetIdfor retrieving the original content,report.findingsfor evidence, andreport.manifestas a candidate baseline. Save that manifest only after the destination update succeeds, then pass it aspreviousManifestnext time. Keep the previous data and baseline on BLOCK.
The workflow requires a successful audit run and a matching, non-demo report
with the boolean gatePassed: true. It stops on failed/unfinished runs, API
errors and malformed reports. The audit is capped at $0.10 and 60 seconds, with
no automatic retries. After an ambiguous timeout, inspect the existing Apify
run before starting another: retrying can create another paid audit.
The source dataset must remain immutable until your refresh finishes. The example does not include a destination connector, embeddings, notifications, scheduling or persistent baseline storage. Test those parts of your own workflow before enabling a schedule. n8n hosting, your crawler/destination and post-run storage/download costs may be separate. A PASS does not authorize deletion or prove content accuracy, freshness or whole-site completeness.
Tested integration: n8n 2.6.4 with the live Apify API, October 3, 2026. Complete-export PASS and controlled missing-page BLOCK were verified on three public documentation pages. Failed-source and unreviewed-source checks stop before starting an audit. These are owner tests, not customer results. Other n8n versions and customer-specific destination updates have not been verified.
Copy the importable n8n workflow JSON
{"name": "Crawl Quality Gate - check a finished crawl before refresh","nodes": [{"id": "start-manually","name": "Start manually","type": "n8n-nodes-base.manualTrigger","typeVersion": 1,"position": [0,0],"parameters": {}},{"id": "configure-source","name": "Configure source","type": "n8n-nodes-base.code","typeVersion": 2,"position": [220,0],"parameters": {"jsCode": "// Replace these values with your completed crawl and required pages.\n// Review the crawler's failed-request count yourself; this is not inferred.\nreturn [{ json: {\n sourceRunId: 'REPLACE_WITH_FINISHED_CRAWL_RUN_ID',\n namespace: 'client-help-center',\n expectedUrls: ['https://example.com/help/returns'],\n upstreamFailedRequests: null,\n previousManifest: null\n} }];"}},{"id": "validate-configuration","name": "Validate configuration","type": "n8n-nodes-base.code","typeVersion": 2,"position": [440,0],"parameters": {"jsCode": "const c = $input.first().json;\nif (typeof c.sourceRunId !== 'string' || !/^[A-Za-z0-9]{17}$/.test(c.sourceRunId)) throw new Error('Set sourceRunId to the finished crawler run ID.');\nif (typeof c.namespace !== 'string' || !/^[A-Za-z0-9_-]{1,80}$/.test(c.namespace)) throw new Error('Set a stable namespace.');\nif (!Number.isInteger(c.upstreamFailedRequests) || c.upstreamFailedRequests < 0) throw new Error('Review the source failed-request count and set upstreamFailedRequests explicitly.');\nif (c.upstreamFailedRequests !== 0) throw new Error('Source has failed requests. Keep the current knowledge base.');\nif (!Array.isArray(c.expectedUrls) || c.expectedUrls.length > 1000 || c.expectedUrls.some(u => typeof u !== 'string' || !/^https?:\\/\\/[^\\s]+$/.test(u))) throw new Error('Provide valid expectedUrls.');\nif (!c.expectedUrls.length && !c.previousManifest) throw new Error('Supply required URLs or a previously committed manifest to detect missing pages.');\nreturn [{json:c}];"}},{"id": "read-source-run","name": "Read source run","type": "n8n-nodes-base.httpRequest","typeVersion": 4.2,"position": [660,0],"parameters": {"authentication": "genericCredentialType","genericAuthType": "httpHeaderAuth","url": "={{ 'https://api.apify.com/v2/actor-runs/' + $json.sourceRunId }}","options": {"timeout": 90000,"response": {"response": {"responseFormat": "json"}}}},"retryOnFail": false,"onError": "stopWorkflow","notes": "Select your Apify Header Auth credential. Name: Authorization; value: Bearer followed by your token. Never paste a token into this workflow JSON.","notesInFlow": true},{"id": "require-completed-source","name": "Require completed source","type": "n8n-nodes-base.code","typeVersion": 2,"position": [880,0],"parameters": {"jsCode": "const run = $input.first().json.data;\nconst c = $('Validate configuration').first().json;\nif (!run || run.id !== c.sourceRunId || run.status !== 'SUCCEEDED' || !run.finishedAt || !/^[A-Za-z0-9]{17}$/.test(run.defaultDatasetId ?? '')) throw new Error('Source is not a completed successful crawl. Keep the current knowledge base.');\nconst auditInput = {mode:'dataset', namespace:c.namespace, datasetId:run.defaultDatasetId, sourceFinished:true, upstreamFailedRequests:c.upstreamFailedRequests, expectedUrls:c.expectedUrls};\nif (c.previousManifest) auditInput.previousManifest = c.previousManifest;\nreturn [{json:{auditInput,sourceRunId:run.id}}];"}},{"id": "run-crawl-audit","name": "Run crawl audit","type": "n8n-nodes-base.httpRequest","typeVersion": 4.2,"position": [1100,0],"parameters": {"authentication": "genericCredentialType","genericAuthType": "httpHeaderAuth","method": "POST","url": "https://api.apify.com/v2/acts/obeying_laureate~crawl-quality-gate/runs","sendQuery": true,"queryParameters": {"parameters": [{"name": "waitForFinish","value": "60"},{"name": "timeout","value": "60"},{"name": "memory","value": "256"},{"name": "maxTotalChargeUsd","value": "0.10"},{"name": "restartOnError","value": "false"},{"name": "forcePermissionLevel","value": "LIMITED_PERMISSIONS"}]},"sendBody": true,"specifyBody": "json","jsonBody": "={{ $json.auditInput }}","options": {"timeout": 90000,"response": {"response": {"responseFormat": "json"}}}},"retryOnFail": false,"onError": "stopWorkflow","notes": "Select your Apify Header Auth credential. Name: Authorization; value: Bearer followed by your token. Never paste a token into this workflow JSON.","notesInFlow": true},{"id": "require-successful-audit","name": "Require successful audit","type": "n8n-nodes-base.code","typeVersion": 2,"position": [1320,0],"parameters": {"jsCode": "const run = $input.first().json.data;\nif (!run || run.status !== 'SUCCEEDED' || !run.finishedAt || !/^[A-Za-z0-9]{17}$/.test(run.defaultKeyValueStoreId ?? '') || !/^[A-Za-z0-9]{17}$/.test(run.id ?? '')) throw new Error('Audit has not finished successfully. Keep the current knowledge base. Inspect the existing run before retrying: ' + (run?.id ?? 'unknown'));\nreturn [{json:{auditRunId:run.id,storeId:run.defaultKeyValueStoreId}}];"}},{"id": "read-audit-report","name": "Read audit report","type": "n8n-nodes-base.httpRequest","typeVersion": 4.2,"position": [1540,0],"parameters": {"authentication": "genericCredentialType","genericAuthType": "httpHeaderAuth","url": "={{ 'https://api.apify.com/v2/key-value-stores/' + $json.storeId + '/records/REPORT' }}","options": {"timeout": 90000,"response": {"response": {"responseFormat": "json"}}}},"retryOnFail": false,"onError": "stopWorkflow","notes": "Select your Apify Header Auth credential. Name: Authorization; value: Bearer followed by your token. Never paste a token into this workflow JSON.","notesInFlow": true},{"id": "validate-decision","name": "Validate decision","type": "n8n-nodes-base.code","typeVersion": 2,"position": [1760,0],"parameters": {"jsCode": "const report = $input.first().json;\nconst source = $('Require completed source').first().json;\nconst audit = $('Require successful audit').first().json;\nif (report.schemaVersion !== 1 || report.demo !== false || report.namespace !== source.auditInput.namespace || report.source?.type !== 'dataset' || report.source?.id !== source.auditInput.datasetId || typeof report.gatePassed !== 'boolean' || !['PASS','BLOCK'].includes(report.status) || report.gatePassed !== (report.status === 'PASS') || report.deletionAuthorized !== false || report.manifest?.acceptable !== report.gatePassed || !Number.isInteger(report.summary?.errorCount) || report.summary.errorCount < 0 || report.gatePassed !== (report.summary.errorCount === 0)) throw new Error('Unexpected audit report. Keep the current knowledge base.');\nreturn [{json:{gatePassed:report.gatePassed,auditRunId:audit.auditRunId,sourceRunId:source.sourceRunId,sourceDatasetId:source.auditInput.datasetId,report,reportUrl:'https://console.apify.com/actors/runs/' + audit.auditRunId + '#output'}}];"}},{"id": "gate-passed-","name": "Gate passed?","type": "n8n-nodes-base.if","typeVersion": 2.2,"position": [1980,0],"parameters": {"conditions": {"options": {"caseSensitive": true,"leftValue": "","typeValidation": "strict","version": 2},"conditions": [{"id": "gate-true","leftValue": "={{ $json.gatePassed }}","rightValue": true,"operator": {"type": "boolean","operation": "true","singleValue": true}}],"combinator": "and"},"options": {}}},{"id": "pass-connect-your-refresh-here","name": "PASS - connect your refresh here","type": "n8n-nodes-base.noOp","typeVersion": 1,"position": [2200,-120],"parameters": {},"notes": "Only attach the destination refresh to this branch. Read warnings. Persist report.manifest only after the refresh succeeds. PASS is not permission to delete records.","notesInFlow": true},{"id": "block-keep-existing-knowledge-base","name": "BLOCK - keep existing knowledge base","type": "n8n-nodes-base.noOp","typeVersion": 1,"position": [2200,140],"parameters": {},"notes": "Inspect report.findings. Leave the destination and committed manifest unchanged. This branch intentionally ends without writes.","notesInFlow": true},{"id": "read-before-connecting","name": "Read before connecting","type": "n8n-nodes-base.stickyNote","typeVersion": 1,"position": [480,-430],"parameters": {"height": 350,"width": 660,"content": "## Crawl Quality Gate — stop partial exports before a refresh\n1. Enter a completed crawler run ID, required page URLs and reviewed failed-request count.\n2. Choose the same Apify Header Auth credential in all three HTTP Request nodes. Keep credentials out of the JSON.\n3. A real completed audit (PASS or BLOCK) costs $0.10. The run cap is $0.10. No automatic retry.\n4. Both branches currently end without changing anything. Connect only the PASS branch to your existing destination update.\n5. Save report.manifest only after that update succeeds. A PASS never authorizes deletion.\n6. API errors, invalid reports and incomplete runs stop execution. Inspect an ambiguous run before retrying to avoid another charge."}}],"connections": {"Start manually": {"main": [[{"node": "Configure source","type": "main","index": 0}]]},"Configure source": {"main": [[{"node": "Validate configuration","type": "main","index": 0}]]},"Validate configuration": {"main": [[{"node": "Read source run","type": "main","index": 0}]]},"Read source run": {"main": [[{"node": "Require completed source","type": "main","index": 0}]]},"Require completed source": {"main": [[{"node": "Run crawl audit","type": "main","index": 0}]]},"Run crawl audit": {"main": [[{"node": "Require successful audit","type": "main","index": 0}]]},"Require successful audit": {"main": [[{"node": "Read audit report","type": "main","index": 0}]]},"Read audit report": {"main": [[{"node": "Validate decision","type": "main","index": 0}]]},"Validate decision": {"main": [[{"node": "Gate passed?","type": "main","index": 0}]]},"Gate passed?": {"main": [[{"node": "PASS - connect your refresh here","type": "main","index": 0}],[{"node": "BLOCK - keep existing knowledge base","type": "main","index": 0}]]}},"settings": {"executionOrder": "v1"},"active": false,"pinData": {},"tags": []}
Interpreting checks
- A required URL must have one valid row. Missing, failed or duplicate rows do not satisfy it. Default baseline tolerance is zero missing URLs.
maxMissingRatiocan allow intentional baseline changes; it does not waive required URLs and never proves that absent pages were deleted on the website.- Fragments are removed from URL identity; query strings, paths and trailing slashes are preserved. Canonical tags are not used to silently merge pages.
- A title resembling an error page is a warning. Such a title can also belong to legitimate documentation, so this heuristic alone does not block.
- A
PASSmeans configured checks passed. Without required URLs or a prior manifest, omitted pages are undetectable. Even with them, newly added pages missed by the source are unknown. Content truth, freshness, semantic quality, hidden errors and whole-site completeness are not verified.
Limits and data handling
Up to 1,000 rows, 8 MB serialized source JSON and 200,000 content characters per document. Oversized input is rejected rather than truncated. Datasets are read with read-only permissions and checked for changes during the read; use a finished dataset. No website URL is fetched by this Actor.
Reports contain URLs, counts and hashes, not source text. Pasted input remains in the run's INPUT storage under Apify's retention and access settings. Process only data you may use, and keep confidential inputs and URLs private.
Pilot pricing
US$0.10 per completed real audit, up to the limits
above. A useful BLOCK report is billable, as is PASS. The synthetic demo has
no custom event charge. Invalid configuration or a failed source read does not
produce an audit event. Check the actual pricing tab before running. Run-time
platform usage is included in this event price. Post-run storage and downloads
follow Apify's standard rules. There is no separate start or automatic row fee.
A real export check
In our October 2026 owner test, Website Content Crawler retrieved three public
Apify documentation pages. All three passed when listed as required URLs. A
controlled copy with one row removed returned BLOCK, identifying the missing
URL and the one-third coverage loss against the accepted manifest. This is a
technical example, not evidence of customer demand or prevented financial loss.
Development
Requires Node.js 22 or newer. Install with npm ci --ignore-scripts. Run
npm test. For a local sample, create the results parent directory and run
npm run demo. The local command writes JSON and HTML to a new directory and
refuses to overwrite an existing result. Cloud execution uses npm start.
Report reproducible issues with sanitized example input. Do not post tokens, private content or customer URLs publicly.