RAG Content Quality Auditor avatar

RAG Content Quality Auditor

Pricing

from $3.00 / 1,000 page auditeds

Go to Apify Store
RAG Content Quality Auditor

RAG Content Quality Auditor

Audit websites or Apify datasets before RAG/LLM ingestion. Score content quality, detect duplicates and low-value pages, estimate tokens, and identify what to keep for AI knowledge bases, embeddings, and vector databases.

Pricing

from $3.00 / 1,000 page auditeds

Rating

0.0

(0)

Developer

Flavian COMBES

Flavian COMBES

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

Audit website or dataset content before it reaches your RAG, LLM, AI agent, embeddings, or vector-database pipeline.

RAG Content Quality Auditor scores every successfully analyzed document, identifies exact/near duplicates and low-value pages, estimates tokens, suggests chunking defaults, and tells you which documents are worth keeping for ingestion.

Use it as the quality gate between crawling and RAG ingestion. Review what you collected before spending time and money on chunking, embeddings, indexing, and vector storage. The Actor is deterministic, explainable, and does not call an external LLM. It measures content quality signals, not retrieval accuracy or business relevance.

1. What it does

  • Recommends which accessible documents to include or review before ingestion.
  • Flags exact and likely near-duplicate main content within the run.
  • Measures content size, structure, boilerplate and machine readability.
  • Identifies probable JavaScript shells that need a different upstream crawler.
  • Estimates tokens and suggests starting chunk sizes.
  • Produces page-level explanations and a reconciled global audit report.

2. Why audit before RAG ingestion

A crawler can collect pages that your AI knowledge base does not need: navigation pages, repeated content, empty shells or pages with little useful text. An audit makes those properties visible before you embed the corpus. Keep your existing crawler and use this Actor as a content quality check between collection and ingestion.

3. What you get

OutputWhere to find itWhat to do next
Page analysesDefault dataset / Page analysesFilter recommended_for_rag = true, inspect reasons and join to your source data
AUDIT_REPORTDefault key-value store / Global audit reportReview corpus tokens, duplicate counts and top issues
DIAGNOSTICSDefault key-value store / Skipped and failed resourcesFix or exclude unprocessed resources; no successful-analysis event is requested

Illustrative output from the repository's synthetic documentation fixture:

{
"url": "https://example.com/document-0",
"rag_readiness_score": 91.46,
"rag_readiness_level": "excellent",
"rag_value": "medium",
"recommended_for_rag": true,
"estimated_tokens": 239,
"duplicate_status": "unique",
"content_status": "ok",
"recommended_chunk_size": 239,
"recommended_chunk_overlap": 0
}

This is an example, not a measured performance claim. examples/page-results.json include an exact duplicate and a JavaScript shell, with explanations for excluding both.

4. Quick start

In Apify, choose Website, enter your website URL and set Maximum pages. Run the Actor, then open Page analyses and Global audit report.

For a first run, the input form suggests eight pages from the public Python tutorial, link depth 1, and include_patterns = ["https://docs.python.org/3/tutorial/*"], matching .actor/INPUT.json. Replace the pattern when auditing another section, or clear it to crawl the whole hostname within the configured limits. API calls that omit max_pages use 50 and omitted include_patterns remains unrestricted within the hostname. Start small and inspect results before expanding an audit.

5. Example website input

{
"mode": "website",
"start_url": "https://docs.python.org/3/tutorial/",
"max_pages": 8,
"max_depth": 1,
"include_patterns": ["https://docs.python.org/3/tutorial/*"]
}

Website mode visits only the exact starting hostname. It follows bounded sitemap and link discovery, respects robots.txt by default and fetches server-rendered content. It does not execute JavaScript.

6. Example dataset input

{
"mode": "dataset",
"dataset_id": "YOUR_SOURCE_DATASET_ID",
"max_pages": 100
}

Use the Source dataset picker in Apify Console to grant this limited-permissions Actor read-only access to the dataset you want to audit. API callers can still pass the dataset ID or unique name directly. This mode analyzes existing content without recrawling its URLs. Text-only records are supported even when their source URL is missing; use source_id to find the original zero-based row.

7. Example output: interpreting a recommendation

The sample documentation page receives these components:

{
"score_breakdown": {
"content": 22.96,
"structure": 18.5,
"cleanliness": 15.0,
"uniqueness": 15.0,
"metadata": 10.0,
"accessibility": 10.0
}
}

A duplicate can still contain well-structured, readable information. Its uniqueness score and recommendation change because an earlier document already represents that content. duplicate_of identifies that source document. Review near-duplicates before discarding them: small differences can matter.

8. Global audit report

The report includes discovery/selection/processing counts, successful/failed/skipped totals, duplicate counts, recommended pages, top issues and estimated corpus tokens. It also records runtime, website HTTP request count and optional llms.txt observations.

estimated_corpus_tokens includes recommended documents only. estimated_analyzed_tokens includes every successfully analyzed document, including duplicates and low-value pages. overall_score averages unique or duplicate-unchecked documents while excluding navigation/index profiles; score_denominator makes the population explicit. It is null when no eligible document remains.

examples/AUDIT_REPORT.json. Include/exclude pattern suggestions are empty in V1 because reliable path-level inference is not implemented; page-level recommendations remain actionable.

9. Use cases

  • Prepare website content before embeddings.
  • Clean documentation before RAG ingestion.
  • Remove repeated pages from a crawler dataset.
  • Estimate the size of an AI knowledge base.
  • Identify low-value pages for manual review.
  • Prepare a corpus for Pinecone, Qdrant, Weaviate or another vector database through your own pipeline. No direct database integrations are implemented.
  • Give an AI agent a machine-readable quality gate before knowledge ingestion.

10. Scoring methodology

ComponentMaximumSignals
Content usefulness30Logarithmic word-count signal, capped at 600 words
Structure20Title, headings, paragraphs, lists/tables/code; mild heading-jump penalty
Cleanliness15Main-content share with an allowance for ordinary navigation/footer text
Uniqueness15Unique: 15; near duplicate: 5; exact duplicate: 0; unchecked: 7.5
Metadata10Title: 4; description, canonical and JSON-LD presence: 2 each
Machine accessibility10Meaningful parseable text: 10; short: 7; empty/JS shell: 2

Levels: excellent ≥85, good ≥70, mixed ≥50, poor <50.

recommended_for_rag is true only for a successful analysis with a score ≥70, content_status = ok, at least 60 word units, a unique or duplicate-unchecked status, and a profile other than navigation/index. Exact and near duplicates are excluded. Disabling duplicate detection explicitly marks uniqueness as unverified.

rag_value is a separate content/profile proxy, not a judgment of business relevance. No rule understands whether a document answers your specific questions. Missing llms.txt has zero effect on scoring. All formulas and thresholds are in ARCHITECTURE.md.

11. Input reference

FieldDefaultMeaning / limits
modewebsitewebsite or dataset
start_urlPython tutorialPublic HTTP(S); required in website mode; no credentials, ports 80/443 only
dataset_idemptyRequired in dataset mode; Console uses an Apify dataset picker with read-only access; API callers may pass an ID or unique name
max_pages501–1,000 page attempts or source rows, including failures; form suggests 8
max_depth30–10; form suggests 1; start=0, sitemap pages=1; 0 disables sitemap discovery
include_patterns[]Full normalized URL globs; form suggests https://docs.python.org/3/tutorial/*; empty includes all eligible same-host URLs
exclude_patterns[]Exclusions win, including on page redirects
respect_robots_txttrueCrawl-rule checks for website requests
detect_duplicatestrueMain-content comparisons within this run
check_ai_metadatatrueObserve /llms.txt and /llms-full.txt in website mode
max_concurrency41–8 simultaneous page fetches
request_timeout_secs255–60 seconds per HTTP attempt

Patterns use shell-style globs, not regular expressions. They are case-sensitive after URL normalization. At most 20 patterns of 200 characters per list. Robots/sitemap/llms discovery files are outside page include/exclude filters, but remain subject to network, host and applicable robots checks. Website-specific settings are ignored in dataset mode.

12. Output reference

FieldsPurpose
source_type, source_idOrigin and stable join key
url, requested_url, final_url, canonical_urlSupplied/requested/final URLs; canonical is an untrusted signal
status, http_status, content_status, error_codeAnalysis/HTTP/content diagnostics; unavailable HTTP values are null
rag_readiness_score, rag_readiness_level, score_breakdownExplainable quality score
rag_value, page_profile, recommended_for_ragRetrieval-potential proxy, content profile and include/review decision
estimated_tokens, token_estimation_methodApproximate token size
raw_html_chars, raw_text_chars, main_content_chars, word_count, boilerplate_ratioExtraction measurements
Heading/paragraph/list/table/code/link countsStructure signals; navigation links are counted for discovery
duplicate_status, duplicate_of, similarity, duplicate_matchDuplicate classification and representative
has_canonical, has_json_ld, has_meta_description, metadata_availableObserved HTML metadata availability
robots_checked, robots_allowed, robots_urlFactual crawl-policy decision, not legal consent
recommended_chunk_size, recommended_chunk_overlap, chunking_reasonStarting guidance in estimated tokens
reasons, issues, recommendations, collected_atShort explanations and collection timestamp

The .actor/dataset_schema.json is the complete field contract. Raw HTML and full document text are intentionally not included. Join analyses back to your source dataset to build the retained corpus. Diagnostics use the same shape but are saved in DIAGNOSTICS, with recommended_for_rag = false.

13. Using with Website Content Crawler / RAG Web Browser

Run your existing crawler, then pass its dataset ID to this Actor. Content precedence is markdown → cleanText → text → content → html. Nonempty strings with at least 40 characters are considered first; shorter fields are fallback candidates. HTML is parsed before scoring. URL precedence is loadedUrl → url → canonicalUrl.

HTML metadata is not inferred from cleaned text, and missing metadata is explicitly reported. Empty/unusable rows are diagnostic records, not paid analyses. These are field-compatible workflows, not claims of official integration or support for every upstream schema.

14. RAG / embeddings / vector database workflow

Crawler or existing dataset
→ RAG Content Quality Auditor
→ join recommended source documents
→ review near-duplicates and domain relevance
→ chunk and embed in your own pipeline
→ vector database / retrieval augmented generation / AI agent

Chunk guidance ranges from a single short chunk to 768 tokens with 96-token overlap for long structured documents. FAQ guidance favors question/answer pairs. These are starting recommendations, not optimal chunking guarantees. This Actor does not chunk, embed or upload your documents to a vector database.

15. MCP / AI-agent usage

The same JSON inputs work through the Actor API and through Apify's Actor tooling where available. No custom MCP server is included. An agent can:

  1. Submit website/dataset input with a small limit.
  2. Read the dataset plus AUDIT_REPORT and DIAGNOSTICS output links.
  3. Filter recommended_for_rag, join by source ID/URL and pass the retained original content to its ingestion tool.
  4. Present uncertain or near-duplicate cases for review.

16. Pricing

Launch pricing is $0.003 per successfully analyzed document using the custom page-audited event — about $3 per 1,000 successful page/document audits. Failed, blocked, unsupported or unusable resources do not request this custom billing event.

An analyzed duplicate, short page or JavaScript shell still represents analysis work and is eligible for the event because the content was fetched/parsed/scored and a result was produced. Diagnostics are stored separately from the dataset. The report distinguishes billable eligibility, requested events and actual platform event counts.

Cloud benchmark evidence

Single-run Apify Cloud checks on 2026-09-30 used 512 MB memory, concurrency 4 and public Python documentation. They validate scale/cost behavior for this build; they are not a universal performance guarantee because websites differ in latency, page size and content type.

Selected pagesSuccessful auditsSkippedRuntimeApify run cost
10100~7 s UI / 2.5 s enginedisplayed as $0.000
50437~1 min UI / 46.6 s engine$0.001
1009281 min 56 s UI / 112.5 s engine$0.003
500491916 min 51 s UI / 1005.3 s engine$0.031

The 500-page run processed all 500 selected URLs with zero page-processing failures; nine resources were skipped because they were unsupported or exceeded configured safety limits. Start with a small audit on your own corpus before increasing the limit.

Platform usage is intended to be included in PPE pricing when the Actor is published. Check the Actor's current Pricing tab for the live commercial configuration. BENCHMARK_PLAN.md records the measurement methodology.

17. Limitations

  • Static public HTML, XHTML, plain text and Markdown only. No browser rendering, login, cookies, proxy rotation, CAPTCHA bypass, PDF extraction, OCR or multimedia analysis.
  • Exact-host bounded discovery; no promise of exhaustive crawling, sitemap coverage or semantic relevance.
  • Heuristic extraction can misidentify unusual templates. Markdown analysis covers common structure, not the full CommonMark specification.
  • Near-duplicate detection is lexical, with bounded candidate comparisons. It can miss matches, especially short or boilerplate-heavy documents. Review version-specific content.
  • Token estimates are Unicode characters divided by four and vary by language/model. CJK word units count characters; other word boundaries use Unicode word groups.
  • Structured-data presence is measured; its semantic correctness is not validated. A canonical URL is not proof of the preferred source.
  • Dataset mode analyzes content as supplied; it does not check whether the original site is currently accessible or whether its source metadata is accurate.
  • V1 does not resume partial runs. A nonempty output dataset is rejected on startup to prevent accidental duplicate publication/charging. Start a fresh run after interruption; previously saved rows remain available.

18. Responsible crawling

Robots.txt is checked by default through the same protected transport. A missing file (404/410) allows crawling with a warning. Network/server errors, access denials or unrecognizable robots content cause conservative skipping. Parser decisions are recorded. These observations do not establish copyright rights, legal authorization or AI-training consent.

All resolved IP addresses must be public. Each redirect is revalidated, cross-host redirects and HTTPS downgrades are refused, and sockets connect to a validated numeric address while retaining hostname/TLS verification. Small request, body, retry and discovery limits constrain work. The implementation does not use environmental HTTP proxies.

19. Technical notes

Python 3.12, Apify SDK, HTTPX, Beautiful Soup and Protego. No external AI API or model download. Apify's SDK has transitive dependencies including Crawlee; this Actor does not use its crawling framework.

python -m venv .venv
# Activate the environment, then:
python -m pip install -e ".[test]"
python -m pytest -q
python scripts/local_smoke.py --mode website
python scripts/local_smoke.py --mode dataset

For the real entry point, place input at storage/key_value_stores/default/INPUT.json, then run python -m my_actor. VALIDATION.md covers Windows commands, live tests, Docker and Cloud validation. ARCHITECTURE.md contains the security model and decision log.

20. FAQ

What is RAG readiness? A set of measurable ingestion-quality signals: useful extracted text, structure, cleanliness, uniqueness, metadata and machine readability. It is not measured answer accuracy.

Does this Actor use an LLM? No. Core analysis uses deterministic local rules.

Does it crawl JavaScript websites? It fetches the initial server response and diagnoses probable JS shells. Use a browser-capable crawler upstream when necessary.

Can I use an existing Apify dataset? Yes. Dataset mode supports common text, Markdown, HTML and URL fields without recrawling.

Does it support Website Content Crawler outputs? It supports common output fields. Verify your actual dataset's field mapping with a small run first.

Are token counts exact? No. They are estimates, not model-specific tokenizer results.

Does a low score mean a page is bad? No. A short answer, a product page or a specialized format may still be valuable for your application. Review the reasons and your retrieval requirements.

Does llms.txt affect the score? No. Availability is optional report metadata.

Does it support PDFs? No. Convert them upstream to cleaned text/Markdown, then audit that dataset.

Does it respect robots.txt? Yes, by default, using a real parser and a documented conservative fallback.

What counts as a paid result? One successfully analyzed page/document. Duplicate and low-value analyses count; failures, blocked resources and unusable rows do not request the custom billing event.

Why are no pages recommended? The corpus may be inaccessible, too short, duplicated or unsuitable under the explicit rules. Read the report and diagnostics. A valid audit with zero recommended pages is not automatically a failed run.