Detect duplicate embedding spend
Created by
Sebastián S
Actor
RAG Dataset Linter
Find exact duplicate RAG chunks and excessive overlap, estimate wasted embedding tokens, and stop redundant vectors before indexing.
RAG Dataset Lintersebastian-actors/rag-dataset-linter
Gate
Action
Source type
Source ID
+9 fieldsTextNumberBooleanListObject
Input
Minimum tokens:5
Maximum tokens:2000
Near-duplicate similarity:0.95
Maximum adjacent overlap ratio:0.35
Inline records
chunkId:pricing-0+2
documentId:pricing+2
sourceUrl:https://example.com/pricing+2
chunkIndex:0+2
title:Pricing+2
headingPath
chunkText:Embedding duplicate chunks wastes tokens and creates redundant vectors that can crowd retrieval results.+2
Maximum chunks:1000
Output fields
Gate
Action
Source type
Source ID
Audited
Findings
Errors
Warnings
Exact duplicate rate
Near duplicate rate
Estimated wasted tokens
Truncated
Finished
Sign up on Apify01
Create your Apify account to access the RAG Dataset Linter.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
