Dataset Filter - keep only the rows that match your conditions avatar

Dataset Filter - keep only the rows that match your conditions

Under maintenance

Pricing

$0.50 / 1,000 record filtereds

Go to Apify Store
Dataset Filter - keep only the rows that match your conditions

Dataset Filter - keep only the rows that match your conditions

Under maintenance

Point it at any scraper's dataset, write your conditions in plain language, and get back only the rows that match - with a reason on every row. Never deletes a row it could not read. Pay only for records it actually judged.

Pricing

$0.50 / 1,000 record filtereds

Rating

0.0

(0)

Developer

Deric Rifqi

Deric Rifqi

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Dataset Filter — keep only the rows that match your conditions

Your scraper returned 8,000 rows. You need the 300 that actually match. Write your conditions as plain yes/no statements, point this at the dataset, and get back only those rows — each with a score, a verdict and a reason.

criteria:
- key: is_hiring question: The company is currently hiring engineers.
- key: in_europe question: The company operates in Europe.
- key: not_agency question: This is an operating company, not an agency or consultancy.
mode: all

Every returned row keeps all your original fields and gains a _filter object:

field
verdictkeep · review · drop
reasonall_criteria_met, partial_match, insufficient_data, or failed_<your key>
score0–10, the weighted average of your conditions
rationalethe per-condition probabilities behind it
rules_appliedwhich safety gates fired

Three ways to combine conditions

  • all — every condition must hold. Each one becomes its own gate, so when a row is dropped you are told which condition it failed, not just that it failed.
  • any — at least one must hold. Useful for "anything mentioning X, Y or Z".
  • score — the weighted average must clear a threshold. Give each condition a weight. Rows between dropBelow and threshold land in review rather than vanishing.

It will not silently lose your rows

A filter that deletes something you needed has done damage you cannot see: the row is gone and nothing tells you it was there. So the whole thing is asymmetric.

  • A row it could not read — empty, truncated, or about something else — is never dropped. It comes back as insufficient_data for you to look at. Every condition gate is guarded on this, in all three modes.
  • A row it could not judge (an API failure) is kept, flagged, and not billed.
  • An unconfident drop becomes review.
  • Only your own conditions can drop a row. Our uncertainty never can.
  • If more than a fifth of the rows cannot be judged, the run fails rather than handing you a filtered list that is quietly incomplete.

Leave review switched on in What to return unless you have a reason not to. That bucket is where the borderline rows go.

What you pay

Billed per record judged — not per row returned, and never for a row that could not be judged. maxItems is a hard ceiling, so you know the cost of a run before you start it.

Note that everything is judged and billed, including rows you filter out of the output: judging a row is the work. If you only want to pay for a subset, cut it down with maxItems or filter upstream.

Honest about the limits

The conditions are yours, so there is nothing to calibrate against: the score is a plain weighted average of the per-condition probabilities, not a fitted model, and it is presented as one. Sharp, narrow, factual conditions work; vague ones ("this is a good fit") do not — one broad question is exactly what this approach exists to avoid.

Around 13 records a second at the default concurrency.