# Dataset Filter - keep only the rows that match your conditions (`jev-tools/dataset-filter`) Actor

Point it at any scraper's dataset, write your conditions in plain language, and get back only the rows that match - with a reason on every row. Never deletes a row it could not read. Pay only for records it actually judged.

- **URL**: https://apify.com/jev-tools/dataset-filter.md
- **Developed by:** [Deric Rifqi](https://apify.com/jev-tools) (community)
- **Categories:** AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.50 / 1,000 record filtereds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Dataset Filter — keep only the rows that match your conditions

Your scraper returned 8,000 rows. You need the 300 that actually match. Write
your conditions as plain yes/no statements, point this at the dataset, and get
back only those rows — each with a score, a verdict and a reason.

```
criteria:
  - key: is_hiring     question: The company is currently hiring engineers.
  - key: in_europe     question: The company operates in Europe.
  - key: not_agency    question: This is an operating company, not an agency or consultancy.
mode: all
```

Every returned row keeps all your original fields and gains a `_filter` object:

| field | |
| --- | --- |
| `verdict` | `keep` · `review` · `drop` |
| `reason` | `all_criteria_met`, `partial_match`, `insufficient_data`, or `failed_<your key>` |
| `score` | 0–10, the weighted average of your conditions |
| `rationale` | the per-condition probabilities behind it |
| `rules_applied` | which safety gates fired |

### Three ways to combine conditions

- **all** — every condition must hold. Each one becomes its own gate, so when a
  row is dropped you are told *which* condition it failed, not just that it
  failed.
- **any** — at least one must hold. Useful for "anything mentioning X, Y or Z".
- **score** — the weighted average must clear a threshold. Give each condition a
  `weight`. Rows between `dropBelow` and `threshold` land in `review` rather
  than vanishing.

### It will not silently lose your rows

A filter that deletes something you needed has done damage you cannot see: the
row is gone and nothing tells you it was there. So the whole thing is
asymmetric.

- A row it **could not read** — empty, truncated, or about something else — is
  never dropped. It comes back as `insufficient_data` for you to look at. Every
  condition gate is guarded on this, in all three modes.
- A row it **could not judge** (an API failure) is **kept**, flagged, and **not
  billed**.
- An **unconfident** drop becomes `review`.
- Only **your own conditions** can drop a row. Our uncertainty never can.
- If more than a fifth of the rows cannot be judged, the run **fails** rather
  than handing you a filtered list that is quietly incomplete.

Leave `review` switched on in **What to return** unless you have a reason not
to. That bucket is where the borderline rows go.

### What you pay

Billed **per record judged** — not per row returned, and never for a row that
could not be judged. `maxItems` is a hard ceiling, so you know the cost of a run
before you start it.

Note that everything is judged and billed, including rows you filter out of the
output: judging a row is the work. If you only want to pay for a subset, cut it
down with `maxItems` or filter upstream.

### Honest about the limits

The conditions are yours, so there is nothing to calibrate against: the score is
a **plain weighted average** of the per-condition probabilities, not a fitted
model, and it is presented as one. Sharp, narrow, factual conditions work; vague
ones ("this is a good fit") do not — one broad question is exactly what this
approach exists to avoid.

Around 13 records a second at the default concurrency.

# Actor input Schema

## `criteria` (type: `array`):

One yes/no statement per condition, written about a single record. Each needs a short key and the statement itself; weight is only used in score mode. Write statements that are true of the rows you WANT.

## `mode` (type: `string`):

all: every condition must hold. any: at least one must. score: the weighted average must clear the threshold.

## `threshold` (type: `number`):

In all/any mode: how certain a single condition must be to count as true (0-1, default 0.5). In score mode: the 0-10 score a record must reach to be kept (default 5).

## `dropBelow` (type: `number`):

Records scoring under this are dropped; between this and the threshold they are marked for review. Defaults to half the threshold.

## `sourceDatasetId` (type: `string`):

The dataset ID of any scraper run. Leave empty and paste rows into 'records' instead.

## `records` (type: `array`):

Rows to filter, if you are not reading them from a dataset. Any shape is accepted.

## `keepVerdicts` (type: `array`):

Everything is judged and billed; this only decides what lands in the output. 'review' holds the rows it was not sure about - leaving it on is how you avoid silently losing borderline rows.

## `maxItems` (type: `integer`):

A hard ceiling on the run, and therefore on the cost.

## `idField` (type: `string`):

Which field to use as the record id in the output and the logs. Falls back to the row number when the field is missing.

## `concurrency` (type: `integer`):

Parallel judgements. 12 handles about 13 records a second.

## `includeSignals` (type: `boolean`):

Adds each condition's calibrated probability to every row, so you can re-threshold your own results without paying again.

## Actor input object example

```json
{
  "criteria": [
    {
      "key": "is_hiring",
      "question": "The company is currently hiring engineers."
    },
    {
      "key": "in_europe",
      "question": "The company operates in Europe."
    }
  ],
  "mode": "all",
  "keepVerdicts": [
    "keep",
    "review"
  ],
  "maxItems": 5000,
  "idField": "id",
  "concurrency": 12,
  "includeSignals": false
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "criteria": [
        {
            "key": "is_hiring",
            "question": "The company is currently hiring engineers."
        },
        {
            "key": "in_europe",
            "question": "The company operates in Europe."
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("jev-tools/dataset-filter").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "criteria": [
        {
            "key": "is_hiring",
            "question": "The company is currently hiring engineers.",
        },
        {
            "key": "in_europe",
            "question": "The company operates in Europe.",
        },
    ] }

# Run the Actor and wait for it to finish
run = client.actor("jev-tools/dataset-filter").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "criteria": [
    {
      "key": "is_hiring",
      "question": "The company is currently hiring engineers."
    },
    {
      "key": "in_europe",
      "question": "The company operates in Europe."
    }
  ]
}' |
apify call jev-tools/dataset-filter --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,jev-tools/dataset-filter"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/lIVwLDCKFXMb5c20y/builds/TTCfXasXKixOnZ8Bg/openapi.json
