# File-to-Dataset Ingestion & Normalization Pipeline (`quanmatrix/file-to-dataset-processing-pipeline`) Actor

A mixed-format ingestion layer that converts external structured files into normalized Apify Dataset rows, not merely a profiler or dataset-to-dataset transformer.

- **URL**: https://apify.com/quanmatrix/file-to-dataset-processing-pipeline.md
- **Developed by:** [Rafael Barreto Haddad](https://apify.com/quanmatrix) (community)
- **Categories:** Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.17 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## File-to-Dataset Processing Pipeline

A mixed-format ingestion layer that converts external structured files into normalized Apify Dataset rows, not merely a profiler or dataset-to-dataset transformer.

A mixed-format ingestion layer that converts external structured files into normalized Apify Dataset rows, not merely a profiler or dataset-to-dataset transformer.

A mixed-format ingestion layer that converts external structured files into normalized Apify Dataset rows, not merely a profiler or dataset-to-dataset transformer.

Convert CSV, JSON, JSONL and XML into normalized Apify Dataset rows with one lightweight Actor. The pipeline accepts inline content or public HTTP(S) URLs, parses supported structured formats, records source provenance, infers a practical field/type schema, and writes ready-to-use rows to the default Dataset.

### Why use this Actor

Real automation rarely receives perfectly standardized input. Partners send CSV exports, APIs expose JSON, event streams arrive as JSONL, old systems still publish XML, and research workflows often mix several formats at once. Maintaining separate import scripts for every source creates brittle glue code, duplicated validation and avoidable operational cost.

This Actor provides one deterministic ingestion layer before RAG, CRM, analytics, catalog, job-feed, lead-enrichment or agent workflows. It is intentionally HTTP-first and runs with a small 256 MB default footprint.

### Key features

- Parse CSV, JSON, JSONL and XML in the same run.
- Accept inline payloads or public HTTP(S) URLs.
- Reject private, loopback, link-local and otherwise unsafe network targets.
- Combine multiple sources into one normalized Dataset.
- Add `_sourceIndex` and `_rowIndex` provenance fields when requested.
- Infer a field/type schema for the accepted output rows.
- Produce per-source reports with parsed and accepted row counts.
- Store `PIPELINE_SUMMARY` in the default key-value store.
- Cap remote downloads and total output rows to keep runs predictable.
- Work without a browser, login, API key or external LLM.

### Output

Each accepted source row is written to the default Apify Dataset. When provenance is enabled, two metadata fields are added:

- `_sourceIndex`: zero-based source number from the `documents` input array.
- `_rowIndex`: zero-based row number within that parsed source.

The `PIPELINE_SUMMARY` record contains total output rows, successful and failed source counts, an inferred schema and a source-level diagnostic report. This makes the Actor useful both as a converter and as an auditable ingestion gate.

Example normalized Dataset row:

```json
{
  "id": "SKU-42",
  "name": "Example product",
  "price": "19.90",
  "_sourceIndex": 0,
  "_rowIndex": 3
}
```

### Example

Input:

```json
{
  "documents": [
    {
      "name": "catalog",
      "format": "csv",
      "content": "id,name,price\n1,Alpha,10.00\n2,Beta,20.00"
    },
    {
      "name": "status-feed",
      "format": "jsonl",
      "content": "{\"id\":3,\"status\":\"new\"}\n{\"id\":4,\"status\":\"active\"}"
    }
  ],
  "maxRows": 1000,
  "addProvenance": true
}
```

The run writes four normalized rows and a summary describing both sources and the combined field schema.

### Use cases

- **RAG ingestion:** normalize heterogeneous exports before chunking or embedding.
- **CRM migration:** convert partner or legacy CSV/JSON files into consistent Dataset rows.
- **E-commerce catalogs:** ingest supplier feeds and preserve source provenance.
- **Job feeds:** normalize recurring JSONL, CSV or XML job exports.
- **Lead pipelines:** standardize inbound lead lists before enrichment and routing.
- **Research data:** combine structured public files into a single auditable Dataset.
- **Agent workflows:** give AI agents a predictable Dataset endpoint instead of format-specific parsing logic.
- **Migration validation:** inspect parsed row counts and schema evidence before downstream delivery.

### Pricing

The product is designed as pay per normalized output row. Pricing is tiered by the factory according to observed cost and margin. The Actor does not need a separate artificial start event to create billable value.

### Reliability and safety

Remote URLs must resolve to public HTTP(S) targets. Private/local addresses are blocked before fetching. Remote response size is bounded, and the run stops emitting rows after the configured `maxRows` limit. A malformed source is reported at source level; other valid sources can still be processed. If no source yields usable rows, the run fails explicitly instead of returning fabricated data.

### Limitations

- Remote sources must use public HTTP(S).
- Each fetched source is capped at 10 MB.
- Complex XML is normalized conservatively rather than attempting domain-specific mapping.
- Binary formats such as XLSX, Parquet and PDF are outside the initial scope.
- CSV values remain strings unless a downstream transformation explicitly changes them.
- Schema inference describes observed JSON value types; it is not a substitute for a business-domain contract.

The Actor is meant to be a dependable intake layer, not a magical data-cleaning oracle. Humans already invented enough of those.

# Changelog

This Actor's version history is a separate document: https://apify.com/quanmatrix/file-to-dataset-processing-pipeline/changelog.md

# Actor input Schema

## `documents` (type: `array`):

CSV, JSON, JSONL or XML sources. Use either content or a public HTTP(S) URL.

## `maxRows` (type: `integer`):

Maximum number of normalized rows to write to the default Dataset across all supplied documents.

## `addProvenance` (type: `boolean`):

Add \_sourceIndex and \_rowIndex fields so downstream workflows can trace each output row to its input source and position.

## `previousAnalysis` (type: `object`):

Optional prior Gen2 output used to calculate decision-metric deltas and regression.

## `valuePerImpactUnitUsd` (type: `number`):

Optional user-supplied economic value per impact unit. Leave empty to avoid monetary estimation.

## `monthlyRuns` (type: `integer`):

Optional expected monthly run count used only with valuePerImpactUnitUsd for economic impact estimation.

## `mcpConnectors` (type: `array`):

Optional MCP connectors authorized in your Apify account. Use them to send or write this Actor result to tools such as Slack, Notion, GitHub, Sentry, Supabase, or another compatible MCP service.

## `mcpToolName` (type: `string`):

Optional exact MCP tool name. Leave blank to let the selected MCP action preset discover a compatible tool automatically.

## `mcpToolArguments` (type: `object`):

JSON object passed to the selected MCP tool. String values may use {{actor\_title}}, {{result\_summary}}, or {{result\_json}} placeholders.

## `mcpFailOnError` (type: `boolean`):

When enabled, an MCP delivery error fails the Actor run. Disabled by default so data extraction and intelligence results remain available even if the external destination is unavailable.

## `mcpActionPreset` (type: `string`):

Choose a safe action pattern. AUTO\_SAFE\_WRITE discovers a compatible non-destructive write tool automatically; use a specific preset for Slack, GitHub, Notion, or database delivery.

## Actor input object example

```json
{
  "documents": [
    {
      "name": "sample",
      "format": "json",
      "content": "[{\"id\":1,\"name\":\"alpha\"},{\"id\":2,\"name\":\"beta\"}]"
    }
  ],
  "maxRows": 10000,
  "addProvenance": true,
  "monthlyRuns": 1,
  "mcpToolName": "",
  "mcpToolArguments": {},
  "mcpFailOnError": false,
  "mcpActionPreset": "AUTO_SAFE_WRITE"
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("quanmatrix/file-to-dataset-processing-pipeline").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("quanmatrix/file-to-dataset-processing-pipeline").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call quanmatrix/file-to-dataset-processing-pipeline --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,quanmatrix/file-to-dataset-processing-pipeline"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/JIPcc25nIYbykRO6X/builds/eCK8GlzOHOhov1PHc/openapi.json
