File-to-Dataset Ingestion & Normalization Pipeline avatar

File-to-Dataset Ingestion & Normalization Pipeline

Pricing

from $1.17 / 1,000 results

Go to Apify Store
File-to-Dataset Ingestion & Normalization Pipeline

File-to-Dataset Ingestion & Normalization Pipeline

A mixed-format ingestion layer that converts external structured files into normalized Apify Dataset rows, not merely a profiler or dataset-to-dataset transformer.

Pricing

from $1.17 / 1,000 results

Rating

0.0

(0)

Developer

Rafael Barreto Haddad

Rafael Barreto Haddad

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Share

File-to-Dataset Processing Pipeline

A mixed-format ingestion layer that converts external structured files into normalized Apify Dataset rows, not merely a profiler or dataset-to-dataset transformer.

A mixed-format ingestion layer that converts external structured files into normalized Apify Dataset rows, not merely a profiler or dataset-to-dataset transformer.

A mixed-format ingestion layer that converts external structured files into normalized Apify Dataset rows, not merely a profiler or dataset-to-dataset transformer.

Convert CSV, JSON, JSONL and XML into normalized Apify Dataset rows with one lightweight Actor. The pipeline accepts inline content or public HTTP(S) URLs, parses supported structured formats, records source provenance, infers a practical field/type schema, and writes ready-to-use rows to the default Dataset.

Why use this Actor

Real automation rarely receives perfectly standardized input. Partners send CSV exports, APIs expose JSON, event streams arrive as JSONL, old systems still publish XML, and research workflows often mix several formats at once. Maintaining separate import scripts for every source creates brittle glue code, duplicated validation and avoidable operational cost.

This Actor provides one deterministic ingestion layer before RAG, CRM, analytics, catalog, job-feed, lead-enrichment or agent workflows. It is intentionally HTTP-first and runs with a small 256 MB default footprint.

Key features

  • Parse CSV, JSON, JSONL and XML in the same run.
  • Accept inline payloads or public HTTP(S) URLs.
  • Reject private, loopback, link-local and otherwise unsafe network targets.
  • Combine multiple sources into one normalized Dataset.
  • Add _sourceIndex and _rowIndex provenance fields when requested.
  • Infer a field/type schema for the accepted output rows.
  • Produce per-source reports with parsed and accepted row counts.
  • Store PIPELINE_SUMMARY in the default key-value store.
  • Cap remote downloads and total output rows to keep runs predictable.
  • Work without a browser, login, API key or external LLM.

Output

Each accepted source row is written to the default Apify Dataset. When provenance is enabled, two metadata fields are added:

  • _sourceIndex: zero-based source number from the documents input array.
  • _rowIndex: zero-based row number within that parsed source.

The PIPELINE_SUMMARY record contains total output rows, successful and failed source counts, an inferred schema and a source-level diagnostic report. This makes the Actor useful both as a converter and as an auditable ingestion gate.

Example normalized Dataset row:

{
"id": "SKU-42",
"name": "Example product",
"price": "19.90",
"_sourceIndex": 0,
"_rowIndex": 3
}

Example

Input:

{
"documents": [
{
"name": "catalog",
"format": "csv",
"content": "id,name,price\n1,Alpha,10.00\n2,Beta,20.00"
},
{
"name": "status-feed",
"format": "jsonl",
"content": "{\"id\":3,\"status\":\"new\"}\n{\"id\":4,\"status\":\"active\"}"
}
],
"maxRows": 1000,
"addProvenance": true
}

The run writes four normalized rows and a summary describing both sources and the combined field schema.

Use cases

  • RAG ingestion: normalize heterogeneous exports before chunking or embedding.
  • CRM migration: convert partner or legacy CSV/JSON files into consistent Dataset rows.
  • E-commerce catalogs: ingest supplier feeds and preserve source provenance.
  • Job feeds: normalize recurring JSONL, CSV or XML job exports.
  • Lead pipelines: standardize inbound lead lists before enrichment and routing.
  • Research data: combine structured public files into a single auditable Dataset.
  • Agent workflows: give AI agents a predictable Dataset endpoint instead of format-specific parsing logic.
  • Migration validation: inspect parsed row counts and schema evidence before downstream delivery.

Pricing

The product is designed as pay per normalized output row. Pricing is tiered by the factory according to observed cost and margin. The Actor does not need a separate artificial start event to create billable value.

Reliability and safety

Remote URLs must resolve to public HTTP(S) targets. Private/local addresses are blocked before fetching. Remote response size is bounded, and the run stops emitting rows after the configured maxRows limit. A malformed source is reported at source level; other valid sources can still be processed. If no source yields usable rows, the run fails explicitly instead of returning fabricated data.

Limitations

  • Remote sources must use public HTTP(S).
  • Each fetched source is capped at 10 MB.
  • Complex XML is normalized conservatively rather than attempting domain-specific mapping.
  • Binary formats such as XLSX, Parquet and PDF are outside the initial scope.
  • CSV values remain strings unless a downstream transformation explicitly changes them.
  • Schema inference describes observed JSON value types; it is not a substitute for a business-domain contract.

The Actor is meant to be a dependable intake layer, not a magical data-cleaning oracle. Humans already invented enough of those.