File-to-Dataset Ingestion & Normalization Pipeline
Pricing
from $1.17 / 1,000 results
File-to-Dataset Ingestion & Normalization Pipeline
A mixed-format ingestion layer that converts external structured files into normalized Apify Dataset rows, not merely a profiler or dataset-to-dataset transformer.
Pricing
from $1.17 / 1,000 results
Rating
0.0
(0)
Developer
Rafael Barreto Haddad
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
File-to-Dataset Processing Pipeline
A mixed-format ingestion layer that converts external structured files into normalized Apify Dataset rows, not merely a profiler or dataset-to-dataset transformer.
A mixed-format ingestion layer that converts external structured files into normalized Apify Dataset rows, not merely a profiler or dataset-to-dataset transformer.
A mixed-format ingestion layer that converts external structured files into normalized Apify Dataset rows, not merely a profiler or dataset-to-dataset transformer.
Convert CSV, JSON, JSONL and XML into normalized Apify Dataset rows with one lightweight Actor. The pipeline accepts inline content or public HTTP(S) URLs, parses supported structured formats, records source provenance, infers a practical field/type schema, and writes ready-to-use rows to the default Dataset.
Why use this Actor
Real automation rarely receives perfectly standardized input. Partners send CSV exports, APIs expose JSON, event streams arrive as JSONL, old systems still publish XML, and research workflows often mix several formats at once. Maintaining separate import scripts for every source creates brittle glue code, duplicated validation and avoidable operational cost.
This Actor provides one deterministic ingestion layer before RAG, CRM, analytics, catalog, job-feed, lead-enrichment or agent workflows. It is intentionally HTTP-first and runs with a small 256 MB default footprint.
Key features
- Parse CSV, JSON, JSONL and XML in the same run.
- Accept inline payloads or public HTTP(S) URLs.
- Reject private, loopback, link-local and otherwise unsafe network targets.
- Combine multiple sources into one normalized Dataset.
- Add
_sourceIndexand_rowIndexprovenance fields when requested. - Infer a field/type schema for the accepted output rows.
- Produce per-source reports with parsed and accepted row counts.
- Store
PIPELINE_SUMMARYin the default key-value store. - Cap remote downloads and total output rows to keep runs predictable.
- Work without a browser, login, API key or external LLM.
Output
Each accepted source row is written to the default Apify Dataset. When provenance is enabled, two metadata fields are added:
_sourceIndex: zero-based source number from thedocumentsinput array._rowIndex: zero-based row number within that parsed source.
The PIPELINE_SUMMARY record contains total output rows, successful and failed source counts, an inferred schema and a source-level diagnostic report. This makes the Actor useful both as a converter and as an auditable ingestion gate.
Example normalized Dataset row:
{"id": "SKU-42","name": "Example product","price": "19.90","_sourceIndex": 0,"_rowIndex": 3}
Example
Input:
{"documents": [{"name": "catalog","format": "csv","content": "id,name,price\n1,Alpha,10.00\n2,Beta,20.00"},{"name": "status-feed","format": "jsonl","content": "{\"id\":3,\"status\":\"new\"}\n{\"id\":4,\"status\":\"active\"}"}],"maxRows": 1000,"addProvenance": true}
The run writes four normalized rows and a summary describing both sources and the combined field schema.
Use cases
- RAG ingestion: normalize heterogeneous exports before chunking or embedding.
- CRM migration: convert partner or legacy CSV/JSON files into consistent Dataset rows.
- E-commerce catalogs: ingest supplier feeds and preserve source provenance.
- Job feeds: normalize recurring JSONL, CSV or XML job exports.
- Lead pipelines: standardize inbound lead lists before enrichment and routing.
- Research data: combine structured public files into a single auditable Dataset.
- Agent workflows: give AI agents a predictable Dataset endpoint instead of format-specific parsing logic.
- Migration validation: inspect parsed row counts and schema evidence before downstream delivery.
Pricing
The product is designed as pay per normalized output row. Pricing is tiered by the factory according to observed cost and margin. The Actor does not need a separate artificial start event to create billable value.
Reliability and safety
Remote URLs must resolve to public HTTP(S) targets. Private/local addresses are blocked before fetching. Remote response size is bounded, and the run stops emitting rows after the configured maxRows limit. A malformed source is reported at source level; other valid sources can still be processed. If no source yields usable rows, the run fails explicitly instead of returning fabricated data.
Limitations
- Remote sources must use public HTTP(S).
- Each fetched source is capped at 10 MB.
- Complex XML is normalized conservatively rather than attempting domain-specific mapping.
- Binary formats such as XLSX, Parquet and PDF are outside the initial scope.
- CSV values remain strings unless a downstream transformation explicitly changes them.
- Schema inference describes observed JSON value types; it is not a substitute for a business-domain contract.
The Actor is meant to be a dependable intake layer, not a magical data-cleaning oracle. Humans already invented enough of those.