XML to CSV and Excel Converter
Pricing
Pay per event
XML to CSV and Excel Converter
Flatten XML text, uploaded files, and public XML URLs into CSV- and Excel-ready rows with record selection, namespaces, attributes, nested fields, columns, and parse errors.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Stas Persiianenko
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
Turn XML documents into clean, spreadsheet-ready rows without maintaining a conversion script. This XML to CSV converter accepts pasted XML, uploaded files, public XML URLs, and mixed batches. It detects repeated records or follows your chosen record-node path, then flattens attributes and nested values into dataset columns.
The default Apify dataset can be downloaded as CSV, Excel, JSON, XML, and other supported formats. The same rows are available through the Dataset API for scheduled pipelines.
What does this XML converter do?
For each source, the Actor:
- validates the XML before conversion;
- selects records from
recordPath, or detects the first repeated object node; - optionally removes namespace prefixes;
- includes or excludes XML attributes;
- flattens nested elements into columns;
- joins repeated scalar values;
- applies an ordered column list when supplied;
- saves successful records to the default dataset; and
- saves a diagnostic error row when one source cannot be parsed or downloaded.
Valid sources continue even when another source in the same batch fails. A run in which every source fails exits with a non-zero status.
Who is it for?
- Analysts converting vendor or ERP XML exports before opening them in Excel.
- Data engineers normalizing XML feeds for warehouses, ETL tools, or scheduled imports.
- Operations teams turning recurring catalog, inventory, order, or report files into stable columns.
- Developers who need an API-based XML to CSV conversion step without hosting parser code.
- Automation builders connecting XML-producing systems to Make, Zapier, webhooks, or Apify schedules.
Why use this Actor?
Unlike a one-off browser converter, the Actor supports repeatable runs, mixed batches, schedules, API access, and dataset integrations. You can control the row node and output columns instead of accepting an opaque automatic mapping. Source metadata stays attached to every row, which helps trace batch results.
No browser or proxy is used. Conversion happens in the Actor container, while public or uploaded files are downloaded directly over HTTP or HTTPS.
Supported XML inputs
Choose any combination of these routes:
| Input route | Best for |
|---|---|
xmlText | One pasted document or an API request containing XML |
xmlFile | One file uploaded through Apify Console or a public file URL |
xmlUrls | Several public XML files in one run |
sources | A named mixed batch of inline XML and public URLs |
Each object in sources must contain exactly one of xml or url.
Use name to give its output rows a recognizable _sourceName.
Public downloads are limited to 20 MB per source. Only HTTP and HTTPS URLs that resolve to public addresses are accepted. Authenticated URLs, embedded URL credentials, local hosts, and private network addresses are not supported.
Select the XML nodes that become rows
Set recordPath to a dot- or slash-separated path such as:
catalog.productorders/orderfeed.entries.entry
When namespace removal is enabled, use names without prefixes.
For example, erp:orders/erp:order becomes orders.order.
If recordPath is blank, the Actor selects the first repeated object node it finds.
If there is no repeated object node, the root value becomes one row.
For production pipelines, set recordPath explicitly so a source-structure change cannot alter row selection silently.
Flatten attributes, nested fields, and arrays
With the default settings, this XML:
<product sku="P-100"><name>Travel Mug</name><price currency="USD">18.5</price><tags><tag>travel</tag><tag>kitchen</tag></tags></product>
produces columns like:
{"@sku": "P-100","name": "Travel Mug","price.@currency": "USD","price.#text": 18.5,"tags.tag": "travel | kitchen"}
Change attributePrefix, nestedSeparator, or arraySeparator to match your downstream naming convention.
Set includeAttributes to false when attributes are not needed.
Keep a stable column set
Use columns to define the order and names of XML-derived columns:
{"columns": ["@id","customer.name","customer.region","total.@currency","total.#text"]}
A missing selected value becomes null.
Leaving columns empty preserves every detected flattened field.
Names beginning with _ are reserved for source metadata and cannot be selected as XML columns.
Input parameters
| Field | Type | Default | Description |
|---|---|---|---|
xmlText | string | — | One inline XML document |
xmlFile | string | — | Uploaded file URL or public XML URL |
xmlUrls | string[] | [] | Public XML files to process |
sources | object[] | [] | Named objects containing exactly one xml or url |
recordPath | string | automatic | Dot or slash path to the node used as each row |
maxRows | integer | 1000 | Successful-row limit across all sources, from 1 to 100,000 |
removeNamespaces | boolean | true | Remove prefixes such as ns: from element and attribute names |
includeAttributes | boolean | true | Include XML attributes as flattened columns |
attributePrefix | string | @ | Prefix applied to attribute names |
nestedSeparator | string | . | Separator used in nested column names |
arraySeparator | string | ` | ` |
columns | string[] | [] | Optional ordered output column list |
At least one XML source is required.
Output fields
XML-derived fields vary by document. Every row also includes these stable metadata fields:
| Field | Meaning |
|---|---|
_sourceName | Input label or generated source name |
_sourceType | text, file, or url |
_sourceUrl | Download URL, otherwise null |
_recordPath | Selected or detected row path |
_recordIndex | One-based row number within the source |
_error | Parse or download error, otherwise null |
Error rows contain metadata and _error, have a null _recordIndex, and are not charged as converted items.
Example output
{"@sku": "P-100","name": "Travel Mug","category": "Kitchen","price.#text": 18.5,"price.@currency": "USD","_sourceName": "Inline XML","_sourceType": "text","_sourceUrl": null,"_recordPath": "catalog.product","_recordIndex": 1,"_error": null}
Open the default dataset to download all dynamic XML columns. The overview view focuses on source metadata and errors.
How much does it cost to convert XML to spreadsheet rows?
Pricing has two parts: one start event per run and one item event per successfully converted row.
Error rows have no item charge.
The current event prices and volume tiers are shown in Apify Console before every run.
At the Bronze rate of $0.00005 per run plus $0.00018 per converted row:
| Successful rows | Example cost |
|---|---|
| 10 | $0.00185 |
| 100 | $0.01805 |
| 1,000 | $0.18005 |
Larger usage tiers reduce the per-row price. Actual charges follow the active pricing shown on the Actor page; the examples exclude unrelated platform storage or compute charges that may apply under your Apify plan.
Get started
- Open the Actor in Apify Console.
- Paste XML into XML text, upload a file, or add public XML URLs.
- Leave Record node path blank for automatic detection, or set an explicit path.
- Adjust namespace, attribute, and flattening options.
- Add
columnsif your spreadsheet import requires a stable schema. - Set
maxRowsfor the run. - Click Start.
- Open the default dataset and export it as CSV or Excel.
The prefilled catalog input is ready to run and returns two product rows.
Schedule recurring XML to CSV conversion
Create an Apify Schedule using the same input whenever a public XML feed updates.
Explicit recordPath and columns settings keep the destination schema stable between runs.
Common workflows include:
- nightly supplier catalog conversion;
- weekly inventory or price imports;
- recurring ERP order extracts;
- XML feed normalization before warehouse loading; and
- conversion followed by dataset webhooks.
For change detection, pass the resulting exports to a downstream comparison step rather than treating the converter itself as a monitoring service.
Integrations
The default dataset works with:
- Apify schedules and webhooks;
- Make and Zapier;
- Google Sheets and Microsoft Excel exports;
- Python, JavaScript, or shell scripts using the Dataset API;
- cloud storage and database loaders; and
- Apify's MCP server.
Use _sourceName and _sourceUrl to preserve lineage when merging several inputs.
Run with the API using cURL
Replace YOUR_APIFY_TOKEN with your token:
curl -X POST \"https://api.apify.com/v2/acts/automation-lab~xml-to-csv-excel-converter/runs?token=YOUR_APIFY_TOKEN&waitForFinish=300" \-H "Content-Type: application/json" \-d '{"xmlText": "<catalog><item id=\"1\"><name>Sample item</name></item></catalog>","recordPath": "catalog.item","maxRows": 100}'
Read the run's defaultDatasetId, then request /v2/datasets/DATASET_ID/items?format=csv or format=xlsx.
Run with JavaScript
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: process.env.APIFY_TOKEN });const run = await client.actor('automation-lab/xml-to-csv-excel-converter').call({xmlUrls: ['https://www.w3schools.com/xml/plant_catalog.xml'],recordPath: 'CATALOG.PLANT',columns: ['COMMON', 'BOTANICAL', 'ZONE', 'PRICE'],maxRows: 100,});const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(items);
Run with Python
import osfrom apify_client import ApifyClientclient = ApifyClient(os.environ["APIFY_TOKEN"])run = client.actor("automation-lab/xml-to-csv-excel-converter").call(run_input={"xmlText": "<orders><order id='1'><total>42</total></order></orders>","recordPath": "orders.order","columns": ["@id", "total"],})items = client.dataset(run["defaultDatasetId"]).list_items().itemsprint(items)
Use with MCP and AI assistants
Add this Actor to Claude Code:
claude mcp add --transport http apify \"https://mcp.apify.com?tools=automation-lab/xml-to-csv-excel-converter"
Claude Desktop, Cursor, and VS Code setup
Claude Desktop, Cursor, and VS Code can use this MCP configuration:
{"mcpServers": {"apify": {"url": "https://mcp.apify.com?tools=automation-lab/xml-to-csv-excel-converter"}}}
Example prompts:
- “Convert this product XML into rows and keep SKU, name, category, and price columns.”
- “Download this public XML feed, use
catalog.itemas the record path, and return 500 rows.” - “Flatten these namespaced order documents and explain any parse-error rows.”
Never place secrets or private authenticated URLs in prompts sent to an AI assistant.
Limits and failure behavior
- Each downloaded file is limited to 20 MB.
- A run returns at most 100,000 successful rows.
- Public URL downloads time out after 30 seconds per attempt.
- Network failures, HTTP 429, and HTTP 5xx responses receive bounded retries.
- Stable client errors are not retried.
- DTD validation and XSD schema validation are not provided.
- The Actor does not fetch authenticated, private-network, FTP, or local files.
- Complex mixed-content XML is flattened according to the parser's object representation.
- Repeated objects become indexed nested columns when they are inside, rather than equal to, the selected record node.
- Automatic record detection chooses the first repeated object node; use
recordPathwhen ambiguity matters.
A failed source writes an error row so batch diagnostics remain inspectable. When all sources fail, the run also exits non-zero.
Legality and responsible use
Only process XML that you are authorized to access and transform. Respect source terms, privacy rules, retention requirements, and intellectual-property rights. Do not use signed file URLs beyond their intended audience or lifetime.
The Actor does not need source credentials and rejects URLs with embedded credentials. Apify stores run inputs and datasets according to your platform storage settings, so choose retention settings appropriate for sensitive business exports.
Troubleshooting
Why did I get “record path was not found”?
Check capitalization and namespace handling.
XML names are case-sensitive.
When removeNamespaces is enabled, omit prefixes from the path.
Run once without recordPath and inspect _recordPath to see the automatically detected value.
Why are my desired fields missing?
Remove columns to inspect every detected field, then copy the exact flattened names into your ordered list.
Attributes use attributePrefix, and element text paired with attributes commonly appears under #text.
Why is there one row instead of many?
The selected node may be the container rather than the repeated child.
For <orders><order>...</order></orders>, use orders.order, not orders.
Why did a URL fail?
Confirm it is a public HTTP or HTTPS URL, resolves to public addresses, returns the XML directly, is no larger than 20 MB, and does not require cookies or login.
The _error field contains the specific download or parse message.
Frequently asked questions
Does it create a physical .csv or .xlsx file?
The Actor writes normalized rows to the default Apify dataset. Use the dataset Export button or API format parameter to download CSV or Excel, avoiding duplicate stored output files.
Can I convert several XML documents in one run?
Yes. Use xmlUrls or sources, and use meaningful source names for lineage.
Are parse errors charged as items?
No. Error rows are preserved for diagnosis but only successfully converted rows emit the item charge event.
Can I preserve namespace prefixes?
Yes. Set removeNamespaces to false, then use prefixed element names in recordPath and columns.
Can I choose my own delimiter?
Yes. nestedSeparator controls nested column names and arraySeparator controls repeated scalar values.
The final CSV delimiter is selected when exporting the Apify dataset.
Related automation-lab Actors
- XML JSON Converter for bidirectional XML and JSON conversion.
- JSON to CSV Converter when the source format is already JSON.
- CSV Diff Tool for comparing two tabular exports after conversion.
Support
If a valid XML structure does not flatten as expected, include a small anonymized XML sample, the input options, and the expected row columns in your Apify issue. Remove personal data, credentials, private URLs, and confidential business values before sharing a sample.

