# Headerless CSV Batch Import (`sapph1re/headerless-csv-batch-import`) Actor

Import headerless CSV from pasted text or public URLs. Preserve first records, leading zeros and literal strings, with source and line provenance, explicit file errors and row limits.

- **URL**: https://apify.com/sapph1re/headerless-csv-batch-import.md
- **Developed by:** [Roman V](https://apify.com/sapph1re) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Headerless CSV Batch Import

Import a small batch of authorized headerless CSV files using one ordered column
list. The first record stays data. Each delivered row contains literal string
values, its source URL, its record number and the physical lines it occupied.
Separate file receipts explain empty files, excluded widths and failures.

Existing CSV Actors also support batching and custom headers. This utility
focuses on literal strings, source positions and separate file failure receipts.
No competitor reliability, price advantage or superiority is claimed.

### Try the synthetic example

An API request with no body or `{}` uses two synthetic inline records. The Console
prefill and `example-input.json` use the same data:

```json
{
  "inlineCsv": "0001,North,\"first record\"\n0002,South,\n",
  "fieldNames": ["id", "region", "note"],
  "separator": ",",
  "maxRows": 100,
  "maxChargeUsd": 0.05
}
```

For your files, replace `inlineCsv` with `csvUrls`, an array of one to ten exact
public HTTP(S) CSV URLs, and set your column names and separator explicitly.
Remove the inline property completely when selecting URL mode. URL mode does not
supply a fabricated public demonstration file. Use URLs whose content and access rights you have checked.

### What the data means

Each dataset row has these fields:

| Field | Meaning |
| --- | --- |
| `values` | Object mapping your column names to strings, including empty strings |
| `sourceUrl` | Exact input URL, or null for inline input |
| `fileIndex` | Zero-based original input position; exact duplicate URLs use their first position |
| `recordIndex` | One-based logical CSV record, including excluded width records but skipping empty physical lines |
| `physicalStartLine`, `physicalEndLine` | One-based inclusive physical lines, with CRLF counted as one line break |
| `fileId` | SHA-256 identity of the exact source URL or the inline-source marker |
| `rowId` | SHA-256 identity of source, exact source bytes, column list, delimiter and logical position |

Values are never inferred as numbers, dates, booleans or nulls. Leading zeros,
spaces, empty fields, quoted separators, quoted line breaks and doubled quotes
are preserved. A leading UTF-8 BOM is removed from parsing. UTF-8 is the only
supported encoding; invalid bytes fail the file. An entirely empty physical line
is skipped and counted. A delimiter-only row or quoted empty string is a record.
A whitespace-only line is also a record and must have the correct width.

Repeated values at different positions have different row identities. Repeating
an unchanged input and source yields the same identities. Changing the source
bytes, including line endings, can change every row identity in that file.
Identical URLs are deduplicated by exact string; query parameters are neither
reordered nor removed. A redirected file retains its original input URL identity.

Only complete rows enter the dataset. Width mismatches are excluded, with their
record number, line span and expected and actual widths in `OUTPUT`. No padding,
truncation or shifting occurs. Invalid quoting or encoding rejects the entire
file, including any otherwise valid prefix. Healthy files in the batch survive.

### File outcomes and delivery

`OUTPUT` in the default key-value store contains no source rows or offending cell
contents. Each file has a stable ID, original input index, duplicate-input count,
record counts, selected and delivered counts, up to 20 error positions and a count
of omitted errors.

- `COMPLETE`: the observed file was fully parsed and all valid rows selected.
- `EMPTY`: a genuinely empty file, including files containing only empty lines.
- `PARTIAL`: widths, row size, row count or serialized-output limits excluded rows.
- `FAILED`: fetch, robots, content, encoding, quoting or time failure.
- `SKIPPED`: the batch reached a limit before that file was observed.

The batch is `PARTIAL` if any file is partial, failed or unobserved, and `FAILED`
if every attempted file failed. Diagnostics are separate from dataset rows.
Check `billing.deliveryStatus` and the delivered count before using results.
`selectedRows` describes selection, not proven delivery. Invalid input exits 10;
failed runtime or all-file failure exits 11. Complete, empty and explicit partial
results exit 0.

There is at most one dataset append. Before it, the Actor writes delivery intent.
After a confirmed append, it saves a separate delivery receipt and completion
state. If acknowledgement is lost, delivery and charges are reported as unknown,
not zero. If the dataset confirms delivery but charge bookkeeping fails, the
confirmed delivered count remains available while charge count is unknown.

A completed restart replays its saved receipt without fetching or appending.
An incomplete or inconsistent restart refuses to retry. It preserves available
checkpoint facts and labels them as lower bounds when later work is unknown.
Use a new run and storage for a fresh attempt after reviewing an incomplete
receipt. Automatic retries of ambiguous appends are disabled; exactly-once
execution through arbitrary distributed failures is not promised.

### Limits

| Resource | Bound |
| --- | --- |
| Sources | 10 exact URLs, or one inline CSV |
| Column names | 1 to 100 unique names; 64 ASCII characters each |
| Concurrency | 1 |
| Requests | 40 total, including robots requests, redirects and failed attempts |
| Redirects | At most 3 per fetch, including robots fetches |
| Time | 15 seconds per HTTP request; 120-second entrypoint deadline, with time reserved for delivery |
| Encoded response body | 2 MiB per response; 8 MiB across the batch |
| Decoded response body | 2 MiB per response; 8 MiB across the batch |
| Accepted rows | 1,000 maximum, or a lower `maxRows` or affordable-event limit |
| Complete serialized row | 8 KiB, including provenance and names |
| Serialized SDK append | 512 KiB for the entire one-array append |
| Diagnostics | At most 20 issue positions per file, plus omitted-issue count |

The append cap uses the pinned SDK's actual indented item JSON format. It can
stop selection before the row-count limit. No second chunk is appended. Limits
preserve an ordered prefix of qualifying rows and report skipped files or omitted
rows explicitly. Oversize rows are excluded whole; later valid rows may survive.

Byte metrics report bytes actually observed, including the read or decoder output
that revealed overflow. They are not clamped to the configured limits. The pinned
Brotli decoder uses an advisory output-buffer size and can emit a larger block;
that block is counted and rejected on overflow. Valid buffered Brotli output is
fully drained, and truncated streams, trailing corruption and excessive gzip
members fail. No overflowing file is delivered.

HTTP transport accepts `text/csv`, `application/csv`, `text/plain` and
`application/octet-stream`, with UTF-8 or no declared charset. Identity, gzip,
zlib-wrapped deflate and Brotli are supported. HTML, XML, known spreadsheet/binary
signatures and unsupported encodings are rejected. Raw DEFLATE is not inferred.

### Access and input safety

Use only authorized nonsensitive CSV files. Every request and redirect is guarded
against non-public and transition-network addresses. DNS answers are checked, one
checked address is selected, and socket creation checks it again. Robots rules
are observed per origin. Robots 404/410 allows access; other robots failures fail
that file. Environment proxies, cookies, authentication, custom headers, browser
sessions and source discovery are not used. Reserved `.test`, `.invalid` and
`.example` labels cannot be requested, though offline evaluator transports may
use them as fixture identifiers.

Credential-like column names and recognized token/key/password patterns are
rejected with generic outcomes. Signed and credential-bearing URLs are rejected.
Pattern checks cannot identify every confidential value. Use nonsensitive input
rather than treating these checks as a universal secret detector.

Formula-like strings such as `=1+1` remain literal strings and are never executed.
Opening exported data in a spreadsheet can cause that spreadsheet to interpret
formulas. Review or escape such values for the destination application.

No Excel or Sheets integration, formula evaluation, CSV repair, merge-by-key,
remote schema, login, page crawling or link discovery is provided.

### Metering and qualification

When paid pricing is enabled, one confirmed delivered dataset row is the billable unit. Diagnostics,
empty results and rejected rows create no row events. The only supported paid
event is `apify-default-dataset-item`, with a positive flat price supplied by the
platform. The runtime refuses mismatched, unsupported or extra-event pricing.
Check the current platform price before running. Local SDK billing tests use
synthetic prices and do not establish customer charges or owner costs.

`maxChargeUsd` limits row-event spending together with the platform run cap.
Fractional budgets are rounded down to whole affordable rows. Use a supported positive platform run cap. Setting the input maxRows or
maxChargeUsd to zero stops source requests and delivery; a literal zero platform
cap is not a no-spend guarantee. An event cap does not cap compute, storage,
transfer or later retention costs.

The implementation uses original CSV parsing and product code. Public-network
and delivery plumbing adapts reviewed clean portfolio patterns. Python 3.12,
Apify SDK 4.0.0, Apify Client 3.2.0 and all transitive runtime packages are pinned
with hashes in `requirements.txt`. The base Docker image tag is not an immutable
hosted-build receipt.

# Actor input Schema

## `csvUrls` (type: `array`):

One to ten exact HTTP(S) URLs on standard ports, without credentials or signed access parameters. Exact duplicates are fetched once. Clear inlineCsv when using URLs. Excel and Sheets sources are unsupported.

## `inlineCsv` (type: `string`):

Literal headerless UTF-8 text, at most 2 MiB. Remove this property when using csvUrls. This prefill is synthetic sample data, not a public source.

## `fieldNames` (type: `array`):

One to 100 unique nonempty names, at most 64 ASCII letters, digits, spaces, underscores, periods or hyphens, starting with a letter or underscore. Credential-like names are rejected. Names and separator are required for an explicit source.

## `separator` (type: `string`):

Choose comma or semicolon. Delimiter, headers, numbers, dates, nulls and encodings are never inferred.

## `maxRows` (type: `integer`):

Zero to 1000 whole rows across all files. The 512 KiB serialized append cap may stop earlier. Zero performs no source requests.

## `maxChargeUsd` (type: `number`):

Additional cap for dataset row events, combined with the platform run cap. Zero performs no source requests. This does not cap compute, storage, transfers or other platform costs. Check the live listing for the active price.

## Actor input object example

```json
{
  "inlineCsv": "0001,North,\"first record\"\n0002,South,\n",
  "fieldNames": [
    "id",
    "region",
    "note"
  ],
  "separator": ",",
  "maxRows": 100,
  "maxChargeUsd": 0.05
}
```

# Actor output Schema

## `records` (type: `string`):

No description

## `outcomes` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "inlineCsv": `0001,North,"first record"
0002,South,`,
    "fieldNames": [
        "id",
        "region",
        "note"
    ],
    "separator": ","
};

// Run the Actor and wait for it to finish
const run = await client.actor("sapph1re/headerless-csv-batch-import").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "inlineCsv": """0001,North,\"first record\"
0002,South,
""",
    "fieldNames": [
        "id",
        "region",
        "note",
    ],
    "separator": ",",
}

# Run the Actor and wait for it to finish
run = client.actor("sapph1re/headerless-csv-batch-import").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "inlineCsv": "0001,North,\\"first record\\"\\n0002,South,\\n",
  "fieldNames": [
    "id",
    "region",
    "note"
  ],
  "separator": ","
}' |
apify call sapph1re/headerless-csv-batch-import --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,sapph1re/headerless-csv-batch-import"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/vAUOSZTcohakKicOx/builds/5QwqUaxpa8g20jtaG/openapi.json
