Headerless CSV Batch Import
Pricing
$1.00 / 1,000 results
Headerless CSV Batch Import
Import headerless CSV from pasted text or public URLs. Preserve first records, leading zeros and literal strings, with source and line provenance, explicit file errors and row limits.
Pricing
$1.00 / 1,000 results
Rating
0.0
(0)
Developer
Roman V
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
6 days ago
Last modified
Categories
Share
Import a small batch of authorized headerless CSV files using one ordered column list. The first record stays data. Each delivered row contains literal string values, its source URL, its record number and the physical lines it occupied. Separate file receipts explain empty files, excluded widths and failures.
Existing CSV Actors also support batching and custom headers. This utility focuses on literal strings, source positions and separate file failure receipts. No competitor reliability, price advantage or superiority is claimed.
Try the synthetic example
An API request with no body or {} uses two synthetic inline records. The Console
prefill and example-input.json use the same data:
{"inlineCsv": "0001,North,\"first record\"\n0002,South,\n","fieldNames": ["id", "region", "note"],"separator": ",","maxRows": 100,"maxChargeUsd": 0.05}
For your files, replace inlineCsv with csvUrls, an array of one to ten exact
public HTTP(S) CSV URLs, and set your column names and separator explicitly.
Remove the inline property completely when selecting URL mode. URL mode does not
supply a fabricated public demonstration file. Use URLs whose content and access rights you have checked.
What the data means
Each dataset row has these fields:
| Field | Meaning |
|---|---|
values | Object mapping your column names to strings, including empty strings |
sourceUrl | Exact input URL, or null for inline input |
fileIndex | Zero-based original input position; exact duplicate URLs use their first position |
recordIndex | One-based logical CSV record, including excluded width records but skipping empty physical lines |
physicalStartLine, physicalEndLine | One-based inclusive physical lines, with CRLF counted as one line break |
fileId | SHA-256 identity of the exact source URL or the inline-source marker |
rowId | SHA-256 identity of source, exact source bytes, column list, delimiter and logical position |
Values are never inferred as numbers, dates, booleans or nulls. Leading zeros, spaces, empty fields, quoted separators, quoted line breaks and doubled quotes are preserved. A leading UTF-8 BOM is removed from parsing. UTF-8 is the only supported encoding; invalid bytes fail the file. An entirely empty physical line is skipped and counted. A delimiter-only row or quoted empty string is a record. A whitespace-only line is also a record and must have the correct width.
Repeated values at different positions have different row identities. Repeating an unchanged input and source yields the same identities. Changing the source bytes, including line endings, can change every row identity in that file. Identical URLs are deduplicated by exact string; query parameters are neither reordered nor removed. A redirected file retains its original input URL identity.
Only complete rows enter the dataset. Width mismatches are excluded, with their
record number, line span and expected and actual widths in OUTPUT. No padding,
truncation or shifting occurs. Invalid quoting or encoding rejects the entire
file, including any otherwise valid prefix. Healthy files in the batch survive.
File outcomes and delivery
OUTPUT in the default key-value store contains no source rows or offending cell
contents. Each file has a stable ID, original input index, duplicate-input count,
record counts, selected and delivered counts, up to 20 error positions and a count
of omitted errors.
COMPLETE: the observed file was fully parsed and all valid rows selected.EMPTY: a genuinely empty file, including files containing only empty lines.PARTIAL: widths, row size, row count or serialized-output limits excluded rows.FAILED: fetch, robots, content, encoding, quoting or time failure.SKIPPED: the batch reached a limit before that file was observed.
The batch is PARTIAL if any file is partial, failed or unobserved, and FAILED
if every attempted file failed. Diagnostics are separate from dataset rows.
Check billing.deliveryStatus and the delivered count before using results.
selectedRows describes selection, not proven delivery. Invalid input exits 10;
failed runtime or all-file failure exits 11. Complete, empty and explicit partial
results exit 0.
There is at most one dataset append. Before it, the Actor writes delivery intent. After a confirmed append, it saves a separate delivery receipt and completion state. If acknowledgement is lost, delivery and charges are reported as unknown, not zero. If the dataset confirms delivery but charge bookkeeping fails, the confirmed delivered count remains available while charge count is unknown.
A completed restart replays its saved receipt without fetching or appending. An incomplete or inconsistent restart refuses to retry. It preserves available checkpoint facts and labels them as lower bounds when later work is unknown. Use a new run and storage for a fresh attempt after reviewing an incomplete receipt. Automatic retries of ambiguous appends are disabled; exactly-once execution through arbitrary distributed failures is not promised.
Limits
| Resource | Bound |
|---|---|
| Sources | 10 exact URLs, or one inline CSV |
| Column names | 1 to 100 unique names; 64 ASCII characters each |
| Concurrency | 1 |
| Requests | 40 total, including robots requests, redirects and failed attempts |
| Redirects | At most 3 per fetch, including robots fetches |
| Time | 15 seconds per HTTP request; 120-second entrypoint deadline, with time reserved for delivery |
| Encoded response body | 2 MiB per response; 8 MiB across the batch |
| Decoded response body | 2 MiB per response; 8 MiB across the batch |
| Accepted rows | 1,000 maximum, or a lower maxRows or affordable-event limit |
| Complete serialized row | 8 KiB, including provenance and names |
| Serialized SDK append | 512 KiB for the entire one-array append |
| Diagnostics | At most 20 issue positions per file, plus omitted-issue count |
The append cap uses the pinned SDK's actual indented item JSON format. It can stop selection before the row-count limit. No second chunk is appended. Limits preserve an ordered prefix of qualifying rows and report skipped files or omitted rows explicitly. Oversize rows are excluded whole; later valid rows may survive.
Byte metrics report bytes actually observed, including the read or decoder output that revealed overflow. They are not clamped to the configured limits. The pinned Brotli decoder uses an advisory output-buffer size and can emit a larger block; that block is counted and rejected on overflow. Valid buffered Brotli output is fully drained, and truncated streams, trailing corruption and excessive gzip members fail. No overflowing file is delivered.
HTTP transport accepts text/csv, application/csv, text/plain and
application/octet-stream, with UTF-8 or no declared charset. Identity, gzip,
zlib-wrapped deflate and Brotli are supported. HTML, XML, known spreadsheet/binary
signatures and unsupported encodings are rejected. Raw DEFLATE is not inferred.
Access and input safety
Use only authorized nonsensitive CSV files. Every request and redirect is guarded
against non-public and transition-network addresses. DNS answers are checked, one
checked address is selected, and socket creation checks it again. Robots rules
are observed per origin. Robots 404/410 allows access; other robots failures fail
that file. Environment proxies, cookies, authentication, custom headers, browser
sessions and source discovery are not used. Reserved .test, .invalid and
.example labels cannot be requested, though offline evaluator transports may
use them as fixture identifiers.
Credential-like column names and recognized token/key/password patterns are rejected with generic outcomes. Signed and credential-bearing URLs are rejected. Pattern checks cannot identify every confidential value. Use nonsensitive input rather than treating these checks as a universal secret detector.
Formula-like strings such as =1+1 remain literal strings and are never executed.
Opening exported data in a spreadsheet can cause that spreadsheet to interpret
formulas. Review or escape such values for the destination application.
No Excel or Sheets integration, formula evaluation, CSV repair, merge-by-key, remote schema, login, page crawling or link discovery is provided.
Metering and qualification
When paid pricing is enabled, one confirmed delivered dataset row is the billable unit. Diagnostics,
empty results and rejected rows create no row events. The only supported paid
event is apify-default-dataset-item, with a positive flat price supplied by the
platform. The runtime refuses mismatched, unsupported or extra-event pricing.
Check the current platform price before running. Local SDK billing tests use
synthetic prices and do not establish customer charges or owner costs.
maxChargeUsd limits row-event spending together with the platform run cap.
Fractional budgets are rounded down to whole affordable rows. Use a supported positive platform run cap. Setting the input maxRows or
maxChargeUsd to zero stops source requests and delivery; a literal zero platform
cap is not a no-spend guarantee. An event cap does not cap compute, storage,
transfer or later retention costs.
The implementation uses original CSV parsing and product code. Public-network
and delivery plumbing adapts reviewed clean portfolio patterns. Python 3.12,
Apify SDK 4.0.0, Apify Client 3.2.0 and all transitive runtime packages are pinned
with hashes in requirements.txt. The base Docker image tag is not an immutable
hosted-build receipt.