Headerless CSV Batch Import avatar

Headerless CSV Batch Import

Pricing

$1.00 / 1,000 results

Go to Apify Store
Headerless CSV Batch Import

Headerless CSV Batch Import

Import headerless CSV from pasted text or public URLs. Preserve first records, leading zeros and literal strings, with source and line provenance, explicit file errors and row limits.

Pricing

$1.00 / 1,000 results

Rating

0.0

(0)

Developer

Roman V

Roman V

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

6 days ago

Last modified

Categories

Share

Import a small batch of authorized headerless CSV files using one ordered column list. The first record stays data. Each delivered row contains literal string values, its source URL, its record number and the physical lines it occupied. Separate file receipts explain empty files, excluded widths and failures.

Existing CSV Actors also support batching and custom headers. This utility focuses on literal strings, source positions and separate file failure receipts. No competitor reliability, price advantage or superiority is claimed.

Try the synthetic example

An API request with no body or {} uses two synthetic inline records. The Console prefill and example-input.json use the same data:

{
"inlineCsv": "0001,North,\"first record\"\n0002,South,\n",
"fieldNames": ["id", "region", "note"],
"separator": ",",
"maxRows": 100,
"maxChargeUsd": 0.05
}

For your files, replace inlineCsv with csvUrls, an array of one to ten exact public HTTP(S) CSV URLs, and set your column names and separator explicitly. Remove the inline property completely when selecting URL mode. URL mode does not supply a fabricated public demonstration file. Use URLs whose content and access rights you have checked.

What the data means

Each dataset row has these fields:

FieldMeaning
valuesObject mapping your column names to strings, including empty strings
sourceUrlExact input URL, or null for inline input
fileIndexZero-based original input position; exact duplicate URLs use their first position
recordIndexOne-based logical CSV record, including excluded width records but skipping empty physical lines
physicalStartLine, physicalEndLineOne-based inclusive physical lines, with CRLF counted as one line break
fileIdSHA-256 identity of the exact source URL or the inline-source marker
rowIdSHA-256 identity of source, exact source bytes, column list, delimiter and logical position

Values are never inferred as numbers, dates, booleans or nulls. Leading zeros, spaces, empty fields, quoted separators, quoted line breaks and doubled quotes are preserved. A leading UTF-8 BOM is removed from parsing. UTF-8 is the only supported encoding; invalid bytes fail the file. An entirely empty physical line is skipped and counted. A delimiter-only row or quoted empty string is a record. A whitespace-only line is also a record and must have the correct width.

Repeated values at different positions have different row identities. Repeating an unchanged input and source yields the same identities. Changing the source bytes, including line endings, can change every row identity in that file. Identical URLs are deduplicated by exact string; query parameters are neither reordered nor removed. A redirected file retains its original input URL identity.

Only complete rows enter the dataset. Width mismatches are excluded, with their record number, line span and expected and actual widths in OUTPUT. No padding, truncation or shifting occurs. Invalid quoting or encoding rejects the entire file, including any otherwise valid prefix. Healthy files in the batch survive.

File outcomes and delivery

OUTPUT in the default key-value store contains no source rows or offending cell contents. Each file has a stable ID, original input index, duplicate-input count, record counts, selected and delivered counts, up to 20 error positions and a count of omitted errors.

  • COMPLETE: the observed file was fully parsed and all valid rows selected.
  • EMPTY: a genuinely empty file, including files containing only empty lines.
  • PARTIAL: widths, row size, row count or serialized-output limits excluded rows.
  • FAILED: fetch, robots, content, encoding, quoting or time failure.
  • SKIPPED: the batch reached a limit before that file was observed.

The batch is PARTIAL if any file is partial, failed or unobserved, and FAILED if every attempted file failed. Diagnostics are separate from dataset rows. Check billing.deliveryStatus and the delivered count before using results. selectedRows describes selection, not proven delivery. Invalid input exits 10; failed runtime or all-file failure exits 11. Complete, empty and explicit partial results exit 0.

There is at most one dataset append. Before it, the Actor writes delivery intent. After a confirmed append, it saves a separate delivery receipt and completion state. If acknowledgement is lost, delivery and charges are reported as unknown, not zero. If the dataset confirms delivery but charge bookkeeping fails, the confirmed delivered count remains available while charge count is unknown.

A completed restart replays its saved receipt without fetching or appending. An incomplete or inconsistent restart refuses to retry. It preserves available checkpoint facts and labels them as lower bounds when later work is unknown. Use a new run and storage for a fresh attempt after reviewing an incomplete receipt. Automatic retries of ambiguous appends are disabled; exactly-once execution through arbitrary distributed failures is not promised.

Limits

ResourceBound
Sources10 exact URLs, or one inline CSV
Column names1 to 100 unique names; 64 ASCII characters each
Concurrency1
Requests40 total, including robots requests, redirects and failed attempts
RedirectsAt most 3 per fetch, including robots fetches
Time15 seconds per HTTP request; 120-second entrypoint deadline, with time reserved for delivery
Encoded response body2 MiB per response; 8 MiB across the batch
Decoded response body2 MiB per response; 8 MiB across the batch
Accepted rows1,000 maximum, or a lower maxRows or affordable-event limit
Complete serialized row8 KiB, including provenance and names
Serialized SDK append512 KiB for the entire one-array append
DiagnosticsAt most 20 issue positions per file, plus omitted-issue count

The append cap uses the pinned SDK's actual indented item JSON format. It can stop selection before the row-count limit. No second chunk is appended. Limits preserve an ordered prefix of qualifying rows and report skipped files or omitted rows explicitly. Oversize rows are excluded whole; later valid rows may survive.

Byte metrics report bytes actually observed, including the read or decoder output that revealed overflow. They are not clamped to the configured limits. The pinned Brotli decoder uses an advisory output-buffer size and can emit a larger block; that block is counted and rejected on overflow. Valid buffered Brotli output is fully drained, and truncated streams, trailing corruption and excessive gzip members fail. No overflowing file is delivered.

HTTP transport accepts text/csv, application/csv, text/plain and application/octet-stream, with UTF-8 or no declared charset. Identity, gzip, zlib-wrapped deflate and Brotli are supported. HTML, XML, known spreadsheet/binary signatures and unsupported encodings are rejected. Raw DEFLATE is not inferred.

Access and input safety

Use only authorized nonsensitive CSV files. Every request and redirect is guarded against non-public and transition-network addresses. DNS answers are checked, one checked address is selected, and socket creation checks it again. Robots rules are observed per origin. Robots 404/410 allows access; other robots failures fail that file. Environment proxies, cookies, authentication, custom headers, browser sessions and source discovery are not used. Reserved .test, .invalid and .example labels cannot be requested, though offline evaluator transports may use them as fixture identifiers.

Credential-like column names and recognized token/key/password patterns are rejected with generic outcomes. Signed and credential-bearing URLs are rejected. Pattern checks cannot identify every confidential value. Use nonsensitive input rather than treating these checks as a universal secret detector.

Formula-like strings such as =1+1 remain literal strings and are never executed. Opening exported data in a spreadsheet can cause that spreadsheet to interpret formulas. Review or escape such values for the destination application.

No Excel or Sheets integration, formula evaluation, CSV repair, merge-by-key, remote schema, login, page crawling or link discovery is provided.

Metering and qualification

When paid pricing is enabled, one confirmed delivered dataset row is the billable unit. Diagnostics, empty results and rejected rows create no row events. The only supported paid event is apify-default-dataset-item, with a positive flat price supplied by the platform. The runtime refuses mismatched, unsupported or extra-event pricing. Check the current platform price before running. Local SDK billing tests use synthetic prices and do not establish customer charges or owner costs.

maxChargeUsd limits row-event spending together with the platform run cap. Fractional budgets are rounded down to whole affordable rows. Use a supported positive platform run cap. Setting the input maxRows or maxChargeUsd to zero stops source requests and delivery; a literal zero platform cap is not a no-spend guarantee. An event cap does not cap compute, storage, transfer or later retention costs.

The implementation uses original CSV parsing and product code. Public-network and delivery plumbing adapts reviewed clean portfolio patterns. Python 3.12, Apify SDK 4.0.0, Apify Client 3.2.0 and all transitive runtime packages are pinned with hashes in requirements.txt. The base Docker image tag is not an immutable hosted-build receipt.