# Persistent Dataset Appender for Apify (`quartz_apple_dnx/persistent-dataset-appender`) Actor

Append Apify dataset rows into one persistent named dataset with retry protection and webhook-ready automation.

- **URL**: https://apify.com/quartz\_apple\_dnx/persistent-dataset-appender.md
- **Developed by:** [Austin DeMoss](https://apify.com/quartz_apple_dnx) (community)
- **Categories:** Automation, Developer tools, Integrations
- **Stats:** 1 total users, 0 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.50 / 1,000 dataset row appendeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Persistent Dataset Appender for Apify

Append rows from one Apify dataset into a **persistent named dataset** — manually or automatically after another Actor run succeeds.

Use this when repeated Actor runs should accumulate into one long-lived dataset instead of leaving results split across many run datasets.

### What you get

- **One persistent destination** for results produced across repeated runs
- **Webhook-ready automation** using the source run's `defaultDatasetId`
- **Retry protection** that skips completed source datasets and resumes interrupted copies
- **Limited permissions** — the Actor requests READ access only to the selected source dataset
- **Predictable pay-per-event pricing** based on rows actually appended

### Basic input

```json
{
  "targetDatasetName": "accumulated-results",
  "sourceDatasetId": "YOUR_SOURCE_DATASET_ID"
}
```

In Apify Console, `sourceDatasetId` is a dataset picker. API callers can pass a dataset ID or unique name.

### Automatic webhook / Actor integration

Trigger this Actor after an upstream Actor succeeds and pass the upstream run's dataset into the input:

```json
{
  "targetDatasetName": "accumulated-results",
  "sourceDatasetId": "{{resource.defaultDatasetId}}",
  "sourceActorRunId": "{{resource.id}}"
}
```

`sourceActorRunId` is optional trace metadata. The permission-bearing input is `sourceDatasetId`.

### Pricing

The production pay-per-event event is `dataset-row-appended`.

- **$0.0005 per successfully appended row**
- **$0.50 per 1,000 appended rows**

The Actor charges only after a target batch is successfully written when PPE pricing is active. It checks the run's remaining charge budget before copying so a max-charge limit can stop the run gracefully.

Private cloud validation measured about **$0.00663** of Apify platform usage for a successful 1,000-row append. Actual platform usage can vary by run.

### Retry behavior

`skipPreviouslyAppended` defaults to `true`.

Progress is stored using the resolved target dataset ID and source dataset ID. A normal retry:

- skips a source dataset that already completed, or
- resumes from the last recorded source offset after an interrupted run.

This is retry protection, **not transactional exactly-once delivery**. If the process stops after a target batch is written but before its checkpoint is stored, a retry can duplicate at most that uncheckpointed batch. The current batch size is 50 rows.

### Output

The persistent named target dataset contains the copied source rows.

This Actor's default dataset contains one summary row with:

- status
- target dataset name and ID
- source dataset ID
- optional source Actor run ID
- source row count
- rows appended in this run
- rows already appended before this run
- duplicate-skip status
- budget-limit status

### What this Actor deliberately does not do

It does not transform, join, clean, filter, crawl, call an LLM, or act as a general ETL toolkit.

Its job is intentionally small: **reliably accumulate Apify dataset rows into one persistent dataset.**

# Actor input Schema

## `targetDatasetName` (type: `string`):

Persistent named dataset that receives the rows. It is created if it does not already exist.

## `sourceDatasetId` (type: `string`):

Dataset to append. In the UI, select a dataset. API and webhook integrations may pass its dataset ID or unique name.

## `sourceActorRunId` (type: `string`):

Optional trace metadata identifying the Actor run that produced the source dataset. It is not used to obtain storage permission.

## `skipPreviouslyAppended` (type: `boolean`):

Keeps a persistent progress checkpoint so normal retries resume or skip instead of re-appending completed source rows.

## Actor input object example

```json
{
  "targetDatasetName": "accumulated-results",
  "skipPreviouslyAppended": true
}
```

# Actor output Schema

## `summary` (type: `string`):

Status, source/target identifiers, appended row counts, duplicate guard result, and budget-limit status.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("quartz_apple_dnx/persistent-dataset-appender").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("quartz_apple_dnx/persistent-dataset-appender").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call quartz_apple_dnx/persistent-dataset-appender --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,quartz_apple_dnx/persistent-dataset-appender"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ojddBS7kjaHjwtYAY/builds/hyskpt5VETH9KVTg3/openapi.json
