# Docs to Pinecone Sync (`human_intelligence/docs-to-pinecone-sync`) Actor

Keep public HTML documentation in sync with Pinecone. Detect changed pages, embed updated content using your OpenAI key, and update a dedicated Pinecone namespace. Unchanged pages are skipped.

- **URL**: https://apify.com/human\_intelligence/docs-to-pinecone-sync.md
- **Developed by:** [human intelligence](https://apify.com/human_intelligence) (community)
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.50 / 1,000 page checkeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Docs to Pinecone Sync

Keep a small public documentation site synchronized with a dedicated Pinecone namespace. One Actor fetches HTML, extracts the main content, checks for page changes, and updates the corresponding vectors. Unchanged pages do not generate new embeddings.

**Version 1.0.0-rc1 — prepared for private Apify acceptance testing.** Local tests and an actual Apify SDK persistence check have been run. No paid cloud run, live OpenAI embedding or live Pinecone write has been completed in the delivery environment. Do not describe this candidate as production-validated or claim a measured saving. See `TEST_REPORT.md` and `RELEASE_CHECKLIST.md`.

### What it does

- Reads an exact URL list or one simple XML `urlset` sitemap on one public host.
- Extracts HTML content using an optional CSS selector, otherwise `main`, `[role=main]`, `article`, then `body`.
- Tracks up to 100 active sources. Removed sitemap entries remain monitored; omission never deletes vectors.
- Reprocesses an entire changed page using OpenAI `text-embedding-3-small` at 1536 dimensions.
- Updates title metadata without re-embedding unchanged text.
- Stores a durable pending operation so interrupted writes can resume.
- Uses deterministic vector IDs, upserts new revisions first, checks Pinecone log sequence numbers, then removes replaced IDs.
- Deletes an imported page only after consecutive confirmed 404/410 observations; large deletion waves require explicit plan approval.
- Provides one page-results dataset, `SUMMARY` JSON and a readable `REPORT`.

It is not a chatbot. It does not support JavaScript-only pages, authenticated content, PDF/OCR, sitemap indexes, query-string URLs, arbitrary database providers or existing foreign vectors. At most one writer may manage a monitor/namespace. Pinecone updates are eventually consistent; old and new revisions can briefly coexist during replacement.

### Start with the offline demo

Set input to:

```json
{"mode":"demo"}
```

The demo performs an import, an unchanged check, a change, and two missing-page observations with in-memory fixtures. Its vectors are fake and it makes no external website, OpenAI or Pinecone calls. It produces no custom `page-checked` events. Apify platform/start costs may still apply.

### Preview your source

```json
{
  "mode": "preview",
  "monitorName": "my-docs",
  "urls": ["https://YOUR-PUBLIC-DOCS-HOST/docs/getting-started"],
  "maxPages": 10
}
```

Replace the placeholder with your actual public documentation URL. Use the final URL after any cross-domain redirect. For a sitemap, replace `urls` with `sitemapUrl`; optionally narrow `pathPrefix` to `/docs/`. Do not supply both source types.

Preview reads sources and, if available, saved monitor state. It never calls the embedding API, writes to Pinecone, advances the applied baseline or confirms deletions. Its completed checks can incur Apify/event charges when monetized. A preview without previous state describes a first import, not an inspection of arbitrary existing Pinecone data.

### Synchronize

You need:

1. An existing **serverless dense-vector Pinecone index, dimension 1536**, configured for external vectors, not integrated embedding. Cosine similarity is recommended for the initial index. Index creation and its costs stay under your control.
2. A new empty named namespace dedicated to this Actor/monitor. The Actor validates emptiness and claims that target; it will not take over a shared namespace.
3. Your own Pinecone API key with the necessary index read/write operations and an OpenAI API key permitted to create embeddings. Their provider bills are separate from Apify.

Use the same source settings and monitor name as your preview, select `sync`, then fill the encrypted key fields in Apify's input form. Example shape:

```json
{
  "mode": "sync",
  "monitorName": "my-docs",
  "urls": ["https://YOUR-PUBLIC-DOCS-HOST/docs/getting-started"],
  "pineconeIndex": "your-index-name",
  "pineconeNamespace": "docs-sync-first-test",
  "pineconeApiKey": "ENTER-IN-SECRET-FIELD",
  "openaiApiKey": "ENTER-IN-SECRET-FIELD",
  "maxPages": 10,
  "maxEmbeddingTokens": 10000
}
```

Never commit real keys to example files or source code. The Actor does not create indexes or automatically switch models. Reuse the same Actor, monitor, source scope, selector, index and namespace. Changing a bound setting requires a new monitor and an unused namespace. API keys may be rotated without changing the binding.

Once the acceptance checks pass, save the input as an Apify task and schedule that same task at an interval appropriate to your source. Avoid overlapping schedules and manual runs. A cloud Request Queue lock serializes each monitor and target; local development uses a file lock. After an abrupt termination a cloud lock can remain for up to 15 minutes.

### Outputs and recovery

Page results include `url`, `status`, `plannedAction`, `appliedAction`, `error`, and `checkedAt`. `SUMMARY.status` is the authoritative business result:

- `COMPLETE`: every selected source checked and all required actions applied.
- `PREVIEW`: a read-only result; inspect completeness and per-page errors.
- `PARTIAL`: some checks or writes remain. Retry with the same settings after correcting the reported issue.
- `BLOCKED`: configuration, ownership, budget or preflight prevented normal completion.
- `FAILED`: no successful usable scan or an unexpected failure.

Partial/blocked/failed sync runs exit nonzero. A completed process alone is not proof of a complete index. The `pendingAction` field identifies a saved operation. The next sync resumes it before starting a fresh comparison. Do not delete named state storage or coordination queues to make an error disappear.

No rollback/history interface is included. Only the current applied state, up to 100 recent compact deletion records and an unfinished operation are retained. Successful temporary embedding vectors are removed with the committed operation. If an embedding response was lost before it could be saved, repeating that request can cost tokens again. Exact-once external billing is not promised.

### Deletion rules

Timeouts, rate limits, 5xx, robots restrictions, missing links, incomplete runs and content extraction errors are never interpreted as deletion. A 304 is positive evidence the stored page still exists. A full fetch is forced after seven days or 20 cache confirmations.

Imported pages require two valid 404/410 observations in consecutive sync cycles. A failed/missed cycle breaks that sequence; previews do not advance it. A wave exceeding both five pages and 20% of the imported pages is held back. Review the affected URLs and copy the exact `SUMMARY.deletionPlan` into `approveDeletionPlan` to approve that specific wave. A fresh scan rechecks the candidates; the approval does not disable safeguards globally.

Removing a still-live page from the input is not a deletion command. Automatic intentional removal of a live page is outside this version. No full-namespace or whole-index delete endpoint is used.

### Bounds and costs

Hard ceilings: 100 active sources, 400 source HTTP requests, 50 MiB decoded source bytes, 2 MiB per response, 20,000 extracted tokens per page, 500 Pinecone requests and 10 minutes. Defaults: 100,000 embedding tokens per run (maximum 250,000), one page at a time with at least 0.5 seconds between source requests, respecting longer robots intervals. API retries also consume budgets. Limits stop new work and keep incomplete operations recoverable.

**Monetization hook:** `page-checked`, one custom event per successfully evaluated URL and run. It is intended to include fetch, cleaning, comparison and the bounded sync work. A 304 or definite 404/410 is a completed check. Technical fetch failures and demo rows are not custom paid results. Interrupted sync application is conservatively left uncharged by this candidate; its resource costs still belong in the publisher's margin calculation. Pure pending-operation replay is not charged again. Preview completed checks use the same event definition.

Do not enable `apify-default-dataset-item`: it would charge again for diagnostic output. The Actor blocks that configuration. A documented platform start event may be used. The SDK's remaining charge allowance reduces the page budget before scanning; state is saved before custom charges. Ambiguous charges are not blindly repeated after a restart.

This candidate has **no hard-coded sale price and does not configure a paid plan**. Run private acceptance tests without monetization first. Then measure the complete normal/error workload and set a price that covers the underlying Apify costs. OpenAI and Pinecone are always paid through the customer's own accounts. Ongoing named-state storage and retained outputs should be included in the full cost estimate. Manage run-output retention in Apify; the Actor never deletes its needed baseline on an age timer.

### Common actionable errors

| Error | Action |
|---|---|
| `ONE_SOURCE_REQUIRED` | Supply URL list or sitemap, not both. |
| `REDIRECT_HOST_BLOCKED` | Use the actual final host in your source input. |
| `ROBOTS_DENIED` / `ROBOTS_UNAVAILABLE` | Use permitted sources or wait for robots availability; no bypass is performed. |
| `SITEMAP_INDEX_UNSUPPORTED` | Use one actual urlset file or exact URLs. |
| `TOO_MANY_SOURCES` | Narrow the sitemap path or start a smaller monitor. |
| `SELECTOR_NOT_FOUND` | Correct the selector; a new bound selector requires a new monitor. |
| `SUSPICIOUS_CONTENT_LOSS` | Inspect the source; do not bypass the guard blindly. A deliberate large restructuring may need a fresh monitor/namespace. |
| `INCOMPATIBLE_INDEX` | Use a serverless dense 1536-dimensional index for external vectors. |
| `TARGET_NOT_EMPTY` / `TARGET_ALREADY_OWNED` | Select a genuinely unused dedicated namespace. |
| `STATE_MISSING` / `TARGET_CLAIM_MISSING` | Restore the matching state/claim or use a new empty namespace; never overwrite existing unknown vectors. |
| `VISIBILITY_PENDING` | Retry the same monitor later; old vectors are retained until confirmed replacement. |
| `CONCURRENT_RUN` | Wait for the active run/lock expiry. |
| `CHARGE_LIMIT` / `EMBEDDING_TOKEN_LIMIT` | Check budget and pending work before retrying. |
| `API_HTTP_401` / `API_HTTP_403` | Check the relevant provider key and permissions. |
| `MISSING_WRITE_LSN` | Verify the deployed Pinecone API contract before release; no unverified delete follows. |

### Local development

Python 3.12 on Linux/macOS:

```bash
python3 -m venv .venv
.venv/bin/pip install -r requirements-dev.txt
.venv/bin/python -m pytest -q
```

For a local Actor demo, put `examples/demo.json` at `storage/key_value_stores/default/INPUT.json`, then run `.venv/bin/python -m src`. The Docker image preloads the tokenizer data at build time. Outside Docker its first load may download that public tokenizer asset. No API key is needed for the demo.

`tools/build_web_ide.py` regenerates a self-contained launcher from the exact same source modules. `START_HIER.html` contains the German browser-only deployment instructions. No Apify CLI is required by that path.

### Privacy and permissions

Keep Apify permissions at **Limited**. The Actor uses its own named storage, a coordination queue and default output storage. It does not read your unrelated Apify datasets. Keys use secret input fields and are excluded from state and generated reports. Website content sent for embedding goes to OpenAI; text and vectors go to your Pinecone namespace. You must have permission to crawl and process your chosen source.

# Actor input Schema

## `mode` (type: `string`):

Preview checks sources without embedding or external writes. Sync uses your keys. Demo is offline.

## `monitorName` (type: `string`):

Reuse this name and Actor for subsequent runs; source settings and target are bound to it.

## `urls` (type: `array`):

Exact URLs on one host without query parameters. Provide this OR sitemapUrl.

## `sitemapUrl` (type: `string`):

A single urlset sitemap on the same host as the pages. Sitemap indexes are unsupported.

## `pathPrefix` (type: `string`):

For example /docs/. At most 100 active sources after filtering.

## `contentSelector` (type: `string`):

Optional. Otherwise main, role=main, article, then body. Missing selectors are errors.

## `pineconeIndex` (type: `string`):

Sync: existing serverless dense-vector index, dimension 1536, no integrated embedding.

## `pineconeNamespace` (type: `string`):

Sync: an empty namespace exclusively managed by this monitor. Do not use a shared or default namespace.

## `pineconeApiKey` (type: `string`):

Your own key, sent only to Pinecone and the verified index host.

## `openaiApiKey` (type: `string`):

Your own API key for text-embedding-3-small. External usage is billed by OpenAI.

## `maxPages` (type: `integer`):

Unchecked sources stay intact and are prioritized on later runs.

## `maxEmbeddingTokens` (type: `integer`):

Includes retries initiated by this Actor, not other use of your account.

## `maxRunSeconds` (type: `integer`):

Stops new work and preserves unfinished operations.

## `requestDelaySeconds` (type: `number`):

Longer robots.txt limits take precedence.

## `approveDeletionPlan` (type: `string`):

Normally empty. Paste SUMMARY deletionPlan only after reviewing the listed URLs.

## Actor input object example

```json
{
  "mode": "preview",
  "monitorName": "my-docs",
  "pathPrefix": "/",
  "maxPages": 100,
  "maxEmbeddingTokens": 100000,
  "maxRunSeconds": 600,
  "requestDelaySeconds": 0.5
}
```

# Actor output Schema

## `pages` (type: `string`):

No description

## `summary` (type: `string`):

No description

## `report` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("human_intelligence/docs-to-pinecone-sync").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("human_intelligence/docs-to-pinecone-sync").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call human_intelligence/docs-to-pinecone-sync --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,human_intelligence/docs-to-pinecone-sync"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/CLZYvlVthZoK409Ob/builds/pRLrzhh31tgRQN1P0/openapi.json
