Docs to Pinecone Sync
Under maintenancePricing
from $4.50 / 1,000 page checkeds
Docs to Pinecone Sync
Under maintenanceKeep public HTML documentation in sync with Pinecone. Detect changed pages, embed updated content using your OpenAI key, and update a dedicated Pinecone namespace. Unchanged pages are skipped.
Pricing
from $4.50 / 1,000 page checkeds
Rating
0.0
(0)
Developer
human intelligence
Maintained by CommunityActor stats
1
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
Keep a small public documentation site synchronized with a dedicated Pinecone namespace. One Actor fetches HTML, extracts the main content, checks for page changes, and updates the corresponding vectors. Unchanged pages do not generate new embeddings.
Version 1.0.0-rc1 — prepared for private Apify acceptance testing. Local tests and an actual Apify SDK persistence check have been run. No paid cloud run, live OpenAI embedding or live Pinecone write has been completed in the delivery environment. Do not describe this candidate as production-validated or claim a measured saving. See TEST_REPORT.md and RELEASE_CHECKLIST.md.
What it does
- Reads an exact URL list or one simple XML
urlsetsitemap on one public host. - Extracts HTML content using an optional CSS selector, otherwise
main,[role=main],article, thenbody. - Tracks up to 100 active sources. Removed sitemap entries remain monitored; omission never deletes vectors.
- Reprocesses an entire changed page using OpenAI
text-embedding-3-smallat 1536 dimensions. - Updates title metadata without re-embedding unchanged text.
- Stores a durable pending operation so interrupted writes can resume.
- Uses deterministic vector IDs, upserts new revisions first, checks Pinecone log sequence numbers, then removes replaced IDs.
- Deletes an imported page only after consecutive confirmed 404/410 observations; large deletion waves require explicit plan approval.
- Provides one page-results dataset,
SUMMARYJSON and a readableREPORT.
It is not a chatbot. It does not support JavaScript-only pages, authenticated content, PDF/OCR, sitemap indexes, query-string URLs, arbitrary database providers or existing foreign vectors. At most one writer may manage a monitor/namespace. Pinecone updates are eventually consistent; old and new revisions can briefly coexist during replacement.
Start with the offline demo
Set input to:
{"mode":"demo"}
The demo performs an import, an unchanged check, a change, and two missing-page observations with in-memory fixtures. Its vectors are fake and it makes no external website, OpenAI or Pinecone calls. It produces no custom page-checked events. Apify platform/start costs may still apply.
Preview your source
{"mode": "preview","monitorName": "my-docs","urls": ["https://YOUR-PUBLIC-DOCS-HOST/docs/getting-started"],"maxPages": 10}
Replace the placeholder with your actual public documentation URL. Use the final URL after any cross-domain redirect. For a sitemap, replace urls with sitemapUrl; optionally narrow pathPrefix to /docs/. Do not supply both source types.
Preview reads sources and, if available, saved monitor state. It never calls the embedding API, writes to Pinecone, advances the applied baseline or confirms deletions. Its completed checks can incur Apify/event charges when monetized. A preview without previous state describes a first import, not an inspection of arbitrary existing Pinecone data.
Synchronize
You need:
- An existing serverless dense-vector Pinecone index, dimension 1536, configured for external vectors, not integrated embedding. Cosine similarity is recommended for the initial index. Index creation and its costs stay under your control.
- A new empty named namespace dedicated to this Actor/monitor. The Actor validates emptiness and claims that target; it will not take over a shared namespace.
- Your own Pinecone API key with the necessary index read/write operations and an OpenAI API key permitted to create embeddings. Their provider bills are separate from Apify.
Use the same source settings and monitor name as your preview, select sync, then fill the encrypted key fields in Apify's input form. Example shape:
{"mode": "sync","monitorName": "my-docs","urls": ["https://YOUR-PUBLIC-DOCS-HOST/docs/getting-started"],"pineconeIndex": "your-index-name","pineconeNamespace": "docs-sync-first-test","pineconeApiKey": "ENTER-IN-SECRET-FIELD","openaiApiKey": "ENTER-IN-SECRET-FIELD","maxPages": 10,"maxEmbeddingTokens": 10000}
Never commit real keys to example files or source code. The Actor does not create indexes or automatically switch models. Reuse the same Actor, monitor, source scope, selector, index and namespace. Changing a bound setting requires a new monitor and an unused namespace. API keys may be rotated without changing the binding.
Once the acceptance checks pass, save the input as an Apify task and schedule that same task at an interval appropriate to your source. Avoid overlapping schedules and manual runs. A cloud Request Queue lock serializes each monitor and target; local development uses a file lock. After an abrupt termination a cloud lock can remain for up to 15 minutes.
Outputs and recovery
Page results include url, status, plannedAction, appliedAction, error, and checkedAt. SUMMARY.status is the authoritative business result:
COMPLETE: every selected source checked and all required actions applied.PREVIEW: a read-only result; inspect completeness and per-page errors.PARTIAL: some checks or writes remain. Retry with the same settings after correcting the reported issue.BLOCKED: configuration, ownership, budget or preflight prevented normal completion.FAILED: no successful usable scan or an unexpected failure.
Partial/blocked/failed sync runs exit nonzero. A completed process alone is not proof of a complete index. The pendingAction field identifies a saved operation. The next sync resumes it before starting a fresh comparison. Do not delete named state storage or coordination queues to make an error disappear.
No rollback/history interface is included. Only the current applied state, up to 100 recent compact deletion records and an unfinished operation are retained. Successful temporary embedding vectors are removed with the committed operation. If an embedding response was lost before it could be saved, repeating that request can cost tokens again. Exact-once external billing is not promised.
Deletion rules
Timeouts, rate limits, 5xx, robots restrictions, missing links, incomplete runs and content extraction errors are never interpreted as deletion. A 304 is positive evidence the stored page still exists. A full fetch is forced after seven days or 20 cache confirmations.
Imported pages require two valid 404/410 observations in consecutive sync cycles. A failed/missed cycle breaks that sequence; previews do not advance it. A wave exceeding both five pages and 20% of the imported pages is held back. Review the affected URLs and copy the exact SUMMARY.deletionPlan into approveDeletionPlan to approve that specific wave. A fresh scan rechecks the candidates; the approval does not disable safeguards globally.
Removing a still-live page from the input is not a deletion command. Automatic intentional removal of a live page is outside this version. No full-namespace or whole-index delete endpoint is used.
Bounds and costs
Hard ceilings: 100 active sources, 400 source HTTP requests, 50 MiB decoded source bytes, 2 MiB per response, 20,000 extracted tokens per page, 500 Pinecone requests and 10 minutes. Defaults: 100,000 embedding tokens per run (maximum 250,000), one page at a time with at least 0.5 seconds between source requests, respecting longer robots intervals. API retries also consume budgets. Limits stop new work and keep incomplete operations recoverable.
Monetization hook: page-checked, one custom event per successfully evaluated URL and run. It is intended to include fetch, cleaning, comparison and the bounded sync work. A 304 or definite 404/410 is a completed check. Technical fetch failures and demo rows are not custom paid results. Interrupted sync application is conservatively left uncharged by this candidate; its resource costs still belong in the publisher's margin calculation. Pure pending-operation replay is not charged again. Preview completed checks use the same event definition.
Do not enable apify-default-dataset-item: it would charge again for diagnostic output. The Actor blocks that configuration. A documented platform start event may be used. The SDK's remaining charge allowance reduces the page budget before scanning; state is saved before custom charges. Ambiguous charges are not blindly repeated after a restart.
This candidate has no hard-coded sale price and does not configure a paid plan. Run private acceptance tests without monetization first. Then measure the complete normal/error workload and set a price that covers the underlying Apify costs. OpenAI and Pinecone are always paid through the customer's own accounts. Ongoing named-state storage and retained outputs should be included in the full cost estimate. Manage run-output retention in Apify; the Actor never deletes its needed baseline on an age timer.
Common actionable errors
| Error | Action |
|---|---|
ONE_SOURCE_REQUIRED | Supply URL list or sitemap, not both. |
REDIRECT_HOST_BLOCKED | Use the actual final host in your source input. |
ROBOTS_DENIED / ROBOTS_UNAVAILABLE | Use permitted sources or wait for robots availability; no bypass is performed. |
SITEMAP_INDEX_UNSUPPORTED | Use one actual urlset file or exact URLs. |
TOO_MANY_SOURCES | Narrow the sitemap path or start a smaller monitor. |
SELECTOR_NOT_FOUND | Correct the selector; a new bound selector requires a new monitor. |
SUSPICIOUS_CONTENT_LOSS | Inspect the source; do not bypass the guard blindly. A deliberate large restructuring may need a fresh monitor/namespace. |
INCOMPATIBLE_INDEX | Use a serverless dense 1536-dimensional index for external vectors. |
TARGET_NOT_EMPTY / TARGET_ALREADY_OWNED | Select a genuinely unused dedicated namespace. |
STATE_MISSING / TARGET_CLAIM_MISSING | Restore the matching state/claim or use a new empty namespace; never overwrite existing unknown vectors. |
VISIBILITY_PENDING | Retry the same monitor later; old vectors are retained until confirmed replacement. |
CONCURRENT_RUN | Wait for the active run/lock expiry. |
CHARGE_LIMIT / EMBEDDING_TOKEN_LIMIT | Check budget and pending work before retrying. |
API_HTTP_401 / API_HTTP_403 | Check the relevant provider key and permissions. |
MISSING_WRITE_LSN | Verify the deployed Pinecone API contract before release; no unverified delete follows. |
Local development
Python 3.12 on Linux/macOS:
python3 -m venv .venv.venv/bin/pip install -r requirements-dev.txt.venv/bin/python -m pytest -q
For a local Actor demo, put examples/demo.json at storage/key_value_stores/default/INPUT.json, then run .venv/bin/python -m src. The Docker image preloads the tokenizer data at build time. Outside Docker its first load may download that public tokenizer asset. No API key is needed for the demo.
tools/build_web_ide.py regenerates a self-contained launcher from the exact same source modules. START_HIER.html contains the German browser-only deployment instructions. No Apify CLI is required by that path.
Privacy and permissions
Keep Apify permissions at Limited. The Actor uses its own named storage, a coordination queue and default output storage. It does not read your unrelated Apify datasets. Keys use secret input fields and are excluded from state and generated reports. Website content sent for embedding goes to OpenAI; text and vectors go to your Pinecone namespace. You must have permission to crawl and process your chosen source.