Web Page Text Extractor for AI & RAG
Pricing
from $1.00 / 1,000 page extracteds
Web Page Text Extractor for AI & RAG
Extract clean text from up to 10 public web pages for AI and RAG pipelines. Get source URLs, page titles, UTC retrieval times, character counts and SHA-256 content fingerprints, with bounded timeouts and per-page diagnostics. Static HTML only; no browser rendering or login.
Pricing
from $1.00 / 1,000 page extracteds
Rating
0.0
(0)
Developer
Riparazione Computer&Cellulari
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
20 hours ago
Last modified
Categories
Share
Clean Web Text with Source Evidence
Extract the main readable text from a small batch of public, static HTML pages. Receive structured JSON with the source URL, page title, extracted text, character count, retrieval timestamp and a SHA-256 content fingerprint.
Use it in RAG ingestion, research agents and n8n workflows when you already know the URLs and need a simple extraction step. The fingerprint helps detect changes in the extracted text; it does not certify that the source is accurate.
Quick start
Provide between 1 and 10 URLs:
{"urls":["https://example.com/"]}
Successful pages appear in the default dataset. The DIAGNOSTICS key in the default key-value store lists failed URLs and reasons. Always inspect diagnostics: a completed run can contain partial results or no results.
Output
| Field | Meaning |
|---|---|
| url | Final source URL |
| title | Extracted page title, when available |
| text | Main static page text |
| characters | Length of extracted text |
| sha256 | SHA-256 fingerprint of extracted text |
| retrievedAt | Retrieval timestamp in UTC |
| status | success |
Limits
Each page has a hard 20-second processing deadline and a 2 MB response limit. Repeated URLs differing only by a fragment are processed once. URLs must be public HTTP(S), use standard ports and contain no embedded credentials. Local and private network destinations are rejected.
This Actor checks robots.txt and stops when permission cannot be established. It does not render JavaScript, use residential proxies, bypass access controls, process PDFs, perform OCR, sign in to websites or crawl links recursively. Compressed responses are currently unsupported. Cross-host redirects are not accepted as successful input: use the final public URL directly.
Intended use
- Add source URLs and timestamps to extracted text for a research workflow.
- Compare content fingerprints between runs to detect text changes.
- Fetch up to 10 known static pages for an ingestion pipeline.
Pricing
The price is $0.005 per successfully extracted page ($5 per 1,000 pages), plus a $0.001 run-start event at the default 512 MB memory. The start event is billed once per GB of memory, with a minimum of one event. Failed pages do not generate the page-extracted charge; the run-start charge still applies. Platform usage is included in these event prices. Check the store's current pricing before using the Actor. At the default memory, 10 successful pages cost $0.051. A $0.05 spending ceiling may stop the run before all 10 pages are returned.
Troubleshooting
A page deadline, robots restriction, HTTP error or lack of readable static content is reported in DIAGNOSTICS. Use a browser crawler when the target requires JavaScript. An HTTP 200 response is not sufficient: the Actor requires an HTML content type and at least 80 characters of extracted text.
This is an early product with limited real-site validation. No claim of universal coverage or commercial success is made.
Python integration
Use apify-client and store your own Apify token in an environment variable. Do not put it in a public URL or commit it to source code. The accompanying example_client.py includes a $0.05 execution ceiling, a timeout and diagnostics retrieval. The example has been syntax checked against Apify Client 3.2.1; an authenticated end-to-end client run has not been tested.
n8n integration
With the official Apify node, use Run Actor, select this Actor and pass the urls JSON input. Wait for the run to finish, then use Get Dataset Items with the returned defaultDatasetId. Also retrieve DIAGNOSTICS from the returned defaultKeyValueStoreId through the key-value store API. Review failures before feeding the dataset to another service. Set the workflow's own timeout above the Actor timeout.
Validation snapshot
A private cloud test on 1 October 2026 processed five URL inputs: four returned extracted text and one was stopped by the 20-second page deadline. Total run duration was 83 seconds at 512 MB; the displayed platform execution cost was about $0.002, excluding timed storage. This is a small test, not a general success-rate estimate or a guaranteed cost.
A subsequent private development run with pricing active completed in 9 seconds: one successful page and one HTTP error. Apify recorded one apify-actor-start event and one page-extracted event. This validates event counting; it is not a customer sale or proof of commercial profit.