# Web Page Text Extractor for AI & RAG (`riparazionecomputerecellulari/my-actor`) Actor

Extract clean text from up to 10 public web pages for AI and RAG pipelines. Get source URLs, page titles, UTC retrieval times, character counts and SHA-256 content fingerprints, with bounded timeouts and per-page diagnostics. Static HTML only; no browser rendering or login.

- **URL**: https://apify.com/riparazionecomputerecellulari/my-actor.md
- **Developed by:** [Riparazione Computer\&Cellulari](https://apify.com/riparazionecomputerecellulari) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 page extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Clean Web Text with Source Evidence

Extract the main readable text from a small batch of public, static HTML pages. Receive structured JSON with the source URL, page title, extracted text, character count, retrieval timestamp and a SHA-256 content fingerprint.

Use it in RAG ingestion, research agents and n8n workflows when you already know the URLs and need a simple extraction step. The fingerprint helps detect changes in the extracted text; it does not certify that the source is accurate.

### Quick start

Provide between 1 and 10 URLs:

```json
{"urls":["https://example.com/"]}
```

Successful pages appear in the default dataset. The `DIAGNOSTICS` key in the default key-value store lists failed URLs and reasons. Always inspect diagnostics: a completed run can contain partial results or no results.

### Output

| Field | Meaning |
| --- | --- |
| url | Final source URL |
| title | Extracted page title, when available |
| text | Main static page text |
| characters | Length of extracted text |
| sha256 | SHA-256 fingerprint of extracted text |
| retrievedAt | Retrieval timestamp in UTC |
| status | `success` |

### Limits

Each page has a hard 20-second processing deadline and a 2 MB response limit. Repeated URLs differing only by a fragment are processed once. URLs must be public HTTP(S), use standard ports and contain no embedded credentials. Local and private network destinations are rejected.

This Actor checks robots.txt and stops when permission cannot be established. It does not render JavaScript, use residential proxies, bypass access controls, process PDFs, perform OCR, sign in to websites or crawl links recursively. Compressed responses are currently unsupported. Cross-host redirects are not accepted as successful input: use the final public URL directly.

### Intended use

- Add source URLs and timestamps to extracted text for a research workflow.
- Compare content fingerprints between runs to detect text changes.
- Fetch up to 10 known static pages for an ingestion pipeline.

### Pricing

The price is $0.005 per successfully extracted page ($5 per 1,000 pages), plus a $0.001 run-start event at the default 512 MB memory. The start event is billed once per GB of memory, with a minimum of one event. Failed pages do not generate the page-extracted charge; the run-start charge still applies. Platform usage is included in these event prices. Check the store's current pricing before using the Actor. At the default memory, 10 successful pages cost $0.051. A $0.05 spending ceiling may stop the run before all 10 pages are returned.

### Troubleshooting

A page deadline, robots restriction, HTTP error or lack of readable static content is reported in `DIAGNOSTICS`. Use a browser crawler when the target requires JavaScript. An HTTP 200 response is not sufficient: the Actor requires an HTML content type and at least 80 characters of extracted text.

This is an early product with limited real-site validation. No claim of universal coverage or commercial success is made.

### Python integration

Use `apify-client` and store your own Apify token in an environment variable. Do not put it in a public URL or commit it to source code. The accompanying `example_client.py` includes a $0.05 execution ceiling, a timeout and diagnostics retrieval. The example has been syntax checked against Apify Client 3.2.1; an authenticated end-to-end client run has not been tested.

### n8n integration

With the official Apify node, use **Run Actor**, select this Actor and pass the `urls` JSON input. Wait for the run to finish, then use **Get Dataset Items** with the returned `defaultDatasetId`. Also retrieve `DIAGNOSTICS` from the returned `defaultKeyValueStoreId` through the key-value store API. Review failures before feeding the dataset to another service. Set the workflow's own timeout above the Actor timeout.

### Validation snapshot

A private cloud test on 1 October 2026 processed five URL inputs: four returned extracted text and one was stopped by the 20-second page deadline. Total run duration was 83 seconds at 512 MB; the displayed platform execution cost was about $0.002, excluding timed storage. This is a small test, not a general success-rate estimate or a guaranteed cost.

A subsequent private development run with pricing active completed in 9 seconds: one successful page and one HTTP error. Apify recorded one `apify-actor-start` event and one `page-extracted` event. This validates event counting; it is not a customer sale or proof of commercial profit.

# Actor input Schema

## `urls` (type: `array`):

1-10 public static HTML page URLs.

## Actor input object example

```json
{
  "urls": [
    "https://example.com/",
    "https://docs.apify.com/actors/publishing/monetize"
  ]
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://example.com/",
        "https://docs.apify.com/actors/publishing/monetize"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("riparazionecomputerecellulari/my-actor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://example.com/",
        "https://docs.apify.com/actors/publishing/monetize",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("riparazionecomputerecellulari/my-actor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://example.com/",
    "https://docs.apify.com/actors/publishing/monetize"
  ]
}' |
apify call riparazionecomputerecellulari/my-actor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,riparazionecomputerecellulari/my-actor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/zXVnrcVKBJQppabb4/builds/0mw1yqtkZgKub0uVN/openapi.json
