# Article Summarizer & Web Page Summarizer - AI Text Summarizer (`tidytools/web-page-summarizer`) Actor

AI summarizer for URLs, articles or your own text: TL;DR, bullets or executive summary, plus key points, topics, sentiment and keywords. Any language, JavaScript sites. No API key. $6/1,000.

- **URL**: https://apify.com/tidytools/web-page-summarizer.md
- **Developed by:** [Yukai Lin](https://apify.com/tidytools) (community)
- **Categories:** AI, Automation, News
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Article Summarizer & Text Summarizer do?

Give it URLs or your own texts and get a **clear, factual summary of each one**: one line, a paragraph, bullet points, a TL;DR or an executive summary, in the language you choose. Optionally get **key points, topic tags, sentiment and keywords** in the same call, at no extra cost. Works on articles, blogs, documentation, product and company pages, and JavaScript-heavy sites. No API key needed.

**Real run (October 3, 2026, shortened).** Input: `{"urlsText": "https://go.dev/blog/go1.21", "style": "executive", "extras": ["keyPoints", "topics", "sentiment"], "evidence": true}`. Output for that article (3,811 characters, one row, $0.006):

```json
{
    "title": "Go 1.21 is released! - The Go Programming Language",
    "summary": "Go 1.21 is released with new features and improvements. 
- The Profile Guided Optimization feature is now generally available. 
- New built-in functions include min, max, and clear. 
- ...",
    "keyPoints": ["Go 1.21 is released with new features and improvements.", "New built-in functions include min, max, and clear.", "..."],
    "topics": ["go release", "programming language", "webassembly", "performance improvements"],
    "sentiment": "positive",
    "evidence": [{ "quote": "We’ve measured the impact of PGO on a wide set of Go programs and see performance improvements of 2-7%.", "verified": true, "position": 777 }],
    "evidenceDropped": 0,
    "charged": true
}
```

- ✍️ **Five styles**: one line, paragraph, 4–7 bullet points, TL;DR, executive summary
- 🔗 **Supporting quotes checked against the page (optional, included in the price)**: 1–3 short quotes from the page, each one found in the page text; quotes the AI made up are removed
- 🔍 **Insights included in the price**: `keyPoints` (list), `topics` (tags), `sentiment` (positive / neutral / negative / mixed) and `keywords`
- 📝 **URLs or texts**: summarize pages, or paste articles, reviews, transcripts; or chain the output of another Actor
- 📚 **Long pages in full**: *Deep* mode summarizes pages longer than one AI pass section by section, then combines them
- 🌍 **Any output language**: English, Spanish, German, Japanese, Traditional Chinese…
- 🎯 **Focus option**: e.g. "pricing and plans", "key findings", "risks"
- 🧾 **Grounded in the text**: the AI is instructed to use only facts from the page and to ignore menus, cookie banners and footers
- 🌐 **Works on protected and JavaScript sites**: plain HTTP first, a second network if a site blocks data centers, a real browser when needed
- 💸 **Pay only for summaries delivered**: blocked or empty pages, login walls, paywalled teasers, bot checks, "enable JavaScript" and cookie-consent pages are free (the AI is not even called)

### How much does it cost?

| Event | Price |
|---|---|
| Summarized page or text | **$6.00 / 1,000** (insights included) |
| Long page summarized in full (*Deep* mode, only for pages over 30,000 characters) | **$12.00 / 1,000** |

**Example:** summarizing 500 news articles costs 500 × $0.006 = **$3.00** on the Free plan.

**No start fee.** With *Deep* mode on, pages that fit in one pass keep the standard price. Higher Apify plans get volume discounts.

Unlike general "web page → GPT" tools, you don't need an OpenAI key: the price per page covers everything. Some summarizers on Apify Store add a start fee of $0.09 per run (checked September 2026), which is more than summarizing 15 pages here.

#### Control your cost

- **What is charged**: one event per page or text that got a summary. The deep price applies only to pages that were really too long for one pass.
- **What is free**: lines that are not a URL, duplicate URLs, blocked pages, errors, time-outs, login walls, paywalled pages (only the start of the article is visible), bot checks, "enable JavaScript" and cookie-consent pages and pages with less than about 200 characters of text. Every row has `charged: true/false`, and failed rows say "(not charged)".
- **Before the run starts**, the log and the status message show the plan: how many items, the worst-case cost and your max charge per run (also saved as `costPlan` in `SUMMARY`).
- **Max charge per run**: set *Maximum cost per run* in the run options (or `maxTotalChargeUsd` in the API). When it is reached, the run stops starting new items and finishes cleanly: `SUMMARY.status` is `LIMIT_REACHED` and `SUMMARY.notProcessed` lists the items that were not started (count and up to 100 inputs).
- **If Apify restarts the run** (server migration or Resurrect), items already finished are skipped and not charged again (`SUMMARY.resumedSkipped`).

### How to use it

1. Paste your **URLs or domains** (one per line) and/or **texts**. A line that is not a URL becomes one free `invalid_input` row; the rest of the run continues. Every input gets exactly one row (summary or error). The same page written differently (`example.com`, `https://example.com/`, `www.example.com`) is summarized and charged once. An empty input ends with one free `invalid_input` row (the run does not fail).
2. Pick a **style** and **max words**; optionally **insights**, an **output language** and a **focus**.
3. Click **Start** and export the summaries as JSON, CSV or Excel.

#### Input example: executive summary with insights

```json
{
    "urlsText": "https://blog.cloudflare.com/",
    "texts": ["I bought the X200 headphones last month. The noise cancelling is excellent on flights and the battery easily lasts 30 hours, but the ear cushions get hot after an hour and the app keeps disconnecting on my Android phone. Support replied within a day and sent a firmware fix, which solved the app issue."],
    "style": "executive",
    "extras": ["keyPoints", "topics", "sentiment", "keywords"]
}
```

#### Output example (real result for the text above)

```json
{
    "source": "text-0",
    "success": true,
    "summary": "The X200 headphones have both positive and negative aspects.\n- The noise cancelling is excellent.\n- The battery lasts 30 hours.\n- The ear cushions get hot after an hour.\n- The app had connectivity issues on Android phones, but support provided a fix.",
    "keyPoints": ["The X200 headphones have excellent noise cancelling.", "The battery lasts 30 hours.", "The ear cushions get hot after an hour.", "The app keeps disconnecting on Android phones.", "Support replied within a day with a firmware fix."],
    "topics": ["headphones", "noise cancelling", "battery life"],
    "sentiment": "mixed",
    "keywords": ["X200 headphones", "noise cancelling", "battery", "ear cushions", "app", "Android", "firmware fix", "support"],
    "style": "executive",
    "sourceCharacters": 302,
    "analysisDepth": "standard"
}
```

For the Cloudflare blog home page the same run returned the topics `cloudflare`, `ai`, `development`, `security`, `performance` and sentiment `neutral`.

#### Output example: bullet points of a long page, Deep mode (real result)

```json
{
    "url": "https://en.wikipedia.org/wiki/Cloudflare",
    "title": "Cloudflare - Wikipedia",
    "success": true,
    "summary": "- Cloudflare, Inc. is an American technology company that provides a range of internet services, including content delivery network services, cloud cybersecurity, DDoS mitigation, and ICANN-accredited domain registration.\n- The company was founded by Matthew Prince, Lee Holloway, and Michelle Zatlyn, and it went public on the New York Stock Exchange in 2019 under the ticker symbol NET.\n- ...\n- In May 2026, Cloudflare announced the elimination of approximately 1,100 positions, around 20 percent of its workforce, in a restructuring the company attributed to the rapid adoption of artificial intelligence tools.",
    "style": "bullets",
    "contentTruncated": false,
    "sourceCharacters": 158584,
    "analysisDepth": "deep",
    "chunksSummarized": 9,
    "mode": "fast",
    "via": "backend"
}
```

In standard mode the same 158,584-character page is summarized from its first ~32,000 characters (`contentTruncated: true`).

#### Summarize another Actor's output

Set **Dataset ID** to a run's dataset and **Dataset field** to the field to use: URLs are read as pages, other values (article bodies, reviews) are summarized as text. Empty field = the first URL field, else the first long text field.

### Output fields

| Field | Meaning |
|---|---|
| `summary` | The summary in the chosen style |
| `keyPoints`, `topics`, `sentiment`, `keywords` | Only when requested in *Insights* |
| `evidence` | Only with *Supporting quotes* on: up to 3 `{ quote, verified, position }` items. Each quote was found in the page text the AI read (ignoring case, extra spaces, curly vs straight quotes, dashes and Markdown formatting); `position` is its character offset in that text. Quotes longer than 300 characters are shortened and end with "…" |
| `evidenceDropped` | How many quotes from the AI were removed because they are not in the page text (made up, paraphrased, translated or too short) |
| `evidenceNote` | Why no quotes were returned, when that is not the page's fault (e.g. long pages in *Deep* mode) |
| `sourceCharacters` | Length of the page text or your text |
| `contentTruncated` | `true` when only the first ~32,000 characters were read |
| `analysisDepth`, `chunksSummarized` | `deep` pages list how many sections were summarized |
| `mode`, `via` | `fast` or `browser`; fetched by `backend` (our servers), `direct` (Apify's network) or `browser` |
| `input`, `inputIndex` | For URLs: the line you entered and its position in your list; for texts: `inputIndex` is the position of the text |
| `source`, `textPreview` | For texts: `text-0`, `dataset-3`… and the start of the text |
| `datasetIndex` | With *Dataset ID*: position of the value in the source dataset, to join the summary back to the original item |
| `charged`, `chargedEvent` | Whether this row was charged, and with which event |
| `error`, `errorType` | For items that failed (not charged): `invalid_input`, `login_required`, `paywalled`, `blocked`, `no_text`, `not_found`, `unreachable` (unknown domain or no answer), `timeout`… |

### Use with AI agents (MCP)

Connect Apify's MCP server (https://mcp.apify.com?tools=tidytools/web-page-summarizer) to Claude, Cursor or any MCP client, then ask e.g. "Give me a TL;DR and the key points of these five articles".

```json
{ "urlsText": "https://en.wikipedia.org/wiki/Cloudflare", "style": "tldr", "extras": ["keyPoints"] }
```

Failed items are not charged and carry an `errorType` (`blocked`, `not_found`, `timeout`, `no_text`…). The key-value store record `SUMMARY` has an overall `status` (`SUCCESS`, `PARTIAL_RESULTS`, `FAILED`, `NO_RESULTS`, `LIMIT_REACHED`).

### Use it from code and integrations

Run it from your own code with the Apify API. This call waits for the run and returns the results as JSON (replace `YOUR_TOKEN` with your Apify API token):

```bash
curl -X POST "https://api.apify.com/v2/acts/tidytools~web-page-summarizer/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"urlsText":"https://en.wikipedia.org/wiki/Cloudflare"}'
```

The synchronous endpoint waits up to 5 minutes. For bigger runs, start the run with `POST https://api.apify.com/v2/acts/tidytools~web-page-summarizer/runs` and read the dataset when it finishes, or use the `apify-client` package for JavaScript or Python.

**Schedules and integrations:** run it daily or weekly with Apify Schedules, get a webhook when a run finishes, or send the results to Zapier, Make, n8n, Google Sheets, Slack and other apps with Apify integrations. Results can be exported as JSON, CSV, Excel or XML.

### Tips

- For news sites, a **paragraph** summary of the home page gives a quick digest of the day's headlines.
- Use **focus** to pull one topic out of long pages (e.g. "refund policy").
- **Sentiment** is most useful for reviews, posts and news; factual pages are usually `neutral`.
- For very short texts, pick *One line* or *TL;DR*: longer styles can pad a two-sentence text with restatements.

### Limitations

- AI summaries can occasionally miss nuance; verify important facts.
- Images and other files without text come back as a free `no_text` row. PDFs and plain-text files are read as text; a PDF's summary may describe its metadata when it holds little text.
- Only public pages; no logins. For sites that block data centers, add a proxy under *Advanced settings* (Apify proxy usage is billed to your Apify account).
- *Deep* mode reads up to about 240,000 characters (10 sections) per page and takes longer.
- *Supporting quotes* come from the first ~32,000 characters the AI read, stay in the page's language even with an *Output language*, and are not available for long pages summarized in *Deep* mode. A quote shows that the passage is on the page; it does not prove that every sentence of the summary is correct. The list can be empty when no quote from the AI could be found on the page.
- *Time limit per item* (Advanced settings) stops a page or text that takes too long (default 180 seconds, 600 with *Deep*); it comes back as a free `timeout` row.
- Pages with less than about 200 characters of text, login pages, paywalled teasers (a short page that says to subscribe to continue reading), bot checks and "enable JavaScript" or cookie-consent pages are not summarized (free `login_required`, `paywalled`, `blocked` or `no_text` rows).

### FAQ

**Does the Actor collect usage data?** The Actor sends one anonymous usage counter per run (input type, item count range, outcome) to improve the tool; no URLs, content or account data.

**Do I need an OpenAI API key to summarize a web page or text?** No. The price per page covers everything, so you don't need an OpenAI key. Paste URLs, or your own articles, reviews or transcripts in **Texts to summarize (optional)**.

**Which summary formats can I get, and is there a TL;DR?** **Summary style** offers One line, Paragraph, Bullet points (4–7), TL;DR (1–2 sentences) and Executive summary, with **Max words** and **Output language (optional)** to set length and language. **Insights (included in the price)** adds key points, topic tags, sentiment and keywords in the same AI call at no extra cost.

**Does it work on JavaScript sites and very long articles?** Yes. With **Page fetching (HTTP or browser)** on Auto, it reads pages over plain HTTP and switches to a real browser for sites that block it or need JavaScript. For long pages, set **Long pages** to Deep: pages longer than one AI pass are summarized section by section and combined (up to about 240,000 characters), while Standard reads the first 32,000 characters. PDF links and Office documents are converted to text and summarized like web pages (a scanned PDF without a text layer comes back as a free `no_text` row).

**What is charged and what is free?** Each summarized page or text costs $6.00 / 1,000 (insights included), and $12.00 / 1,000 only for pages over 30,000 characters summarized in full in Deep mode; there is no start fee. Lines that are not a URL, duplicate URLs, blocked pages, errors, time-outs, login walls, paywalled pages, bot checks, "enable JavaScript" and cookie-consent pages, and pages with less than about 200 characters of text are free. Wrong input never fails the run: it ends with a free `invalid_input` row that says what is wrong (no URL or text given, a blank or non-string text, an unknown style that falls back to Paragraph, a `maxWords` that is not a number), and `SUMMARY.inputWarnings` lists options that were replaced by their defaults.

### Related Actors

More tools by TidyTools that work well with this one:

- [Website Markdown Crawler](https://apify.com/tidytools/website-markdown-crawler): crawl a whole site or docs portal into clean Markdown for LLMs and RAG
- [AI Web Data Extractor](https://apify.com/tidytools/ai-web-data-extractor): turn any page into structured JSON with the fields you choose
- [Web Page Translator](https://apify.com/tidytools/web-page-translator): translate pages, text, JSON and subtitles with AI

### Support

Open an issue in the **Issues** tab with the URL and your input. Issues are checked regularly.

# Actor input Schema

## `urlsText` (type: `string`):

Main input (fill this, or `urls` / `texts` / `datasetId`). Articles, blog posts, product pages, documentation or any web page. One per line (bare domains such as stripe.com work too). A line that is not a URL becomes one free error row; the rest of the run continues.

## `urls` (type: `array`):

Same as "URLs or domains", in Apify's request list format (also accepts a link to a text file with URLs). Kept for existing integrations; new integrations should use "urlsText".

## `texts` (type: `array`):

Your own texts instead of (or besides) URLs: articles, reviews, transcripts, emails. Each text is one item, charged like one page.

## `style` (type: `string`):

One sentence, a short paragraph, 4–7 bullet points, a TL;DR (1–2 sentences) or an executive summary (bottom line + key facts). Values: one-line = One line; paragraph = Paragraph; bullets = Bullet points; tldr = TL;DR; executive = Executive summary.

## `maxWords` (type: `integer`):

Upper limit for the summary length (one-line summaries are capped at 40 words, TL;DR at 50).

## `extras` (type: `array`):

Extra fields returned with the summary in the same AI call: key points (list), topic tags, sentiment (positive / neutral / negative / mixed) and keywords.

## `evidence` (type: `boolean`):

Also return 1–3 short quotes from the page that support the summary, in the same AI call. Every quote is checked against the page text; quotes that are not really on the page are removed (counted in `evidenceDropped`). Not available for long pages summarized in deep mode.

## `outputLanguage` (type: `string`):

Write the summary in this language, e.g. English, Spanish, Japanese, Traditional Chinese. Empty = same language as the page.

## `focus` (type: `string`):

What the summary should concentrate on, e.g. "pricing and plans" or "key findings".

## `analysisDepth` (type: `string`):

Standard reads the first 32,000 characters of each page (about 5,000 words). Deep summarizes longer pages in full: section by section, then combined (up to about 240,000 characters). Deep costs $12 / 1,000 only for pages that are actually longer than one pass; shorter pages keep the standard price.

## `datasetId` (type: `string`):

Summarize the output of another Actor run: URLs are read as pages, other values (e.g. article bodies, reviews) are summarized as text.

## `datasetField` (type: `string`):

Field to use, e.g. "url" or "text" (dot paths work). Empty = the first URL field found, else the first long text field.

## `mode` (type: `string`):

"Auto" reads pages over plain HTTP and switches to a real browser for sites that block it or need JavaScript. Values: auto = Auto (recommended); fast = Fast (HTTP only); browser = Browser (JavaScript rendering).

## `maxConcurrency` (type: `integer`):

How many pages are summarized at the same time.

## `httpVia` (type: `string`):

Some sites block requests from data centers. Auto retries from a second network before falling back to a real browser. If both are refused, Auto also tries our second server (Oracle, different IP) before a real browser.

## `itemTimeoutSecs` (type: `integer`):

A page or text that takes longer (reading + AI) is stopped and returned as a free error row with errorType "timeout". Empty = 180 s (600 s with deep analysis); 0 = no limit.

## `proxyConfiguration` (type: `object`):

Only used for requests sent from Apify's network. Apify proxy usage is billed to your Apify account.

## Actor input object example

```json
{
  "urlsText": "https://en.wikipedia.org/wiki/Cloudflare",
  "style": "paragraph",
  "maxWords": 80,
  "evidence": false,
  "analysisDepth": "standard",
  "mode": "auto",
  "maxConcurrency": 5,
  "httpVia": "auto"
}
```

# Actor output Schema

## `summaries` (type: `string`):

No description

## `full` (type: `string`):

No description

## `details` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urlsText": "https://en.wikipedia.org/wiki/Cloudflare"
};

// Run the Actor and wait for it to finish
const run = await client.actor("tidytools/web-page-summarizer").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urlsText": "https://en.wikipedia.org/wiki/Cloudflare" }

# Run the Actor and wait for it to finish
run = client.actor("tidytools/web-page-summarizer").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urlsText": "https://en.wikipedia.org/wiki/Cloudflare"
}' |
apify call tidytools/web-page-summarizer --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tidytools/web-page-summarizer"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/UK839QHsFspWRaI7d/builds/bejdXyZokWdldDwuk/openapi.json
