# Rule-Guided Web Excerpts to Word (`automation-lab/rule-guided-web-excerpts-to-word`) Actor

Select verbatim paragraphs from public web pages by inclusion and exclusion phrases, preserving headings and source URLs in one cited Word document.

- **URL**: https://apify.com/automation-lab/rule-guided-web-excerpts-to-word.md
- **Developed by:** [Automation Lab](https://apify.com/automation-lab) (community)
- **Categories:** Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.77 / 1,000 contributing source pages

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Rule-Guided Web Excerpts to Word

Turn a list of public HTML pages into **one cited Word document**. This web page to word converter selects paragraphs that contain your inclusion phrases, excludes unwanted paragraphs, keeps their nearest preceding heading, and places the exact supplied source URL immediately after every excerpt. It never paraphrases the page text.

### Who is it for?

Researchers assembling reading packs, editors collecting source quotations, and analysts preparing traceable excerpts from multiple public pages. Supply the pages you are authorized to use; this is not a web search engine or a whole-site crawler.

### Why use this Actor?

Unlike whole-page DOCX conversion, the output contains only paragraphs that match your rules. Unlike a summary, the paragraph text is not rewritten. The Word file combines sources into a single document; the dataset gives one row per selected paragraph so you can audit the selection.

### Getting started

1. Add one to 25 public page URLs in `startUrls`.
2. Add at least one literal `includeTerms` phrase. A paragraph is selected if it contains **any** include phrase (case-insensitive).
3. Optionally provide `excludeTerms`: a match to **any** exclude phrase discards the paragraph even if it also matches an include phrase.
4. Choose `maxItems` and run the Actor. Download `EXCERPTS.docx` from the run's output/key-value store. Inspect `SUMMARY` for page failures.

### Input parameters

| Field | Meaning |
| --- | --- |
| `startUrls` | 1–25 public, anonymous HTML pages; processed in supplied order. |
| `includeTerms` | Required literal phrases; at least one. |
| `excludeTerms` | Optional literal phrases; exclusion wins. |
| `maxItems` | Maximum matching paragraphs across the run, 1–1000; Console default 10 (omitted API value 100). |
| `documentTitle` | Heading printed at the beginning of the Word file. |

Rules apply identically to every supplied page. They operate on the rendered HTML paragraph text, not on metadata, tags, images, PDF text, or headings. Headings are preserved as context, not filtered excerpts. A paragraph's outer whitespace is trimmed; internal text and punctuation are not paraphrased.

### Example: web scraping source notes

```json
{
  "startUrls": [{"url": "https://en.wikipedia.org/wiki/Web_scraping"}],
  "includeTerms": ["web scraping"],
  "excludeTerms": ["automated access"],
  "maxItems": 5,
  "documentTitle": "Web scraping source notes"
}
```

This is a bounded selection of paragraphs on a real public HTML article, not a promise to find every related page on the internet.

### What comes out?

The default dataset contains `sourceUrl`, `pageTitle`, `heading` (nullable when no heading precedes the paragraph), `paragraph`, `documentUrl`, and `scrapedAt`. For example, a local run on NASA's Earth page with `includeTerms: ["earth"]` and `excludeTerms: ["planet"]` selected this paragraph under “Science in Action for Society”:

> Learn how NASA’s studies of Earth bring benefits to the nation and world.

Its source URL was `https://science.nasa.gov/earth/`. The run stores a single `EXCERPTS.docx` containing the selected paragraphs and citations, plus `SUMMARY` with selected count and page errors. Every selected paragraph is followed by its source URL in the DOCX. The same DOCX download URL appears on every output row.

### How much does it cost to compile cited web excerpts?

Pay per event: one $0.0015 `start` event per run plus one `item` event for each supplied public page that contributes at least one selected paragraph. The number of excerpts on that page does not change its item charge. Empty-match and failed pages incur no item event. At BRONZE ($0.00128 per contributing page), one page costs $0.00278, five pages cost $0.00790, and ten pages cost $0.01430; these are Actor event prices, not platform compute charges. Other tiers: FREE $0.001472; SILVER $0.0009984; GOLD/PLATINUM/DIAMOND $0.000768 per contributing page. This Actor does not enable residential proxy or paid browser fallback automatically.

### Accuracy review tips

In a local test on NASA's Earth page, a literal `earth` inclusion and `planet` exclusion returned three selected rows at `maxItems: 3`; the first row retained the heading “Science in Action for Society.” In a separate two-page local test, seven rows came from NASA and twelve from Wikipedia without a per-page cap. If you need both sources represented in a bounded sample, order the pages and set a maximum larger than the first page's matching count. Always compare a few extracted paragraphs against the page before quoting them elsewhere. HTML can change between runs.

### Integrations

Schedule the same input to refresh a bounded reference pack, or export the dataset to a spreadsheet for a human accuracy review. Use the key-value-store DOCX URL to download the combined deliverable; use the dataset's source URL and heading to trace excerpts before publishing them elsewhere. Store the document securely if source material is sensitive.

### API usage

Launch with cURL (replace `TOKEN` with your own Apify API token):

```bash
curl -X POST 'https://api.apify.com/v2/acts/automation-lab~rule-guided-web-excerpts-to-word/runs?token=TOKEN' \
  -H 'Content-Type: application/json' \
  -d '{"startUrls":[{"url":"https://science.nasa.gov/earth/"}],"includeTerms":["earth"],"maxItems":3}'
```

In JavaScript, call `client.actor('automation-lab/rule-guided-web-excerpts-to-word').call(input)` with `apify-client`, then read the run's default dataset and default key-value store. In Python, use `ApifyClient(token).actor('automation-lab/rule-guided-web-excerpts-to-word').call(input)` and download `EXCERPTS.docx` from the returned default key-value-store ID. Keep tokens in secret storage, never in public Task inputs.

### MCP

Expose the Actor through Apify MCP in Claude Code:

```bash
claude mcp add --transport http apify \
  'https://mcp.apify.com?tools=automation-lab/rule-guided-web-excerpts-to-word'
```

For Claude Desktop, Cursor, or VS Code HTTP MCP clients, configure `{ "mcpServers": { "apify": { "url": "https://mcp.apify.com?tools=automation-lab/rule-guided-web-excerpts-to-word" } } }` in the client's MCP settings. Example prompts: Ask: “Extract NASA Earth paragraphs containing earth but not planet and give me the cited DOCX.” Or ask: “Compile web scraping research paragraphs from my Wikipedia URL into a Word file, at most five excerpts.” Each request must supply real URLs; the MCP tool cannot search the web for you.

### Limitations and troubleshooting

Only anonymously available server-rendered HTML paragraphs are inspected. JavaScript-only content, sign-in walls, CAPTCHA, PDFs, and pages that redirect are not automatically bypassed. Follow a redirect yourself and supply its final public URL. A page failure is logged and listed in `SUMMARY`; all-page failure fails the run. No match on a valid page produces an empty dataset and a title-only DOCX. If text is missing, inspect page HTML and try a phrase appearing inside an actual `<p>` element. The Actor does not recursively crawl links or discover pages from search terms.

### Legality and responsible use

Respect source copyright, terms, robots guidance, and applicable privacy rules. A citation does not grant permission to republish copyrighted content. Supply only publicly accessible pages and review excerpts before distributing the compiled document.

### Related Actors

[HTML Readability to Markdown Converter](https://apify.com/automation-lab/html-readability-markdown-converter) converts page content into Markdown rather than compiling rule-filtered citations. [Schema-Guided Web Data to Excel](https://apify.com/automation-lab/schema-guided-web-data-to-excel) extracts structured fields for spreadsheets, not a prose Word document.

### FAQ

**Are excerpts rewritten?** No: text is taken from HTML paragraph nodes; the Actor trims outer whitespace but does not summarize or rewrite. Text generated by browser JavaScript is outside this route.

**Can I provide a search query?** No. Provide exact public page URLs. Search results and whole-site discovery require a separate workflow.

**Why is there no output row?** None of the paragraphs matched, all matching paragraphs were excluded, or the page did not expose usable HTML paragraphs. Check `SUMMARY` and the run log; all failed pages cause a nonzero run exit.

# Changelog

This Actor's version history is a separate document: https://apify.com/automation-lab/rule-guided-web-excerpts-to-word/changelog.md

# Actor input Schema

## `startUrls` (type: `array`):

1–25 anonymous HTTP(S) HTML pages; redirects, login pages and private addresses are not supported.

## `includeTerms` (type: `array`):

Keep paragraphs containing any phrase (case-insensitive literal text). At least one is required.

## `excludeTerms` (type: `array`):

Reject paragraphs containing any phrase; exclusion wins over inclusion.

## `maxItems` (type: `integer`):

Maximum selected paragraphs across all pages in document order.

## `documentTitle` (type: `string`):

Title printed at the top of the single DOCX.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://en.wikipedia.org/wiki/Web_scraping"
    }
  ],
  "includeTerms": [
    "web scraping"
  ],
  "excludeTerms": [],
  "maxItems": 10,
  "documentTitle": "Cited web excerpts"
}
```

# Actor output Schema

## `overview` (type: `string`):

One row per selected paragraph with heading and source URL

## `document` (type: `string`):

Download the single cited DOCX

## `summary` (type: `string`):

Selected count and skipped page errors

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://en.wikipedia.org/wiki/Web_scraping"
        }
    ],
    "includeTerms": [
        "web scraping"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("automation-lab/rule-guided-web-excerpts-to-word").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://en.wikipedia.org/wiki/Web_scraping" }],
    "includeTerms": ["web scraping"],
}

# Run the Actor and wait for it to finish
run = client.actor("automation-lab/rule-guided-web-excerpts-to-word").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://en.wikipedia.org/wiki/Web_scraping"
    }
  ],
  "includeTerms": [
    "web scraping"
  ]
}' |
apify call automation-lab/rule-guided-web-excerpts-to-word --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,automation-lab/rule-guided-web-excerpts-to-word"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/th2PhSekAueXBe8tR/builds/uVyncIrCcc8rXVanH/openapi.json
