Rule-Guided Web Excerpts to Word avatar

Rule-Guided Web Excerpts to Word

Pricing

from $0.77 / 1,000 contributing source pages

Go to Apify Store
Rule-Guided Web Excerpts to Word

Rule-Guided Web Excerpts to Word

Select verbatim paragraphs from public web pages by inclusion and exclusion phrases, preserving headings and source URLs in one cited Word document.

Pricing

from $0.77 / 1,000 contributing source pages

Rating

0.0

(0)

Developer

Automation Lab

Automation Lab

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

7 days ago

Last modified

Categories

Share

Turn a list of public HTML pages into one cited Word document. This web page to word converter selects paragraphs that contain your inclusion phrases, excludes unwanted paragraphs, keeps their nearest preceding heading, and places the exact supplied source URL immediately after every excerpt. It never paraphrases the page text.

Who is it for?

Researchers assembling reading packs, editors collecting source quotations, and analysts preparing traceable excerpts from multiple public pages. Supply the pages you are authorized to use; this is not a web search engine or a whole-site crawler.

Why use this Actor?

Unlike whole-page DOCX conversion, the output contains only paragraphs that match your rules. Unlike a summary, the paragraph text is not rewritten. The Word file combines sources into a single document; the dataset gives one row per selected paragraph so you can audit the selection.

Getting started

  1. Add one to 25 public page URLs in startUrls.
  2. Add at least one literal includeTerms phrase. A paragraph is selected if it contains any include phrase (case-insensitive).
  3. Optionally provide excludeTerms: a match to any exclude phrase discards the paragraph even if it also matches an include phrase.
  4. Choose maxItems and run the Actor. Download EXCERPTS.docx from the run's output/key-value store. Inspect SUMMARY for page failures.

Input parameters

FieldMeaning
startUrls1–25 public, anonymous HTML pages; processed in supplied order.
includeTermsRequired literal phrases; at least one.
excludeTermsOptional literal phrases; exclusion wins.
maxItemsMaximum matching paragraphs across the run, 1–1000; Console default 10 (omitted API value 100).
documentTitleHeading printed at the beginning of the Word file.

Rules apply identically to every supplied page. They operate on the rendered HTML paragraph text, not on metadata, tags, images, PDF text, or headings. Headings are preserved as context, not filtered excerpts. A paragraph's outer whitespace is trimmed; internal text and punctuation are not paraphrased.

Example: web scraping source notes

{
"startUrls": [{"url": "https://en.wikipedia.org/wiki/Web_scraping"}],
"includeTerms": ["web scraping"],
"excludeTerms": ["automated access"],
"maxItems": 5,
"documentTitle": "Web scraping source notes"
}

This is a bounded selection of paragraphs on a real public HTML article, not a promise to find every related page on the internet.

What comes out?

The default dataset contains sourceUrl, pageTitle, heading (nullable when no heading precedes the paragraph), paragraph, documentUrl, and scrapedAt. For example, a local run on NASA's Earth page with includeTerms: ["earth"] and excludeTerms: ["planet"] selected this paragraph under “Science in Action for Society”:

Learn how NASA’s studies of Earth bring benefits to the nation and world.

Its source URL was https://science.nasa.gov/earth/. The run stores a single EXCERPTS.docx containing the selected paragraphs and citations, plus SUMMARY with selected count and page errors. Every selected paragraph is followed by its source URL in the DOCX. The same DOCX download URL appears on every output row.

How much does it cost to compile cited web excerpts?

Pay per event: one $0.0015 start event per run plus one item event for each supplied public page that contributes at least one selected paragraph. The number of excerpts on that page does not change its item charge. Empty-match and failed pages incur no item event. At BRONZE ($0.00128 per contributing page), one page costs $0.00278, five pages cost $0.00790, and ten pages cost $0.01430; these are Actor event prices, not platform compute charges. Other tiers: FREE $0.001472; SILVER $0.0009984; GOLD/PLATINUM/DIAMOND $0.000768 per contributing page. This Actor does not enable residential proxy or paid browser fallback automatically.

Accuracy review tips

In a local test on NASA's Earth page, a literal earth inclusion and planet exclusion returned three selected rows at maxItems: 3; the first row retained the heading “Science in Action for Society.” In a separate two-page local test, seven rows came from NASA and twelve from Wikipedia without a per-page cap. If you need both sources represented in a bounded sample, order the pages and set a maximum larger than the first page's matching count. Always compare a few extracted paragraphs against the page before quoting them elsewhere. HTML can change between runs.

Integrations

Schedule the same input to refresh a bounded reference pack, or export the dataset to a spreadsheet for a human accuracy review. Use the key-value-store DOCX URL to download the combined deliverable; use the dataset's source URL and heading to trace excerpts before publishing them elsewhere. Store the document securely if source material is sensitive.

API usage

Launch with cURL (replace TOKEN with your own Apify API token):

curl -X POST 'https://api.apify.com/v2/acts/automation-lab~rule-guided-web-excerpts-to-word/runs?token=TOKEN' \
-H 'Content-Type: application/json' \
-d '{"startUrls":[{"url":"https://science.nasa.gov/earth/"}],"includeTerms":["earth"],"maxItems":3}'

In JavaScript, call client.actor('automation-lab/rule-guided-web-excerpts-to-word').call(input) with apify-client, then read the run's default dataset and default key-value store. In Python, use ApifyClient(token).actor('automation-lab/rule-guided-web-excerpts-to-word').call(input) and download EXCERPTS.docx from the returned default key-value-store ID. Keep tokens in secret storage, never in public Task inputs.

MCP

Expose the Actor through Apify MCP in Claude Code:

claude mcp add --transport http apify \
'https://mcp.apify.com?tools=automation-lab/rule-guided-web-excerpts-to-word'

For Claude Desktop, Cursor, or VS Code HTTP MCP clients, configure { "mcpServers": { "apify": { "url": "https://mcp.apify.com?tools=automation-lab/rule-guided-web-excerpts-to-word" } } } in the client's MCP settings. Example prompts: Ask: “Extract NASA Earth paragraphs containing earth but not planet and give me the cited DOCX.” Or ask: “Compile web scraping research paragraphs from my Wikipedia URL into a Word file, at most five excerpts.” Each request must supply real URLs; the MCP tool cannot search the web for you.

Limitations and troubleshooting

Only anonymously available server-rendered HTML paragraphs are inspected. JavaScript-only content, sign-in walls, CAPTCHA, PDFs, and pages that redirect are not automatically bypassed. Follow a redirect yourself and supply its final public URL. A page failure is logged and listed in SUMMARY; all-page failure fails the run. No match on a valid page produces an empty dataset and a title-only DOCX. If text is missing, inspect page HTML and try a phrase appearing inside an actual <p> element. The Actor does not recursively crawl links or discover pages from search terms.

Legality and responsible use

Respect source copyright, terms, robots guidance, and applicable privacy rules. A citation does not grant permission to republish copyrighted content. Supply only publicly accessible pages and review excerpts before distributing the compiled document.

HTML Readability to Markdown Converter converts page content into Markdown rather than compiling rule-filtered citations. Schema-Guided Web Data to Excel extracts structured fields for spreadsheets, not a prose Word document.

FAQ

Are excerpts rewritten? No: text is taken from HTML paragraph nodes; the Actor trims outer whitespace but does not summarize or rewrite. Text generated by browser JavaScript is outside this route.

Can I provide a search query? No. Provide exact public page URLs. Search results and whole-site discovery require a separate workflow.

Why is there no output row? None of the paragraphs matched, all matching paragraphs were excluded, or the page did not expose usable HTML paragraphs. Check SUMMARY and the run log; all failed pages cause a nonzero run exit.