Rule-Guided Web Excerpts to Word
Pricing
from $0.77 / 1,000 contributing source pages
Rule-Guided Web Excerpts to Word
Select verbatim paragraphs from public web pages by inclusion and exclusion phrases, preserving headings and source URLs in one cited Word document.
Pricing
from $0.77 / 1,000 contributing source pages
Rating
0.0
(0)
Developer
Automation Lab
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
7 days ago
Last modified
Categories
Share
Turn a list of public HTML pages into one cited Word document. This web page to word converter selects paragraphs that contain your inclusion phrases, excludes unwanted paragraphs, keeps their nearest preceding heading, and places the exact supplied source URL immediately after every excerpt. It never paraphrases the page text.
Who is it for?
Researchers assembling reading packs, editors collecting source quotations, and analysts preparing traceable excerpts from multiple public pages. Supply the pages you are authorized to use; this is not a web search engine or a whole-site crawler.
Why use this Actor?
Unlike whole-page DOCX conversion, the output contains only paragraphs that match your rules. Unlike a summary, the paragraph text is not rewritten. The Word file combines sources into a single document; the dataset gives one row per selected paragraph so you can audit the selection.
Getting started
- Add one to 25 public page URLs in
startUrls. - Add at least one literal
includeTermsphrase. A paragraph is selected if it contains any include phrase (case-insensitive). - Optionally provide
excludeTerms: a match to any exclude phrase discards the paragraph even if it also matches an include phrase. - Choose
maxItemsand run the Actor. DownloadEXCERPTS.docxfrom the run's output/key-value store. InspectSUMMARYfor page failures.
Input parameters
| Field | Meaning |
|---|---|
startUrls | 1–25 public, anonymous HTML pages; processed in supplied order. |
includeTerms | Required literal phrases; at least one. |
excludeTerms | Optional literal phrases; exclusion wins. |
maxItems | Maximum matching paragraphs across the run, 1–1000; Console default 10 (omitted API value 100). |
documentTitle | Heading printed at the beginning of the Word file. |
Rules apply identically to every supplied page. They operate on the rendered HTML paragraph text, not on metadata, tags, images, PDF text, or headings. Headings are preserved as context, not filtered excerpts. A paragraph's outer whitespace is trimmed; internal text and punctuation are not paraphrased.
Example: web scraping source notes
{"startUrls": [{"url": "https://en.wikipedia.org/wiki/Web_scraping"}],"includeTerms": ["web scraping"],"excludeTerms": ["automated access"],"maxItems": 5,"documentTitle": "Web scraping source notes"}
This is a bounded selection of paragraphs on a real public HTML article, not a promise to find every related page on the internet.
What comes out?
The default dataset contains sourceUrl, pageTitle, heading (nullable when no heading precedes the paragraph), paragraph, documentUrl, and scrapedAt. For example, a local run on NASA's Earth page with includeTerms: ["earth"] and excludeTerms: ["planet"] selected this paragraph under “Science in Action for Society”:
Learn how NASA’s studies of Earth bring benefits to the nation and world.
Its source URL was https://science.nasa.gov/earth/. The run stores a single EXCERPTS.docx containing the selected paragraphs and citations, plus SUMMARY with selected count and page errors. Every selected paragraph is followed by its source URL in the DOCX. The same DOCX download URL appears on every output row.
How much does it cost to compile cited web excerpts?
Pay per event: one $0.0015 start event per run plus one item event for each supplied public page that contributes at least one selected paragraph. The number of excerpts on that page does not change its item charge. Empty-match and failed pages incur no item event. At BRONZE ($0.00128 per contributing page), one page costs $0.00278, five pages cost $0.00790, and ten pages cost $0.01430; these are Actor event prices, not platform compute charges. Other tiers: FREE $0.001472; SILVER $0.0009984; GOLD/PLATINUM/DIAMOND $0.000768 per contributing page. This Actor does not enable residential proxy or paid browser fallback automatically.
Accuracy review tips
In a local test on NASA's Earth page, a literal earth inclusion and planet exclusion returned three selected rows at maxItems: 3; the first row retained the heading “Science in Action for Society.” In a separate two-page local test, seven rows came from NASA and twelve from Wikipedia without a per-page cap. If you need both sources represented in a bounded sample, order the pages and set a maximum larger than the first page's matching count. Always compare a few extracted paragraphs against the page before quoting them elsewhere. HTML can change between runs.
Integrations
Schedule the same input to refresh a bounded reference pack, or export the dataset to a spreadsheet for a human accuracy review. Use the key-value-store DOCX URL to download the combined deliverable; use the dataset's source URL and heading to trace excerpts before publishing them elsewhere. Store the document securely if source material is sensitive.
API usage
Launch with cURL (replace TOKEN with your own Apify API token):
curl -X POST 'https://api.apify.com/v2/acts/automation-lab~rule-guided-web-excerpts-to-word/runs?token=TOKEN' \-H 'Content-Type: application/json' \-d '{"startUrls":[{"url":"https://science.nasa.gov/earth/"}],"includeTerms":["earth"],"maxItems":3}'
In JavaScript, call client.actor('automation-lab/rule-guided-web-excerpts-to-word').call(input) with apify-client, then read the run's default dataset and default key-value store. In Python, use ApifyClient(token).actor('automation-lab/rule-guided-web-excerpts-to-word').call(input) and download EXCERPTS.docx from the returned default key-value-store ID. Keep tokens in secret storage, never in public Task inputs.
MCP
Expose the Actor through Apify MCP in Claude Code:
claude mcp add --transport http apify \'https://mcp.apify.com?tools=automation-lab/rule-guided-web-excerpts-to-word'
For Claude Desktop, Cursor, or VS Code HTTP MCP clients, configure { "mcpServers": { "apify": { "url": "https://mcp.apify.com?tools=automation-lab/rule-guided-web-excerpts-to-word" } } } in the client's MCP settings. Example prompts: Ask: “Extract NASA Earth paragraphs containing earth but not planet and give me the cited DOCX.” Or ask: “Compile web scraping research paragraphs from my Wikipedia URL into a Word file, at most five excerpts.” Each request must supply real URLs; the MCP tool cannot search the web for you.
Limitations and troubleshooting
Only anonymously available server-rendered HTML paragraphs are inspected. JavaScript-only content, sign-in walls, CAPTCHA, PDFs, and pages that redirect are not automatically bypassed. Follow a redirect yourself and supply its final public URL. A page failure is logged and listed in SUMMARY; all-page failure fails the run. No match on a valid page produces an empty dataset and a title-only DOCX. If text is missing, inspect page HTML and try a phrase appearing inside an actual <p> element. The Actor does not recursively crawl links or discover pages from search terms.
Legality and responsible use
Respect source copyright, terms, robots guidance, and applicable privacy rules. A citation does not grant permission to republish copyrighted content. Supply only publicly accessible pages and review excerpts before distributing the compiled document.
Related Actors
HTML Readability to Markdown Converter converts page content into Markdown rather than compiling rule-filtered citations. Schema-Guided Web Data to Excel extracts structured fields for spreadsheets, not a prose Word document.
FAQ
Are excerpts rewritten? No: text is taken from HTML paragraph nodes; the Actor trims outer whitespace but does not summarize or rewrite. Text generated by browser JavaScript is outside this route.
Can I provide a search query? No. Provide exact public page URLs. Search results and whole-site discovery require a separate workflow.
Why is there no output row? None of the paragraphs matched, all matching paragraphs were excluded, or the page did not expose usable HTML paragraphs. Check SUMMARY and the run log; all failed pages cause a nonzero run exit.