Webpage Text & Markdown Extractor
Pricing
Pay per usage
Webpage Text & Markdown Extractor
Convert up to 1,000 webpage URLs into clean readable text, Markdown, metadata, canonical URLs, images, and deduplicated links for AI and content workflows.
Pricing
Pay per usage
Rating
0.0
(0)
Developer
snapperwapper
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Convert public webpage URLs into clean readable text and Markdown for RAG, LLM, research, content migration, and data-enrichment workflows.
Output
- Readability-focused text and Markdown
- Title, description, canonical URL, Open Graph image, and language
- Final URL and HTTP status
- Word and character counts
- Optional deduplicated absolute links with anchor text
- Structured HTTP, network, timeout, and unsupported-content errors
Example input
{"urls":["https://example.com"],"maxConcurrency":10,"timeoutSecs":20,"maxChars":100000,"includeLinks":true}
Good uses
- Preparing webpage content for RAG and embeddings
- Feeding readable articles to LLM agents
- Content migration and archiving
- Research and monitoring pipelines
- Extracting text without browser-rendering expense
Scope and limitations
The Actor uses normal HTTP requests and readability extraction. It does not execute client-side JavaScript, bypass logins, or defeat anti-bot controls. Pages whose useful content is rendered only in a browser may return limited text. Only process pages you may access, and respect site terms and applicable law.
Pricing
Published as standard Apify pay-per-usage while traffic is validated. It requires no browser, proxy, login, or paid external API by default.