AI Web Extractor: URL → Clean Markdown + JSON for LLM/RAG
Pricing
from $3.00 / 1,000 results
AI Web Extractor: URL → Clean Markdown + JSON for LLM/RAG
Turn any URL into clean, LLM-ready Markdown + structured JSON (title, headings, main content, links, metadata, token count). Perfect for RAG pipelines, AI agents, and LLM context.
Pricing
from $3.00 / 1,000 results
Rating
0.0
(0)
Developer
Marvin Eguilos
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
AI Web Extractor — URL → Clean Markdown + JSON for LLM & RAG
Give it a URL. Get back clean, LLM-ready Markdown and structured JSON. Title, headings, main content, links, metadata, and an accurate token count — every page, one tidy result. Built for RAG pipelines, AI agents, and anyone who needs the content of a page without the navbars, ads, cookie banners, and boilerplate.
Feed the open web to your LLM the way it wants to be fed: as clean Markdown, with the token budget already counted.
✨ What it does
- Main-content extraction — Mozilla Readability strips nav, sidebars, ads, and footers so you keep just the article body.
- HTML → Markdown — high-fidelity conversion (GitHub-Flavored Markdown: tables, code blocks, lists, links) via Turndown.
- Structured JSON —
title,description,siteName,lang,byline,headings[],links[],wordCount,tokenCount,fetchedAt. - Accurate token counts — counted with the GPT/
cl100k-family tokenizer so you know exactly how much context each page costs before you send it to a model. - Token budgeting — optional
maxTokenstruncates output to fit your context window. - Robust by design — one bad URL never kills the run. Failed pages return a clean error record (and are never charged).
- Polite crawling — respects
robots.txtby default, sends a real User-Agent, and retries transient errors. - JS rendering when you need it — flip
renderJs: trueto render client-side pages with a headless browser (opt-in, higher compute).
🎯 Use cases
| You want to… | This Actor gives you… |
|---|---|
| Build a RAG knowledge base | Clean Markdown chunks with token counts, ready to embed. |
| Give an AI agent web context | Structured JSON your agent can reason over — no HTML noise. |
| Scrape docs / blogs to Markdown | Publishable Markdown you can drop straight into a repo or wiki. |
| Feed an LLM prompt | Pre-counted tokens so you never blow the context window. |
| Archive / snapshot pages | Portable, diff-friendly Markdown + metadata. |
📥 Input
| Field | Type | Default | Description |
|---|---|---|---|
urls | string[] | — (required) | One or more URLs. Each successful page is one result. |
outputFormat | both | markdown | json | both | Include Markdown, the JSON fields, or both. |
onlyMainContent | boolean | true | Strip nav/ads/sidebars with Readability. |
includeLinks | boolean | true | Include extracted absolute links + anchor text. |
maxPagesPerRun | integer | 1000 | Safety cap on URLs processed per run. |
renderJs | boolean | false | Render JS-heavy pages with a headless browser (higher cost). |
respectRobotsTxt | boolean | true | Skip URLs disallowed by the site's robots.txt. |
maxTokens | integer | 0 | Truncate Markdown to ~N tokens (0 = no limit). |
Example input
{"urls": ["https://example.com","https://en.wikipedia.org/wiki/Markdown"],"outputFormat": "both","onlyMainContent": true,"includeLinks": true,"maxTokens": 0}
📤 Output
One dataset item per URL. Successful example:
{"url": "https://example.com/","finalUrl": "https://example.com/","statusCode": 200,"title": "Example Domain","description": null,"siteName": null,"lang": "en","byline": null,"excerpt": "This domain is for use in documentation examples...","headings": [],"wordCount": 17,"tokenCount": 29,"fetchedAt": "2026-07-19T02:17:19.954Z","markdown": "This domain is for use in documentation examples without needing permission. Avoid use in operations.\n\n[Learn more](https://iana.org/domains/example)","links": [{ "url": "https://iana.org/domains/example", "text": "Learn more" }]}
Failed URL (returned, not charged):
{"url": "https://not-a-real-domain-xyz.com/","finalUrl": "https://not-a-real-domain-xyz.com/","statusCode": null,"error": "getaddrinfo ENOTFOUND not-a-real-domain-xyz.com","fetchedAt": "2026-07-19T02:17:19.831Z"}
🔌 Call it from code (Apify API)
curl -X POST "https://api.apify.com/v2/acts/YOUR_USERNAME~ai-web-extractor/run-sync-get-dataset-items?token=YOUR_TOKEN" \-H "Content-Type: application/json" \-d '{"urls":["https://en.wikipedia.org/wiki/Markdown"]}'
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: 'YOUR_TOKEN' });const { defaultDatasetId } = await client.actor('YOUR_USERNAME/ai-web-extractor').call({ urls: ['https://en.wikipedia.org/wiki/Markdown'] });const { items } = await client.dataset(defaultDatasetId).listItems();console.log(items[0].markdown);
💸 Pricing (Pay-Per-Event)
| Event | Price |
|---|---|
| Actor start | $0.05 per run |
| Extracted page | $0.003 per successful page |
- You only pay for pages that succeed — failed URLs are never charged.
- 🎁 Free tier: free-plan users' platform usage is covered by Apify, so you can try it and run small jobs at no cost before scaling up.
- Cleaner structured JSON + accurate token counts than typical single-purpose "URL to Markdown" tools — at the same market price point.
⚖️ Acceptable use
This is a general-purpose format-conversion tool: you supply the URLs and are responsible for having the right to crawl and use the content you submit. By default the Actor respects robots.txt and identifies itself with a descriptive User-Agent. It does not target any single platform's private API and does not harvest personal data as a feature. Please crawl responsibly and comply with each site's terms of service and applicable law.
🧱 Under the hood
Node.js · Crawlee (Cheerio + optional Playwright) · @mozilla/readability · Turndown (+ GFM) · gpt-tokenizer · Apify SDK. Stateless — nothing is stored between runs.