AI Web Extractor: URL → Clean Markdown + JSON for LLM/RAG avatar

AI Web Extractor: URL → Clean Markdown + JSON for LLM/RAG

Pricing

from $3.00 / 1,000 results

Go to Apify Store
AI Web Extractor: URL → Clean Markdown + JSON for LLM/RAG

AI Web Extractor: URL → Clean Markdown + JSON for LLM/RAG

Turn any URL into clean, LLM-ready Markdown + structured JSON (title, headings, main content, links, metadata, token count). Perfect for RAG pipelines, AI agents, and LLM context.

Pricing

from $3.00 / 1,000 results

Rating

0.0

(0)

Developer

Marvin Eguilos

Marvin Eguilos

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

AI Web Extractor — URL → Clean Markdown + JSON for LLM & RAG

Give it a URL. Get back clean, LLM-ready Markdown and structured JSON. Title, headings, main content, links, metadata, and an accurate token count — every page, one tidy result. Built for RAG pipelines, AI agents, and anyone who needs the content of a page without the navbars, ads, cookie banners, and boilerplate.

Feed the open web to your LLM the way it wants to be fed: as clean Markdown, with the token budget already counted.


✨ What it does

  • Main-content extraction — Mozilla Readability strips nav, sidebars, ads, and footers so you keep just the article body.
  • HTML → Markdown — high-fidelity conversion (GitHub-Flavored Markdown: tables, code blocks, lists, links) via Turndown.
  • Structured JSONtitle, description, siteName, lang, byline, headings[], links[], wordCount, tokenCount, fetchedAt.
  • Accurate token counts — counted with the GPT/cl100k-family tokenizer so you know exactly how much context each page costs before you send it to a model.
  • Token budgeting — optional maxTokens truncates output to fit your context window.
  • Robust by design — one bad URL never kills the run. Failed pages return a clean error record (and are never charged).
  • Polite crawling — respects robots.txt by default, sends a real User-Agent, and retries transient errors.
  • JS rendering when you need it — flip renderJs: true to render client-side pages with a headless browser (opt-in, higher compute).

🎯 Use cases

You want to…This Actor gives you…
Build a RAG knowledge baseClean Markdown chunks with token counts, ready to embed.
Give an AI agent web contextStructured JSON your agent can reason over — no HTML noise.
Scrape docs / blogs to MarkdownPublishable Markdown you can drop straight into a repo or wiki.
Feed an LLM promptPre-counted tokens so you never blow the context window.
Archive / snapshot pagesPortable, diff-friendly Markdown + metadata.

📥 Input

FieldTypeDefaultDescription
urlsstring[](required)One or more URLs. Each successful page is one result.
outputFormatboth | markdown | jsonbothInclude Markdown, the JSON fields, or both.
onlyMainContentbooleantrueStrip nav/ads/sidebars with Readability.
includeLinksbooleantrueInclude extracted absolute links + anchor text.
maxPagesPerRuninteger1000Safety cap on URLs processed per run.
renderJsbooleanfalseRender JS-heavy pages with a headless browser (higher cost).
respectRobotsTxtbooleantrueSkip URLs disallowed by the site's robots.txt.
maxTokensinteger0Truncate Markdown to ~N tokens (0 = no limit).

Example input

{
"urls": [
"https://example.com",
"https://en.wikipedia.org/wiki/Markdown"
],
"outputFormat": "both",
"onlyMainContent": true,
"includeLinks": true,
"maxTokens": 0
}

📤 Output

One dataset item per URL. Successful example:

{
"url": "https://example.com/",
"finalUrl": "https://example.com/",
"statusCode": 200,
"title": "Example Domain",
"description": null,
"siteName": null,
"lang": "en",
"byline": null,
"excerpt": "This domain is for use in documentation examples...",
"headings": [],
"wordCount": 17,
"tokenCount": 29,
"fetchedAt": "2026-07-19T02:17:19.954Z",
"markdown": "This domain is for use in documentation examples without needing permission. Avoid use in operations.\n\n[Learn more](https://iana.org/domains/example)",
"links": [
{ "url": "https://iana.org/domains/example", "text": "Learn more" }
]
}

Failed URL (returned, not charged):

{
"url": "https://not-a-real-domain-xyz.com/",
"finalUrl": "https://not-a-real-domain-xyz.com/",
"statusCode": null,
"error": "getaddrinfo ENOTFOUND not-a-real-domain-xyz.com",
"fetchedAt": "2026-07-19T02:17:19.831Z"
}

🔌 Call it from code (Apify API)

curl -X POST "https://api.apify.com/v2/acts/YOUR_USERNAME~ai-web-extractor/run-sync-get-dataset-items?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"urls":["https://en.wikipedia.org/wiki/Markdown"]}'
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const { defaultDatasetId } = await client
.actor('YOUR_USERNAME/ai-web-extractor')
.call({ urls: ['https://en.wikipedia.org/wiki/Markdown'] });
const { items } = await client.dataset(defaultDatasetId).listItems();
console.log(items[0].markdown);

💸 Pricing (Pay-Per-Event)

EventPrice
Actor start$0.05 per run
Extracted page$0.003 per successful page
  • You only pay for pages that succeed — failed URLs are never charged.
  • 🎁 Free tier: free-plan users' platform usage is covered by Apify, so you can try it and run small jobs at no cost before scaling up.
  • Cleaner structured JSON + accurate token counts than typical single-purpose "URL to Markdown" tools — at the same market price point.

⚖️ Acceptable use

This is a general-purpose format-conversion tool: you supply the URLs and are responsible for having the right to crawl and use the content you submit. By default the Actor respects robots.txt and identifies itself with a descriptive User-Agent. It does not target any single platform's private API and does not harvest personal data as a feature. Please crawl responsibly and comply with each site's terms of service and applicable law.


🧱 Under the hood

Node.js · Crawlee (Cheerio + optional Playwright) · @mozilla/readability · Turndown (+ GFM) · gpt-tokenizer · Apify SDK. Stateless — nothing is stored between runs.