Web Fetch
Pricing
from $1.50 / 1,000 fetches
Web Fetch
Real-time web fetch API that turns any URL into clean Markdown, HTML, or links, with automatic JavaScript rendering and anti-bot protection.
Pricing
from $1.50 / 1,000 fetches
Rating
0.0
(0)
Developer
Apify
Maintained by ApifyActor stats
0
Bookmarked
8
Total users
3
Monthly active users
3 days ago
Last modified
Categories
Share
Turn any URL into clean, LLM-ready Markdown, HTML, or links with a single API call. Web Fetch automatically bypasses anti-bot protection, rate limits, and browser-based challenges, then strips out navigation, ads, and boilerplate so you get clean content back - no browser automation infrastructure, no proxy management, no CAPTCHA-solving code to maintain yourself.
Web Fetch is the unblocker your AI agent needs to reach the web.
Because it runs as an always-on Standby Actor on the Apify platform, there's no run to start and no result to poll for - you call one HTTP endpoint and get your answer back in the response, like calling any other API.
Why use Web Fetch to scrape websites?
- Feed LLMs and RAG pipelines clean content. Markdown is far cheaper (in tokens) and easier for a model to parse than raw HTML, making it the ideal input for retrieval-augmented generation, AI agents, and fine-tuning datasets.
- Get past anti-bot walls without building your own bypass logic. Web Fetch automatically handles common blocking mechanisms, including browser-based challenges, so you don't need to reverse-engineer them yourself.
- No infrastructure to run or scale. Because it's a Standby Actor, there's no cold start for a new run and no server to provision - just call the endpoint and get a response.
- Automate anywhere. Call it directly via HTTP, through the Apify API client libraries for Python and JavaScript, or wire it into Apify integrations like Make, Zapier, and n8n.
How to use Web Fetch
-
Authenticate with your Apify API token, either as a
tokenquery parameter or as anAuthorization: Bearer <token>header (the header is more secure since URLs can end up in logs or browser history) - the platform uses it to identify and bill the calling user for each successful fetch. -
Send a GET or POST request to the Actor's Standby URL,
https://web-fetch.apify.actor, with the URL you want to convert - both behave identically:$curl 'https://web-fetch.apify.actor/?url=https://apify.com&formats=markdown,links&token=***'curl -X POST 'https://web-fetch.apify.actor/' \-H 'Content-Type: application/json' \-H 'Authorization: Bearer ***' \-d '{"url": "https://apify.com","formats": ["markdown", "links"]}' -
Read the
text/markdown/html/raw/links/fetch/metadatafields straight out of the JSON response - no polling, no separate results endpoint.
Prefer a regular Actor run instead? Fill in the same fields on the Input tab and hit Start - Web Fetch performs the single fetch, saves the result to the run's dataset, and exits, just like any other Actor.
Running the Actor in standard mode gives you access to all of Apify's native integrations. It also lets you schedule Web Fetch to run on a schedule.
Input
The fields below are exactly what the Standby URL accepts, either as GET query params or as a POST JSON body; the same fields (minus unwrap, which is Standby/MCP only) also appear on the Input tab for a regular Actor run.
| Field | Description |
|---|---|
url | The URL of a web page or resource to fetch (required). |
formats | Which outputs to return, case-insensitive: text, markdown, html, raw, links. If omitted, a single best-effort format is picked based on the content type (see below). |
unwrap | Standby/MCP only - not available (and not needed) for a regular Actor run, which always returns the JSON envelope. Return the first requested format that can actually be produced (or the best-effort default) as a plain HTTP body instead of the JSON envelope; raw is the only format that never fails, regardless of content type. Composes with formats - not mutually exclusive. Default: false. Accepts true/false, 1/0, case-insensitive. |
headers | Extra HTTP headers to send with the request. |
{"url": "https://apify.com","formats": ["markdown", "links"]}
The links format will return a list of all links found on the page, without any deduplication. This is a useful output format for building a crawl queue that fetches more content.
Output
Every request (other than unwrap=true) returns a JSON envelope with fetch (how the page was fetched), metadata (what's on the page), and one key per requested format. If formats is omitted, a single best-effort format is picked based on the fetched content type, so you never need to know it up front: HTML or PDF defaults to markdown, plain text (text/plain) defaults to text, and any other binary content (images, archives, etc.) defaults to raw. raw itself always has to be requested explicitly via formats - it's the heaviest payload of the bunch, so it's never included unless asked for:
{"url": "https://apify.com","fetch": {"loadedUrl": "https://apify.com","loadedTime": "2026-07-27T12:41:41.064Z","httpStatusCode": 200,"contentLengthBytes": 45210,"contentType": "text/html; charset=utf-8"},"metadata": {"canonicalUrl": "https://apify.com","title": "Apify: Full-stack web scraping and data extraction platform","description": "Cloud platform for web scraping, browser automation, AI agents, and data for AI.","author": null,"keywords": null,"languageCode": "en","openGraph": [{ "property": "og:title", "content": "Apify" },{ "property": "og:site_name", "content": "Apify" }],"jsonLd": null,"headers": {"content-type": "text/html; charset=utf-8","content-length": "45210"}},"markdown": "# Apify: Full-stack web scraping and data extraction platform\n\n...","links": ["https://apify.com/pricing", "https://apify.com/store"]}
Each requested format is best-effort: if it can't be produced for the fetched resource's content type (e.g. markdown for an image), it's returned as null rather than failing the whole request. raw can be produced for any content type, so requesting formats: ["raw"] always succeeds regardless of what the URL returns. Beyond that, if the fetch fails outright (bad input, unreachable site, none of the requested formats could be produced, or a timeout), you'll get a flat error with a stable machine-readable code and a matching HTTP status instead:
{"code": "UNSUPPORTED_CONTENT_TYPE","error": "The URL returned a content type (image/jpeg) that cannot be converted to any of the requested formats: markdown. Add \"raw\" to formats (works for any content type) to fetch it as-is."}
Unwrapped response: unwrap=true
Setting unwrap=true returns a single format as a plain HTTP body instead of the JSON envelope, with an appropriate Content-Type (text/markdown, text/html, text/plain, or - for raw - the upstream resource's own content type).
It composes with formats - there's no mutual exclusion:
unwrap=truealone → the best-effort format as a plain body (same content-type-based defaulting as above: HTML/PDF →markdown,text/plain→text, other binary →raw).formats=["markdown"]&unwrap=true→ the Markdown astext/markdown.formats=["raw"]&unwrap=true→ the raw original body, verbatim (works for any content type, including images and archives).
If multiple formats are requested with unwrap=true, they're tried in the given order and the first one that can actually be produced for the fetched content type is returned as the body - e.g. formats=["markdown","raw"] returns Markdown when possible, or the raw body when it's not. raw is treated like any other requested format in that order - it's just the one format that never fails, regardless of content type, so listing it gives you a guaranteed response if everything listed before it comes back empty.
# Returns the PDF's extracted text as markdown (the best-effort default for PDF)curl 'https://web-fetch.apify.actor/?url=https://example.com/file.pdf&unwrap=true&token=***'# Returns the same thing, requested explicitlycurl 'https://web-fetch.apify.actor/?url=https://example.com/file.pdf&unwrap=true&formats=markdown&token=***'# Returns the original PDF file, verbatimcurl 'https://web-fetch.apify.actor/?url=https://example.com/file.pdf&unwrap=true&formats=raw&token=***' -o file.pdf
You only get an UNSUPPORTED_CONTENT_TYPE JSON error once every requested format has failed for the fetched content type (e.g. unwrap=true&formats=markdown on an image, with no raw fallback requested) - add raw to formats to guarantee a response regardless of content type.
Every successful fetch is also saved as an item in the run's dataset - unwrap=true requests get the same shape as any other request (fetch, metadata, and the produced format's field), just with only one format populated - which you can download in JSON, HTML, CSV, or Excel format from the Output tab.
Response
| Field | Type | Description |
|---|---|---|
url | string | The URL you requested. |
text | string | Plain-text rendering of the cleaned article content (if requested). |
markdown | string | Page content converted to clean Markdown (if requested). |
html | string | Cleaned HTML - navigation, footers, and ads stripped, main article content only (if requested). |
raw | string | Original raw HTTP response body (if requested): as-is for HTML/text, base64-encoded for binary. |
links | string[] | Absolute URLs of every link found on the page (if requested). |
fetch.loadedUrl | string | The final URL, after any redirects. |
fetch.loadedTime | string | When the fetch completed, in ISO 8601. |
fetch.httpStatusCode | number | HTTP status code returned by the target site. |
fetch.contentLengthBytes | number | Size of the response body, in bytes. |
fetch.contentType | string | The Content-Type returned by the target site. |
metadata.canonicalUrl | string | The page's <link rel="canonical"> URL, or the loaded URL if absent. |
metadata.title | string | Page <title>. |
metadata.description | string | The <meta name="description"> content, if present. |
metadata.languageCode | string | IETF language tag from the page's <html lang> attribute. |
metadata.openGraph | array | Open Graph, Twitter Card, and related meta tags as { property, content } pairs. |
metadata.jsonLd | array | Parsed <script type="application/ld+json"> blocks, if any. |
metadata.headers | object | All HTTP response headers from the target site. |
Use Web Fetch via MCP
Web Fetch also runs a Model Context Protocol server at /mcp, exposing a single web-fetch tool with the same parameters and JSON output as the main API. Add it to any MCP-compatible client (Claude Code, Cursor, etc.) pointed at the Actor's Standby URL:
$claude mcp add web-fetch https://web-fetch.apify.actor/mcp -t http
Or, since this Actor is published on Apify Store, get the tool automatically by adding apify/web-fetch through the Apify MCP server instead of connecting directly.
How much does it cost to run Web Fetch?
Web Fetch uses pay-per-event pricing: you're charged one fetch event for every successful request, and nothing for failed requests. Check the Actor's page on Apify Store for the current price per event. There's no separate compute or proxy bill to account for.
Tips for getting the best results
- Request only the
formatsyou actually need (e.g. justmarkdown) to keep responses smaller and faster to parse. - For file types other than HTML and PDF (images, archives, etc.), only
formats: ["raw"]will return anything -text/markdown/html/linkscome backnullsince there's no text or markup to extract; this is also why these content types default torawwhenformatsis omitted. - The target URL's response is capped at 10 MB and an overall fetch timeout of 60 seconds (covering both fetching and reading the response) - a response over either limit fails with
UPSTREAM_FETCH_ERROR/FETCH_TIMEOUTinstead of being partially processed.
FAQ
Does Web Fetch handle sites with anti-bot protection?
Yes - it's built to automatically get past common anti-bot mechanisms, rate limits, and browser-based challenges without any extra configuration on your end.
Does Web Fetch render JavaScript-heavy pages like a browser would?
Yes. Web Fetch renders JavaScript, so content that is generated or loaded client-side can be included in the output.
Does Web Fetch support PDFs?
Yes - text, markdown, and html all work for PDF URLs, using extracted text (no layout, images, or tables); raw returns the original PDF bytes, base64-encoded. Link extraction (links) is not yet supported for PDFs.
Is web scraping legal?
Scraping publicly available, non-personal data is generally legal, but what you do with the data afterward matters, and some content is protected by copyright or a site's Terms of Service. If you're unsure, seek legal advice. Read more in this blog post.
Why not use Claude native web fetch?
Native LLM web fetches often rely on limited fetch infrastructure that can't reliably get past anti-bot walls or JavaScript-heavy pages at scale. Web Fetch is built for that job.
What is a web fetch tool?
A web fetch tool retrieves content from a URL on behalf of an AI agent or LLM app, usually converting it to a format the model can parse like Markdown.
Why convert web content to Markdown?
Markdown is the perfect format to feed an LLM. It's lighter than HTML but preserves text structure like titles and headings. Using Markdown instead of HTML lowers your token usage and your AI costs.
Something's not working, or I need a custom feature.
Please open an issue on the Actor's Issues tab with the URL and formats you tried. Custom extraction pipelines (screenshots, structured/LLM extraction, table extraction from PDFs, and more) can be built as a bespoke solution - reach out via Apify's custom solutions page.