AI Agent Web Fetcher avatar

AI Agent Web Fetcher

Pricing

from $1.40 / 1,000 results

Go to Apify Store
AI Agent Web Fetcher

AI Agent Web Fetcher

An advanced web fetcher that can fetch almost all websites and convert them to LLM-friendly Markdown format. Perfect for AI agents, RAG systems, and integration with search actors.

Pricing

from $1.40 / 1,000 results

Rating

0.0

(0)

Developer

Abot API

Abot API

Maintained by Community

Actor stats

1

Bookmarked

23

Total users

4

Monthly active users

2 days ago

Last modified

Share

AI Web Fetcher: Any Page to Clean, LLM-Ready Markdown

AI Web Fetcher renders any URL in a full browser and returns clean, LLM-ready Markdown, plus the page's title, description and author when the page publishes them. Fetch one URL on demand, or point it at another actor's output dataset and it works through every link in a field you choose, up to a cap you set. Read results back from the dataset (with an Overview or Markdown Content view), the key-value store, the API, or call the actor's own persistent Standby HTTP endpoint for real-time fetches from your own backend or agent.

Why This Scraper?

  • Real browser rendering. Every page loads in a full headless browser, so JavaScript-rendered content, not just the raw HTML response, ends up in the output.
  • Clean Markdown, not raw HTML. Scripts, navigation, ads, forms and hidden elements are stripped before conversion, so downstream LLMs and agents get readable text instead of markup noise.
  • Automatic retry on a fresh connection. A page that fails or times out on the first attempt is retried in a second browser session, up to two more times within the same per-page time budget; with Apify Proxy on, each retry runs on a newly established proxy session.
  • Two ways to feed it URLs. Fetch one URL directly, or point it at another actor's output dataset and it pulls the link field from every item, up to a cap you set.
  • Built for agent pipelines. Chain it after a search actor from the Integrations tab (it picks up that run's dataset automatically) or call the persistent Standby HTTP endpoint for real-time, low-latency fetches from your own backend or agent framework.
  • Metadata alongside the content. Title, description and author, when the page publishes them, come back with every fetched page.
  • Fails loudly, not quietly. A page that cannot be fetched within its time budget is reported as a failure in the run's status and log, never as a dataset item with empty content.

Use Cases

  • AI agent and RAG builders: convert any page into LLM-ready Markdown as live context for an agent, chatbot or retrieval index.
  • Search-augmented apps: chain this actor after a web search actor so every result link is fetched and converted automatically, without writing glue code.
  • Research and content teams: pull readable text from articles, docs and product pages without manually stripping ads and navigation.
  • Backend and automation builders: call the Standby HTTP endpoint directly from your own service for on-demand, real-time page fetches.

Data You Get

Sample shape: values are illustrative placeholders, not from a live record.

FieldExample
url"https://example.com/articles/global-markets-update"
markdown (page content)"# Global Markets Update\n\nA summary of today's session..."
markdown (headings)"## Today's Highlights"
markdown (paragraphs)"Plain paragraph text carried over from the page body."
markdown (lists)"- First point\n- Second point"
markdown (links)"[Read the full report](#)"
markdown (tables)"| Metric | Value |\n| --- | --- |\n| Sample | 42 |"
markdown (images)"![Chart of quarterly results](#)"
markdown (emphasis)"**bold** and *italic* text kept as written"
markdown (blockquotes and code)"> A quoted passage" and `inline code`
markdown (noise removed)scripts, navigation, ads, forms and hidden elements never reach this field
markdown (empty string)possible when a page loads but nothing readable is left after cleanup
metadata{ "title": "...", "description": "...", "author": "..." }
metadata.title"Global Markets Update"
metadata.description"A daily summary of major market moves."
metadata.author"Jane Analyst"

metadata.title, metadata.description and metadata.author come back as empty strings, not missing keys, when the page does not publish that piece of information.

How to Use

  1. Choose a mode: fetch a single URL, or point at a dataset from another actor's run to fetch every URL listed in it.
  2. Fill in the URL (or the dataset ID and the field name that holds the URL), and turn on a connection for pages that block plain requests.
  3. Set Max URLs to cap a batch run, then click Start.
  4. Read the Markdown and metadata from the dataset (Overview or Markdown Content view), the key-value store (single URL mode), or the API.

Fetch a single URL:

{
"url": "https://example.com/articles/global-markets-update"
}

A JavaScript-heavy page, waiting for the network to settle:

{
"url": "https://example.com/dashboard",
"waitUntil": "networkidle"
}

A page that needs a residential connection:

{
"url": "https://example.com/product/4821",
"waitUntil": "load",
"proxy": {
"useApifyProxy": true,
"apifyProxyGroups": ["RESIDENTIAL"]
}
}

Batch mode from another actor's dataset:

{
"datasetId": "YOUR_DATASET_ID",
"urlField": "link",
"maxUrls": 25
}

Run it from your code

Python:

from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("abotapi/ai-fetch-python").call(
run_input={"url": "https://example.com/articles/global-markets-update"}
)
for page in client.dataset(run["defaultDatasetId"]).iterate_items():
print(page["url"], page["metadata"]["title"])

JavaScript:

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' });
const run = await client.actor('abotapi/ai-fetch-python').call({
url: 'https://example.com/articles/global-markets-update',
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();

Or connect it to Make, Zapier, n8n, Google Sheets or webhooks from the Integrations tab. Chained after another actor's run from the Integrations tab, this actor automatically receives that run's dataset ID, so you usually only need to set Url Field if the URLs sit in a field other than link.

How a run handles a slow or blocked page

Each URL gets its own 90 second budget. The actor first renders the page in a full headless browser; if that attempt fails or times out, it retries the page in a second headless browser session, up to two times, as long as the 90 second budget allows. With Apify Proxy on, each retry uses a newly established proxy session. A dataset item is written only for a successful fetch, so a failed URL never creates a billable dataset row, it is reported through the run's status message and log instead. Max URLs (1 to 100) caps how many links a batch run reads from the dataset; it has no effect in single URL mode.

Input Parameters

ParameterTypeDefaultDescription
urlstringhttps://www.google.comThe webpage to fetch (single URL mode).
datasetIdstring(none)Dataset ID to pull URLs from (batch mode); leave empty for single URL mode.
urlFieldstringlinkField in each dataset item that holds the URL.
maxUrlsinteger10Maximum URLs to fetch from the dataset, from 1 to 100.
proxyobject(none)Optional connection settings for the fetch; enable Apify Proxy with a residential group for pages that block plain requests.
waitUntilstringdomcontentloadedNavigation wait condition: load, domcontentloaded, networkidle, or commit.
debugbooleanfalseReserved for debug output; has no effect on the current run or its output.

Output Example

Sample shape: values are illustrative placeholders, not from a live record.

{
"url": "https://example.com/articles/global-markets-update",
"markdown": "# Global Markets Update\n\nA summary of today's session, including sector moves and a look ahead.\n\n## Today's Highlights\n\n- First point\n- Second point\n\nRead the [full report](https://example.com/reports/2026-q3).",
"metadata": {
"title": "Global Markets Update",
"description": "A daily summary of major market moves.",
"author": "Jane Analyst"
}
}

Plan Requirement

Without a proxy, the actor fetches over a direct connection, which works for most publicly reachable pages. For sites that block plain requests, turn on Apify Proxy with a residential group under Proxy Configuration; the actor then retries a blocked or slow page on a fresh proxy session before giving up. Standby mode needs enough concurrent capacity to hold open the number of requests you expect, adjust memory under the actor's Standby settings if you plan sustained real-time traffic.

FAQ

How much does it cost?

You pay for what your run uses; the Pricing tab on the Store page shows the current rate. Batch runs are capped by Max URLs, and a failed fetch never creates a dataset item, so a blocked or slow page does not add a billable row.

This actor renders whatever page you point it at, the same way a browser would; it does not target one site's data specifically. You are responsible for how you use the fetched content: check the target site's terms and the laws that apply to you, especially before republishing paywalled or personal content.

Can I run it on a schedule?

Yes, schedule it from the Schedules tab like any other actor. Every run is independent: it always (re)fetches the URLs or dataset you give it, there is no built-in change tracking or incremental memory, so add your own de-duplication if you only want new pages.

Why is the Markdown mostly empty or missing text I can see in my browser?

The default wait condition stops as soon as the initial page loads, so content a page injects later (infinite scroll, chat widgets, some single-page apps) may not be present yet; switch Wait Until to networkidle for pages like that. Scripts, navigation, ads, forms and hidden elements are also deliberately stripped before conversion, so a navigation-heavy page will look sparse in Markdown by design.

Why did my run fail instead of returning an empty dataset?

If a page cannot be fetched within its time budget, through both the primary and fallback browser attempts, that URL fails on its own without writing anything to the dataset, and the run's status message reports how many URLs succeeded and failed. If every URL in a run fails, the whole run fails rather than finishing with an empty dataset, so "nothing could be read" is never confused with "no data on this page". One exception: in batch mode, if no dataset item has a value in the Url Field you set, there is nothing to fetch, so the run finishes with an empty dataset and a warning in the log; check that Url Field matches your dataset.

Can I use it with AI agents or MCP?

Yes. Call it like any other Apify actor from your agent framework or MCP client through the API or client libraries, or enable Standby mode and call its /fetch endpoint directly (GET or POST, with url and optional waitUntil) for real-time, low-latency fetches without starting a new run each time. Standby requests use a direct connection; for pages that need Apify Proxy, start a regular run instead.

๐Ÿ”— Want more AI data?

Pair this actor with these related scrapers from the same team:

๐Ÿงฉ Doc To Markdown MCP Server
An MCP server that converts documents to clean Markdown. Convert PDFs, Word docs, Excel...
๐Ÿค– AI Search Tool
Give your AI agents real-world knowledge. This Actor provides high-quality web, news...
๐Ÿค– Web Scraper For Llms
Stealth web scraping engine built for LLMs. Converts any web page to clean markdown or...
๐Ÿค– Doc To Markdown
Convert documents (PDF, Word, PowerPoint, Excel, HTML, images) to clean Markdown...
๐Ÿ“ฑ Universal Media Extractor
Extract videos, audio, and metadata from 1000+ websites including YouTube, TikTok...
๐Ÿ“‡ Yellow Pages NZ
From $0.8/1K. Scrapes business listings from Yellow.co.nz (New Zealand Yellow Pages)...

๐Ÿ‘‰ Browse all abotapi scrapers

๐Ÿ’ฌ Support & custom scrapers

  • ๐Ÿž Found a bug or a missing field? Open a ticket on the Issues tab. We usually reply within hours.
  • ๐Ÿ› ๏ธ Need another site, extra fields or a private build? Email abotapi@proton.me or message Telegram @abotapi.
  • โญ Enjoying it? A quick review on the actor page helps other users find it.