MCP Web Research Agent: Search to JSON Schema
Pricing
from $100.00 / 1,000 web searches
MCP Web Research Agent: Search to JSON Schema
Ask a research question and name your fields. The agent searches the web, reads the top pages and returns every item it finds as JSON in your schema, with source URL and a per-field confidence label. Callable as a tool via the Apify MCP server.
Pricing
from $100.00 / 1,000 web searches
Rating
0.0
(0)
Developer
Turgay NANTA
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
MCP Web Research Agent: Web Search to Structured JSON
AI agents are good at reasoning and bad at reading raw web pages. When an assistant searches the web, it usually receives long markdown or HTML, spends tokens reading it, and then gives you facts in a different shape every time. That is hard to put into a table, a database or the next step of a workflow.
Ask a research question and name the fields you want. This Actor searches the web, reads the top result pages, and returns every item it finds as JSON in exactly your schema, with a source URL and a confidence label for every field. One call, one predictable object, ready for an agent, a RAG pipeline or a spreadsheet.
- Question in, table out: "Top 5 open-source vector databases" with fields
name,license,star_count,featuresreturns one record per database found. - Many items per page: a "top 10" article becomes 10 records, not one blurb.
- Your schema, your types:
string,number,boolean,array,object, with a description per field to guide extraction. - Grounded, not invented: values must come from the source page; every field is labelled
high,grounded,unverifiedornull. - Merged across sources: the same item found on several pages is combined into one record, gaps filled from the other pages.
- Built for agents: callable as a tool through the Apify MCP server, or through the API from any workflow.
- Clear cost: a 5-source research run costs up to $0.3501 in Actor fees (see Pricing).
Example output (shortened)
Input: query = "Top 5 open-source vector databases", schema name, license, star_count (number), features (array).
{"query": "Top 5 open-source vector databases","schema": ["name", "license", "star_count", "features"],"result_count": 2,"results": [{"name": "Milvus","license": "Apache 2.0","star_count": 20000,"features": ["hybrid search", "GPU indexing"],"_source_url": "https://example-blog.com/vector-databases","_confidence": {"name": "high", "license": "high", "star_count": "high", "features": "grounded"}},{"name": "Qdrant","license": "Apache 2.0","star_count": null,"features": ["payload filtering"],"_source_url": "https://example-blog.com/vector-databases","_confidence": {"name": "high", "license": "high", "star_count": "null", "features": "grounded"}}],"sources": [{"title": "Best vector databases", "url": "https://example-blog.com/vector-databases"}]}
Illustrative example showing the format; the values and URL are made up. A real run returns what its sources say, and null where they say nothing.
At a glance
| Input | A research question (query) + the fields you want (schema) + how many sources to read (maxResults, 1-50) |
| Output | One dataset item: query, schema, result_count, results (your records), sources |
| Per record | Your fields + _source_url + _confidence per field |
| Search | Apify RAG Web Browser (web search + page content as markdown) |
| Extraction | Language model, grounded in each source page, one object per distinct item |
| Languages | Text values in English, Turkish, German, Spanish or French |
| Pricing | Per search + per source that yields data; see Pricing |
| Runs as | A classic Actor (Console, API, schedules, n8n, Make, Zapier) and an MCP tool through the Apify MCP server |
Quick start
- Type a Research question, for example
Best project management tools for small teams. - Edit Fields to extract: a list of field names, or objects with
name,typeanddescription. The form starts with an example schema. - Set Maximum results (how many search result pages to read; default 5).
- Click Start. The Dataset tab has one item with all records in
results.
An empty run (no question) searches nothing, charges nothing and returns one row telling you what to fill in.
Why structured research
| Plain web search tool for an agent | This Actor | |
|---|---|---|
| What the agent receives | Page text or markdown | JSON records in your schema |
| Shape of the answer | Different every time | Same keys, same types, every run |
| Many items on one page | Agent must parse them | One record per item |
| Duplicates across sources | Agent must merge | Merged by name, gaps filled |
| Can the agent trust a value? | Unknown | Per-field label: found in source or not |
| Token cost for the agent | Full pages | Only the extracted fields |
Use cases
AI agent builders
Give your assistant one tool, "structured web research", instead of search plus scraping plus parsing. The agent passes a question and a schema, and gets a JSON object it can reason over or store.
RAG and knowledge base teams
Turn web research into clean rows before indexing: each record carries its _source_url, so answers can cite where a fact came from.
Market and competitor research
"Pricing of CRM tools for small businesses" with name, starting_price (number), free_plan (boolean), integrations (array). Read 10 sources, get one merged comparison table.
Sales and business development
"Logistics software companies in Germany" with company, city, website, focus. A quick first list from public articles and directories, each row with its source.
Product managers and analysts
"Open-source alternatives to X" with name, license, language, last_release. Fast landscape tables for a decision document.
Journalists and content teams
Collect facts on a topic into a table with source links, then verify the unverified fields before publishing.
Procurement and operations
"Industrial label printers under 2000 euro" with model, brand, price (number), print_width_mm (number). A first shortlist with sources for each value.
Worked example: a tool comparison in one call
You need a comparison of project management tools for a team decision.
- Question:
best project management tools for small teams 2026 - Schema:
[{"name": "tool", "type": "string", "description": "product name"},{"name": "starting_price_usd", "type": "number", "description": "lowest paid plan per user per month"},{"name": "free_plan", "type": "boolean"},{"name": "key_features", "type": "array"}]
- maxResults: 8, so the Actor reads eight result pages.
- Run. Each page is read separately; a "10 best tools" article produces up to ten records.
- Merging. "Asana" found on three pages becomes one record; if one page had no price but another did, the price is filled in.
- Review. Sort by
tool, check any field whose_confidenceisunverified, and open_source_urlfor the ones you want to confirm.
Actor fees if all 8 sources yield data: $0.0001 + $0.10 + 8 x $0.05 = $0.5001, plus the RAG Web Browser usage on your account.
How it works
- Search. The Actor runs Apify's RAG Web Browser with your question and
maxResults. It returns the top search results together with each page's content as markdown. - Read each source. For every source page, the first 12,000 characters of its content, your question and your field list go to the language model.
- Extract many items. The model is told to return a JSON array with one object per distinct relevant item on the page, using only the page content, and
nullfor anything the page does not state. Text values are written in your chosen language; field names are never translated. - Parse robustly. The first complete JSON value is taken from the reply, even if the model wrapped it in code fences or added text after it. A reply without valid JSON is logged and that source is skipped.
- Coerce types. Numbers are cleaned (
"20,000 stars"to20000,"1.299,90"to1299.9), booleans recognized from words (yes,available,in stock,sold out...), single values wrapped into arrays where an array is expected. - Label confidence. Each value is checked against the source text (see Confidence labels).
- Merge duplicates. Records with the same value in the first text field (case-insensitive) are merged; empty fields are filled from later copies, with their confidence labels.
- Return one object with
resultsand the list ofsourcesthat were read.
Input
| Field | Type | Default | Description |
|---|---|---|---|
query | string | none | The research question or topic. Written like a web search. |
schema | array or object | example schema | The fields you want (formats below). Up to 25 fields. |
maxResults | integer | 5 | How many search result pages to read and structure, 1-50. |
language | string | English | Language of text values: English, Türkçe, Deutsch, Español, Français. |
model | string | claude-haiku-4-5-20251001 | Advanced. claude-sonnet-4-6 for harder extraction. |
query has no default on purpose: every run with a question starts a web search on your account, so nothing is searched until you ask something.
Schema formats
Names only (all fields become text):
["name", "website", "country"]
Objects with type and description (recommended):
[{"name": "name", "type": "string", "description": "product or company name"},{"name": "price", "type": "number", "description": "monthly price in USD"},{"name": "open_source", "type": "boolean"},{"name": "features", "type": "array"}]
Map:
{"name": "product name", "price": {"type": "number", "description": "monthly price"}}
Allowed types: string, number, boolean, array, object. Any other type (including integer) falls back to string, so use number for counts and years.
Output
The dataset contains one item per run:
| Key | Content |
|---|---|
query | Your question |
schema | The field names used |
result_count | Number of records after merging |
results | The records: your fields + _source_url + _confidence |
sources | Title and URL of every page the search returned |
Confidence labels
| Label | Meaning |
|---|---|
high | A text or number value was found in the source text (numbers compared digit by digit, at least two digits) |
grounded | A boolean or array value is supported by the source: a matching word for true/false (such as "available" or "sold out"), or at least one array element present in the text |
unverified | The value is not found in the source; it may be inferred or paraphrased. object fields are always unverified because free-form objects cannot be checked reliably |
null | The field is empty |
For an agent, a simple rule works well: use high and grounded values directly, and verify or drop unverified ones.
When there is nothing to return
- The search returns no pages: the item is a
_summarywith a note; start and search fees apply. - Pages were read but none contained relevant items: those sources are not charged; if no record at all could be produced, the run ends with an error message.
Pricing
This Actor charges per event. Current prices are on the Pricing tab; at the time of writing:
| Event | When | Price |
|---|---|---|
actor-start | Once per run | $0.0001 |
search-performed | Once per run, after the web search succeeds | $0.10 |
page-structured | Per source page that yielded at least one record | $0.05 |
Source pages with no relevant item, or whose extraction failed, are not charged.
Examples (Actor fees)
| Run | Calculation | Total (maximum) |
|---|---|---|
| 1 source | $0.0001 + $0.10 + 1 x $0.05 | $0.1501 |
| 5 sources (default) | $0.0001 + $0.10 + 5 x $0.05 | $0.3501 |
| 10 sources | $0.0001 + $0.10 + 10 x $0.05 | $0.6001 |
| 50 sources | $0.0001 + $0.10 + 50 x $0.05 | $2.6001 |
| Empty input | nothing searched | $0 |
Search is billed separately: the web search runs as an Apify RAG Web Browser run on your account, and that run's usage is billed to you at its own price.
Using it from an AI agent (MCP)
The Apify MCP server exposes Store Actors as tools for MCP clients such as Claude Desktop, Claude Code, Cursor and other agent frameworks. Add the Apify MCP server to your client with your Apify token and include this Actor (enezli/mcp-research-agent) in the tools it loads. The agent can then call it with:
{"query": "EU regulations on AI in hiring","schema": [{"name": "regulation", "type": "string"},{"name": "year", "type": "number"},{"name": "applies_to", "type": "string"},{"name": "key_obligations", "type": "array"}],"maxResults": 5}
and receive the single JSON object described above. Tips for agent prompts:
- Ask the agent to design a small schema (3-8 fields) that answers the user's question.
- Tell the agent to cite
_source_urlfor each fact it reports. - Tell the agent to treat
unverifiedfields as uncertain.
Integrations
API
curl -X POST "https://api.apify.com/v2/acts/enezli~mcp-research-agent/run-sync-get-dataset-items?token=YOUR_TOKEN" \-H "Content-Type: application/json" \-d '{"query":"Top 5 open-source vector databases","schema":["name","license",{"name":"star_count","type":"number"}],"maxResults":5}'
Python
from apify_client import ApifyClientclient = ApifyClient("YOUR_TOKEN")run = client.actor("enezli/mcp-research-agent").call(run_input={"query": "Top 5 open-source vector databases","schema": [{"name": "name", "type": "string"}, {"name": "license", "type": "string"}],"maxResults": 5,})item = next(client.dataset(run["defaultDatasetId"]).iterate_items())for rec in item.get("results", []):print(rec["name"], rec["license"], rec["_source_url"])
Schedules
Save a question and schema as a task and schedule it weekly to refresh a comparison table or a market list.
n8n, Make, Zapier
Run Actor → Get dataset items → split results into rows → write to Google Sheets, Airtable or Notion. Route rows with unverified fields to a review step.
Workflow recipes
Weekly market table. Save "pricing of [category] tools" with a fixed schema as a task, schedule it weekly, and append results to a Google Sheet. New tools and changed prices show up as new rows you can compare over time.
Agent answer with citations. Let your assistant call the Actor, then answer the user only from high and grounded values, citing _source_url next to each fact. Ask it to say explicitly when a field was null everywhere.
Lead list seed. Ask "[industry] companies in [city]" with company, website, city, focus. Treat the output as a first list to verify, not a finished database; each row tells you where it came from.
Due diligence snapshot. Before a meeting, research "[company] funding, products and customers" with a small schema, and keep the JSON with its sources in your notes.
Tips for better results
- Write the question like a search query that would return list-style pages ("best", "top", "alternatives to", "companies in").
- Add a
descriptionto each field, especially units ("price per user per month in USD"). - Make the first text field the item's name. It is used to merge duplicates across sources.
- Use
numberfor numeric fields, including counts and years. - Start with 5 sources, raise
maxResultsonly if the table is too thin. - Keep schemas short. Many fields and long pages can exceed the model's reply size for one source; that source is then skipped and not charged.
- Use the stronger model for technical or dense sources.
Limitations
- Search results decide coverage. The Actor reads what the web search returns for your question; it does not browse beyond those pages.
- First 12,000 characters per page. Items further down a very long page may be missed.
- Reply size per source. Very long lists on one page may be cut; a cut reply is skipped and not charged.
- Merging uses the first text field. Items with different spellings ("PostgreSQL" vs "Postgres") stay separate.
- Confidence labels are heuristic. A correct paraphrase can be
unverified; they indicate where to look, not truth. - No logins or paywalls. Content behind a login is not available to the search.
- Maximum 25 fields and 50 sources per run.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| One row saying "No research question given" | query was empty | Add a question |
_summary with "no results" | The search returned no pages | Use a broader or differently worded question |
| Run failed: "No source could be structured" | None of the pages contained relevant items, or extraction failed on all of them | Rephrase the question toward list-style pages; simplify the schema |
| Run failed: "Web search failed" | The search step failed | Run again; the search fee is not charged when the search fails |
Many null values | Sources do not state those facts | Adjust fields or add descriptions |
| Duplicate items with slightly different names | Merging is exact on the first text field | Clean up in your sheet, or describe the name field more precisely |
A field you typed as integer comes back as text | integer is not an allowed type here | Use number |
FAQ
Is this an MCP server? It is an Apify Actor that AI agents can call as a tool through the Apify MCP server. It also runs as a normal Actor from the Console, API and automation tools.
Which search engine does it use? It uses Apify's RAG Web Browser, which performs the web search and returns page content.
Can it return several items from one page? Yes. Each source page can produce many records, one per distinct item.
Does it make things up?
The model is instructed to use only the source content and to write null otherwise. Every value is also checked against the source text and labelled, so an agent can tell grounded values from possible inferences.
Can I get the answer in another language? Yes. Text values are written in the chosen output language. Field names and proper nouns stay as they are.
What if a source fails? It is logged and skipped; the other sources continue. Skipped sources are not charged.
How is this different from a normal web scraper? A scraper returns the content of pages you choose. This Actor starts from a question, finds the pages, and returns only the facts you asked for, in your schema.
Can I use it for a RAG pipeline? Yes. Each record has its source URL, so you can store structured facts with citations instead of raw page chunks.
How long does a run take? It depends on the search and on the number of sources; sources are processed one after another.
Is my data stored? Only in your own Apify storage. Source content is sent to the language model provider solely to extract your fields.
What's new
- Better number parsing: values such as
20,000and1,299.90are now read correctly (earlier versions could turn them into 20 or 1.2999). - Empty input is free: a run without a question searches nothing and explains what to fill in.
Trademarks
This Actor is not affiliated with, endorsed by, or connected to Anthropic, the Model Context Protocol project, or any website it reads. All product and company names are trademarks of their respective owners.


