MCP Web Research Agent: Search to JSON Schema avatar

MCP Web Research Agent: Search to JSON Schema

Pricing

from $100.00 / 1,000 web searches

Go to Apify Store
MCP Web Research Agent: Search to JSON Schema

MCP Web Research Agent: Search to JSON Schema

Ask a research question and name your fields. The agent searches the web, reads the top pages and returns every item it finds as JSON in your schema, with source URL and a per-field confidence label. Callable as a tool via the Apify MCP server.

Pricing

from $100.00 / 1,000 web searches

Rating

0.0

(0)

Developer

Turgay NANTA

Turgay NANTA

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

1

Monthly active users

4 days ago

Last modified

Share

MCP Web Research Agent: Web Search to Structured JSON

AI agents are good at reasoning and bad at reading raw web pages. When an assistant searches the web, it usually receives long markdown or HTML, spends tokens reading it, and then gives you facts in a different shape every time. That is hard to put into a table, a database or the next step of a workflow.

Ask a research question and name the fields you want. This Actor searches the web, reads the top result pages, and returns every item it finds as JSON in exactly your schema, with a source URL and a confidence label for every field. One call, one predictable object, ready for an agent, a RAG pipeline or a spreadsheet.

  • Question in, table out: "Top 5 open-source vector databases" with fields name, license, star_count, features returns one record per database found.
  • Many items per page: a "top 10" article becomes 10 records, not one blurb.
  • Your schema, your types: string, number, boolean, array, object, with a description per field to guide extraction.
  • Grounded, not invented: values must come from the source page; every field is labelled high, grounded, unverified or null.
  • Merged across sources: the same item found on several pages is combined into one record, gaps filled from the other pages.
  • Built for agents: callable as a tool through the Apify MCP server, or through the API from any workflow.
  • Clear cost: a 5-source research run costs up to $0.3501 in Actor fees (see Pricing).

Example output (shortened)

Input: query = "Top 5 open-source vector databases", schema name, license, star_count (number), features (array).

{
"query": "Top 5 open-source vector databases",
"schema": ["name", "license", "star_count", "features"],
"result_count": 2,
"results": [
{
"name": "Milvus",
"license": "Apache 2.0",
"star_count": 20000,
"features": ["hybrid search", "GPU indexing"],
"_source_url": "https://example-blog.com/vector-databases",
"_confidence": {"name": "high", "license": "high", "star_count": "high", "features": "grounded"}
},
{
"name": "Qdrant",
"license": "Apache 2.0",
"star_count": null,
"features": ["payload filtering"],
"_source_url": "https://example-blog.com/vector-databases",
"_confidence": {"name": "high", "license": "high", "star_count": "null", "features": "grounded"}
}
],
"sources": [{"title": "Best vector databases", "url": "https://example-blog.com/vector-databases"}]
}

Illustrative example showing the format; the values and URL are made up. A real run returns what its sources say, and null where they say nothing.


At a glance

InputA research question (query) + the fields you want (schema) + how many sources to read (maxResults, 1-50)
OutputOne dataset item: query, schema, result_count, results (your records), sources
Per recordYour fields + _source_url + _confidence per field
SearchApify RAG Web Browser (web search + page content as markdown)
ExtractionLanguage model, grounded in each source page, one object per distinct item
LanguagesText values in English, Turkish, German, Spanish or French
PricingPer search + per source that yields data; see Pricing
Runs asA classic Actor (Console, API, schedules, n8n, Make, Zapier) and an MCP tool through the Apify MCP server

Quick start

  1. Type a Research question, for example Best project management tools for small teams.
  2. Edit Fields to extract: a list of field names, or objects with name, type and description. The form starts with an example schema.
  3. Set Maximum results (how many search result pages to read; default 5).
  4. Click Start. The Dataset tab has one item with all records in results.

An empty run (no question) searches nothing, charges nothing and returns one row telling you what to fill in.


Why structured research

Plain web search tool for an agentThis Actor
What the agent receivesPage text or markdownJSON records in your schema
Shape of the answerDifferent every timeSame keys, same types, every run
Many items on one pageAgent must parse themOne record per item
Duplicates across sourcesAgent must mergeMerged by name, gaps filled
Can the agent trust a value?UnknownPer-field label: found in source or not
Token cost for the agentFull pagesOnly the extracted fields

Use cases

AI agent builders

Give your assistant one tool, "structured web research", instead of search plus scraping plus parsing. The agent passes a question and a schema, and gets a JSON object it can reason over or store.

RAG and knowledge base teams

Turn web research into clean rows before indexing: each record carries its _source_url, so answers can cite where a fact came from.

Market and competitor research

"Pricing of CRM tools for small businesses" with name, starting_price (number), free_plan (boolean), integrations (array). Read 10 sources, get one merged comparison table.

Sales and business development

"Logistics software companies in Germany" with company, city, website, focus. A quick first list from public articles and directories, each row with its source.

Product managers and analysts

"Open-source alternatives to X" with name, license, language, last_release. Fast landscape tables for a decision document.

Journalists and content teams

Collect facts on a topic into a table with source links, then verify the unverified fields before publishing.

Procurement and operations

"Industrial label printers under 2000 euro" with model, brand, price (number), print_width_mm (number). A first shortlist with sources for each value.


Worked example: a tool comparison in one call

You need a comparison of project management tools for a team decision.

  1. Question: best project management tools for small teams 2026
  2. Schema:
    [
    {"name": "tool", "type": "string", "description": "product name"},
    {"name": "starting_price_usd", "type": "number", "description": "lowest paid plan per user per month"},
    {"name": "free_plan", "type": "boolean"},
    {"name": "key_features", "type": "array"}
    ]
  3. maxResults: 8, so the Actor reads eight result pages.
  4. Run. Each page is read separately; a "10 best tools" article produces up to ten records.
  5. Merging. "Asana" found on three pages becomes one record; if one page had no price but another did, the price is filled in.
  6. Review. Sort by tool, check any field whose _confidence is unverified, and open _source_url for the ones you want to confirm.

Actor fees if all 8 sources yield data: $0.0001 + $0.10 + 8 x $0.05 = $0.5001, plus the RAG Web Browser usage on your account.


How it works

  1. Search. The Actor runs Apify's RAG Web Browser with your question and maxResults. It returns the top search results together with each page's content as markdown.
  2. Read each source. For every source page, the first 12,000 characters of its content, your question and your field list go to the language model.
  3. Extract many items. The model is told to return a JSON array with one object per distinct relevant item on the page, using only the page content, and null for anything the page does not state. Text values are written in your chosen language; field names are never translated.
  4. Parse robustly. The first complete JSON value is taken from the reply, even if the model wrapped it in code fences or added text after it. A reply without valid JSON is logged and that source is skipped.
  5. Coerce types. Numbers are cleaned ("20,000 stars" to 20000, "1.299,90" to 1299.9), booleans recognized from words (yes, available, in stock, sold out...), single values wrapped into arrays where an array is expected.
  6. Label confidence. Each value is checked against the source text (see Confidence labels).
  7. Merge duplicates. Records with the same value in the first text field (case-insensitive) are merged; empty fields are filled from later copies, with their confidence labels.
  8. Return one object with results and the list of sources that were read.

Input

FieldTypeDefaultDescription
querystringnoneThe research question or topic. Written like a web search.
schemaarray or objectexample schemaThe fields you want (formats below). Up to 25 fields.
maxResultsinteger5How many search result pages to read and structure, 1-50.
languagestringEnglishLanguage of text values: English, Türkçe, Deutsch, Español, Français.
modelstringclaude-haiku-4-5-20251001Advanced. claude-sonnet-4-6 for harder extraction.

query has no default on purpose: every run with a question starts a web search on your account, so nothing is searched until you ask something.

Schema formats

Names only (all fields become text):

["name", "website", "country"]

Objects with type and description (recommended):

[
{"name": "name", "type": "string", "description": "product or company name"},
{"name": "price", "type": "number", "description": "monthly price in USD"},
{"name": "open_source", "type": "boolean"},
{"name": "features", "type": "array"}
]

Map:

{"name": "product name", "price": {"type": "number", "description": "monthly price"}}

Allowed types: string, number, boolean, array, object. Any other type (including integer) falls back to string, so use number for counts and years.


Output

The dataset contains one item per run:

KeyContent
queryYour question
schemaThe field names used
result_countNumber of records after merging
resultsThe records: your fields + _source_url + _confidence
sourcesTitle and URL of every page the search returned

Confidence labels

LabelMeaning
highA text or number value was found in the source text (numbers compared digit by digit, at least two digits)
groundedA boolean or array value is supported by the source: a matching word for true/false (such as "available" or "sold out"), or at least one array element present in the text
unverifiedThe value is not found in the source; it may be inferred or paraphrased. object fields are always unverified because free-form objects cannot be checked reliably
nullThe field is empty

For an agent, a simple rule works well: use high and grounded values directly, and verify or drop unverified ones.

When there is nothing to return

  • The search returns no pages: the item is a _summary with a note; start and search fees apply.
  • Pages were read but none contained relevant items: those sources are not charged; if no record at all could be produced, the run ends with an error message.

Pricing

This Actor charges per event. Current prices are on the Pricing tab; at the time of writing:

EventWhenPrice
actor-startOnce per run$0.0001
search-performedOnce per run, after the web search succeeds$0.10
page-structuredPer source page that yielded at least one record$0.05

Source pages with no relevant item, or whose extraction failed, are not charged.

Examples (Actor fees)

RunCalculationTotal (maximum)
1 source$0.0001 + $0.10 + 1 x $0.05$0.1501
5 sources (default)$0.0001 + $0.10 + 5 x $0.05$0.3501
10 sources$0.0001 + $0.10 + 10 x $0.05$0.6001
50 sources$0.0001 + $0.10 + 50 x $0.05$2.6001
Empty inputnothing searched$0

Search is billed separately: the web search runs as an Apify RAG Web Browser run on your account, and that run's usage is billed to you at its own price.


Using it from an AI agent (MCP)

The Apify MCP server exposes Store Actors as tools for MCP clients such as Claude Desktop, Claude Code, Cursor and other agent frameworks. Add the Apify MCP server to your client with your Apify token and include this Actor (enezli/mcp-research-agent) in the tools it loads. The agent can then call it with:

{
"query": "EU regulations on AI in hiring",
"schema": [
{"name": "regulation", "type": "string"},
{"name": "year", "type": "number"},
{"name": "applies_to", "type": "string"},
{"name": "key_obligations", "type": "array"}
],
"maxResults": 5
}

and receive the single JSON object described above. Tips for agent prompts:

  • Ask the agent to design a small schema (3-8 fields) that answers the user's question.
  • Tell the agent to cite _source_url for each fact it reports.
  • Tell the agent to treat unverified fields as uncertain.

Integrations

API

curl -X POST "https://api.apify.com/v2/acts/enezli~mcp-research-agent/run-sync-get-dataset-items?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"query":"Top 5 open-source vector databases","schema":["name","license",{"name":"star_count","type":"number"}],"maxResults":5}'

Python

from apify_client import ApifyClient
client = ApifyClient("YOUR_TOKEN")
run = client.actor("enezli/mcp-research-agent").call(run_input={
"query": "Top 5 open-source vector databases",
"schema": [{"name": "name", "type": "string"}, {"name": "license", "type": "string"}],
"maxResults": 5,
})
item = next(client.dataset(run["defaultDatasetId"]).iterate_items())
for rec in item.get("results", []):
print(rec["name"], rec["license"], rec["_source_url"])

Schedules

Save a question and schema as a task and schedule it weekly to refresh a comparison table or a market list.

n8n, Make, Zapier

Run Actor → Get dataset items → split results into rows → write to Google Sheets, Airtable or Notion. Route rows with unverified fields to a review step.


Workflow recipes

Weekly market table. Save "pricing of [category] tools" with a fixed schema as a task, schedule it weekly, and append results to a Google Sheet. New tools and changed prices show up as new rows you can compare over time.

Agent answer with citations. Let your assistant call the Actor, then answer the user only from high and grounded values, citing _source_url next to each fact. Ask it to say explicitly when a field was null everywhere.

Lead list seed. Ask "[industry] companies in [city]" with company, website, city, focus. Treat the output as a first list to verify, not a finished database; each row tells you where it came from.

Due diligence snapshot. Before a meeting, research "[company] funding, products and customers" with a small schema, and keep the JSON with its sources in your notes.


Tips for better results

  • Write the question like a search query that would return list-style pages ("best", "top", "alternatives to", "companies in").
  • Add a description to each field, especially units ("price per user per month in USD").
  • Make the first text field the item's name. It is used to merge duplicates across sources.
  • Use number for numeric fields, including counts and years.
  • Start with 5 sources, raise maxResults only if the table is too thin.
  • Keep schemas short. Many fields and long pages can exceed the model's reply size for one source; that source is then skipped and not charged.
  • Use the stronger model for technical or dense sources.

Limitations

  • Search results decide coverage. The Actor reads what the web search returns for your question; it does not browse beyond those pages.
  • First 12,000 characters per page. Items further down a very long page may be missed.
  • Reply size per source. Very long lists on one page may be cut; a cut reply is skipped and not charged.
  • Merging uses the first text field. Items with different spellings ("PostgreSQL" vs "Postgres") stay separate.
  • Confidence labels are heuristic. A correct paraphrase can be unverified; they indicate where to look, not truth.
  • No logins or paywalls. Content behind a login is not available to the search.
  • Maximum 25 fields and 50 sources per run.

Troubleshooting

SymptomLikely causeWhat to do
One row saying "No research question given"query was emptyAdd a question
_summary with "no results"The search returned no pagesUse a broader or differently worded question
Run failed: "No source could be structured"None of the pages contained relevant items, or extraction failed on all of themRephrase the question toward list-style pages; simplify the schema
Run failed: "Web search failed"The search step failedRun again; the search fee is not charged when the search fails
Many null valuesSources do not state those factsAdjust fields or add descriptions
Duplicate items with slightly different namesMerging is exact on the first text fieldClean up in your sheet, or describe the name field more precisely
A field you typed as integer comes back as textinteger is not an allowed type hereUse number

FAQ

Is this an MCP server? It is an Apify Actor that AI agents can call as a tool through the Apify MCP server. It also runs as a normal Actor from the Console, API and automation tools.

Which search engine does it use? It uses Apify's RAG Web Browser, which performs the web search and returns page content.

Can it return several items from one page? Yes. Each source page can produce many records, one per distinct item.

Does it make things up? The model is instructed to use only the source content and to write null otherwise. Every value is also checked against the source text and labelled, so an agent can tell grounded values from possible inferences.

Can I get the answer in another language? Yes. Text values are written in the chosen output language. Field names and proper nouns stay as they are.

What if a source fails? It is logged and skipped; the other sources continue. Skipped sources are not charged.

How is this different from a normal web scraper? A scraper returns the content of pages you choose. This Actor starts from a question, finds the pages, and returns only the facts you asked for, in your schema.

Can I use it for a RAG pipeline? Yes. Each record has its source URL, so you can store structured facts with citations instead of raw page chunks.

How long does a run take? It depends on the search and on the number of sources; sources are processed one after another.

Is my data stored? Only in your own Apify storage. Source content is sent to the language model provider solely to extract your fields.


What's new

  • Better number parsing: values such as 20,000 and 1,299.90 are now read correctly (earlier versions could turn them into 20 or 1.2999).
  • Empty input is free: a run without a question searches nothing and explains what to fill in.

Trademarks

This Actor is not affiliated with, endorsed by, or connected to Anthropic, the Model Context Protocol project, or any website it reads. All product and company names are trademarks of their respective owners.