Baidu Search Scraper - Web & News SERP Results avatar

Baidu Search Scraper - Web & News SERP Results

Pricing

$2.00 / 1,000 result scrapeds

Go to Apify Store
Baidu Search Scraper - Web & News SERP Results

Baidu Search Scraper - Web & News SERP Results

Scrape Baidu web and news search results with the real destination URL, title, snippet, source, publish date and rank. No redirect resolving, no login, no cookies. Independent tool, not affiliated with Baidu.

Pricing

$2.00 / 1,000 result scrapeds

Rating

0.0

(0)

Developer

Scrape Sage

Scrape Sage

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

5 days ago

Last modified

Share

Disclaimer: This Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by Baidu, Inc. or any of its subsidiaries. All trademarks mentioned are the property of their respective owners. "Baidu" is referenced only to describe the publicly available website this Actor collects data from.

Scrape Baidu search results at scale - web search and Baidu News - and get back clean rows with the real destination URL, title, snippet, source, publish date and ranking position. No login, no cookies, no redirect-resolving.

Built for China market research, Baidu SEO tracking, competitor monitoring and feeding Chinese-language search results to an AI agent.


The thing that makes this different: real URLs

Baidu's visible result links are www.baidu.com/link?url=... redirects. They expire, and resolving them costs you an extra HTTP request per result.

This actor reads Baidu's own destination attribute, so every row already carries the final URL - ready to crawl, dedupe or feed to an agent. You also get the redirect link in redirectUrl if you want it.

What you get

One row per result, every row carrying the same fields:

FieldWhat it is
titleResult title, highlight markup stripped
urlReal destination URL
redirectUrlBaidu's own /link?url= form
descriptionResult snippet
sourceName, sourceDomainPublisher name and domain
publishedText, publishedAtDate as Baidu shows it, plus ISO 8601 when unambiguous
rank, rankOnPage, pagePosition overall and within the page
resultKindorganic or featured (Baidu's rich blocks)
queryMatchedWhether the result actually mentions your query
searchTypeweb or news
query, resultKeyWhich query found it; a stable unique key

Search modes

  • Web - the main Baidu index
  • News - Baidu News results, with publisher and date
  • Both - a page can rank in both indexes; you get both rows, each with its own ranking

Plus optional related searches (相关搜索) - the query suggestions Baidu prints under the results, which are useful keyword-research seeds.

Filters

onlyOrganic · excludeBaiduModules · maxPagesPerQuery · resultsPerPage · maxResults


Example input

{
"queries": ["人工智能", "电动汽车"],
"searchType": "both",
"maxPagesPerQuery": 3,
"maxResults": 100,
"proxyConfiguration": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"] }
}

Example output

{
"type": "result",
"query": "人工智能",
"searchType": "web",
"page": 1,
"rank": 1,
"rankOnPage": 1,
"title": "人工智能(中国普通高等学校本科专业) - 百度百科",
"url": "https://baike.baidu.com/item/%E4%BA%BA%E5%B7%A5%E6%99%BA%E8%83%BD/24604211",
"description": "人工智能(Artificial Intelligence),是一门普通高等学校本科专业...",
"sourceName": "百度百科",
"sourceDomain": "baike.baidu.com",
"resultKind": "featured",
"queryMatched": true
}

Pricing

$0.002 per result. You are charged only for rows actually delivered to your dataset.

You are not billed for junk results

Baidu never returns an empty page. Give it a term it cannot match and it serves a page of unrelated filler - we measured three nonsense queries returning 3-10 results each, none of which mentioned the term.

By default (requireQueryMatch) this actor recognises that: if a query's first page contains no result mentioning your term, nothing is stored and nothing is billed, and the run tells you so. Turn it off if you want everything Baidu returns.

Honest limits

  • Baidu's index is overwhelmingly Chinese. English queries return far fewer, weaker results. Query in Chinese for real coverage.
  • Dates are sparse on web results (~40%) because Baidu only prints a date when it has one; news results carry them far more often. publishedAt is filled only when the date is unambiguous - a bare "8月5日" with no year is returned in publishedText rather than guessed into a year.
  • Use the residential proxy (the default). Baidu throttles per IP.

Automate & schedule

Run this Actor on autopilot and pull results into your own stack:

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'MY_APIFY_TOKEN' });
const run = await client.actor('scrapesage/baidu-search-scraper').call({
"queries": [
"人工智能"
],
"searchType": "web",
"maxPagesPerQuery": 2,
"maxResults": 50,
"proxyConfiguration": {
"useApifyProxy": true,
"apifyProxyGroups": [
"RESIDENTIAL"
]
}
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(`Got ${items.length} records`);

Integrate with any app

Connect the dataset to thousands of apps - no code required:

  • Make - multi-step automation scenarios.
  • Zapier - push new records straight into your CRM or spreadsheet.
  • Slack - get notified when a scheduled run finds something new.
  • Google Drive / Sheets - auto-export every run to a spreadsheet.
  • Airbyte - pipe results into your data warehouse.
  • GitHub - trigger runs from commits or releases.

Use with AI assistants (MCP)

The output is clean, LLM-ready JSON. You can call this Actor from Claude, ChatGPT or any agent framework through the Apify MCP server - describe the data you need and let the assistant run this Actor for you.

Agent-ready: autonomous payments (x402 & Skyfire)

This actor is agent-ready — AI agents can discover it, run it, and pay for it autonomously, with no Apify account and no human in the loop. It uses pay-per-event pricing and limited permissions, so it qualifies for Apify's agentic-payment standards:

  • x402 — an open, HTTP-native payment protocol. Agents pay per run in USDC on the Base network directly through the Apify MCP server — no account, no API key.
  • Skyfire — agent-to-service payments for fully autonomous AI-agent workflows.

Building an AI agent, MCP tool, or autonomous data pipeline? This scraper is ready to plug in and pay as it goes.

FAQ

How is this Actor billed? Pay-per-event: you pay only for the results it delivers, with no monthly rental and no start fee. The per-event price is shown on the Pricing tab.

Can I schedule it and get results automatically? Yes - create a Schedule and add a webhook or an integration (Google Sheets, Slack, Make, Zapier) to push each run's dataset wherever you need it.

Which export formats are available? Every run's dataset can be downloaded as JSON, CSV, Excel (XLSX), XML, HTML or RSS from the Apify Console or the API.

Can I run it from code or an AI agent? Yes - through the Apify API and client libraries, or from Claude, ChatGPT and other assistants via the Apify MCP server.

Is it legal to scrape Baidu? This Actor collects publicly available data only. You are responsible for using the output in compliance with applicable laws (including data-protection law where personal data is involved) and the source's terms. See the Disclaimer below.

Is this an official Baidu tool? No. It is an independent, third-party Actor with no affiliation to, endorsement by or sponsorship from Baidu, Inc. See the Disclaimer below.

Tips

  • Set searchType: "both" and dedupe on resultKey to track a page's web and news ranking in one job.
  • rank is the position across all pages for that query - chart it over scheduled runs to track ranking movement.
  • Turn on includeRelatedSearches for keyword research; the terms come back as rows with type: "relatedSearch".
  • Naver Scraper - the same job for Korea (places, blogs, news, cafe)

Disclaimer

This Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by Baidu, Inc. or any of its subsidiaries. All trademarks mentioned are the property of their respective owners.

"Baidu" and any related marks are the property of their respective owners and are used here only in a descriptive, nominative sense - to identify the publicly accessible website from which this Actor collects data. This Actor is not an official Baidu product, is not authorised or certified by Baidu, Inc., and does not distribute Baidu software. It collects only publicly available information; you are responsible for ensuring your use of that data complies with applicable laws, regulations and the terms of the source website.

Need help?

Open an issue on the Actor's Issues tab, or visit the Apify help center. Feature requests are welcome - this Actor is actively maintained.