Baidu Search Scraper - Web & News SERP Results avatar

Baidu Search Scraper - Web & News SERP Results

Pricing

$2.00 / 1,000 result scrapeds

Go to Apify Store
Baidu Search Scraper - Web & News SERP Results

Baidu Search Scraper - Web & News SERP Results

Scrape Baidu web and news search results with the real destination URL, title, snippet, source, publish date and rank. No redirect resolving, no login, no cookies.

Pricing

$2.00 / 1,000 result scrapeds

Rating

0.0

(0)

Developer

Scrape Sage

Scrape Sage

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

Scrape Baidu search results at scale - web search and Baidu News - and get back clean rows with the real destination URL, title, snippet, source, publish date and ranking position. No login, no cookies, no redirect-resolving.

Built for China market research, Baidu SEO tracking, competitor monitoring and feeding Chinese-language search results to an AI agent.


The thing that makes this different: real URLs

Baidu's visible result links are www.baidu.com/link?url=... redirects. They expire, and resolving them costs you an extra HTTP request per result.

This actor reads Baidu's own destination attribute, so every row already carries the final URL - ready to crawl, dedupe or feed to an agent. You also get the redirect link in redirectUrl if you want it.

What you get

One row per result, every row carrying the same fields:

FieldWhat it is
titleResult title, highlight markup stripped
urlReal destination URL
redirectUrlBaidu's own /link?url= form
descriptionResult snippet
sourceName, sourceDomainPublisher name and domain
publishedText, publishedAtDate as Baidu shows it, plus ISO 8601 when unambiguous
rank, rankOnPage, pagePosition overall and within the page
resultKindorganic or featured (Baidu's rich blocks)
queryMatchedWhether the result actually mentions your query
searchTypeweb or news
query, resultKeyWhich query found it; a stable unique key

Search modes

  • Web - the main Baidu index
  • News - Baidu News results, with publisher and date
  • Both - a page can rank in both indexes; you get both rows, each with its own ranking

Plus optional related searches (相关搜索) - the query suggestions Baidu prints under the results, which are useful keyword-research seeds.

Filters

onlyOrganic · excludeBaiduModules · maxPagesPerQuery · resultsPerPage · maxResults


Example input

{
"queries": ["人工智能", "电动汽车"],
"searchType": "both",
"maxPagesPerQuery": 3,
"maxResults": 100,
"proxyConfiguration": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"] }
}

Example output

{
"type": "result",
"query": "人工智能",
"searchType": "web",
"page": 1,
"rank": 1,
"rankOnPage": 1,
"title": "人工智能(中国普通高等学校本科专业) - 百度百科",
"url": "https://baike.baidu.com/item/%E4%BA%BA%E5%B7%A5%E6%99%BA%E8%83%BD/24604211",
"description": "人工智能(Artificial Intelligence),是一门普通高等学校本科专业...",
"sourceName": "百度百科",
"sourceDomain": "baike.baidu.com",
"resultKind": "featured",
"queryMatched": true
}

Pricing

$0.002 per result. You are charged only for rows actually delivered to your dataset.

You are not billed for junk results

Baidu never returns an empty page. Give it a term it cannot match and it serves a page of unrelated filler - we measured three nonsense queries returning 3-10 results each, none of which mentioned the term.

By default (requireQueryMatch) this actor recognises that: if a query's first page contains no result mentioning your term, nothing is stored and nothing is billed, and the run tells you so. Turn it off if you want everything Baidu returns.

Honest limits

  • Baidu's index is overwhelmingly Chinese. English queries return far fewer, weaker results. Query in Chinese for real coverage.
  • Dates are sparse on web results (~40%) because Baidu only prints a date when it has one; news results carry them far more often. publishedAt is filled only when the date is unambiguous - a bare "8月5日" with no year is returned in publishedText rather than guessed into a year.
  • Use the residential proxy (the default). Baidu throttles per IP.

Tips

  • Set searchType: "both" and dedupe on resultKey to track a page's web and news ranking in one job.
  • rank is the position across all pages for that query - chart it over scheduled runs to track ranking movement.
  • Turn on includeRelatedSearches for keyword research; the terms come back as rows with type: "relatedSearch".
  • Naver Scraper - the same job for Korea (places, blogs, news, cafe)