Baidu Images Scraper avatar

Baidu Images Scraper

Pricing

from $1.99 / 1,000 search results

Go to Apify Store
Baidu Images Scraper

Baidu Images Scraper

Search Baidu Images and extract normalized image URLs, source pages, dimensions, formats, and listing metadata with API pagination and DOM fallback.

Pricing

from $1.99 / 1,000 search results

Rating

0.0

(0)

Developer

Search API

Search API

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

5 days ago

Last modified

Share

Search Baidu Images by keyword and collect normalized image records. The Actor calls Baidu's structured data.images response from the loaded browser session first, validates status and content type before parsing, and supplements incomplete or unavailable API pages with a rendered result-grid fallback.

Inputs

Use query for one phrase or queries for a deduplicated batch of up to 25 phrases. Output for relevance is interleaved round-robin across queries. maxItems is the global cap, maxItemsPerQuery is the per-query cap, and maxPages is the maximum 30-result API pages per query.

Minimal input

{
"query": "熊猫",
"maxItems": 10,
"maxPages": 1,
"proxyConfiguration": { "useApifyProxy": false }
}

Batch, pagination, filters, and developer options

{
"queries": ["猫", "狗"],
"maxItems": 20,
"maxItemsPerQuery": 10,
"maxPages": 2,
"sortBy": "widthDesc",
"minWidth": 800,
"minHeight": 600,
"fileFormat": "jpg",
"sourceDomainContains": "example.com",
"requestDelayMs": 300,
"maxConcurrency": 2,
"maxRequestRetries": 1,
"navigationTimeoutSecs": 60,
"requestHandlerTimeoutSecs": 120,
"blockMedia": true,
"market": "zh-CN",
"proxyConfiguration": { "useApifyProxy": false }
}

No-results input

{
"query": "qzxv9918273645neverexists",
"maxItems": 5,
"maxPages": 1,
"proxyConfiguration": { "useApifyProxy": false }
}

sortBy supports relevance, titleAsc, titleDesc, widthDesc, heightDesc, and sizeDesc. Width, height, format, and source-domain filters are applied after normalized records are built; missing values do not pass a configured filter.

Output

The Actor writes one JSON item per unique image URL. The dataset JSON files are available in the run's default dataset storage and through the Apify dataset API/Console links after the run completes. The OUTPUT key-value record contains per-query counts, API/DOM fallback usage, page failures, filters, and the final status.

Each record can include:

  • stable SHA-256 identity, query/global positions, query/page context, and extraction method;
  • image, thumbnail, preview, large/original, host-page, and source URLs;
  • source domain/name, title/alt text, dimensions, aspect ratio, pixel count, format, MIME type, and optional file size;
  • animation, video, copyright, watermark, AI-edit, commodity, clarity, aurora, and source-type indicators when Baidu publishes them;
  • validated API status/content type/attempt metadata, locale, search metadata, and UTC timestamp.

The dataset contract requires at least 20 populated fields per record, including the image and source-page URLs, host domains, query/page positions, and extraction metadata. Optional measurements and labels are omitted when Baidu does not provide them.

No example dataset record is included because Chrome could not load the live Baidu Images search page (ERR_CONNECTION_TIMED_OUT) for the required source comparison. No synthetic record is presented as a live extraction. The site and evidence are recorded in the repository's AGENTS.md.

Reliability and access boundaries

The Actor uses bounded retries and exponential backoff for temporary API/browser failures, a shared persistent Playwright context for cookies, optional authorized Apify/custom proxies, and a consistent browser locale. If the host system DNS cannot resolve image.baidu.com, it performs a bounded A-record lookup through public DNS and maps only that official hostname inside the browser; if both paths fail, the run reports a failure instead of fabricating records. Media and fonts can be blocked to reduce load; images remain available for DOM fallback and metadata extraction. Verification and CAPTCHA pages fail closed and are never stored as records. The Actor does not solve challenges, bypass authentication, or log cookies, authorization headers, proxy credentials, tokens, or raw API payloads.

Validation

npm test
npm run check
npm run validate:readme
npm run validate:dataset
npm run validate:run
apify validate-schema
apify run --purge --input-file qa-inputs/one.json

The schema validator checks the contract shape. The run validator checks required fields and types, safe URLs, unique IDs, sequential positions, and the OUTPUT summary from a completed local run.