WordPress Posts & Pages Scraper avatar

WordPress Posts & Pages Scraper

Pricing

from $1.00 / 1,000 content items

Go to Apify Store
WordPress Posts & Pages Scraper

WordPress Posts & Pages Scraper

Export public WordPress posts and pages as clean text and original HTML for content migration, audit, and research.

Pricing

from $1.00 / 1,000 content items

Rating

0.0

(0)

Developer

Akshay Aggarwal

Akshay Aggarwal

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Categories

Share

Export a site's public WordPress posts and pages for a content migration, audit, or research project. Each result has the original rendered HTML, readable text, permalink, WordPress ID, type, publication and modification dates, and public author/category/tag details when available.

Quick start

{"siteUrls":["https://wordpress.org/news"],"contentTypes":["posts","pages"],"maxItemsPerSite":2}

Example public post from WordPress.org News (preview shortened for readability; dataset rows contain the full rendered HTML and text):

{"id":21731,"type":"post","site_url":"https://wordpress.org/news","url":"https://wordpress.org/news/2026/09/wordpress-7-1-2-release/","title":"WordPress 7.1.2 Release","content_html":"<p class=\"wp-block-paragraph\">This security release features a fix for a critical severity security vulnerability.</p>...","content_text":"This security release features a fix for a critical severity security vulnerability. ...","excerpt":"","published_at":"2026-09-22T14:01:20","modified_at":"2026-09-23T14:04:10"}

See example-output.json for a complete public page row.

Scope and limits

  • Reads only anonymous GET /wp-json/wp/v2/posts and pages with status=publish, through HTTPS. A site must expose those standard REST routes at the supplied site URL. Custom content types and REST-disabled sites are outside scope.
  • Accepts 1–10 sites, posts/pages selection, a total saved-item cap of 1–5000 per site (default 10), and optional after/before ISO 8601 publication filters. WordPress may apply site-specific filtering, access controls, and rate limits. The Actor stops at its item, request, time, or spending limit; large requests may be partial.
  • Skips items with protected content or excerpts, non-published status, invalid links, or duplicate type/ID. It never authenticates, requests drafts/private items, supplies passwords, or downloads images/media. It preserves the source's date and modified strings without timezone conversion.
  • Author and taxonomy names depend on public embedded metadata. Where public IDs exist but names do not, IDs are retained. Missing fields are omitted or null. content_html is original rendered HTML; content_text, title, and excerpt remove tags and decode entities. The Actor does not execute HTML or fetch linked assets.
  • The dataset contains only saved result rows. OUTPUT in the default key-value store records fetched, saved, skipped, errors, limits, request count, and unprocessed sites. A missing REST route or other source error appears there; results already saved remain available.

Billing

One content-item event represents one saved post or page in the dataset. The Actor stops adding rows when the event spending limit is reached. Fetches, skips, errors, and summary entries are not billable result rows. Check the Actor's current price in Apify before running it.

Output fields

id, type, site_url, url, title, content_html, content_text, excerpt, published_at, modified_at; optional author, categories, and tags when public. IDs are unique within a site and content type; use (site_url, type, id) as a stable cross-site key.

This Actor is independent of WordPress and uses the public Posts, Pages, and Pagination API contracts.