WordPress Posts & Pages Scraper
Pricing
from $1.00 / 1,000 content items
WordPress Posts & Pages Scraper
Export public WordPress posts and pages as clean text and original HTML for content migration, audit, and research.
Pricing
from $1.00 / 1,000 content items
Rating
0.0
(0)
Developer
Akshay Aggarwal
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Export a site's public WordPress posts and pages for a content migration, audit, or research project. Each result has the original rendered HTML, readable text, permalink, WordPress ID, type, publication and modification dates, and public author/category/tag details when available.
Quick start
{"siteUrls":["https://wordpress.org/news"],"contentTypes":["posts","pages"],"maxItemsPerSite":2}
Example public post from WordPress.org News (preview shortened for readability; dataset rows contain the full rendered HTML and text):
{"id":21731,"type":"post","site_url":"https://wordpress.org/news","url":"https://wordpress.org/news/2026/09/wordpress-7-1-2-release/","title":"WordPress 7.1.2 Release","content_html":"<p class=\"wp-block-paragraph\">This security release features a fix for a critical severity security vulnerability.</p>...","content_text":"This security release features a fix for a critical severity security vulnerability. ...","excerpt":"","published_at":"2026-09-22T14:01:20","modified_at":"2026-09-23T14:04:10"}
See example-output.json for a complete public page row.
Scope and limits
- Reads only anonymous
GET /wp-json/wp/v2/postsandpageswithstatus=publish, through HTTPS. A site must expose those standard REST routes at the supplied site URL. Custom content types and REST-disabled sites are outside scope. - Accepts 1–10 sites, posts/pages selection, a total saved-item cap of 1–5000 per site (default 10), and optional
after/beforeISO 8601 publication filters. WordPress may apply site-specific filtering, access controls, and rate limits. The Actor stops at its item, request, time, or spending limit; large requests may be partial. - Skips items with protected content or excerpts, non-published status, invalid links, or duplicate type/ID. It never authenticates, requests drafts/private items, supplies passwords, or downloads images/media. It preserves the source's
dateandmodifiedstrings without timezone conversion. - Author and taxonomy names depend on public embedded metadata. Where public IDs exist but names do not, IDs are retained. Missing fields are omitted or null.
content_htmlis original rendered HTML;content_text, title, and excerpt remove tags and decode entities. The Actor does not execute HTML or fetch linked assets. - The dataset contains only saved result rows.
OUTPUTin the default key-value store records fetched, saved, skipped, errors, limits, request count, and unprocessed sites. A missing REST route or other source error appears there; results already saved remain available.
Billing
One content-item event represents one saved post or page in the dataset. The Actor stops adding rows when the event spending limit is reached. Fetches, skips, errors, and summary entries are not billable result rows. Check the Actor's current price in Apify before running it.
Output fields
id, type, site_url, url, title, content_html, content_text, excerpt, published_at, modified_at; optional author, categories, and tags when public. IDs are unique within a site and content type; use (site_url, type, id) as a stable cross-site key.
This Actor is independent of WordPress and uses the public Posts, Pages, and Pagination API contracts.