GOV.UK Content Search Scraper avatar

GOV.UK Content Search Scraper

Pricing

from $22.87 / 1,000 results

Go to Apify Store
GOV.UK Content Search Scraper

GOV.UK Content Search Scraper

Scrape GOV.UK: search the entire UK government publications catalogue (policies, guidance, news, statistics). Filter by query, organisation, format or date. Returns titles, descriptions, URLs, organisations and publication dates.

Pricing

from $22.87 / 1,000 results

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

2

Monthly active users

6 days ago

Last modified

Share

ParseForge Banner

πŸ‡¬πŸ‡§ GOV.UK Content Search Scraper

πŸš€ Search the entire UK government catalogue in seconds. Pull 600,000+ pages, publications, news stories, statistics, consultations, and services straight from GOV.UK. Filter by query, organisation, format, or date. No sign-up, no manual paging, no parser to maintain.

The GOV.UK Content Search Scraper queries the official UK government content index and returns up to 11 structured fields per record, including titles, descriptions, URLs, document formats, publication dates, and the publishing organisations and world locations. GOV.UK is the canonical home for almost every UK government publication, news story, statistical release, consultation, and citizen service.

The catalogue covers the entire central UK government estate, from the Cabinet Office and HM Treasury to HMRC, the Ministry of Defence, the Department for Transport, and over a thousand executive agencies, arms-length bodies, and tribunals. This Actor makes that data downloadable as CSV, Excel, JSON, or XML in under five minutes. Filters run server-side, so you skip the parser engineering entirely.

🎯 Target AudienceπŸ’‘ Primary Use Cases
Policy analysts, regulatory and compliance teams, journalists, lobbyists, GovTech vendors, academic researchers, market-intelligence firmsPolicy monitoring, regulatory horizon-scanning, consultation tracking, FOI release feeds, ministerial speech analysis, statistics release pipelines, press monitoring

πŸ“‹ What the GOV.UK Content Search Scraper does

Six filtering workflows in a single run:

  • πŸ”Ž Free-text search. Query any keyword or phrase across titles, descriptions, and body text of every GOV.UK page.
  • πŸ“‚ Format filter. Restrict to a single GOV.UK document format from a list of 130+ (news_story, press_release, guidance, official_statistics, consultation_outcome, statutory_guidance, FOI release, and many more).
  • πŸ›οΈ Organisation filter. Restrict to a single department or agency by slug (e.g. cabinet-office, hm-revenue-customs, department-for-transport).
  • πŸ“… Date range filter. publishedAfter and publishedBefore scope to any window on the public timestamp.
  • πŸ”’ Sort order. Relevance (default), newest first, oldest first, title A-Z, or most popular.
  • 🌍 World locations and topical events. Returned per record so you can pivot by country or government event.

Each record includes the document title, description, canonical GOV.UK URL, content ID, format and document type, publication date, every publishing organisation (with slug and acronym), associated world locations, and topical events.

πŸ’‘ Why it matters: UK government publications drive regulation, market opportunities, public-sector procurement, and the news cycle. Building your own GOV.UK pipeline means writing a paginated search client, mapping 130+ formats, joining organisation slugs, and refreshing daily. This Actor skips all of that and gives you a clean refreshed snapshot on every run.

πŸ“Š Data fields

Each record includes: contentId, description, documentType, documentTypeRaw, format, organisations, publishedAt, publishedDate, relevanceScore, scrapedAt, title, topicalEvents, totalAvailable, url, worldLocations. These field names come straight from the actor's dataset schema, so what you see here is what lands in your dataset.

πŸš€ How to use

  1. πŸ“ Sign up. Create a free account w/ $5 credit (takes 2 minutes).
  2. 🌐 Open the Actor. Go to the GOV.UK Content Search Scraper page on the Apify Store.
  3. 🎯 Set input. Type a query (or leave empty), pick an organisation or format, set a date window if you need one, and set maxItems.
  4. πŸš€ Run it. Click Start and let the Actor collect your results.
  5. πŸ“₯ Download. Grab your dataset in the Dataset tab as CSV, Excel, JSON, or XML.

⏱️ Total time from signup to a downloaded GOV.UK dataset: 3-5 minutes. No coding required.

πŸ’‘ Pro Tip: browse the complete ParseForge collection for more reference-data scrapers.

⚠️ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by the UK Government, the Government Digital Service, or any UK department. All trademarks mentioned are the property of their respective owners. Only publicly available open-government content is collected, under the Open Government Licence.

πŸ†˜ Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.