Hansard UK Parliament Debates Scraper avatar

Hansard UK Parliament Debates Scraper

Pricing

from $23.63 / 1,000 results

Go to Apify Store
Hansard UK Parliament Debates Scraper

Hansard UK Parliament Debates Scraper

Export the official transcripts of UK Parliament debates and speeches from Hansard. Filter by House (Commons or Lords), search term, member, and date range. Each record includes the full speech text, speaker, debate section, and a permalink to the official Hansard transcript.

Pricing

from $23.63 / 1,000 results

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

4

Total users

1

Monthly active users

6 days ago

Last modified

Share

ParseForge Banner

πŸ—£οΈ Hansard UK Parliament Debates Scraper

πŸš€ Export UK Parliament debate transcripts in seconds. Pull every spoken contribution from the House of Commons and House of Lords, filtered by topic, member, date, or department. Each record is a clean structured speech with full text, speaker, debate section, and a permalink to the official Hansard transcript. No sign-up, no manual paging, no parser to maintain.

The Hansard UK Parliament Debates Scraper queries the official Hansard transcript catalogue and returns up to 17 structured fields per record, including the contribution ID, speaker name and member ID, House, debate section, sitting date, full speech text, word count, ordering metadata, and a deep permalink back to the official Hansard page.

The catalogue covers the official record of every spoken contribution in the UK Parliament, including ministerial statements, backbench speeches, oral questions, urgent questions, statements, and full debates. Hansard has tracked the proceedings of the UK Parliament since 1803 and is the canonical record cited by historians, journalists, and political researchers.

🎯 Target AudienceπŸ’‘ Primary Use Cases
Political analysts and researchers, journalists, NLP and machine-learning teams, public-affairs and lobbying firms, civic-tech projects, academic political scientists, content creatorsSpeech and rhetoric analysis, member voting-context research, topic mining, NLP training corpora, ministerial statement monitoring, lobbyist due diligence, civic-tech transparency tools

πŸ“‹ What the Hansard UK Debates Scraper does

Six filtering workflows in a single run:

  • πŸ”Ž Free-text search. Match a keyword or phrase across every spoken contribution (e.g. "climate change", "NHS funding", "AUKUS").
  • πŸ›οΈ House filter. Restrict to House of Commons, House of Lords, or both.
  • πŸ‘€ Member filter. Substring match on the speaker name (e.g. "Keir Starmer", "Lord Hannan").
  • πŸ“… Date range. Scope to any sitting-date window with startDate and endDate.
  • πŸ›οΈ Department filter. Substring match on the responsible government department (e.g. "Treasury", "Department for Education").
  • πŸ”’ Page-driven sample. Pull the latest contributions across all topics when no query is set.

Each record includes the contribution ID, the member's name and (when available) their member ID, the House, the section ("Commons Chamber", "Westminster Hall", etc.), the debate section title, the Hansard internal section code, the sitting date, the timecode of the contribution, the full speech text (HTML preserved), a word count, the order in the debate, the paragraph tag, and a deep permalink back to the official Hansard transcript page.

πŸ’‘ Why it matters: Hansard transcripts power policy analysis, NLP corpora, civic transparency, and political journalism. Building your own pipeline means writing a paginated client, mapping debate identifiers to permalinks, normalising HTML across sittings, and refreshing daily. This Actor skips all of that and gives you a clean refreshed snapshot on every run.

πŸ“Š Data fields

Each record includes: attributedTo, contributionId, debateSection, debateSectionId, hansardSection, house, memberId, memberName, orderInDebate, paragraphTag, scrapedAt, section, sittingDate, text, timecode, url, wordCount. These field names come straight from the actor's dataset schema, so what you see here is what lands in your dataset.

πŸš€ How to use

  1. πŸ“ Sign up. Create a free account w/ $5 credit (takes 2 minutes).
  2. 🌐 Open the Actor. Go to the Hansard UK Parliament Debates Scraper page on the Apify Store.
  3. 🎯 Set input. Type a search term or member name, optionally pick a House and date range, and set maxItems.
  4. πŸš€ Run it. Click Start and let the Actor collect your contributions.
  5. πŸ“₯ Download. Grab your dataset in the Dataset tab as CSV, Excel, JSON, or XML.

⏱️ Total time from signup to a downloaded Hansard dataset: 3-5 minutes. No coding required.

πŸ’‘ Pro Tip: browse the complete ParseForge collection for more reference-data scrapers.

⚠️ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by the UK Parliament, the House of Commons, the House of Lords, or the Hansard Society. All trademarks mentioned are the property of their respective owners. Only publicly available open Hansard transcript data is collected, under the Open Parliament Licence.

πŸ†˜ Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.