German Imprint Scraper (Contact+Social Links)
Pricing
$1.50/month + usage
German Imprint Scraper (Contact+Social Links)
An ultra-fast, AI-enhanced scraper that automatically finds and extracts Impressum data from German websites. Get perfectly structured company names, addresses, emails, phones, tax IDs, and social links for B2B enrichment.
Pricing
$1.50/month + usage
Rating
5.0
(2)
Developer
CodeScraper
Maintained by CommunityActor stats
3
Bookmarked
70
Total users
2
Monthly active users
11 days ago
Last modified
Categories
Share
⭐ AI-Powered German Impressum Scraper – B2B Legal Data Extraction
This Apify Actor is a blazing-fast, AI-enhanced domain enrichment tool designed exclusively to automatically discover, clean, and structure Impressum (Legal Notice) data from German websites.
Say goodbye to messy, unstructured legal pages. By combining lightning-fast DOM parsing with an advanced LLM processing engine (Groq), this scraper cuts through website clutter to deliver pristine, ready-to-use B2B contact and compliance data.
🔥 Why choose this scraper?
- 🤖 Smart AI Structuring: The built-in AI translates complex, unstructured legal paragraphs into perfectly formatted JSON (e.g., extracting exact legal entity names and standardizing international phone numbers).
- ⚡ Hybrid Reliability: Seamlessly toggles between ultra-fast Regex pattern matching and Premium AI parsing.
- 🎯 "Never Miss" Routing: If a homepage lacks a clear legal link, the scraper's intelligent routing automatically hunts down hidden
/impressumpages. - 🛡️ Zero Maintenance: Operates entirely without browser overhead (Puppeteer/Playwright), keeping your runs incredibly fast and memory-efficient.
💲 Pricing Note: This Actor operates on a predictable, risk-free Pay-Per-Event (PPE) model. You ONLY pay for successful extractions!
- Standard Extraction (Regex): $0.003 per domain processed.
- Premium Extraction (AI Enabled): $0.0045 per domain processed.
- No API Keys Required: We cover the AI token costs! You do not need to provide your own Groq or OpenAI API keys. Just input your URLs and let the AI do the heavy lifting.
- Volume discounts are automatically applied for Bronze, Silver, and Gold tier users.
🚀 What It Does
For every website URL or domain provided, the actor locates the Impressum page, extracts the core legal text, and outputs a perfectly structured, standardized JSON profile of the business.
🏢 Extracted Data
The actor outputs a clean dataset containing exactly what you need for compliance and B2B enrichment:
- 🔗 URLs: The original target website and the exact Impressum URL found.
- 🏷️ Company Name: The extracted legal name of the business (e.g., GmbH, AG).
- 📍 Address: The full physical address of the company.
- 📧 Emails: All detected contact email addresses.
- 📞 Phones: All detected phone numbers.
- 🏛️ Legal IDs: The Commercial Register Number (HRB/HRA) and Tax ID (Umsatzsteuer-ID).
- 🌐 Social Links: Detected social media profiles (Facebook, Instagram, LinkedIn, X, TikTok, YouTube, Pinterest).
- 🧠 Metadata: Page Title, Description, Keywords, and H1 Header Text.
🧠 Smart AI Models Included
If you enable the Premium AI Extraction, you can choose the specific AI model that powers your data structuring directly from the input settings. No API key required.
- 🥇 openai/gpt-oss-120b (Most Capable): Highest reasoning capability. Best for complex or confusing legal pages.
- 🥈 openai/gpt-oss-20b (Recommended): Blazing fast. Perfectly balances speed and accuracy for mass processing.
- 🥉 qwen/qwen3.6-27b (Alternative): Excellent alternative with very strong multilingual support for non-German domains.
- 🚀 groq/compound & compound-mini: High-speed experimental models optimized for short text bursts.
⚡ Core Features
- 🤖 Hybrid Extraction — Seamlessly toggle between ultra-fast Regex pattern matching or highly accurate AI parsing. If the AI fails, it automatically falls back to Regex.
- 🎯 "Never Miss" Fallback Routing — If the homepage has no clear Impressum link, the scraper dynamically guesses the route (e.g.,
/impressum) and analyzes that instead. - 🚀 Extreme Speed — Uses a
CheerioCrawlercombined with strict DOM-stripping (removing heavy images, navs, and scripts) to keep memory usage low and speed blazing fast. - 🔗 Smart URL Parsing — No need to format your lists. Just paste raw domains like
marley.deordr-johanna-budwig.deand the Actor automatically normalizes them. - 🧼 Clean Output — Automatically strips out embedded SVGs, payment badges, and code artifacts from metadata and company names.
⚙️ Input Configuration
| Field | Type | Required | Description | Default |
|---|---|---|---|---|
startUrls | Array | Yes | List of website URLs to scrape (http, https, with/without www). | ["dr-johanna-budwig.de"] |
enableAi | Boolean | No | Use AI to flawlessly extract and format data fields. Falls back to Regex if disabled. | false |
aiModel | String | No | Select the AI model to use for extraction (if enableAi is true). | openai/gpt-oss-20b |
maxConcurrency | Number | No | How many pages to process at the same time. Maximum is 10 to ensure stability. | 10 |
🧩 Example Input
{"startUrls": ["https://www.dr-johanna-budwig.de/", "https://www.marley.de"],"enableAi": true,"aiModel": "openai/gpt-oss-20b","maxConcurrency": 10}
📊 Example Output
{"originalUrl": "https://www.dr-johanna-budwig.de/","impressumUrl": "https://www.dr-johanna-budwig.de/policies/legal-notice","address": "An den Kolonaten 2-4, 26160 Bad Zwischenahn, Germany","emails": ["kontakt@dr-johanna-budwig.de"],"phones": ["0441 390 630 0"],"registerNumber": "HRB209987","taxId": "DE300469959","socialLinks": {"pinterest": "https://www.pinterest.com/drjohannabudwig","tiktok": "https://www.tiktok.com/@dr.johannabudwig","facebook": "https://web.facebook.com/Dr.Johanna.Budwig/?locale=de_DE&_rdc=1&_rdr#","instagram": "https://www.instagram.com/dr.johannabudwig/?hl=de","youtube": "https://www.youtube.com/user/DrJohannaBudwig"},"metadata": {"title": "Dr. Johanna Budwig","description": "Dr. Johanna Budwig","h1": "Frühstücksöl Limette & Orange"},"companyName": "Dr. Johanna Budwig GmbH"}
💡 Use Cases
- B2B Lead Generation: Enrich datasets with verified German company contact data (Emails & Phones).
- Compliance Checks: Validate if a target website meets German Impressum law requirements.
- CRM Database Enrichment: Clean up missing legal entity names and tax IDs in Salesforce or HubSpot.
- Competitive Market Research: Analyze contact and social media details from thousands of competitor domains in minutes.
❓ FAQs
1. Do I need my own API key for the AI? No! Unlike other scrapers, our Actor has advanced, load-balanced Groq API keys built directly into the backend. You just pay the flat Apify PPE rate per domain.
2. How am I charged for this Actor? This is a Pay-Per-Event (PPE) actor. You are charged a flat rate per domain successfully processed. If the scraper cannot find an Impressum page, you are not charged anything for that URL. You only pay for what you scrape!
3. Will the AI scrape Impressum data hidden in an image? No. Under German law (DDG), Impressum data must be easily readable text. Our scraper strips images to optimize for extreme speed and cost-efficiency. It will only extract text-based legal notices.
🧑💻 Developer Info
Author: codescraper Email: codescraper011@gmail.com Platform: Apify Language: TypeScript (Cheerio Crawler)
🏷️ Tags
German Impressum Detector · impressum · germany · legal-scraper · company-data · contact-scraper · cheerio · apify · automation · b2b-leads · lead-generation