FAQ Question Answer Extractor
Pricing
from $6.50 / 1,000 faq source page scanneds
FAQ Question Answer Extractor
Extract FAQ question-answer pairs from visible HTML and FAQPage schema, reconcile duplicates, and flag schema/visible mismatches
FAQ Question Answer Extractor
Pricing
from $6.50 / 1,000 faq source page scanneds
Extract FAQ question-answer pairs from visible HTML and FAQPage schema, reconcile duplicates, and flag schema/visible mismatches
Public URLs containing visible FAQs or FAQPage structured data.
[]Optional public page URLs; pageUrls is preferred for role clarity.
[]Optional public XML sitemap URLs. Accepted pages remain bounded by maxPages.
[]Optional records with sourceUrl and html or currentHtml for deterministic extraction.
[]Extract FAQPage structured data when present.
Compare visible FAQ text against FAQPage structured data.
Deduplicate repeated questions across scanned pages.
Optional case-insensitive phrases. Only questions containing at least one phrase are retained.
[]Maximum FAQ question-answer pairs to emit per page.
Optional hostname allowlist for fetched pages.
[]Maximum pages to fetch in one run.
General link discovery is disabled.
Include short source evidence snippets in output rows.
Opt in to raw HTML artifacts in key-value storage.
Delay in milliseconds between outbound page requests.
Maximum time in milliseconds to wait for a page request.
User agent profile to use for public page requests.
Maximum estimated PPE spend before the actor exits gracefully.