Quora Dataset Builder
Pricing
from $1.00 / 1,000 results
Go to Apify Store
Quora Dataset Builder
Build bounded normalized Quora datasets from public URLs or native search seeds.
Quora Dataset Builder
Pricing
from $1.00 / 1,000 results
Build bounded normalized Quora datasets from public URLs or native search seeds.
Overview emits page-level summaries without answer/comment/activity/feed expansion; detail preserves the current bounded traversal.
Question, answer, profile, topic, Space, or native search URLs.
[]Native Quora search seeds. Search cards are followed only when their visible URLs are supported.
[]Explicit question-page seeds. Only question routes are accepted.
[]Explicit answer-page seeds. Parent question joins are emitted only from observed evidence.
[]Explicit Topic-page seeds for bounded visible question discovery.
[]Explicit Space-page seeds for bounded visible feed discovery.
[]Explicit profile-page seeds. Only profile routes are accepted; emitted as normalized profile rows exactly like a profile route discovered from startUrls.
[]Explicit standalone Post-page seeds. Only post routes are accepted; emitted with the same topic_post field family topic-discovered posts use (post_title, post_body, post_author_*).
[]Select normalized entities, explicit graph relationships, relationship-collection counts, status rows, and the final run manifest.
[ "question", "answer", "comment", "profile", "profile_activity", "topic", "topic_post", "space", "space_feed_item", "entity_edge", "status", "run_manifest"]Projection after one canonical extraction path. Nested groups emitted answers and comments under explicit question/answer parents; flat keeps entity rows with parent context; edges emits explicit edge rows plus status rows.
canonical suppresses overlapping entity rows by stable ID or canonical URL (default); none preserves every observed non-status row. The run manifest reports duplicate_records_suppressed.
Annotate canonical entity rows with new/changed/unchanged/scope_mismatch state and suppress only unchanged semantic rows when includeUnchanged is false. Requires stateStoreId; status, graph, and manifest rows remain visible.
Retain unchanged semantic entities when incremental state is enabled; output projection does not alter comparison hashes.
Apify key-value store for canonical entity snapshots. Required when onlyNewItems is enabled; never store cookies or credentials here.
Optional allowlist for compact rows, including bounded dotted paths through nested collections (for example answers.answer_text, comments.comment_text, feed.feed_text, or contributors.contributor_name where applicable). Child identity/URL fields and the actor identity, access/status, source, traversal, graph, filter, and parser-envelope fields remain protected.
[]Include field_sources and fields_present maps. Disable only when a smaller payload is required; semantic values and protected access fields remain unchanged.
Retain relationship rows whose endpoint URL was not explicitly observed, with unresolved_reason. Disabling this suppresses only those edge rows and does not imply complete relationships.
Explicit relationship-type allowlist for the newly ported follower/following/want-answers/upvoter/space-role relationship collections. Empty leaves the existing always-on content-membership/authorship edges untouched (question_has_topic, answer_authored_by, comment_belongs_to_answer, profile_has_activity, space_has_feed_item, etc.); this allowlist only gates the additional collections below. Sensitive follower/want-answer/upvoter types also require includeSensitiveRelationshipData and their matching collection opt-in.
[]Enable explicitly labeled profile follower/following/followed-question/Space/Topic relationship collections. This does not prove a complete list.
Enable explicitly labeled question follower or want-answer collections. Sensitive identity rows also require the privacy gate.
High-sensitivity opt-in for explicitly labeled answer upvoter identities; requires includeSensitiveRelationshipData=true.
Explicit privacy gate required for follower, want-answer, and upvoter identity edges.
Retain explicitly visible target link text and image (to_name, to_image_url) on relationship edges. URLs remain evidence; no target navigation or recursive target scraping is performed.
Emit privacy-preserving relationship_collection rows with the real observed total (collection_total_text/collection_total_value) but no target identities. Requires relationship_collection in resultTypes. Observed totals are not always a server-confirmed count.
Enable explicitly labeled Space contributor, moderator, and admin collections.
Retain bounded explicit Space post/question membership edges; this never claims a complete Space inventory.
Bound explicit links emitted from each labeled relationship collection. Zero emits no relationship-collection edges and never means an empty collection.
Maximum question seeds or discovered question pages to process.
Maximum visible answer rows retained from each question page.
Global answer-row cap applied after per-question caps. Zero means no configured global answer cap.
Retain normalized question parent rows when selected in resultTypes.
Retain visible author fields and profile/profile-activity rows. When false, author fields and author rows are removed from the emitted shape; no contact enrichment is performed.
Opt into bounded public profile snapshots for visible author profile links; profile access/status is separate from answer/comment success and never includes contact enrichment.
Retain explicit topic links and topic relationships where observed.
Retain Space rows and feed items where the source route exposes them.
Enable bounded visible comment output when the per-answer cap is positive.
Retain explicitly labeled per-comment reaction breakdowns alongside aggregate reaction labels. Raw labels are preserved and typed companions are added only when the localized number is unambiguous.
Retain links inside an explicitly labeled related/attached region scoped to answer cards; unlabeled navigation and neighboring regions are excluded.
Bound eligible related-content links per answer when enabled; zero emits no related-content items.
Retain only comments with an explicit visible Top/Most helpful/Popular comments label. Unknown comments are excluded and counted; votes and position are not used as top evidence.
Order builder comment rows by observed order or an explicitly typed metric/timestamp. Unknown selected metrics remain in stable fallback order and are counted.
Enable bounded visible nested comment replies and explicit parent paths; this also enables comment rows. Hidden or unexpanded replies are not inferred.
Request explicitly rendered media URL fields scoped to question roots, answer/comment nodes, and Topic post cards where the page exposes them.
Request bounded visible text from explicitly linked Space feed cards; page-level text and complete feed history are excluded.
Request visible non-Quora links scoped to question roots, answer/comment nodes, and Topic post cards; Quora links, comment permalinks, and links outside the owning scope are excluded.
Emit explicit relationship rows even when entity_edge is not listed in resultTypes.
Maximum visible comment rows retained per answer when enabled.
Maximum visible nested reply depth when includeReplyGraph is enabled; reaching the bound is not complete history.
Cap for normalized entity/status rows before graph and manifest rows.
Independent cap for deterministic entity_edge rows.
Bounded lazy-render scroll budget applied during detail traversal; scroll attempts and termination state are reported.
Bound exact visible continuation-control clicks on direct question, answer, topic, and Space collection pages independently from scrolling. Search seeds fan out through their discovered child URLs and retain their own source accounting.
Add bounded document hash and optional HTML bytes to each page-level entity row.
Include bounded raw-document HTML when raw evidence is enabled and bounded answer-card HTML when answer cards are extracted.
Maximum UTF-8 bytes retained for optional raw HTML evidence.
Delay between page batches in seconds; this is not an anti-bot bypass guarantee.
Maximum independent seed pages processed concurrently in deterministic batches. Each seed owns a page; discovered seeds merge in batch order.
Browser language hint. Short language codes and supported regional aliases are accepted and normalize to the short code; auto keeps the actor default. This does not translate content or prove access.
Apify proxy routing. Quora's anti-bot layer blocks the free Datacenter/Automatic tier on roughly half of all real requests (cloud-verified), so this actor always routes through RESIDENTIAL proxies for reliable results regardless of what is selected here — leave this on the default Residential selection unless you need to supply your own proxy URLs.
{ "useApifyProxy": true, "apifyProxyGroups": [ "RESIDENTIAL" ]}Redact narrow credential-like query parameter values in emitted URLs, graph endpoints, redirect chains, and raw HTML. The original document hash remains based on the captured document.
When zero, suppress raw_html snapshots while retaining hash, size, and redaction diagnostics. Positive values permit bounded HTML for the requested run; external Apify dataset/run retention still applies.
Hard upper byte cap for returned raw HTML after the maxHtmlBytes cap.