Quora Dataset Builder
Pricing
from $1.00 / 1,000 results
Quora Dataset Builder
Build bounded normalized Quora datasets from public URLs or native search seeds.
Pricing
from $1.00 / 1,000 results
Rating
0.0
(0)
Developer
Crawler Bros
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
21 days ago
Last modified
Categories
Share
Normalized question, answer, comment, topic-post, and Space-feed rows expose type-specific age fields when explicit machine-readable timestamps are present: question_age_*, answer_age_*, comment_age_*, post_age_*, and feed_age_*. Topic-post rows additionally expose post_updated_at, post_updated_age_* when a scoped update/edit marker survives the payload boundary. Relative, missing, and malformed dates fail closed, and future dates are clamped to zero. The same evidence survives normalized, flat, nested, and edge-oriented output projections.
When enabled, includeMedia and includeOutboundLinks apply to bounded question roots, Topic post cards, answer cards, and comment nodes, emitting combined plus typed image/video channels (question_*, post_*, answer_*, and comment_* media fields) alongside scoped outbound links; includeHtml additionally exposes bounded answer_html for answer-card rows. These are scoped observations, not complete standalone Post or Answer extraction.
Comment rows optionally expose positive-only comment_author_is_verified from an exact accessible verification phrase inside the same comment-author scope; it never inherits parent answer, post, or profile state.
Fixture-backed examples (synthetic)
These examples describe fixture-shaped output, not a live Quora dataset. Graph joins require explicit endpoint evidence; unresolved edges remain unresolved.
{"content_type":"entity_edge","edge_id":"fixture-edge-1","from_content_type":"question","to_content_type":"topic","relationship":"question_has_topic","edge_state":"resolved","access_state":"visible","source_items_emitted":1,"source_cap_reached":false,"field_sources":{"relationship":"derived_signals","from_record_id":"url_or_visible_link","to_record_id":"url_or_visible_link"}}
{"content_type":"status","access_state":"cloudflare","http_status":403,"source_items_emitted":0,"source_cap_reached":false,"field_sources":{"access_state":"runtime_or_access"}}
Authorized session input accepts either structured cookies or the secret cookieString browser-header form (name=value; name2=value2). Structured cookies win on duplicate names; malformed pairs are ignored, the raw header is never emitted, and string cookies have no invented expiry metadata. This is an access input, not a bypass or private-content guarantee.
commentSortOrder supports observed, upvotes_desc, reactions_desc, replies_desc, and newest. Dataset Builder converts only unambiguous localized comment counts into typed values before sorting; missing values remain in stable fallback order and are counted in comment_sort_fallback_count. includeEngagementBreakdown defaults to true for backward compatibility and can disable only the per-reaction decomposition; aggregate reaction labels remain available, while icons and aggregate-only counts are never decomposed. includeRelatedContent independently enables bounded answer-card related/attached links, controlled by maxRelatedContent; zero or disabled emits no related-content items.
maxItemsPerSource is a common semantic cap across direct, query, and typed sources. It is applied after extraction but before output projection; status and run-manifest rows remain visible. Capped rows expose source_cap, source_items_observed, source_items_emitted, and source_cap_reached. A value of 0 preserves unlimited per-source behavior; it never means complete Quora history. Every emitted row also echoes the normalized boundary as max_items_per_source for run-level reconciliation.
Nested projection is supported with bounded dotted paths such as answers.answer_text, comments.comment_text, feed.feed_text, or contributors.contributor_name where that collection exists. Child identity/URL fields remain protected, and an empty list keeps the full backward-compatible row. Answer projections additionally preserve answer_translation_state/answer_translation_label for explicit answer-card translation markers and answer_original_url only for a visible original-answer link; locale and language mismatch are not evidence.
Answer rows also expose answer_is_anonymous only when the scoped author region visibly says Anonymous; absent authors are not treated as anonymous.
Answer rows also expose conditional answer_author_is_verified only for an exact accessible verification marker inside the bounded answer-author scope; credentials, profile URLs, generic icons, and neighboring cards remain unknown.
Answer rows also expose conservative card-state evidence: answer_visibility and answer_text_completeness use only exact owning-card paywall/continuation controls, answer_is_collapsed and answer_is_quora_plus are positive-only markers, and answer_is_top_labeled is emitted only for an exact Top/Best/Most helpful label. answer_rank_source is explicit_label only in that case; otherwise observed_order describes bounded DOM order rather than a global ranking.
Answer rows expose answer_updated_at only when a same-card time/metadata node is explicitly labeled updated, edited, modified, or carries data-updated-at; the first unlabeled time is never reclassified as an edit date.
Answer rows also preserve raw card-local answer_upvote_count, answer_comment_count, answer_view_count, and answer_share_count labels with typed _value companions when localization is unambiguous. Missing controls are omitted, raw labels are never replaced by zero, and answer_comment_count is not the number of rows collected.
When includeEngagementBreakdown is enabled, answer rows additionally expose reaction_breakdown only for explicit labeled like/love/heart/helpful/reaction controls in the owning answer card. Upvotes, aggregate-only counts, icons, comment controls, and neighboring-card controls are excluded; disabling the input removes both answer and comment breakdown arrays while retaining aggregate raw labels.
Answer rows also expose is_ai_answer, ai_label_text, and ai_label_source only for an exact same-card AI, AI-generated answer, or generated by Quora marker. Class names, answer wording, model-like prose, and neighboring labels never create AI state.
Answer rows expose answer_jsonld_type only from an Answer JSON-LD node whose URL exactly matches the answer URL; neighboring structured nodes and route-shaped inference are excluded.
Answer, comment, and bounded topic-post rows optionally expose author_credentials when a credential or tagline is explicitly visible in the local author region; the field is omitted when ambiguous and is not an expertise-verification claim.
Comment rows additionally expose comment_author_credentials from that same bounded comment-author region; it is explicit context, not expertise verification, and remains separate from answer/post author credentials.
Bounded topic-post rows also expose conditional post_author_is_verified only for an exact accessible verification marker inside that post card's author scope; profile URLs, credentials, generic icons, and neighboring cards remain unknown.
Bounded topic-post rows also expose optional post_author_credentials from the local post-author region. Isolated and sequential Dataset Builder paths use the same bounded evaluator and provenance; generic author_credentials remains compatibility context and does not verify expertise.
Bounded topic-post rows expose post_author_is_anonymous only for an exact visible Anonymous/Anonymous User/Anonymous Contributor marker inside that post card's author scope; missing identity is not anonymity evidence.
Exact high-value output names: ../RESEARCH/README_FIELD_REFERENCE.md#dataset-builder.
When cookies are supplied, session_cookie_expiry_state reports only input expiry metadata; it does not prove that Quora accepted the session.
Build a bounded, normalized Quora dataset from direct public URLs and native Quora search seeds. The actor visits supported pages with Playwright, keeps question→answer→comment and profile/topic/Space URL joins, optionally attaches raw-document evidence, and emits deterministic entity_edge rows plus a reconciliation manifest.
Question rows preserve question_created_at and question_updated_at only when an explicitly URL-matched Question JSON-LD node supplies dateCreated, dateModified, or dateUpdated. Neighboring Question nodes, relative labels, and scrape time never populate these fields; absent structured evidence is omitted.
When answer rows contain explicit machine-readable creation dates, question rows also expose last_answered_at as the maximum date observed in the bounded answer set, plus last_answered_at_evidence with the observed/date-bearing counts and bounded_observation completeness state. It is never presented as the global latest answer or complete answer history.
scrapeDepth defaults to detail, preserving the existing bounded traversal. overview emits page-level summaries without answer, comment, profile-activity, or Space-feed expansion. Every row reports requested_depth, observed_depth, and detail_expansion_state; bounded means a configured traversal budget was reached, not complete Quora history, and blocked pages report observed_depth: "status".
The actor is deliberately fail-closed. Cloudflare, login, Quora+, removed pages, and navigation failures become typed status rows; they are never parsed as entity records. Cookies and proxies are optional authorized access controls, not bypass guarantees.
Relative labels remain raw; when a scrape reference is available, date_interpretations exposes conservative bounded intervals for supported prefix, suffix, reversed-order, and non-Latin forms (for example vor 2 Tagen, 3 giorni fa, 3日前, and 3天前) and never claims an invented exact timestamp. Unsupported labels remain raw without a guessed date.
The final run_manifest includes bounded input_accounting entries for every expanded source seed, including zero-row sources; entity/status counts exclude derived graph rows so downstream reconciliation remains additive.
Output projections
outputFormat changes the consumer shape after the same canonical extraction and provenance path:
-
nestedemits each emitted question as one parent row with an orderedanswersarray; each answer carries an orderedcommentsarray when explicit parent URLs resolve. Unresolved children remain top-level withnested_orphan_reason. Requested explicit graph rows remain top-level companions. Counts describe emitted/observed bounded rows, never Quora totals or complete history. -
normalized(default) emits the selected normalized entities, status rows, optional graph rows, and manifest. -
flatemits one row per selected entity while retaining stable identity, parent URL context, access state, source URL, timing, and field-level provenance. -
edgesemits explicitentity_edgerows plus status rows, which is convenient for graph loading without dropping access diagnostics.
includeUnresolvedEdges defaults to true. When an observed relationship has no explicit endpoint URL, the builder can retain it as edge_state: "unresolved" with unresolved_reason: "missing_explicit_target_url"; setting it to false suppresses only those edge rows. It never turns names or IDs into guessed joins and never implies complete relationship history. Every projection carries output_format, record_id, access_state, source_url, scrape_timestamp, and provenance fields when the source row has them.
outputFields is an optional output-field allowlist applied after extraction. An empty list preserves the full row contract. The builder always protects content_type, identity (record_id, url, and canonical/source URL fields when present), access/status fields, depth state, scrape timing, and parser provenance so a small payload cannot become an untraceable success. includeFieldEvidence defaults to true and controls the field_sources and fields_present maps; output_projection reports full or selected, and the run manifest records both projection choices.
Question rows keep the visible question_answer_count label separate from bounded extraction: question_answer_count_value is added only when that raw label parses unambiguously, question_answer_count_scraped reports answer rows extracted for the question, and answer_limit_reached reports a reached per-question cap only when additional visible answer cards were observed. The existing answer_count remains the extracted-row count for compatibility; none of these fields claims complete answer history. Known limitation: Quora's anonymous view can expose noticeably fewer answers than maxAnswersPerQuestion allows on its very highest-traffic questions (observed on a real 9.5K-answer question, which returned as few as 1-2 visible answers across separate cloud runs even with scroll retries and a generous cap) - this reflects a real, variable per-session ceiling on Quora's anonymous answer feed rather than a block, a non-200 response, or a bug in the traversal; scroll_state: "exhausted" with answer_limit_reached: false on such a question is a genuine (if disappointing) result, not a silent failure.
dedupeMode defaults to canonical, suppressing overlapping entity rows by stable entity ID or canonical URL before graph projection. Set it to none to preserve every observed non-status row. The run manifest exposes the selected dedupe_mode and duplicate_records_suppressed; status rows are never suppressed.
Mixed direct, typed, and /search?q= seeds preserve matched_input, matched_input_kind, bounded source_observations, source_observation_count, and (when overlapping search seeds are present) source_queries. The first exact source is retained without changing canonical identity; duplicate observations are merged before projection, and these provenance fields remain protected by outputFields.
Important controls
startUrls: question, answer, profile, topic, Space, or/search?q=...URLs.questionUrls,answerUrls,topicUrls,spaceUrls,profileUrls, andpostUrls: typed seed lists with route validation; use these when mixedstartUrlswould make source accounting ambiguous.profileUrlsseeds a normalprofilerow;postUrlsseeds atopic_post-shaped row for a standalone Post page. Any non-emptyspaceUrlslist automatically enablesincludeSpacesfor that run, since giving an explicit Space seed is itself a request for Space content. Known limitation: Quora has migrated many legacy Topic (and some Question/Profile) URLs to redirect straight to a per-topic<slug>.quora.comSpace subdomain; when that happens to atopicUrls/startUrlsseed andincludeSpacesis not enabled, the actor reports it transparently as astatusrow (redirected_to_space, with the real redirect target infinal_url) instead of silently returning nothing. The same transparentstatusrow (asspace_hidden_by_config) also covers a Space URL pasted directly intostartUrlswithincludeSpacesleft off - sincestartUrlsdocuments accepting Space URLs directly andspaceships in the defaultresultTypes, this is a realistic path to hit, not just the redirect case. Known limitation: aprofileUrls/questionUrls/postUrlsseed can also land on Quora's Topic page for that same named entity (e.g. a well-known person's profile URL redirecting to their Topic page) - the actor always reports this transparently as an additionalstatusrow (redirected_to_topic, with the real redirect target infinal_url) alongside normal extraction, so aresultTypesselection that only lists the original entity (e.g.profile) is never left with a silent zero-row result.- Relationship graph:
relationshipTypesis an explicit allowlist for follower/following/want-answers/upvoter/space-role relationship collections observed on a profile, question, answer, or Space page (e.g.profile_follows_profile,space_has_contributor). Each family has its own opt-in (includeProfileEdges,includeQuestionFollowerEdges,includeAnswerUpvoterEdges,includeSpaceRoleEdges) and identity-revealing families additionally requireincludeSensitiveRelationshipData=true.includeTargetMetadataretains the visible target name/image on each edge;includeRelationshipCounts(withrelationship_collectioninresultTypes) emits privacy-preserving count-only rows (relationship_collection_key,collection_total_text,collection_total_value) with no target identities.maxRelationshipItemsPerCollectionbounds how many links one labeled collection contributes. Existing content-membership/authorship edges (question↔topic, answer↔question, comment↔answer, profile↔activity, Space↔feed item) remain always-on and are not gated by this allowlist. The run manifest'srelationship_collection_diagnosticfield explains, for the types you actually requested, whether real edge rows came back, whether Quora only exposed a count (no identities), whether a matching collection was found but suppressed by a config gate, or whether no matching collection was found at all - so a zero-row result is never silent.- What is real today: every relationship type's member count is real whenever Quora renders it (e.g. a profile's
126.3K Followers/5.3K Followerstab, or a bareFollowingtab) -includeRelationshipCounts+relationship_collectioninresultTypessurfaces these ascollection_total_value/collection_total_textwith no identities. Member identities are real today only for a profile's Followers collection (profile_followed_by_profile): Quora renders identities for arole="tab"Followers control only after it is clicked (a client-side re-render, not present in the initial DOM), so this actor clicks that tab and captures the newly rendered<a href="/profile/...">links asentity_edge/nested edge rows (edge_evidence_field: "clicked_tab_rerender"). - Known limitation: identity lists for Following, followed Questions/Spaces/Topics, Expertise, question Want-Answers/Followers, answer Upvoters, and Space Contributors/Moderators/Admins/content are not yet extracted - Quora renders their member identities only after the same kind of client-side interaction (tab click, "load more", etc.) that Followers now uses, and that click-and-diff behavior has only been implemented and verified for the Followers tab. Requesting those types still yields real counts (where visible) via
relationship_collection, but no member-identity edge rows.
- What is real today: every relationship type's member count is real whenever Quora renders it (e.g. a profile's
searchQueries: native Quora search seeds; visible supported links are queued with bounded caps. Known limitation: Quora frequently redirects an unauthenticated/search?q=...request straight to its (geo-localized) homepage instead of rendering results - cloud-verified across multiple residential exit countries and bothen/default locale. The actor reports this transparently as astatusrow (login_requiredorunsupported_url, matching the real redirect target) rather than fabricating search results; supplyingcookies/cookieStringfor an authorized session is the reliable way to get native search results.maxQuestions,maxAnswersPerQuestion,maxAnswers,maxCommentsPerAnswer, andmaxItems: independent traversal/entity limits.maxAnswersis a global post-extraction cap and its suppression count is reported on the manifest.includeQuestion,includeAuthor,includeAuthorSnapshot,includeTopics,includeSpaces,includeComments,includeReplyGraph,includeMedia,includeFeedText,includeRelatedContent, andincludeEntityGraph: explicit dataset-shape controls;includeReplyGraphenables bounded visible nested comments and explicit parent paths, and automatically enables comment rows.maxCommentDepthbounds that graph; comment rows exposecomment_depth_limit_reachedwhen deeper visible comments were omitted. Comment rows also expose source-scopecomment_items_observed_count,comment_items_emitted_count, andcomment_limit_reached, separating distinct eligible comments seen beforemaxCommentsfrom rows retained after the cap.includeFeedTextindependently enables bounded visible Space feed-card text.includeRelatedContentcollects only links inside explicitly labeled related/attached regions of answer cards and usesmaxRelatedContentas a per-answer cap.includeAuthor=falseremoves author fields and profile/profile-activity rows. When both author controls are enabled,includeAuthorSnapshotqueues only visibly linked Quora profile URLs, emits bounded profile rows with visible/Person-JSON-LD name, bio, and image evidence, and reports separateprofile_snapshot_state; it never performs contact enrichment.- Profile snapshot rows additionally expose conditional
profile_author_idonly from an explicit Person JSON-LD identifier, plusprofile_is_verifiedandprofile_is_restrictedonly for exact visible verification/private markers in the profile shell; route slugs, missing identifiers, and missing markers remain unknown and activity-card content is excluded from the marker scope.
topCommentsOnly retains only comments with an explicit scoped Top/Most helpful/Popular label. Unknown comments are excluded and counted in comment_filter_unknown; comment position, votes, wording, and missing labels never establish top status.
Comment rows expose positive-only comment_is_anonymous evidence for an exact Anonymous/Anonymous User/Anonymous Contributor marker in the scoped author region; missing identity is never treated as anonymous.
scrapeDepth:overviewfor page-level summaries ordetailfor bounded child traversal (default).includeRawEvidenceandincludeHtml: optional bounded full-document hash/byte evidence.redactSensitiveQueryParams: defaults to true and redacts narrow token/session-style query values in emitted URLs, graph endpoints, redirect chains, and returned raw HTML. The original-document hash/byte count remain truthful; returned-byte and redaction-state fields make sanitation explicit. This does not remove arbitrary personal or secret text from page content.- The raw-evidence sanitation pass also redacts narrow assignment/header patterns and exact sensitive keys in raw structured payloads; ordinary semantic text is preserved and arbitrary secrets are not guaranteed to be detected.
rawEvidenceRetentionDays: bounded 0–30 snapshot window requested for the run.0suppressesraw_htmlwhile retaining hash/size and diagnostics (raw_evidence_state: "metadata_only"); positive values permit capped HTML (retained_bounded).maxRawEvidenceBytesadds an independent hard cap aftermaxHtmlBytes. This controls actor output, not external Apify dataset/run retention.maxGraphEdges: independent deterministic relationship cap.requestDelay,maxConcurrency,language,cookies, andproxyConfiguration: explicit batch pacing, bounded independent seed concurrency, locale, authorized session, and routing controls. Above the default concurrency, each seed uses an isolated page; discovered seeds are merged deterministically in batch order.proxyConfigurationalways routes through Apify's RESIDENTIAL proxy tier by default (and falls back to it even if the field is left blank or set to the free Datacenter/Automatic tier), because cloud verification found the free tier gets blocked by Quora's anti-bot layer on roughly half of all requests while RESIDENTIAL reliably reaches the real page. Set an explicitapifyProxyGroupsvalue or your ownproxyUrlsonly if you deliberately want a different routing. ThemaxNavigationRetriesinput (default2) retries transient navigation exceptions up to three times with bounded exponential backoff; attempts, retries, and actual delays are emitted as runtime provenance.
Rows use content_type values such as question, answer, comment, profile, profile_activity, topic, topic_post, space, space_feed_item, entity_edge, status, and run_manifest. Question rows preserve bounded question-root author name/URL/structured ID plus positive-only verification/anonymity states, media, and external links only after nested answer/comment/post regions are removed; answer, comment, Topic post, and Space feed rows preserve their own scoped media/link evidence. Space feed rows additionally expose bounded feed_type, feed_title, optional feed_text, author name/URL/structured ID, positive-only feed_author_is_verified and feed_author_is_anonymous, explicit feed_author_credentials from the local feed-card author region, timestamp, typed media, and external links when the owning feed card renders them. feed_author_is_anonymous requires an exact Anonymous/Anonymous User/Anonymous Contributor marker and no profile link; missing identity is not anonymity evidence. Question author IDs require matched Question JSON-LD and question-author markers never come from answer cards. Comment rows also preserve explicit author IDs, machine/relative timestamps, visibility labels, scoped upvote/reaction/reply-count labels, labeled collection state, and graph fields when those comment-local DOM signals exist. includeMedia and includeOutboundLinks are independent opt-ins; Quora links, comment permalinks, page-level/sidebar assets, and evidence outside the owning scope are excluded. A comment_reply_count is a source label only and is never treated as the number of collected rows. Comment graph rows preserve comment_depth, comment_is_reply, explicit parent IDs/URLs, and comment_reply_path only when all observed parent hops are present; missing parent endpoints remain unresolved rather than fabricated. Topic pages expose topic_post only for explicitly linked visible /post/... records; these rows preserve bounded card body, title, author, timestamp, and link evidence, but do not claim complete standalone-Post extraction. Every row has a typed record_id, access_state, parser version, field-level evidence channels, and scrape timing. Graph rows are only created when endpoint URLs are explicitly present; names are never used as joins.
This is a bounded public-page dataset builder, not a guarantee of complete Quora history. Visible pagination, access state, and Quora layout determine what can be populated.
Question rows also expose optional question_author_image_url when an image is visibly attached to the profile link in the bounded question-author container, and question_author_credentials only when an explicit credential or tagline is visible there. Neither value is inferred from question prose, neighboring answers, profile navigation, or generic author text.
Complete schema input inventory
This generated inventory mirrors every property in .actor/input_schema.json; enum values are retained in the description below.
| Input field | Type | Default | Description |
|---|---|---|---|
scrapeDepth | string | "detail" | Overview emits page-level summaries without answer/comment/activity/feed expansion; detail preserves the current bounded traversal. Allowed values: overview, detail. |
startUrls | array<any> | [] | Question, answer, profile, topic, Space, or native search URLs. |
searchQueries | array<string> | [] | Native Quora search seeds. Search cards are followed only when their visible URLs are supported. |
questionUrls | array<object> | [] | Explicit question-page seeds. Only question routes are accepted. |
answerUrls | array<object> | [] | Explicit answer-page seeds. Parent question joins are emitted only from observed evidence. |
topicUrls | array<object> | [] | Explicit Topic-page seeds for bounded visible question discovery. Quora's current Topic feed links each item straight to its top/featured answer permalink (/<question-slug>/answer/<author>), never to a bare question URL, so question_links is populated by deriving the parent question URL from each visible answer link (deduped against any bare question links also present) rather than by matching bare-question URLs alone - cloud-verified 2026-08-27 across 4 large real topics (Science, Technology, Movies, Travel), each recovering real, followable question URLs. Known limitation: Quora's Topic feed currently surfaces only question/answer items, never standalone /post/... links, so post_links is reliably empty for every topic today - this reflects real current feed composition, not a bug or a block. Occasionally a topic's feed has not finished its client-side render by the time of extraction (intermittent, unrelated to authentication); a topic row with an empty feed and access_state: "public" in that rare case reflects a genuine unauthenticated view of that page at that moment, not a bug. The topic page-level row (title, URL) itself is always reliable. |
spaceUrls | array<object> | [] | Explicit Space-page seeds for bounded visible feed discovery. |
profileUrls | array<object> | [] | Explicit profile-page seeds. Only profile routes are accepted; emitted as normalized profile rows exactly like a profile route discovered from startUrls. |
postUrls | array<object> | [] | Explicit standalone Post-page seeds. Only post routes are accepted; emitted with the same topic_post field family Topic-discovered posts use (post_title, post_body, post_author_*). |
resultTypes | array<string> | ["question","answer","comment","profile","profile_activity","topic","topic_post","space","space_feed_item","entity_edge","status","run_manifest"] | Select normalized entities, explicit graph relationships, relationship_collection counts, status rows, and the final run manifest. relationship_collection is opt-in (also requires includeRelationshipCounts). |
outputFormat | string | "normalized" | Projection after one canonical extraction path. Nested groups explicit answers/comments under safely matched parents; flat keeps entity rows with parent context; edges emits explicit edge rows plus status rows. Allowed values: normalized, nested, flat, edges. |
dedupeMode | string | "canonical" | canonical suppresses overlapping entity rows by stable ID or canonical URL (default); none preserves every observed non-status row. The run manifest reports duplicate_records_suppressed. Allowed values: canonical, none. |
onlyNewItems | boolean | false | Requires stateStoreId; annotates canonical entities with new, changed, unchanged, or scope_mismatch and suppresses only unchanged semantic rows when includeUnchanged is false. A mixed/partial source run is fail-closed and does not commit state; status, graph, and manifest rows remain visible. |
includeUnchanged | boolean | true | Retain unchanged entities when incremental state is enabled; output format and projection do not alter semantic comparison hashes. |
stateStoreId | string | "" | Apify key-value store for canonical entity snapshots; required with onlyNewItems and never used for credentials or cookies. |
outputFields | array<string> | [] | Optional allowlist (up to 3,000 unique names) for compact rows, including bounded dotted paths through nested collections (for example answers.answer_text, comments.comment_text, feed.feed_text, or contributors.contributor_name where applicable). Child identity/URL fields and the actor identity, access/status, source, traversal, graph, filter, and parser-envelope fields remain protected. |
includeFieldEvidence | boolean | true | Include field_sources and fields_present maps. Disable only when a smaller payload is required; semantic values and protected access fields remain unchanged. |
includeUnresolvedEdges | boolean | true | Retain relationship rows whose endpoint URL was not explicitly observed, with unresolved_reason. Disabling this suppresses only those edge rows and does not imply complete relationships. |
relationshipTypes | array<string> | [] | Explicit relationship-type allowlist for the newly ported follower/following/want-answers/upvoter/space-role relationship collections. Empty leaves the existing always-on content-membership/authorship edges untouched; this allowlist only gates the additional collections. Sensitive follower/want-answer/upvoter types also require includeSensitiveRelationshipData and their matching collection opt-in. |
includeProfileEdges | boolean | false | Enable explicitly labeled profile follower/following/followed-question/Space/Topic relationship collections. |
includeQuestionFollowerEdges | boolean | false | Enable explicitly labeled question follower or want-answer collections. Sensitive identity rows also require the privacy gate. |
includeAnswerUpvoterEdges | boolean | false | High-sensitivity opt-in for explicitly labeled answer upvoter identities; requires includeSensitiveRelationshipData=true. |
includeSensitiveRelationshipData | boolean | false | Explicit privacy gate required for follower, want-answer, and upvoter identity edges. |
includeTargetMetadata | boolean | false | Retain explicitly visible target link text and image (to_name, to_image_url) on relationship edges. URLs remain evidence; no target navigation or recursive target scraping is performed. |
includeRelationshipCounts | boolean | false | Emit privacy-preserving relationship_collection rows with the real observed total (collection_total_text/collection_total_value) but no target identities. Requires relationship_collection in resultTypes. |
includeSpaceRoleEdges | boolean | false | Enable explicitly labeled Space contributor, moderator, and admin collections. |
includeSpaceContentEdges | boolean | true | Retain bounded explicit Space post/question membership edges; this never claims a complete Space inventory. |
maxRelationshipItemsPerCollection | integer | 500 | Bound explicit links emitted from each labeled relationship collection. Zero emits no relationship-collection edges and never means an empty collection. |
maxQuestions | integer | 25 | Maximum question seeds or discovered question pages to process. |
maxAnswersPerQuestion | integer | 10 | Maximum visible answer rows retained from each question page. |
maxAnswers | integer | 0 | Global answer-row cap applied after per-question caps. Zero means no configured global answer cap. |
includeQuestion | boolean | true | Retain normalized question parent rows when selected in resultTypes. |
includeAuthor | boolean | true | Retain visible author fields and profile/profile-activity rows. When false, author fields and author rows are removed from the emitted shape; no contact enrichment is performed. |
includeAuthorSnapshot | boolean | false | Opt into bounded public profile snapshots for visible author profile links; profile access/status is separate from answer/comment success and never includes contact enrichment. |
includeTopics | boolean | true | Retain explicit topic links and topic relationships where observed. |
includeSpaces | boolean | false | Retain Space rows and feed items where the source route exposes them. Automatically forced on whenever spaceUrls is non-empty. |
includeComments | boolean | false | Enable bounded visible comment output when the per-answer cap is positive. |
includeEngagementBreakdown | boolean | true | Retain explicitly labeled per-comment reaction breakdowns alongside aggregate reaction labels. Raw labels are preserved and typed companions are added only when the localized number is unambiguous. |
includeRelatedContent | boolean | false | Retain explicit links from visibly labeled related/attached regions inside answer cards. |
maxRelatedContent | integer | 20 | Per-answer cap for eligible related-content links when enabled; zero emits no items. |
topCommentsOnly | boolean | false | Retain only comments with an explicit visible Top/Most helpful/Popular comments label. Unknown comments are excluded and counted; votes and position are not used as top evidence. |
commentSortOrder | string | "observed" | Order builder comment rows by observed order or an explicitly typed metric/timestamp. Unknown selected metrics remain in stable fallback order and are counted. Allowed values: observed, upvotes_desc, reactions_desc, replies_desc, newest. |
includeReplyGraph | boolean | false | Enable bounded visible nested comment replies and explicit parent paths; this also enables comment rows. Hidden or unexpanded replies are not inferred. |
includeMedia | boolean | false | Request visible media URL fields where the page exposes them. |
includeFeedText | boolean | false | Request bounded visible text from explicitly linked Space feed cards; page-level text and complete feed history are excluded. |
includeOutboundLinks | boolean | false | Request visible non-Quora links scoped inside question roots, answer/comment nodes, or Topic post cards; nested/page-level links and comment permalinks are excluded. |
includeEntityGraph | boolean | false | Emit explicit relationship rows even when entity_edge is not listed in resultTypes. |
maxCommentsPerAnswer | integer | 0 | Maximum visible comment rows retained per answer when enabled. |
maxCommentDepth | integer | 20 | Maximum visible nested reply depth when includeReplyGraph is enabled; reaching the bound is not complete history. |
maxItems | integer | 0 | Cap for normalized entity/status rows before graph and manifest rows. |
maxGraphEdges | integer | 5000 | Independent cap for deterministic entity_edge rows. |
maxScrollAttempts | integer | 5 | Bounded lazy-render scroll budget applied during detail traversal; scroll attempts and termination state are reported. |
maxLoadMoreAttempts | integer | 0 | Independent exact visible continuation-control budget for direct question, answer, topic, and Space collection pages; reports clicks/state and does not claim complete history. |
includeRawEvidence | boolean | false | Add bounded document hash and optional HTML bytes to each page-level entity row. |
includeHtml | boolean | false | Include bounded HTML bytes when raw evidence is enabled. |
maxHtmlBytes | integer | 100000 | Maximum UTF-8 bytes retained for optional raw HTML evidence. |
requestDelay | number | 1 | Delay between page batches in seconds; this is not an anti-bot bypass guarantee. |
maxConcurrency | integer | 1 | Maximum independent seed pages processed concurrently in deterministic batches. Each seed owns a page; discovered seeds merge in batch order. |
maxNavigationRetries | integer | 2 | Bounded retries after navigation exceptions (e.g. a transient proxy-tunnel handshake failure); exponential backoff is reported in navigation_retry_delays_seconds. |
language | string | "auto" | Browser language hint. Short language codes and supported regional aliases are accepted and normalize to the short code; auto keeps the actor default. This does not translate content or prove access. Allowed values: auto, en, es, fr, de, pt, it, ja, ko, hi, id, nl, pl, tr, vi, zh, ru, en-US, en-GB, de-DE, fr-FR, es-ES, pt-BR, it-IT, ja-JP, ko-KR, hi-IN, tr-TR, zh-CN, zh-TW, ru-RU. |
cookies | array<object> | — | Optional secret cookies for content the caller is authorized to access. |
cookieString | string | — | Optional browser-exported name=value; name2=value2 header. Parsed locally into authorized session cookies; the raw header is never emitted or logged. |
proxyConfiguration | object | {"useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"]} | Apify proxy routing. Always resolves to the RESIDENTIAL tier by default (and as a fallback for a blank or free-tier selection) for reliable access; set an explicit apifyProxyGroups or proxyUrls value to deliberately override it. |
redactSensitiveQueryParams | boolean | true | Redact narrow credential-like query parameter values in emitted URLs, graph endpoints, redirect chains, and raw HTML. The original document hash remains based on the captured document. |
rawEvidenceRetentionDays | integer | 30 | When zero, suppress raw_html snapshots while retaining hash, size, and redaction diagnostics. Positive values permit bounded HTML for the requested run; external Apify dataset/run retention still applies. |
maxRawEvidenceBytes | integer | 100000 | Hard upper byte cap for returned raw HTML after the maxHtmlBytes cap. |
maxItemsPerSource | integer | 0 | Maximum semantic records emitted from each direct URL, query, or typed source. Status and run-manifest rows are preserved. Use 0 for no per-source cap. |