Quora Search Scraper avatar

Quora Search Scraper

Pricing

from $1.00 / 1,000 results

Go to Apify Store
Quora Search Scraper

Quora Search Scraper

Search Quora by keywords and extract questions, answers, author info, upvotes, and engagement metrics. Find relevant Q&A content across any topic.

Pricing

from $1.00 / 1,000 results

Rating

0.0

(0)

Developer

Crawler Bros

Crawler Bros

Maintained by Community

Actor stats

4

Bookmarked

434

Total users

30

Monthly active users

20 days ago

Last modified

Share

Question rows optionally expose bounded question_age_days and question_age_state, derived only from explicit machine-readable creation timestamps; relative or invalid dates remain unavailable.

Post rows likewise expose bounded post_age_days and post_age_state only when an explicit machine-readable post creation timestamp is present; relative or invalid dates remain unavailable.

Post rows also expose post_updated_age_days and post_updated_age_state when explicit post_updated_at is machine-readable. These fields describe last-edit age separately from post creation age and fail closed for relative, invalid, missing, or future timestamps.

Profile-activity rows expose bounded activity_age_days and activity_age_state only from explicit machine-readable activity creation timestamps; the activity stream remains a bounded sample.

Answer rows with an explicit machine-readable answer creation timestamp also carry derived answer_age_days and answer_age_state. Relative labels, invalid dates, and missing dates fail closed; future timestamps are marked and their numeric age is clamped to zero. This is bounded timestamp arithmetic, not an activity or quality score.

Nested comment rows optionally expose positive-only comment_author_is_verified from an exact accessible verification phrase inside the same comment-author scope; parent answer/post/profile markers and neighboring cards are not reused.

Nested comment rows also expose optional comment_author_image_url only from an image descendant inside the matched comment author profile link, and comment_author_credentials only from explicit text inside that bounded comment-author scope; parent answer/post credentials, neighboring avatars, and expertise inference are excluded.

Search-result rows also optionally expose positive-only search_author_is_verified from an exact accessible verification phrase inside the same search-result card; this field is independent from answer, post, and profile verification.

They also preserve explicit search_author_is_anonymous only for an exact Anonymous marker in that same card; missing profile identity is not treated as anonymous.

Fixture-backed examples (synthetic)

The detailed examples below and the shared field reference use synthetic fixture-shaped contracts, not live Quora records. Search cards, totals, and continuation state require scoped evidence; blocked sources produce typed status rows.

{"content_type":"search_result","search_title":"Fixture result","search_rank":1,"search_view_count_text":"900 views","access_state":"visible","source_items_emitted":1,"source_cap_reached":false,"field_sources":{"search_title":"visible_dom_or_metadata","search_rank":"runtime_or_access"}}
{"content_type":"status","access_state":"cloudflare","http_status":403,"source_items_emitted":0,"source_cap_reached":false,"field_sources":{"access_state":"runtime_or_access"}}

Standalone-post top-comment example (synthetic; requires includePostComments: true):

{"content_type":"post","post_id":"fixture-post","post_body":"Visible post body.","top_comments":[{"comment_id":"comment-1","comment_text":"Visible top comment.","comment_collection_label":"Top comments","comment_is_top_labeled":true,"comment_visibility":"visible","comment_depth":0}],"access_state":"public","field_sources":{"top_comments":"visible_dom_or_metadata"}}

Authorized session input accepts either structured cookies or the secret cookieString browser-header form (name=value; name2=value2). Structured cookies win on duplicate names; malformed pairs are ignored, the raw header is never emitted, and string cookies have no invented expiry metadata. This is an access input, not a bypass or private-content guarantee.

For standalone post rows, explicitly labeled visible post-header and Originally Answered evidence is preserved as post_header and originally_answered; both fields are omitted when that distinct source evidence is absent. Standalone post rows also expose conditional post_author_credentials from the bounded owning post-author scope; this is descriptive context only and is never inferred from post prose or a profile slug.

minAnswerUpvotes is a conservative local answer filter. Positive thresholds retain only answer records with explicit parseable upvote evidence; missing or ambiguous metrics are excluded from the filtered answer set and counted in answer_filter_unknown_upvotes, never treated as zero. The parent question reports observed, emitted, and threshold diagnostics.

answerSortOrder supports observed, upvotes_desc, newest, and views_desc for extracted answer rows. Typed metrics/timestamps determine ordering; records missing the selected metric retain stable fallback order and are counted in answer_sort_fallback_count. filterAiAnswers supports include, exclude, and only: only explicit visible AI attribution qualifies, while missing labels remain unknown and are never classified as human. Legacy boolean API values map to include/exclude.

For compatibility with other Quora actors, sortAnswersBy maps relevance, recency, and upvotes to the canonical sort modes, and minUpvotes maps to minAnswerUpvotes. Canonical names take precedence when both are supplied.

Discovery compatibility aliases are also accepted: searchQuery, queries, and searchKeywords → searchQueries; contentTypes and scrapeType → resultTypes; timeFilter → searchTimeFilter; maxItemsPerQuery → maxItemsPerSource; and maxResultsPerQuery → maxResults. searchType accepts marketplace human labels such as All types, Author, and Question, normalizing them to the bounded canonical filter. sortBy maps only supported values such as recent to sortOrder=newest; unsupported modes such as most_upvoted fail validation rather than being silently mapped to a different metric. includeHtmlContent maps only to bounded answer HTML (includeAnswerHtml); it does not claim raw HTML for every entity type. Canonical names take precedence.

topCommentsOnly retains only comments with an explicit scoped Top/Most helpful/Popular label. Unknown comments are excluded and counted in comment_filter_unknown; position, votes, and wording do not establish that a comment is top.

onlyUnanswered is a fail-closed question opportunity filter. When enabled, question and search-card rows are retained only when question_is_unanswered=true from an explicit typed zero answer count; unknown and positive counts are excluded while status rows remain. It is local payload filtering and does not prove complete answer-history coverage.

Retained rows from that mode expose unanswered_filter_enabled=true and unanswered_filter_state (matched_zero_answer or not_applicable) so downstream consumers can audit why the mode was active.

commentSortOrder supports observed, upvotes_desc, reactions_desc, replies_desc, and newest. Sorting uses only typed comment metrics or parseable machine timestamps; missing values remain in stable fallback order and are counted in comment_sort_fallback_count.

Profile rows preserve explicit profile_joined_date labels and bounded visible profile_active_spaces links when present; missing evidence is omitted and no complete membership claim is made.

maxItemsPerSource is a common semantic cap across direct, query, and typed sources. It is applied after extraction but before output projection; status and run-manifest rows remain visible. Capped rows expose source_cap, source_items_observed, source_items_emitted, and source_cap_reached. Every emitted row also echoes the normalized boundary as max_items_per_source for run-level reconciliation. A value of 0 preserves unlimited per-source behavior; it never means complete Quora history.

Nested projection is supported with bounded dotted paths such as answers.answer_text, comments.comment_text, feed.feed_text, or contributors.contributor_name where that collection exists. Child identity/URL fields remain protected, and an empty list keeps the full backward-compatible row.

Exact high-value output names: ../RESEARCH/README_FIELD_REFERENCE.md#quora-scraper-and-search-scraper.

Standalone-post rows expose post_author_is_anonymous only when the owning post scope contains an exact visible Anonymous/Anonymous User/Anonymous Contributor marker; missing identity is never inferred as anonymous.

outputFields is an optional compact-output allowlist. It is applied after extraction, filtering, deduplication, and provenance; protected identity, access/status, source, traversal, and parser fields remain even when omitted from the allowlist. includeFieldEvidence independently controls field_sources and fields_present; output_projection reports full or selected, and the run manifest records the requested projection.

For Space rows, includeSpaceSettings is opt-in and emits space_settings_state, space_allowed_content_types, space_comment_permission, space_submission_policy, space_content_requirements, space_distribution_state, space_contributor_request_state, and bounded space_settings_text only from explicitly labeled rendered settings/content-permission regions. Missing settings remain unknown; role labels, counts, and public access never infer policy.

The run manifest includes bounded input_accounting entries for each processed source URL: final URL, processed/error outcome, emitted record/type/access counts, navigation totals, and duration. This is reconciliation evidence, not platform billing or cost data.

Question and answer rows preserve raw visible engagement labels alongside the existing typed values (view_count_text, answer_count_text, follow_count_text, upvotes_text, comments_count_text, and shares_count_text). This lets consumers retain localized evidence instead of losing it during numeric normalization. Question rows expose question_is_unanswered only from an explicit typed zero answer count; missing answer counts remain unknown. They may additionally expose answer_to_view_ratio and follow_to_answer_ratio only when their same-scope typed operands are explicit and their denominators are non-zero; search-card rows use their corresponding scoped typed metrics when available. Each companion evidence object records the versioned formula/basis and operands. These are workload/opportunity signals, not quality, popularity, relevance, or ranking claims. Answer rows expose answer_is_collapsed only when an answer-scoped expand/collapsed control explicitly reports aria-expanded="false"; truncation, paywall, and absent-control states remain distinct. Unified Post rows expose post_body_completeness and explicit-label-only post_is_restricted; missing or gated bodies remain conservative rather than being treated as complete or private. Unified Post rows also preserve raw author_content_views and typed author_content_views_value only when an explicit author card/region contains a labeled content-views metric. Post view controls and ordinary body text never populate this author-scoped field.

scrapeDepth defaults to detail, preserving current answer/profile-activity behavior when those controls are enabled. overview disables answer, comment, and profile-activity expansion while retaining discovery/page summaries. Every row reports requested_depth, observed_depth, and detail_expansion_state; bounded indicates a configured traversal cap, not complete history, and gated pages report observed_depth: "status".

Native search rows additionally preserve search_filter_evidence (requested filters plus visibly observed labels), explicit card-scoped search_author_name/search_author_url, search_author_id only when matching JSON-LD identifies the same entity and visible profile URL, and bounded search_author_credentials when a profile link and adjacent credential/tagline evidence are rendered, explicit search_view_count_text, search_answer_count_text, and search_follow_count_text labels with typed companions when parseable, search_total_visible/search_total_visible_text only when Quora visibly labels an explicit total, and search_continuation_state (not_observed, available, exhausted, blocked, or cap_reached). When maxLoadMoreAttempts is positive, the actor separately clicks exact visible search continuation controls and records search_load_more_attempts, search_load_more_clicks, and search_load_more_state; this is independent from scrolling and never claims complete history. The number of extracted cards is never promoted to a search-total estimate, and absence of a continuation control is never treated as exhaustion.

Unified question rows also expose conditional question_author_credentials from the bounded owning Question author scope; this is descriptive context only and is never inferred from question prose, answer cards, or a profile slug. When includeMedia is enabled, each retained search card also exposes combined media_urls plus typed image_urls and video_urls collected only from that card. Page-level/sidebar assets and unrelated result cards are excluded; disabled or unobserved media remains empty/omitted according to the existing row projection. Optional dateFrom and dateTo accept ISO dates or datetimes and apply an inclusive local filter to native search cards with machine-readable timestamps. customDateFrom and customDateTo are backward-compatible aliases, with dateFrom/dateTo taking precedence when both are supplied. Returned rows carry search_date_range_state=matched_machine_datetime; cards without machine timestamps are excluded when a bound is set. This does not claim Quora applied a server-side date filter.

Search-result rows with an explicit machine-readable search_published_at also carry derived search_age_days and search_age_state. Age is calculated against the run's scrape_timestamp; relative labels, invalid dates, and missing dates produce machine_datetime_unavailable, while future timestamps produce machine_datetime_future and a clamped age of zero. These fields are bounded timestamp arithmetic, not a trend score or a claim about complete Quora activity.

When a scoped card visibly exposes Promoted, Sponsored, Ad, or Advertisement evidence, the search row preserves is_promoted, is_ad, promotion_label, promotion_source, promotion_url when that marker is inside a link, promotion_state: promoted, and bounded promotion_evidence metadata. Commercial wording, outbound links, or high rank do not create these fields; unknown cards remain unclassified.

includePromoted defaults to true and can be set to false to remove only explicit positive promoted/ad rows. excludePromoted remains supported and takes precedence when both controls are supplied. includePromotionEvidence can omit promotion fields, and includeUncertain marks accessible cards without explicit evidence as promotion_state: unknown; unknown never means organic. This is bounded evidence extraction, not commercial-content classification.

includePostComments is an opt-in standalone-post mode. When enabled, maxCommentsPerPost bounds the visible comment nodes inspected and top_comments is emitted only for comments inside an explicitly labeled Top, Most helpful, Popular, or supported localized equivalent collection. Ordinary comment regions, position, votes, and reactions never establish a top-comment claim; an absent label omits the field.

Nested answer and standalone-post comments preserve parent_comment_id only from an explicit rendered parent marker. parent_comment_url is emitted separately only when that parent exposes an explicit comment permalink; missing graph endpoints remain unknown.

Search visibly accessible Quora pages by keywords or direct URLs and extract evidence-backed questions, answers, author information, labeled engagement, and access diagnostics. Cookies and proxy configuration are optional authorized inputs, not guaranteed access; Cloudflare, login, Quora+, removed, and network outcomes remain typed status rows.

What can this scraper do?

  • Search by keywords — Enter a search term and the scraper preserves visible result cards, source/query provenance, rank evidence, and observed filter state when the configured route is accessible
  • Scrape direct URLs — Provide Quora question, profile, topic, space, or standalone post URLs to scrape specific pages
  • Extract bounded visible answers — Get rendered answer text, author details, labeled upvotes, comments, and engagement metrics under configured caps; hidden or complete history is not claimed
  • Detect AI answers — Automatically identifies Quora AI-generated answers vs. human answers
  • Multiple content types — Questions, answers, user profiles, topics, and spaces all in one scraper
  • Filter and cap output — Select result types, cap total records, and set an independent answer limit per question
  • Access diagnostics — Cloudflare, login, and Quora+ gates are emitted as explicit status records instead of fake content

Input

FieldTypeRequiredDefaultDescription
Search Keywordsstring[]No—Keywords used by the configured search path; result availability, ordering, and relevance are access-dependent.
Direct Quora URLsstring[]No—Direct Quora URLs to scrape (questions, profiles, topics, spaces, posts, or /search?q=...)
Start URLsrequest listNo—API-compatible URL objects; alternative to Direct Quora URLs
Result Typesenum[]Noallquestion, answer, profile, profile_activity, topic, space, post, comment, or search_result
Scrape DepthenumNodetailoverview keeps page/discovery summaries shallow; detail preserves bounded child traversal
Max ResultsintegerNo50Maximum number of results per search query or answers per question (1–5,000)
Native Search TypeenumNoallFor visible /search cards only: all, question, answer, profile, topic, space, or post
Native Search Time FilterenumNoall_timeFor visible /search cards only: all_time, past_hour, past_day, past_week, past_month, or past_year; non-all filters require a visible machine timestamp
Native Search Sort OrderenumNorelevanceVisible-card ordering: relevance (DOM order), newest (machine timestamp), or most_viewed (visible view count); missing evidence keeps stable source order
Deduplication ModeenumNocanonicalcanonical suppresses overlapping entity rows by stable ID/normalized canonical URL; none preserves every observed non-status row. The run manifest reports duplicate_records_suppressed.
Global OrderenumNoobservedobserved, newest, views_desc, answers_desc, or followers_desc; typed-evidence ordering is buffered and unknown values use stable fallback placement.
Max ItemsintegerNo0Global output cap; 0 means unlimited
Max Items Per SourceintegerNo0Per-query/direct-URL semantic-item cap; status/access rows remain visible; 0 means unlimited
Delay Between RequestsnumberNo2Seconds between search/page navigations (0–30). This controls pacing and cost but is not a bypass guarantee.
Max Answers per QuestionintegerNo50Independent answer cap for each question (0–5,000)
Max Scroll AttemptsintegerNo20Bounded browser pagination attempts per question (0–20)
Include AnswersbooleanNotrueDisable answer expansion when only question metadata is needed
Include Media and LinksbooleanNotrueExtract visible image/video URLs and external links from question and answer content
Include Answer HTMLbooleanNotrueInclude rendered answer HTML; disable to reduce dataset size
Include Engagement BreakdownbooleanNofalseOpt in to explicitly labeled answer-scoped upvote/reaction items in nested answers; aggregate metrics remain separate
Include Visible CommentsbooleanNofalseEmit only comments rendered in visible comment-content nodes
Max Comments per AnswerintegerNo50Independent visible-comment cap per answer
Include Post CommentsbooleanNofalseFor standalone posts, emit only explicitly top-labeled visible comments as top_comments
Max Comments per PostintegerNo50Bound visible post-comment nodes inspected when post-comment mode is enabled
Include Profile ActivitybooleanNofalseExtract public activity visibly rendered on profile pages
Profile Activity Typesenum[]Noanswer/question/postActivity categories to retain
Max Profile Activity ItemsintegerNo50Maximum visible activity rows per profile
Only New or Changed ItemsbooleanNofalseSuppress unchanged records using persistent IDs and content hashes
Incremental State Store IDstringOnly with incremental mode—Existing Apify Key-Value Store ID used to persist state across runs
Proxy ConfigurationobjectNoApify AUTOProxy settings. The default uses Apify's free AUTO pool; paid residential access is optional and may be needed if Quora challenges the request.

At least one of Search Keywords, Direct Quora URLs, or Start URLs is required.

Example input

{
"searchQueries": ["python programming", "machine learning"],
"maxResults": 10,
"proxyConfiguration": {
"useApifyProxy": true,
"apifyProxyGroups": ["RESIDENTIAL"]
}
}
{
"directUrls": [
"https://www.quora.com/What-is-Python-used-for",
"https://www.quora.com/topic/Python-programming-language-1",
"https://www.quora.com/profile/Guido-van-Rossum-1"
],
"maxResults": 20
}

Output

Each run produces a dataset with flat rows. Every row includes a content_type field so you can filter by type.

Every populated field also appears in field_sources, using conservative channel classes such as visible_dom_or_metadata, url_or_visible_link, machine_datetime_or_jsonld, scoped_label_or_structured_data, rendered_dom_or_document, runtime_or_access, or input_provenance. This is channel-level provenance, not a claim that a particular internal CSS selector supplied the value.

Question results

FieldTypeExample
content_typestring"question"
titlestring"What is Python primarily used for?"
urlstring"https://www.quora.com/What-is-Python-primarily-used-for"
answer_countinteger100
follow_countinteger42
topicsstring[]["Python programming language", "Software Development"]
source_urlstring"https://www.quora.com/What-is-Python-used-for"
source_querystring"python programming" (empty if from direct URL)
scrape_timestampstring"2026-03-08T18:28:03.140078+00:00"

Answer results

FieldTypeExample
content_typestring"answer"
titlestring"What is Python primarily used for?"
urlstring"https://www.quora.com/What-is-Python-primarily-used-for/answer/John-Smith"
answer_textstringBounded visibly rendered answer text when available
answer_urlstringDirect link to the answer
author_namestring"John Smith"
author_urlstring"https://www.quora.com/profile/John-Smith"
author_credentialsstring"Software Engineer at Google"
upvotesinteger89
comments_countinteger4
shares_countinteger2
answer_timestampstring"2y"
is_ai_answerbooleanfalse
question_titlestring"What is Python primarily used for?"
question_urlstring"https://www.quora.com/What-is-Python-primarily-used-for"
source_urlstringOriginal input URL
source_querystringSearch keyword (empty if from direct URL)
scrape_timestampstringISO 8601 timestamp

Profile results

FieldTypeExample
content_typestring"profile"
titlestring"Guido van Rossum"
namestring"Guido van Rossum"
urlstring"https://www.quora.com/profile/Guido-van-Rossum-1"
biostringUser biography
credentialsstringProfessional credentials
profile_image_urlstringProfile picture URL
follower_countinteger3100
following_countinteger15
answer_countinteger42
question_countinteger5
total_viewsinteger1200000
content_views_this_monthinteger4200 only when an exact monthly window label is visible
content_views_this_month_textstringRaw monthly content-view label/value
source_urlstringOriginal input URL
scrape_timestampstringISO 8601 timestamp

Topic results

FieldTypeExample
content_typestring"topic"
titlestring"Python (programming language)"
namestring"Python (programming language)"
urlstring"https://www.quora.com/topic/Python-programming-language-1"
descriptionstringTopic description
follower_countinteger1600000
question_countinteger5000
source_urlstringOriginal input URL
scrape_timestampstringISO 8601 timestamp

Space results

FieldTypeExample
content_typestring"space"
titlestring"Data Science"
namestring"Data Science"
urlstring"https://www.quora.com/q/data-science"
descriptionstringSpace description
follower_countinteger104000
follower_count_textstringRaw visible follower label
member_countinteger900 only when explicitly labeled as members
member_count_textstringRaw visible member label
space_visibilitystring"public" only when explicitly labeled
post_countinteger500
contributor_countinteger25
source_urlstringOriginal input URL
scrape_timestampstringISO 8601 timestamp

Sample output

{
"content_type": "answer",
"title": "What is Python primarily used for?",
"url": "https://www.quora.com/What-is-Python-primarily-used-for/answer/Pratima-Yadav-117",
"answer_text": "Python is used for various purposes due to its versatile nature...",
"answer_url": "https://www.quora.com/What-is-Python-primarily-used-for/answer/Pratima-Yadav-117",
"author_name": "Pratima Yadav",
"author_url": "https://www.quora.com/profile/Pratima-Yadav-117",
"author_credentials": "",
"upvotes": 4,
"comments_count": 0,
"shares_count": 2,
"answer_timestamp": "4y",
"is_ai_answer": false,
"question_title": "What is Python primarily used for?",
"question_url": "https://www.quora.com/What-is-Python-primarily-used-for",
"source_url": "https://www.quora.com/What-is-Python-used-for",
"scrape_timestamp": "2026-03-08T18:28:03.140078+00:00"
}

Additional normalized fields

When answers are enabled and machine-readable timestamps are present, last_answered_at is the latest timestamp among answer rows extracted in this run. It is omitted when only relative labels are visible and does not claim hidden/deleted answer history.

Question records may include canonical_url, question_id, question_body, description, view_count, created_at, topic_urls, combined media_urls, typed image_urls/video_urls, and outbound_links when Quora exposes them. The typed channels are scoped to the question root after nested answer regions are removed. The unified output also adds question_view_count_value, question_answer_count_value, and question_follow_count_value only when the corresponding raw label is conservatively parseable. When answers are enabled, answer_count_scraped records the number of answer rows actually extracted under the configured cap; answer_limit_reached reports that the configured cap was reached, and neither field proves that hidden answers exist. This is separate from Quora’s visible answer_count. Answer records may include answer_id, answer_html, answer_created_at, answer_updated_at, answer_upvote_count_value, answer_comment_count_value, answer_share_count_value, media_urls, and outbound_links. Set includeMedia to false to omit media, includeOutboundLinks to false to omit external links, and includeAnswerHtml to false to omit rendered answer HTML. Fields are omitted when unavailable; no placeholder values are emitted. Question rows may additionally include related_questions only when visibly linked cards occur inside an explicitly labeled Related questions region; unlabelled links and utility/answer routes are not treated as related questions.

Answer rows additionally expose answer_translation_state and answer_translation_label only for an explicit translation/original-answer marker scoped to that answer card, plus answer_original_url only for an explicit original-answer link. Browser locale, language mismatch, and translated-looking prose never synthesize these fields. Answer rows additionally expose answer_is_anonymous only for an exact visible Anonymous marker in the scoped author region; missing authors and inaccessible profiles remain unknown.

Every normalized row also includes provenance fields: parser_strategy, parser_version, record_id, requested_result_types, fields_present, scroll_attempts, and scroll_state. The actor additionally reports bounded operational evidence: scrape_duration_ms, navigation_attempts, navigation_retries, ordered navigation_retry_delays_seconds, final_url, http_status when available, proxy_used, proxy_resolution_state (resolved or unavailable — a RESIDENTIAL Apify Proxy endpoint is forced by default), request_delay_seconds, and browser_locale; these values describe this run only. Endpoint resolution is not proof of successful access, and proxy credentials are never exposed. parser_version is currently quora-suite-v1 and identifies the deployed parser contract. record_id prefers a visible entity ID and falls back to canonical URL, providing a stable join key across overlapping seeds. fields_present is calculated from the actual non-empty row before provenance is added, so consumers can distinguish a field that was not rendered from a parser or access failure. For search/question/answer rows, scroll_state is not_requested, max_items_reached, exhausted, or cap_reached; it records whether visible pagination ended naturally or hit the configured budget. Each run also emits one content_type: "run_manifest" row with aggregate emitted-record/type/access counts, attempted source pages, total navigation attempts/retries, run timestamps/duration, requested modes, pacing, locale, and proxy state. Filter this row out when only entity records are desired. Status rows use parser_strategy: "access_gate_or_status"; entity rows identify the visible-DOM/JSON-LD strategy.

When onlyNewItems is enabled, provide a persistent Apify Key-Value Store ID. Emitted entity rows include change_type (new or changed) and a deterministic content_hash; unchanged rows are omitted. Status rows are retained so blocked runs remain visible.

redirect_chain is the ordered URL sequence observed from Playwright request ancestry, bounded to 16 entries and including the requested seed when available. It is diagnostic evidence, not a reconstruction of redirects hidden from the browser.

When includeComments is enabled, visible comments are emitted as content_type: "comment" rows with comment text, author, timestamp, depth, and parent answer/question references. Quora may lazy-load or hide comments; the actor does not infer hidden comments from comments_count.

When includeProfileActivity is enabled, visibly rendered public profile links are emitted as content_type: "profile_activity" rows for selected answer, question, or post categories. This is not a claim of complete historical profile activity; pagination, login-only tabs, and hidden items are not inferred. Each row's activity_text is populated only when includeActivityText is true (default), and a bounded, redacted activity_html (up to 30000 characters) is populated only when includeActivityHtml is opted in. maxItemsPerTab independently caps how many rows are emitted per tab (answers/questions/posts) in addition to the overall maxProfileActivityItems cap. Profile rows may include profile_topics as explicit visible topic URL/name pairs from the profile metadata scope; activity-card topics and text-based topic inference are excluded. When Quora renders a per-topic answer count inside an explicit "Knows About"/"Topic Credentials" card, the matching profile_topics[] entry also carries answer_count_text (raw label) and answer_count_value (typed integer); topics without such a card never get a fabricated count. They may also include employment and education only from explicit public labels or Person JSON-LD (jobTitle, worksFor, alumniOf); no identity or activity inference is performed.

Set includeEntityGraph to true to additionally emit normalized content_type: "entity_edge" rows. Edges have deterministic edge_id, typed from_record_id/to_record_id, URLs, a relationship name, and edge_evidence_field. Supported relationships include question-to-topic/answer, answer-to-author, comment-to-parent/author, profile-to-activity, and space/post-to-feed or author. The actor emits an edge only when the scraped row contains an explicit endpoint URL or ID; it never joins by a display name or guesses a hidden relationship. maxGraphEdges is an independent cap and graph rows can be emitted in addition to maxItems. Graph rows use parser_strategy: "derived_relationships" and retain source/access provenance.

Answer rows also include transparent derived signals: answer_evidence_score / answer_quality_evidence measure observable evidence completeness using quora-answer-evidence-v1, while answer_quality_proxy_score / answer_quality_proxy measure observable structure and evidence signals using quora-answer-quality-proxy-v1. The latter considers only text substance, sentence/formatting signals, visible links/media, identity/date evidence, and labeled metrics. Neither signal is Quora's correctness, authority, relevance, popularity, or top-answer ranking.

Direct /post/... URLs are emitted as content_type: "post" records with visible body, author, timestamp, media, and external-link fields. The actor does not treat an unclassified URL as a post.

Direct /search?q=... URLs use the native Quora search page and emit content_type: "search_result" rows for visibly rendered cards, including rank, title, snippet, and result URL. These are discovery records; the actor does not claim that a search card contains the full question page.

How much does it cost?

The Quora Scraper uses pay-per-event pricing at $5 per 1,000 results. Each question, answer, profile, topic, or space counts as one result.

A typical run scraping 1 search query with 10 answers costs approximately $0.05–0.10 in platform credits (including compute and proxy).

FAQs

Do I need a Quora account or cookies?

No for the default public mode. An optional cookies input is available when an authorized session is needed to pass an access gate; cookies are secret input and may expire. session_cookie_expiry_state reports only expiry metadata present at input time and never proves that Quora accepted the session. The actor never accesses private content by design.

When includeComments is enabled, nested answer comment rows expose comment_limit_reached when the positive maxCommentsPerAnswer cap is reached and comment_depth_limit_reached when visible comments deeper than maxCommentDepth are omitted. They also expose comment_items_observed_count and comment_items_emitted_count, separating distinct eligible comments observed before the extraction cap from retained nested rows. Parent answer rows expose answer_comments_observed_count and answer_comments_emitted_count; the latter is reconciled after collection filters and local caps. maxDepth is accepted as a compatibility alias. They also expose separate reaction/upvote labels, conservative reply status/path fields, and bounded reply-expansion diagnostics including reply_expansion_source. maxReplyExpandAttempts controls visible reply/comment expansion per question page. This is a bounded traversal diagnostic, not a complete comment-history claim. When an explicitly comment-scoped control exposes a typed reaction label, the same rows also preserve comment_reaction_breakdown items with type, raw label, and raw value, plus optional value_numeric for unambiguous localized counts; upvotes, icons, and aggregate-only values are excluded. When Quora exposes a comment-author identifier in an explicit data/structured attribute, nested comment rows preserve comment_author_id; profile slugs are never treated as IDs. Nested comments also expose positive-only comment_is_anonymous evidence when an exact anonymous marker is visible in the scoped author region. Nested rows preserve comment_collection_label and comment_is_top_labeled only for an explicit Top/Most helpful/Popular comments region; ordering and votes are not used.

Post-capable results may additionally expose top_comments, containing only explicitly labeled top/helpful/popular comment objects; it is omitted without label evidence and is not a complete-history claim. For standalone posts, maxCommentDepth applies to these visible nested comments too, and comment_depth_limit_reached is emitted only when a deeper rendered comment was observed and omitted. When enabled, comment-content-scoped media/external links and explicitly labeled comment upvote/reaction totals are retained under the same includeMedia/includeOutboundLinks controls. Nested rows also expose comment_visibility only when explicit metadata or an exact placeholder label identifies visible, deleted, hidden, or restricted state. Breakdown items additionally expose value_numeric when the raw value parses unambiguously; raw labels remain authoritative.

Nested answer rows preserve observed order versus explicit top/best labels with answer_rank_source, and expose conservative answer_text_completeness and explicit-label-only Quora+ state. With includeEngagementBreakdown enabled, nested answer rows also expose reaction_breakdown only for explicitly labeled answer-scoped upvote/reaction controls; items may include value_numeric after unambiguous parsing, and comment-scoped labels and icons are excluded.

How does keyword search work?

The scraper first uses external indexed discovery for keyword inputs, then falls back to native Quora search pages when no indexed URLs are available. Direct /search?q=... inputs always use the native page. Native search cards are represented separately as search_result records so discovery metadata is not confused with scraped question content. Overlapping seeds are deduplicated by typed entity identity, while source_queries retains every query that found the entity. Each canonical row also exposes matched_input, matched_input_kind, bounded source_observations, and source_observation_count for exact mixed URL/query overlap evidence. Set globalOrder to buffer and deterministically order comparable parent rows; typed metrics or machine timestamps are required, unknown parents use stable fallback placement, and children remain adjacent to their parent. maxConcurrency (1–3) applies only to independent indexed discovery queries; page extraction remains sequential so cross-seed deduplication and incremental state updates stay deterministic.

What proxy should I use?

The default is Apify's free AUTO pool. Quora may challenge datacenter traffic; residential access is an optional paid configuration, not a guaranteed bypass. If a challenge remains, the actor emits an access_state status record.

How many results can I get per run?

You can extract up to 5,000 results per search query or direct URL. For questions with many answers, the scraper automatically scrolls to load more content.

Can I scrape private or restricted profiles?

No. The scraper only extracts publicly visible content. Private profiles, restricted answers, and Quora+ paywalled content are not accessible.

What does is_ai_answer mean?

Quora generates AI-powered answers for some questions. The scraper automatically detects these and marks them with is_ai_answer: true and author_name: "Quora AI" so you can filter them.

What memory setting should I use?

Use 1024 MB or higher. Quora's pages are resource-intensive and require more memory than typical websites.

Why are some fields empty?

Some fields like follow_count, author_credentials, or description may be empty when Quora doesn't display that information on the page. This is normal and depends on the specific content being scraped.

What types of Quora URLs are supported?

  • Questions: https://www.quora.com/What-is-Python-used-for
  • Profiles: https://www.quora.com/profile/Username
  • Topics: https://www.quora.com/topic/Topic-Name
  • Spaces: https://www.quora.com/q/space-name or https://spacename.quora.com/

Limitations

  • Requires 1024 MB memory minimum (Quora's pages are JavaScript-heavy)
  • Keyword search returns 1–5 relevant question URLs per query
  • Only publicly visible content can be scraped — no login-restricted or Quora+ content
  • If Cloudflare or a login wall is encountered, the actor emits an access_state status record so an empty result is not mistaken for successful extraction
  • Answer timestamps are relative (e.g., "2y", "6mo") rather than exact dates
  • Quora may occasionally restrict access from data center IPs — residential proxy recommended Relative date labels remain raw; when a scrape reference is available, date_interpretations exposes conservative bounded intervals for supported prefix, suffix, reversed-order, and non-Latin forms (for example vor 2 Tagen, 3 giorni fa, 3日前, and 3天前) and never claims an invented exact timestamp. Unsupported labels remain raw without a guessed date.

includeOutboundLinks independently controls visible external links for question, answer, and post rows; when omitted it follows includeMedia for backward compatibility, enabling links-only and media-only payloads.

maxAnswerHtmlBytes optionally bounds rendered answer HTML by UTF-8 bytes. Positive values emit answer_html_bytes and answer_html_truncated; zero preserves the uncapped legacy behavior, and the cap does not claim complete answer HTML.

includeRelatedQuestions controls explicitly labeled related-question links on question pages and defaults to true; disabling it only removes that context collection and does not change answer or post traversal.

includeQuestionTopics independently controls explicit topic names/URLs on question rows and defaults to true; disabling it only removes topic context and does not change answer or post traversal.

includeQuestionDetails independently controls rendered question description/body text and defaults to true; disabling it removes only those text fields while preserving identity, metrics, answer/post traversal, and status envelopes.

includeQuestionHtml and maxQuestionHtmlBytes independently control optional UTF-8-bounded rendered HTML from the answer-isolated question scope; this is distinct from the legacy answer-HTML alias and from full-document Raw Page evidence.

maxQuestionTopics bounds explicit topic names/URLs on question rows (0 means uncapped). Positive values emit question_topics_scraped and question_topics_limit_reached; these are local observation/cap diagnostics, not complete topic-membership claims.

maxRelatedQuestions bounds the emitted related-question collection on question rows (0 means uncapped). Positive values emit related_questions_scraped and related_questions_limit_reached; these are local observation/cap diagnostics, not completeness claims.

Complete schema input inventory

This generated inventory mirrors every property in .actor/input_schema.json; enum values are retained in the description below.

Input fieldTypeDefaultDescription
scrapeDepthstring"detail"Overview keeps discovery/question summaries shallow; detail preserves answer/profile-activity traversal where enabled and reports its observed expansion state. Allowed values: overview, detail.
searchQueriesarray<string>—Keywords to search for on Quora. Attempts bounded discovery and preserves visible source/query evidence; availability, ordering, and relevance depend on the accessible result page and authorized access configuration.
queriesarray<string>—Compatibility alias for searchQueries. Canonical searchQueries takes precedence when both are supplied.
searchKeywordsarray<string>—Compatibility alias for searchQueries. Canonical searchQueries takes precedence, then queries, when multiple names are supplied.
searchQuerystring—Compatibility alias for searchQueries for a single query. Canonical searchQueries, queries, and searchKeywords take precedence when supplied.
directUrlsarray<string>—Direct Quora URLs to scrape (questions, profiles, topics, spaces). Example: https://www.quora.com/What-is-Python
startUrlsarray<any>[]Alternative to directUrls. Accepts Quora URL objects.
targetsarray<string>[]Mixed-target compatibility alias: URL-shaped values join directUrls and other values join searchQueries; canonical fields and targets are merged, then deduped according to dedupeMode downstream.
resultTypesarray<string>["question","answer","profile","profile_activity","topic","space","post","search_result"]Restrict output to selected entity types. search_result rows are visible Quora discovery cards, not fully scraped question entities.
contentTypesarray<string>—Compatibility alias for resultTypes. Plural values such as questions, answers, profiles, topics, spaces, and posts are normalized to canonical entity types.
scrapeTypestring—Compatibility alias for resultTypes. It is used only when resultTypes and contentTypes are absent; search means the default unified entity set. Allowed values: search, question, answer, profile, author, topic, space, post.
onlyUnansweredbooleanfalseRetain question/search rows only when an explicit typed answer count is exactly zero. Unknown or ambiguous counts are excluded; status and run-manifest rows remain. This does not prove complete answer-history coverage.
excludePromotedbooleanfalseExclude only cards with explicit visible or structured sponsored/promoted/ad evidence; unknown cards remain.
includePromotedbooleantrueInclude cards with explicit promotion evidence. Set false to exclude them; explicit excludePromoted takes precedence when both are supplied.
includePromotionEvidencebooleantrueRetain promotion state, bounded promotion_evidence, and explicit label/source fields on search-result rows.
includeUncertainbooleantrueMark accessible cards without explicit promotion evidence as unknown; this is not a negative classification.
maxResultsinteger50Maximum number of results per search query or answers per question URL.
maxResultsPerQueryinteger50Compatibility alias for maxResults. The canonical maxResults takes precedence.
searchTypestring"all"Fail-closed filter for visible native Quora search cards: all, question, answer, profile, topic, space, or post. Allowed values: all, all types, question, Question, answer, Answer, profile, Profile, Author, topic, Topic, space, Space, post, Post.
searchTimeFilterstring"all_time"Visible-timestamp filter. Non-all filters exclude cards without a machine-readable timestamp; no date is inferred. Allowed values: all_time, past_hour, past_day, past_week, past_month, past_year.
timeFilterstring—Compatibility alias for searchTimeFilter. Values such as 'Past Week' normalize to past_week; the canonical searchTimeFilter takes precedence.
dateFromstring—Optional inclusive ISO date/datetime lower bound applied locally to native search cards with machine-readable timestamps.
dateTostring—Optional inclusive ISO date/datetime upper bound applied locally to native search cards with machine-readable timestamps.
customDateFromstring—Backward-compatible alias for dateFrom. If both are supplied, dateFrom takes precedence.
customDateTostring—Backward-compatible alias for dateTo. If both are supplied, dateTo takes precedence.
sortOrderstring"relevance"Stable DOM relevance order, newest by visible machine timestamp, or most viewed by visible view count. Missing evidence retains original order. Allowed values: relevance, newest, most_viewed.
sortBystring—Compatibility alias for sortOrder. Supported compatibility values include relevance, recent, most_recent, and most_viewed; unsupported values fail validation rather than being guessed.
dedupeModestring"canonical"canonical suppresses overlapping entity rows by stable ID or normalized canonical URL (default). none preserves every observed non-status row; duplicate_records_suppressed in the manifest remains 0 because no suppression is performed. Allowed values: canonical, none.
outputFieldsarray<string>[]Optional allowlist (up to 3,000 unique names) for compact rows, including bounded dotted paths through nested collections (for example answers.answer_text, comments.comment_text, feed.feed_text, or contributors.contributor_name where applicable). Child identity/URL fields and the actor identity, access/status, source, traversal, graph, filter, and parser-envelope fields remain protected.
includeFieldEvidencebooleantrueRetain field_sources and fields_present after projection; disabling removes only these evidence maps, not semantic or protected access fields.
languagestring"auto"Browser language hint. Short language codes and supported regional aliases are accepted and normalize to the short code; auto keeps the actor default. This does not translate content or prove access. Allowed values: auto, en, es, fr, de, pt, it, ja, ko, hi, id, nl, pl, tr, vi, zh, en-US, en-GB, de-DE, fr-FR, es-ES, pt-BR, it-IT, ja-JP, ko-KR, hi-IN, tr-TR, zh-CN, zh-TW.
maxItemsinteger0Maximum total records across all inputs. Use 0 for no total cap.
maxItemsPerSourceinteger0Maximum semantic records emitted from each individual query or direct URL. Status/access rows are preserved; 0 means unlimited.
maxItemsPerQueryinteger0Compatibility alias for maxItemsPerSource. The canonical maxItemsPerSource takes precedence.
requestDelaynumber2Seconds to wait between search/page navigations. Set to 0 for no added delay; this controls pacing but cannot guarantee bypassing Quora access controls.
maxNavigationRetriesinteger2Maximum bounded retries after transient navigation failures (0–3). Access gates, 404/410 responses, and challenge pages are not retried.
maxConcurrencyinteger1Maximum independent indexed discovery queries processed concurrently in deterministic batches. Page extraction remains sequential for shared deduplication and incremental state correctness.
maxAnswersPerQuestioninteger50Independent answer limit for each question URL.
minAnswerUpvotesinteger0When positive, retain only answer records with explicit machine-readable upvote evidence at or above this threshold. Missing or ambiguous upvotes are never treated as zero and are reported in filter diagnostics.
minUpvotesinteger0Compatibility alias for minAnswerUpvotes. The canonical minAnswerUpvotes value takes precedence when both are supplied.
answerSortOrderstring"observed"Order retained answer records by rendered order or explicit typed upvotes, machine timestamps, or view counts. Rows missing the selected metric retain stable fallback order and are counted. Allowed values: observed, upvotes_desc, newest, views_desc.
sortAnswersBystring"relevance"Compatibility alias mapped to answerSortOrder: relevance=observed, recency=newest, upvotes=upvotes_desc. The canonical answerSortOrder takes precedence when both are supplied. Allowed values: relevance, recency, upvotes.
filterAiAnswersstring"include"Include all answers, exclude only explicitly AI-labeled answers, or retain only explicitly AI-labeled answers. Unknown/missing AI labels are never classified and are counted in filter diagnostics. Legacy boolean true/false values are accepted by the API normalizer as exclude/include. Allowed values: include, exclude, only.
topCommentsOnlybooleanfalseRetain only comments with an explicit visible Top/Most helpful/Popular comments label. Unknown comments are excluded and counted; votes and position are not used as top evidence.
commentSortOrderstring"observed"Order comments by observed order or an explicitly typed metric/timestamp. Unknown selected metrics remain in stable fallback order and are counted. Allowed values: observed, upvotes_desc, reactions_desc, replies_desc, newest.
maxScrollAttemptsinteger20Bound the number of browser pagination/scroll attempts per question page. This is capped at 20 to control cost and runaway pages.
includeAnswersbooleantrueExpand and emit answer records for question pages.
includeMediabooleantrueExtract image/video URLs from visible question, answer, and post content. External links are controlled independently by includeOutboundLinks.
includeOutboundLinksbooleantrueIndependently retain visible external anchors from question, answer, and post content; when omitted, follows includeMedia for backward compatibility.
includeAnswerHtmlbooleantrueInclude the rendered HTML of answer content. Disable to reduce dataset size.
includeHtmlContentbooleantrueCompatibility alias for includeAnswerHtml. It controls bounded answer HTML only; canonical includeAnswerHtml takes precedence.
includeRelatedQuestionsbooleantrueRetain explicitly labeled related-question links from question pages. Disable to reduce question-context payload; this does not affect answer extraction.
maxRelatedQuestionsinteger0Bound the emitted related-question collection on question rows. Zero preserves all observed related links; positive values emit related_questions_scraped and related_questions_limit_reached diagnostics.
includeQuestionTopicsbooleantrueRetain explicit topic links associated with question rows. Disable to reduce question-context payload; this does not affect answer or post traversal.
maxQuestionTopicsinteger0Bound retained explicit topic names/URLs on question rows. Zero preserves all observed topic links; positive values emit question_topics_scraped and question_topics_limit_reached diagnostics.
topicFilterarray<string>[]For question and search-result rows, retain only rows with an exact explicitly observed topic label or topic URL matching one of these terms. Unknown topic evidence is excluded and counted; this is local payload filtering, not a server-side Quora filter.
includeSpaceSettingsbooleanfalseOpt in to explicitly labeled Space settings/content-permission evidence for Space rows; missing settings remain unknown and admin-only queues are not scraped.
includeQuestionDetailsbooleantrueRetain rendered question description/body text on question rows. Disable to reduce payload while preserving question identity, metrics, answer/post traversal, and status envelopes.
includeQuestionHtmlbooleanfalseRetain bounded rendered HTML from the question scope after nested answer regions are removed.
maxQuestionHtmlBytesinteger0Optional UTF-8 byte ceiling for question HTML; positive values emit question_html_bytes and question_html_truncated.
maxAnswerHtmlBytesinteger0Optional UTF-8 byte ceiling for rendered answer HTML. Zero preserves uncapped backward-compatible HTML; positive values emit answer_html_bytes and answer_html_truncated.
maxPostHtmlBytesinteger0Optional UTF-8 byte ceiling for rendered standalone Post HTML in unified Post rows. Zero disables HTML retention; positive values emit post_html_bytes and post_html_truncated.
includeCommentsbooleanfalseExtract only comments rendered in visible answer comment-content nodes; hidden comments are not inferred.
includePostCommentsbooleanfalseFor standalone post URLs, retain only comments inside an explicitly labeled Top/Most helpful/Popular collection as top_comments. Ordinary comments, votes, and position never qualify.
maxCommentsPerPostinteger50Bound the visible post-comment nodes inspected when includePostComments is enabled; this is not a complete comment-history guarantee.
includeEngagementBreakdownbooleanfalseOpt in to explicitly labeled answer-scoped upvote/reaction items in nested answer rows. Comment controls and icons are never decomposed.
includePostAuthorContextbooleanfalseOpt in to explicitly labeled About/Bio, Education, and Joined/Member since values inside the owning standalone-post author scope; no profile navigation or inference is performed.
maxPostAuthorContextItemsinteger10Bounds active Spaces and labeled public external links retained from the author scope; observed/emitted/cap diagnostics distinguish truncation.
maxPostAuthorContextItemsinteger10Maximum explicitly labeled active Space links retained per standalone-post author; observed, emitted, and cap diagnostics are preserved.
maxCommentsPerAnswerinteger50Maximum visible comments extracted from each answer when comments are enabled.
maxReplyExpandAttemptsinteger0Bound clicks on clearly labeled visible reply/comment expansion controls per question page.
maxCommentDepthinteger20Maximum explicit comment nesting depth retained from answer pages and standalone-post top-comment collections. Root comments are depth 0; deeper comments are omitted and marked with comment_depth_limit_reached. The maxDepth alias is accepted for compatibility.
includeProfileActivitybooleanfalseExtract public activity links visibly rendered on profile pages.
includeEntityGraphbooleanfalseEmit deterministic entity_edge relationship rows when explicit endpoint URLs or IDs are present. This is opt-in and does not invent relationships.
maxGraphEdgesinteger5000Independent cap for entity_edge rows. Graph rows may be added in addition to maxItems.
profileActivityTypesarray<string>["answer","question","post"]Activity categories to retain: answer, question, or post.
maxProfileActivityItemsinteger50Maximum visibly rendered public activity records per profile.
onlyNewItemsbooleanfalseCompare stable IDs and content hashes against a persistent state store and suppress unchanged rows.
stateStoreIdstring—Existing Apify Key-Value Store ID used for onlyNewItems state. Required when onlyNewItems is enabled.
cookiesarray<object>—Optional authenticated browser cookies. Use only for content you are authorized to view; cookies are stored as secret input and are never logged.
cookieStringstring—Optional browser-exported name=value; name2=value2 header. Parsed locally into authorized session cookies; the raw header is never emitted or logged.
proxyConfigurationobject{"useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"]}Proxy settings. RESIDENTIAL Apify Proxy is used by default and is forced whenever no explicit group is chosen, because Quora's Cloudflare/anti-bot layer blocks the plain datacenter (AUTO) pool on most page types. Provide an explicit apifyProxyGroups or a custom proxyUrl to override.
usernamesarray<string>[]Quora profile usernames/slugs (with or without a leading @ or trailing slash). Normalized to https://www.quora.com/profile/{slug} and merged into the direct-URL queue; also seeds author-discovery and post-discovery when relevant.
includeTopicQuestionsbooleantrueEmit visible topic question links as topic_question rows.
maxTopicQuestionsinteger50Bound topic_question rows emitted per topic page. Zero means unlimited.
includeTopicPostsbooleanfalseEmit explicitly linked standalone Posts from a topic page as topic_post rows.
maxTopicPostsinteger50Bound topic_post rows emitted per topic page. Zero means unlimited.
includeRelatedTopicsbooleantrueEmit visibly linked related topics as topic_related rows.
maxRelatedTopicsinteger50Bound topic_related rows emitted per topic page. Zero means unlimited.
includeAnswerPreviewbooleanfalseOpt in to bounded, already-rendered answer-preview evidence on question/search_result cards, without navigating to the full answer page.
maxAnswerPreviewsinteger3Bound answer_preview_* entries retained per question/search_result card.
minAnswerPreviewUpvotesinteger0Minimum typed upvote threshold an answer preview must meet to be retained. Fails closed: previews with unknown upvotes are excluded once this is positive.
includePostHtmlbooleanfalseLegacy alias: true sets a 100000-byte maxPostHtmlBytes ceiling unless maxPostHtmlBytes is explicitly set.
maxPostsPerSourceinteger0Bound /post/... links discovered (then fully scraped as post rows when resultTypes includes post) from one profile/Space/search source page. Distinct from maxItemsPerSource, which caps already-scraped rows rather than discovery candidates.
includeReactionBreakdownbooleanfalseLegacy alias for includeEngagementBreakdown. The canonical includeEngagementBreakdown takes precedence when both are supplied.
includeRelatedContentbooleanfalseOpt in to an explicit related/attached content region scoped to one answer or post card.
maxRelatedContentinteger20Bound related_content entries retained per answer/post.
includeProfileSensitiveFieldsbooleantrueWhen false, strip location, employment, education, and profile_joined_date from profile/author rows.
includeSensitiveProfileFieldsbooleantrueDeprecated alias for includeProfileSensitiveFields. The canonical field wins on conflict; a conflict is logged to input_normalization_warnings.
includeSocialLinksbooleantrueWhen false, omit social_links and website_url from profile/author rows.
activityTabstring"all"Click the given profile-activity tab before extracting profile_activity rows. Allowed values: all, answers, questions, posts.
includeActivityTextbooleantrueInclude bounded visible activity_text on profile_activity rows.
includeActivityHtmlbooleanfalseInclude bounded, redacted rendered activity_html (up to 30000 bytes) on profile_activity rows.
maxItemsPerTabinteger0Independent cap on profile_activity rows per tab (answers/questions/posts), in addition to maxProfileActivityItems. Zero means unlimited.
maxDetailItemsinteger0When scrapeDepth=detail, bounded number of profile_activity items opened individually for detail_text/detail_title extraction. Zero disables detail expansion.
maxAuthorsinteger1000Maximum unique discovered author rows to emit, independent of maxItems. Requires resultTypes to include author.
maxSourcePagesinteger100Maximum author-discovery source pages (search/question/answer/topic/Space/post/profile) whose already-scraped rows are scanned for author evidence.
maxEvidencePerAuthorinteger10Maximum evidence_urls/evidence_question_urls/evidence_answer_urls/evidence_answer_previews entries retained per discovered author.
includeProfileSnapshotbooleanfalseWhen true and resultTypes includes author, open each discovered author's profile and merge a full profile snapshot onto the author row.
includeGraphEdgesbooleanfalseCompatibility alias for includeEntityGraph. The canonical includeEntityGraph takes precedence when both are supplied.
includeUnchangedbooleantrueWhen onlyNewItems is set, also retain change_type=unchanged rows instead of dropping them.

Unified Search question rows now preserve question_created_at alongside question_updated_at from the scoped Question extractor. The created value is question-scoped visible/structured evidence; both fields are omitted when absent and never use scrape time or unrelated page dates.