Quora Scraper
Pricing
from $3.00 / 1,000 results
Quora Scraper
Scrape Quora questions, answers, user profiles, topics, and spaces. Search by keywords or scrape direct URLs. Extracts full answer text, author info, upvotes, and engagement metrics.
Pricing
from $3.00 / 1,000 results
Rating
0.0
(0)
Developer
Crawler Bros
Maintained by CommunityActor stats
5
Bookmarked
395
Total users
19
Monthly active users
21 days ago
Last modified
Categories
Share
Question rows optionally expose bounded question_age_days and question_age_state, derived only from explicit machine-readable creation timestamps; this is not a complete-history or popularity signal.
Post rows likewise expose bounded post_age_days and post_age_state only when an explicit machine-readable post creation timestamp is present; this is not a complete-activity or popularity claim.
Post rows also expose post_updated_age_days and post_updated_age_state when explicit post_updated_at is machine-readable. These fields describe last-edit age separately from post creation age and fail closed for relative, invalid, missing, or future timestamps.
Profile-activity rows expose bounded activity_age_days and activity_age_state only from explicit machine-readable activity creation timestamps; the activity stream remains a bounded sample, not complete history.
Answer rows with an explicit machine-readable answer creation timestamp also carry derived answer_age_days and answer_age_state. Relative labels, invalid dates, and missing dates fail closed; future timestamps are marked and their numeric age is clamped to zero. This is bounded timestamp arithmetic, not an activity or quality score.
Nested comment rows optionally expose positive-only comment_author_is_verified from an exact accessible verification phrase inside the same comment-author scope; parent answer/post/profile markers and neighboring cards are not reused.
Nested comment rows also expose optional comment_author_image_url only from an image descendant inside the matched comment author profile link, and comment_author_credentials only from explicit text inside that bounded comment-author scope; parent answer/post credentials, neighboring avatars, and expertise inference are excluded.
Search-result rows also optionally expose positive-only search_author_is_verified from an exact accessible verification phrase inside the same search-result card; this field is independent from answer, post, and profile verification.
They also preserve explicit search_author_is_anonymous only for an exact Anonymous marker in that same card; missing profile identity is not treated as anonymous.
Fixture-backed examples (synthetic)
The detailed examples below and the shared field reference use synthetic fixture-shaped contracts, not live Quora records. Populated fields require scoped evidence; blocked sources produce typed status rows.
{"content_type":"question","question_id":"fixture-question-1","question_title":"Fixture question","answer_count_text":"12 answers","answer_count":12,"access_state":"visible","source_items_emitted":1,"source_cap_reached":false,"field_sources":{"question_title":"visible_dom_or_metadata","answer_count":"scoped_label_or_structured_data"}}
{"content_type":"status","access_state":"cloudflare","http_status":403,"source_items_emitted":0,"source_cap_reached":false,"field_sources":{"access_state":"runtime_or_access"}}
Authorized session input accepts either structured cookies or the secret cookieString browser-header form (name=value; name2=value2). Structured cookies win on duplicate names; malformed pairs are ignored, the raw header is never emitted, and string cookies have no invented expiry metadata. This is an access input, not a bypass or private-content guarantee.
For standalone post rows, explicitly labeled visible post-header and Originally Answered evidence is preserved as post_header and originally_answered; both fields are omitted when that distinct source evidence is absent. Standalone post rows also expose conditional post_author_credentials from the bounded owning post-author scope; this is descriptive context only and is never inferred from post prose or a profile slug.
minAnswerUpvotes is a conservative local answer filter. Positive thresholds retain only answer records with explicit parseable upvote evidence; missing or ambiguous metrics are excluded from the filtered answer set and counted in answer_filter_unknown_upvotes, never treated as zero. The parent question reports observed, emitted, and threshold diagnostics.
answerSortOrder supports observed, upvotes_desc, newest, and views_desc for extracted answer rows. Typed metrics/timestamps determine ordering; records missing the selected metric retain stable fallback order and are counted in answer_sort_fallback_count. filterAiAnswers supports include, exclude, and only: only explicit visible AI attribution qualifies, while missing labels remain unknown and are never classified as human. Legacy boolean API values map to include/exclude.
For compatibility with other Quora actors, sortAnswersBy maps relevance, recency, and upvotes to the canonical sort modes, and minUpvotes maps to minAnswerUpvotes. Canonical names take precedence when both are supplied.
Discovery compatibility aliases are also accepted: searchQuery, queries, and searchKeywords → searchQueries; contentTypes and scrapeType → resultTypes; timeFilter → searchTimeFilter; maxItemsPerQuery → maxItemsPerSource; and maxResultsPerQuery → maxResults. searchType accepts marketplace human labels such as All types, Author, and Question, normalizing them to the bounded canonical filter. sortBy maps only supported values such as recent to sortOrder=newest; unsupported modes such as most_upvoted fail validation rather than being silently mapped to a different metric. includeHtmlContent maps only to bounded answer HTML (includeAnswerHtml); it does not claim raw HTML for every entity type. Canonical names take precedence.
topCommentsOnly retains only comments with an explicit scoped Top/Most helpful/Popular label. Unknown comments are excluded and counted in comment_filter_unknown; position, votes, and wording do not establish that a comment is top.
For direct standalone post URLs, includePostComments opts into the bounded top_comments projection, and maxCommentsPerPost (0–5,000; default 50) limits visible comment nodes inspected. Only an explicitly labeled Top, Most helpful, Popular, or supported localized collection qualifies; ordinary comments, position, votes, and reactions never establish a top-comment claim. Overview mode forces this expansion off.
onlyUnanswered is a fail-closed question opportunity filter. When enabled, question and search-card rows are retained only when question_is_unanswered=true from an explicit typed zero answer count; unknown and positive counts are excluded while status rows remain. It is local payload filtering and does not prove complete answer-history coverage.
Retained rows from that mode expose unanswered_filter_enabled=true and unanswered_filter_state (matched_zero_answer or not_applicable) so downstream consumers can audit why the mode was active.
commentSortOrder supports observed, upvotes_desc, reactions_desc, replies_desc, and newest. Sorting uses only typed comment metrics or parseable machine timestamps; missing values remain in stable fallback order and are counted in comment_sort_fallback_count.
Profile rows preserve explicit profile_joined_date labels and bounded visible profile_active_spaces links when present; missing evidence is omitted and no complete membership claim is made.
maxItemsPerSource is a common semantic cap across direct, query, and typed sources. It is applied after extraction but before output projection; status and run-manifest rows remain visible. Capped rows expose source_cap, source_items_observed, source_items_emitted, and source_cap_reached. Every emitted row also echoes the normalized boundary as max_items_per_source for run-level reconciliation. A value of 0 preserves unlimited per-source behavior; it never means complete Quora history.
Nested projection is supported with bounded dotted paths such as answers.answer_text, comments.comment_text, feed.feed_text, or contributors.contributor_name where that collection exists. Child identity/URL fields remain protected, and an empty list keeps the full backward-compatible row.
Exact high-value output names: ../RESEARCH/README_FIELD_REFERENCE.md#quora-scraper-and-search-scraper.
Standalone-post rows expose post_author_is_anonymous only when the owning post scope contains an exact visible Anonymous/Anonymous User/Anonymous Contributor marker; missing identity is never inferred as anonymous.
outputFields is an optional compact-output allowlist. It is applied after extraction, filtering, deduplication, and provenance; protected identity, access/status, source, traversal, and parser fields remain even when omitted from the allowlist. includeFieldEvidence independently controls field_sources and fields_present; output_projection reports full or selected, and the run manifest records the requested projection.
For Space rows, includeSpaceSettings is opt-in and emits space_settings_state, space_allowed_content_types, space_comment_permission, space_submission_policy, space_content_requirements, space_distribution_state, space_contributor_request_state, and bounded space_settings_text only from explicitly labeled rendered settings/content-permission regions. Missing settings remain unknown; role labels, counts, and public access never infer policy.
The run manifest includes bounded input_accounting entries for each processed source URL: final URL, processed/error outcome, emitted record/type/access counts, navigation totals, and duration. This is reconciliation evidence, not platform billing or cost data.
Question and answer rows preserve raw visible engagement labels alongside the existing typed values (view_count_text, answer_count_text, follow_count_text, upvotes_text, comments_count_text, and shares_count_text). This lets consumers retain localized evidence instead of losing it during numeric normalization. Question rows expose question_is_unanswered only from an explicit typed zero answer count; missing answer counts remain unknown. They may additionally expose answer_to_view_ratio and follow_to_answer_ratio only when their same-scope typed operands are explicit and their denominators are non-zero; search-card rows use their corresponding scoped typed metrics when available. Each companion evidence object records the versioned formula/basis and operands. These are workload/opportunity signals, not quality, popularity, relevance, or ranking claims.
Answer rows expose answer_is_collapsed only when an answer-scoped expand/collapsed control explicitly reports aria-expanded="false"; truncation, paywall, and absent-control states remain distinct.
Unified Post rows expose post_body_completeness and explicit-label-only post_is_restricted; missing or gated bodies remain conservative rather than being treated as complete or private.
Unified Post rows also preserve raw author_content_views and typed author_content_views_value only when an explicit author card/region contains a labeled content-views metric. Post view controls and ordinary body text never populate this author-scoped field.
Optional dateFrom and dateTo accept ISO dates or datetimes and apply an inclusive local filter to native search cards with machine-readable timestamps. customDateFrom and customDateTo are backward-compatible aliases, with dateFrom/dateTo taking precedence when both are supplied. Returned rows carry search_date_range_state=matched_machine_datetime; cards without machine timestamps are excluded when a bound is set. This does not claim Quora applied a server-side date filter.
Search-result rows with an explicit machine-readable search_published_at also carry derived search_age_days and search_age_state. Age is calculated against the run's scrape_timestamp; relative labels, invalid dates, and missing dates produce machine_datetime_unavailable, while future timestamps produce machine_datetime_future and a clamped age of zero. These fields are bounded timestamp arithmetic, not a trend score or a claim about complete Quora activity.
Native search-card rows preserve explicit card-scoped search_author_name/search_author_url, search_author_id only from matching JSON-LD entity and visible profile URLs, and bounded search_author_credentials when a profile link and adjacent credential/tagline evidence are rendered, plus search_view_count_text, search_answer_count_text, and search_follow_count_text labels with typed companions when parseable. When maxLoadMoreAttempts is positive, the actor separately clicks exact visible search continuation controls and records search_load_more_attempts, search_load_more_clicks, and search_load_more_state; this is independent from scrolling and never claims complete history. They also use the existing combined media_urls and typed image_urls/video_urls channels when includeMedia is enabled. Collection is limited to explicit media descendants of the owning bounded result card; page-level, sidebar, and neighboring-card assets are excluded. Disabled or unobserved media remains empty/omitted according to the existing row projection.
Unified question rows also expose conditional question_author_credentials from the bounded owning Question author scope; this is descriptive context only and is never inferred from question prose, answer cards, or a profile slug.
When includeMedia is enabled, native search-card rows also expose the existing combined media_urls and typed image_urls/video_urls channels. These are collected only from explicit media descendants of the owning bounded result card; page-level, sidebar, and neighboring-card assets are excluded. Disabled or unobserved media remains empty/omitted according to the existing row projection.
Extract structured evidence from visibly accessible Quora questions, answers, profiles, topics, spaces, and posts. Search by keywords or provide direct URLs; cookies and proxy configuration are optional authorized inputs, not guaranteed access. Cloudflare, login, Quora+, removed, and network outcomes are emitted as typed status rows.
What can this scraper do?
- Search by keywords — Enter a search term and the scraper follows the configured discovery path, preserving source/query and rank evidence when visible cards are accessible
- Scrape direct URLs — Provide Quora question, profile, topic, space, or standalone post URLs to scrape specific pages
- Extract bounded visible answers — Get rendered answer text, author details, labeled upvotes, comments, and engagement metrics under the configured caps; hidden or complete history is not claimed
- Detect AI answers — Automatically identifies Quora AI-generated answers vs. human answers
- Multiple content types — Questions, answers, user profiles, topics, and spaces all in one scraper
- Filter and cap output — Select result types, cap total records, and set an independent answer limit per question
- Access diagnostics — Cloudflare, login, and Quora+ gates are emitted as explicit status records instead of fake content
Input
scrapeDepth accepts overview or detail (default detail). Overview is a bounded initial-render mode and disables answer, comment, and profile-activity expansion even if those toggles are supplied. Detail retains the configured bounded traversal. Every row reports requested_depth, observed_depth, and detail_expansion_state; these fields distinguish requested intent from blocked or bounded work and never claim complete Quora history.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
| Search Keywords | string[] | No | — | Keywords used by the configured discovery path; result availability, ordering, and relevance are access-dependent. |
| Direct Quora URLs | string[] | No | — | Direct Quora URLs to scrape (questions, profiles, topics, spaces, posts, or /search?q=...) |
| Start URLs | request list | No | — | API-compatible URL objects; alternative to Direct Quora URLs |
| Result Types | enum[] | No | all | question, answer, profile, profile_activity, topic, space, post, comment, or search_result |
| Max Results | integer | No | 50 | Maximum number of results per search query or answers per question (1–5,000) |
| Native Search Type | enum | No | all | For visible /search cards only: all, question, answer, profile, topic, space, or post |
| Native Search Time Filter | enum | No | all_time | For visible /search cards only: all_time, past_hour, past_day, past_week, past_month, or past_year; non-all filters require a visible machine timestamp |
| Native Search Sort Order | enum | No | relevance | Visible-card ordering: relevance (DOM order), newest (machine timestamp), or most_viewed (visible view count); missing evidence keeps stable source order |
| Deduplication Mode | enum | No | canonical | canonical suppresses overlapping entity rows by stable ID/normalized canonical URL; none preserves every observed non-status row. The run manifest reports duplicate_records_suppressed. |
| Global Order | enum | No | observed | observed, newest, views_desc, answers_desc, or followers_desc; typed-evidence ordering is buffered and unknown values use stable fallback placement. |
| Max Items | integer | No | 0 | Global output cap; 0 means unlimited |
| Max Items Per Source | integer | No | 0 | Per-query/direct-URL semantic-item cap; status/access rows remain visible; 0 means unlimited |
| Delay Between Requests | number | No | 2 | Seconds between search/page navigations (0–30). This controls pacing and cost but is not a bypass guarantee. |
| Max Answers per Question | integer | No | 50 | Independent answer cap for each question (0–5,000) |
| Max Scroll Attempts | integer | No | 20 | Bounded browser pagination attempts per question (0–20) |
| Include Answers | boolean | No | true | Disable answer expansion when only question metadata is needed |
| Include Media and Links | boolean | No | true | Extract visible image/video URLs and external links from question and answer content |
| Include Answer HTML | boolean | No | true | Include rendered answer HTML; disable to reduce dataset size |
| Include Engagement Breakdown | boolean | No | false | Opt in to explicitly labeled answer-scoped upvote/reaction items in nested answers; aggregate metrics remain separate |
| Include Visible Comments | boolean | No | false | Emit only comments rendered in visible comment-content nodes |
| Max Comments per Answer | integer | No | 50 | Independent visible-comment cap per answer |
| Include Profile Activity | boolean | No | false | Extract public activity visibly rendered on profile pages |
| Profile Activity Types | enum[] | No | answer/question/post | Activity categories to retain |
| Max Profile Activity Items | integer | No | 50 | Maximum visible activity rows per profile |
| Only New or Changed Items | boolean | No | false | Suppress unchanged records using persistent IDs and content hashes |
| Incremental State Store ID | string | Only with incremental mode | — | Name of the Key-Value Store used to persist state across runs (auto-created on first use) |
| Proxy Configuration | object | No | Apify Residential | Proxy settings. Defaults to Apify Residential proxy, which cloud measurement showed clears Quora's anti-bot challenge reliably; the shared AUTO/datacenter pool is challenged on a large fraction of requests and is not used unless you explicitly select a different group. |
At least one of Search Keywords, Direct Quora URLs, or Start URLs is required.
Example input
{"searchQueries": ["python programming", "machine learning"],"maxResults": 10,"proxyConfiguration": {"useApifyProxy": true,"apifyProxyGroups": ["RESIDENTIAL"]}}
{"directUrls": ["https://www.quora.com/What-is-Python-used-for","https://www.quora.com/topic/Python-programming-language-1","https://www.quora.com/profile/Guido-van-Rossum-1"],"maxResults": 20}
Output
Each run produces a dataset with flat rows. Every row includes a content_type field so you can filter by type.
Every populated field also appears in field_sources, using conservative channel classes such as visible_dom_or_metadata, url_or_visible_link, machine_datetime_or_jsonld, scoped_label_or_structured_data, rendered_dom_or_document, runtime_or_access, or input_provenance. This is channel-level provenance, not a claim that a particular internal CSS selector supplied the value.
Question results
| Field | Type | Example |
|---|---|---|
content_type | string | "question" |
title | string | "What is Python primarily used for?" |
url | string | "https://www.quora.com/What-is-Python-primarily-used-for" |
answer_count | integer | 100 |
answer_count_text | string | "100" or "100+" when Quora renders a capped/rounded display |
answer_count_is_lower_bound | boolean | true when the visible label is capped (e.g. Quora's own "100+" style aggregation), meaning the real answer count may be higher |
follow_count | integer | 42 |
topics | string[] | ["Python programming language", "Software Development"] |
source_url | string | "https://www.quora.com/What-is-Python-used-for" |
source_query | string | "python programming" (empty if from direct URL) |
scrape_timestamp | string | "2026-03-08T18:28:03.140078+00:00" |
Answer results
| Field | Type | Example |
|---|---|---|
content_type | string | "answer" |
title | string | "What is Python primarily used for?" |
url | string | "https://www.quora.com/What-is-Python-primarily-used-for/answer/John-Smith" |
answer_text | string | Bounded visibly rendered answer text when available |
answer_url | string | Direct link to the answer |
author_name | string | "John Smith" |
author_url | string | "https://www.quora.com/profile/John-Smith" |
author_credentials | string | "Software Engineer at Google" |
upvotes | integer | 89 |
comments_count | integer | 4 |
shares_count | integer | 2 |
answer_timestamp | string | "2y" |
is_ai_answer | boolean | false |
question_title | string | "What is Python primarily used for?" |
question_url | string | "https://www.quora.com/What-is-Python-primarily-used-for" |
source_url | string | Original input URL |
source_query | string | Search keyword (empty if from direct URL) |
scrape_timestamp | string | ISO 8601 timestamp |
Profile results
| Field | Type | Example |
|---|---|---|
content_type | string | "profile" |
title | string | "Guido van Rossum" |
name | string | "Guido van Rossum" |
url | string | "https://www.quora.com/profile/Guido-van-Rossum-1" |
bio | string | User biography |
credentials | string | Professional credentials |
profile_image_url | string | Profile picture URL |
follower_count | integer | 3100 |
following_count | integer | 15 |
answer_count | integer | 42 |
question_count | integer | 5 |
total_views | integer | 1200000 |
content_views_this_month | integer | 4200 only when an exact monthly window label is visible |
content_views_this_month_text | string | Raw monthly content-view label/value |
source_url | string | Original input URL |
scrape_timestamp | string | ISO 8601 timestamp |
Topic results
| Field | Type | Example |
|---|---|---|
content_type | string | "topic" |
title | string | "Python (programming language)" |
name | string | "Python (programming language)" |
url | string | "https://www.quora.com/topic/Python-programming-language-1" |
description | string | Topic description |
follower_count | integer | 1600000 |
question_count | integer | 5000 |
source_url | string | Original input URL |
scrape_timestamp | string | ISO 8601 timestamp |
Space results
| Field | Type | Example |
|---|---|---|
content_type | string | "space" |
title | string | "Data Science" |
name | string | "Data Science" |
url | string | "https://www.quora.com/q/data-science" |
description | string | Space description |
follower_count | integer | 104000 |
follower_count_text | string | Raw visible follower label |
member_count | integer | 900 only when explicitly labeled as members |
member_count_text | string | Raw visible member label |
space_visibility | string | "public" only when explicitly labeled |
post_count | integer | 500 |
contributor_count | integer | 25 |
source_url | string | Original input URL |
scrape_timestamp | string | ISO 8601 timestamp |
Sample output
{"content_type": "answer","title": "What is Python primarily used for?","url": "https://www.quora.com/What-is-Python-primarily-used-for/answer/Pratima-Yadav-117","answer_text": "Python is used for various purposes due to its versatile nature...","answer_url": "https://www.quora.com/What-is-Python-primarily-used-for/answer/Pratima-Yadav-117","author_name": "Pratima Yadav","author_url": "https://www.quora.com/profile/Pratima-Yadav-117","author_credentials": "","upvotes": 4,"comments_count": 0,"shares_count": 2,"answer_timestamp": "4y","is_ai_answer": false,"question_title": "What is Python primarily used for?","question_url": "https://www.quora.com/What-is-Python-primarily-used-for","source_url": "https://www.quora.com/What-is-Python-used-for","scrape_timestamp": "2026-03-08T18:28:03.140078+00:00"}
Additional normalized fields
When answers are enabled and machine-readable timestamps are present, last_answered_at is the latest timestamp among answer rows extracted in this run. It is omitted when only relative labels are visible and does not claim hidden/deleted answer history.
For keyword inputs, maxConcurrency (1–3) batches independent indexed discovery queries. Result ordering remains deterministic, while Playwright page extraction stays sequential so shared deduplication and incremental state updates remain correct. The run manifest exposes the normalized value as max_concurrency.
Question records may include canonical_url, question_id, question_body, description, view_count, created_at, topic_urls, combined media_urls, typed image_urls/video_urls, and outbound_links when Quora exposes them. The typed channels are scoped to the question root after nested answer regions are removed. The unified output also adds question_view_count_value, question_answer_count_value, and question_follow_count_value only when the corresponding raw label is conservatively parseable. When answers are enabled, answer_count_scraped records the number of answer rows actually extracted under the configured cap; answer_limit_reached reports that the configured cap was reached, and neither field proves that hidden answers exist. This is separate from Quora’s visible answer_count. On popular questions Quora itself renders a rounded/capped answer badge (e.g. "100+"); when that capped form is observed, answer_count_is_lower_bound is true and answer_count_text keeps the trailing +, so answer_count should be read as "at least this many," not exact. Answer records may include answer_id, answer_html, answer_created_at, answer_updated_at, answer_upvote_count_value, answer_comment_count_value, answer_share_count_value, media_urls, and outbound_links. Set includeMedia to false to omit media, includeOutboundLinks to false to omit external links, and includeAnswerHtml to false to omit rendered answer HTML. Fields are omitted when unavailable; no placeholder values are emitted.
Question rows may additionally include related_questions only when visibly linked cards occur inside an explicitly labeled Related questions region; unlabelled links and utility/answer routes are not treated as related questions.
Answer rows additionally expose answer_translation_state and answer_translation_label only for an explicit translation/original-answer marker scoped to that answer card, plus answer_original_url only for an explicit original-answer link. Browser locale, language mismatch, and translated-looking prose never synthesize these fields.
Answer rows additionally expose answer_is_anonymous only for an exact visible Anonymous marker in the scoped author region; missing authors and inaccessible profiles remain unknown.
Every normalized row also includes provenance fields: parser_strategy, parser_version, record_id, requested_result_types, fields_present, scroll_attempts, and scroll_state. The unified actor additionally reports bounded operational evidence: scrape_duration_ms, navigation_attempts, navigation_retries, ordered navigation_retry_delays_seconds, final_url, http_status when available, proxy_used, proxy_resolution_state (resolved, or unavailable if the enforced Residential proxy could not be provisioned), request_delay_seconds, and browser_locale. These values describe this run only; endpoint resolution is not proof of successful access, and the actor never exposes proxy credentials or claims anti-bot bypass. parser_version is currently quora-suite-v1 and identifies the deployed parser contract. record_id prefers a visible entity ID and falls back to canonical URL, providing a stable join key across overlapping seeds. fields_present is calculated from the actual non-empty row before provenance is added, so consumers can distinguish a field that was not rendered from a parser or access failure. For search/question/answer rows, scroll_state is not_requested, max_items_reached, exhausted, or cap_reached; it records whether visible pagination ended naturally or hit the configured budget. Each run also emits one content_type: "run_manifest" row with aggregate emitted-record/type/access counts, attempted source pages, total navigation attempts/retries, run timestamps/duration, requested modes, pacing, locale, and proxy state. Filter this row out when only entity records are desired. Status rows use parser_strategy: "access_gate_or_status"; entity rows identify the visible-DOM/JSON-LD strategy.
When onlyNewItems is enabled, provide a persistent state store name (reuse the same name across runs). Emitted entity rows include change_type (new or changed) and a deterministic content_hash; unchanged rows are omitted. Status rows are retained so blocked runs remain visible.
redirect_chain is the ordered URL sequence observed from Playwright request ancestry, bounded to 16 entries and including the requested seed when available. It is diagnostic evidence, not a reconstruction of redirects hidden from the browser.
When includeComments is enabled, visible comments are emitted as content_type: "comment" rows with comment text, author, timestamp, depth, and parent answer/question references. Quora may lazy-load or hide comments; the actor does not infer hidden comments from comments_count.
When includeProfileActivity is enabled, visibly rendered public profile links are emitted as content_type: "profile_activity" rows for selected answer, question, or post categories. This is not a claim of complete historical profile activity; pagination, login-only tabs, and hidden items are not inferred. includeActivityText (default true) gates the bounded visible activity_text on each row, includeActivityHtml (default false) additionally gates a bounded (≤30000 bytes), redacted activity_html snippet of the same card, and maxItemsPerTab independently caps how many rows are emitted per tab (answers/questions/posts) on top of the overall maxProfileActivityItems cap.
Profile rows may include profile_topics as explicit visible topic URL/name pairs from the profile metadata scope; activity-card topics and text-based topic inference are excluded. When a topic also renders inside an explicit "Knows about"/topic-credentials section with a visible per-topic answer-count label, that entry additionally carries answer_count_text (raw label) and answer_count_value (typed integer, only when the label is conservatively parseable).
They may also include employment and education only from explicit public labels or Person JSON-LD (jobTitle, worksFor, alumniOf); no identity or activity inference is performed.
Set includeEntityGraph to true to additionally emit normalized content_type: "entity_edge" rows. Edges have deterministic edge_id, typed from_record_id/to_record_id, URLs, a relationship name, and edge_evidence_field. Supported relationships include question-to-topic/answer, answer-to-author, comment-to-parent/author, profile-to-activity, and space/post-to-feed or author. The actor emits an edge only when the scraped row contains an explicit endpoint URL or ID; it never joins by a display name or guesses a hidden relationship. maxGraphEdges is an independent cap and graph rows can be emitted in addition to maxItems. Graph rows use parser_strategy: "derived_relationships" and retain source/access provenance.
Answer rows also include transparent derived signals: answer_evidence_score / answer_quality_evidence measure observable evidence completeness using quora-answer-evidence-v1, while answer_quality_proxy_score / answer_quality_proxy measure observable structure and evidence signals using quora-answer-quality-proxy-v1. The latter considers only text substance, sentence/formatting signals, visible links/media, identity/date evidence, and labeled metrics. Neither signal is Quora's correctness, authority, relevance, popularity, or top-answer ranking.
Direct /post/... URLs are emitted as content_type: "post" records with visible body, author, timestamp, media, and external-link fields. The actor does not treat an unclassified URL as a post.
Direct /search?q=... URLs use the native Quora search page and emit content_type: "search_result" rows for visibly rendered cards, including rank, title, snippet, and result URL. These are discovery records; the actor does not claim that a search card contains the full question page.
How much does it cost?
The Quora Scraper uses pay-per-event pricing at $5 per 1,000 results. Each question, answer, profile, profile-activity row, topic, space, post, or comment counts as one result.
A typical run scraping 1 search query with 10 answers costs approximately $0.05–0.10 in platform credits for compute, plus Apify Residential proxy usage (billed separately by Apify at its standard per-GB rate); this actor uses Residential proxy by default because it is required for reliable access.
FAQs
Do I need a Quora account or cookies?
No for the default public mode. An optional cookies input is available when an authorized session is needed to pass an access gate; cookies are secret input and may expire. session_cookie_expiry_state reports only expiry metadata present at input time and never proves that Quora accepted the session. The actor never accesses private content by design.
When includeComments is enabled, nested answer comment rows expose comment_limit_reached when the positive maxCommentsPerAnswer cap is reached and comment_depth_limit_reached when visible comments deeper than maxCommentDepth are omitted. They also expose comment_items_observed_count and comment_items_emitted_count, separating distinct eligible comments observed before the extraction cap from retained nested rows. Parent answer rows expose answer_comments_observed_count and answer_comments_emitted_count; the latter is reconciled after collection filters and local caps. maxDepth is accepted as a compatibility alias. They also expose separately scoped upvote/reaction labels, conservative reply status/path fields, and bounded reply-expansion diagnostics including reply_expansion_source. maxReplyExpandAttempts controls visible reply/comment expansion per question page. This is a bounded traversal diagnostic, not a complete comment-history claim.
When an explicitly comment-scoped control exposes a typed reaction label, the same rows also preserve comment_reaction_breakdown items with type, raw label, and raw value, plus optional value_numeric for unambiguous localized counts; upvotes, icons, and aggregate-only values are excluded.
When Quora exposes a comment-author identifier in an explicit data/structured attribute, nested comment rows preserve comment_author_id; profile slugs are never treated as IDs.
Nested comments also expose positive-only comment_is_anonymous evidence when an exact anonymous marker is visible in the scoped author region.
Nested rows preserve comment_collection_label and comment_is_top_labeled only for an explicit Top/Most helpful/Popular comments region; ordering and votes are not used.
Post-capable results may additionally expose top_comments, containing only explicitly labeled top/helpful/popular comment objects; it is omitted without label evidence and is not a complete-history claim. For standalone posts, maxCommentDepth applies to these visible nested comments too, and comment_depth_limit_reached is emitted only when a deeper rendered comment was observed and omitted. When enabled, comment-content-scoped media/external links and explicitly labeled comment upvote/reaction totals are retained under the same includeMedia/includeOutboundLinks controls.
Nested answer and opt-in standalone-post comments preserve parent_comment_id only from an explicit rendered parent marker. parent_comment_url is emitted separately only when that parent exposes an explicit comment permalink; missing graph endpoints remain unknown.
Nested rows also expose comment_visibility only when explicit metadata or an exact placeholder label identifies visible, deleted, hidden, or restricted state.
Breakdown items additionally expose value_numeric when the raw value parses unambiguously; raw labels remain authoritative.
Nested answer rows preserve observed answer order separately from any explicit top/best label through answer_rank_source, and expose answer_text_completeness plus explicit-label-only Quora+ state. No hidden ranking or complete-history claim is made.
With includeEngagementBreakdown enabled, nested answer rows also expose reaction_breakdown only for explicitly labeled answer-scoped upvote/reaction controls; items may include value_numeric after unambiguous parsing, and comment-scoped labels and icons are excluded.
How does keyword search work?
The scraper first uses external indexed discovery for keyword inputs, then falls back to native Quora search pages when no indexed URLs are available. Direct /search?q=... inputs always use the native page. Native search cards are represented separately as search_result records so discovery metadata is not confused with scraped question content. Overlapping seeds are deduplicated by typed entity identity, while source_queries retains every query that found the entity. Each canonical row also exposes matched_input, matched_input_kind, bounded source_observations, and source_observation_count for exact mixed URL/query overlap evidence. Set globalOrder to buffer and deterministically order comparable parent rows; typed metrics or machine timestamps are required, unknown parents use stable fallback placement, and children remain adjacent to their parent.
Unified native search rows preserve explicit promoted/sponsored/ad evidence when a scoped visible label or structured marker exists. Commercial-looking text and ranking alone are never treated as promotion.
includePromoted defaults to true and can be set to false to remove only explicit positive promotion rows. excludePromoted remains supported and takes precedence when both controls are supplied. includePromotionEvidence controls promotion fields, including bounded promotion_evidence metadata, and includeUncertain can mark accessible unlabeled cards as promotion_state: unknown; unknown is not organic. promotion_url is emitted only when the explicit promotion marker is nested in a link, and is never inferred from the result URL.
What proxy should I use?
Leave the default. The actor defaults to, and internally enforces, Apify Residential proxy: real cloud measurement showed the shared AUTO/datacenter pool gets Cloudflare-challenged on a large share of Quora requests, while Residential reliably clears it. An input that supplies {"useApifyProxy": true} with no group, or omits proxyConfiguration entirely, is upgraded to Residential automatically — this is not a guaranteed bypass, but it is the configuration that measurably works. If you explicitly set a different apifyProxyGroups value, supply your own proxyUrl, or explicitly set useApifyProxy: false, that explicit choice is respected as-is (including running fully proxyless). If a challenge still occurs, the actor emits a typed access_state status record rather than fabricating data.
How many results can I get per run?
You can extract up to 5,000 results per search query or direct URL. For questions with many answers, the scraper automatically scrolls to load more content.
Can I scrape private or restricted profiles?
No. The scraper only extracts publicly visible content. Private profiles, restricted answers, and Quora+ paywalled content are not accessible.
What does is_ai_answer mean?
Quora generates AI-powered answers for some questions. The scraper automatically detects these and marks them with is_ai_answer: true and author_name: "Quora AI" so you can filter them.
What memory setting should I use?
Use 1024 MB or higher. Quora's pages are resource-intensive and require more memory than typical websites.
Why are some fields empty?
Some fields like follow_count, author_credentials, or description may be empty when Quora doesn't display that information on the page. This is normal and depends on the specific content being scraped.
What types of Quora URLs are supported?
- Questions:
https://www.quora.com/What-is-Python-used-for - Profiles:
https://www.quora.com/profile/Username - Topics:
https://www.quora.com/topic/Topic-Name - Spaces:
https://www.quora.com/q/space-nameorhttps://spacename.quora.com/
Limitations
- Requires 1024 MB memory minimum (Quora's pages are JavaScript-heavy)
- Keyword search returns 1–5 relevant question URLs per query
- Only publicly visible content can be scraped — no login-restricted or Quora+ content
- If Cloudflare or a login wall is encountered, the actor emits an
access_statestatus record so an empty result is not mistaken for successful extraction - Answer timestamps are relative (e.g., "2y", "6mo") rather than exact dates
- Quora's anti-bot layer challenges shared/datacenter IPs on a large share of requests; the actor defaults to and internally enforces Apify Residential proxy to keep results reliable
Relative date labels remain raw; when a scrape reference is available,
date_interpretationsexposes conservative bounded intervals for supported prefix, suffix, reversed-order, and non-Latin forms (for examplevor 2 Tagen,3 giorni fa,3日前, and3天前) and never claims an invented exact timestamp. Unsupported labels remain raw without a guessed date.
includeOutboundLinks independently controls visible external links for question, answer, and post rows; when omitted it follows includeMedia for backward compatibility, enabling links-only and media-only payloads.
maxAnswerHtmlBytes optionally bounds rendered answer HTML by UTF-8 bytes. Positive values emit answer_html_bytes and answer_html_truncated; zero preserves the uncapped legacy behavior, and the cap does not claim complete answer HTML.
includeRelatedQuestions controls explicitly labeled related-question links on question pages and defaults to true; disabling it only removes that context collection and does not change answer or post traversal.
includeQuestionTopics independently controls explicit topic names/URLs on question rows and defaults to true; disabling it only removes topic context and does not change answer or post traversal.
includeQuestionDetails independently controls rendered question description/body text and defaults to true; disabling it removes only those text fields while preserving identity, metrics, answer/post traversal, and status envelopes.
includeQuestionHtml and maxQuestionHtmlBytes independently control optional UTF-8-bounded rendered HTML from the answer-isolated question scope; this is distinct from the legacy answer-HTML alias and from full-document Raw Page evidence.
maxQuestionTopics bounds explicit topic names/URLs on question rows (0 means uncapped). Positive values emit question_topics_scraped and question_topics_limit_reached; these are local observation/cap diagnostics, not complete topic-membership claims.
maxRelatedQuestions bounds the emitted related-question collection on question rows (0 means uncapped). Positive values emit related_questions_scraped and related_questions_limit_reached; these are local observation/cap diagnostics, not completeness claims.
Complete schema input inventory
This generated inventory mirrors every property in .actor/input_schema.json; enum values are retained in the description below.
| Input field | Type | Default | Description |
|---|---|---|---|
searchQueries | array<string> | — | Keywords to search for on Quora. Attempts bounded discovery and preserves visible source/query evidence; availability, ordering, and relevance depend on the accessible result page and authorized access configuration. |
queries | array<string> | — | Compatibility alias for searchQueries. Canonical searchQueries takes precedence when both are supplied. |
searchKeywords | array<string> | — | Compatibility alias for searchQueries. Canonical searchQueries takes precedence, then queries, when multiple names are supplied. |
searchQuery | string | — | Compatibility alias for searchQueries for a single query. Canonical searchQueries, queries, and searchKeywords take precedence when supplied. |
directUrls | array<string> | — | Direct Quora URLs to scrape (questions, profiles, topics, spaces). Example: https://www.quora.com/What-is-Python |
startUrls | array<any> | [] | Alternative to directUrls. Accepts Quora URL objects such as {"url":"https://www.quora.com/..."}. |
targets | array<string> | [] | Mixed-target compatibility alias: URL-shaped values join directUrls and other values join searchQueries; canonical fields and targets are merged, then deduped according to dedupeMode downstream. |
resultTypes | array<string> | ["question","answer","profile","profile_activity","topic","space","post","search_result"] | Restrict output to selected entity types. search_result rows are visible Quora discovery cards, not fully scraped question entities. |
contentTypes | array<string> | — | Compatibility alias for resultTypes. Plural values such as questions, answers, profiles, topics, spaces, and posts are normalized to canonical entity types. |
scrapeType | string | — | Compatibility alias for resultTypes. It is used only when resultTypes and contentTypes are absent; search means the default unified entity set. Allowed values: search, question, answer, profile, author, topic, space, post. |
onlyUnanswered | boolean | false | Retain question/search rows only when an explicit typed answer count is exactly zero. Unknown or ambiguous counts are excluded; status and run-manifest rows remain. This does not prove complete answer-history coverage. |
excludePromoted | boolean | false | Exclude only cards with explicit visible or structured sponsored/promoted/ad evidence; unknown cards remain. |
includePromoted | boolean | true | Include cards with explicit promotion evidence. Set false to exclude them; explicit excludePromoted takes precedence when both are supplied. |
includePromotionEvidence | boolean | true | Retain promotion state and explicit label/source fields on native search-result rows. |
includeUncertain | boolean | true | Mark accessible cards without explicit promotion evidence as unknown; this is not a negative classification. |
maxResults | integer | 50 | Maximum number of results per search query or answers per question URL. |
maxResultsPerQuery | integer | 50 | Compatibility alias for maxResults. The canonical maxResults takes precedence. |
searchType | string | "all" | Fail-closed filter for visible native Quora search cards: all, question, answer, profile, topic, space, or post. Allowed values: all, all types, question, Question, answer, Answer, profile, Profile, Author, topic, Topic, space, Space, post, Post. |
searchTimeFilter | string | "all_time" | Visible-timestamp filter. Non-all filters exclude cards without a machine-readable timestamp; no date is inferred. Allowed values: all_time, past_hour, past_day, past_week, past_month, past_year. |
timeFilter | string | — | Compatibility alias for searchTimeFilter. Values such as 'Past Week' normalize to past_week; the canonical searchTimeFilter takes precedence. |
dateFrom | string | — | Optional inclusive ISO date/datetime lower bound applied locally to native search cards with machine-readable timestamps. |
dateTo | string | — | Optional inclusive ISO date/datetime upper bound applied locally to native search cards with machine-readable timestamps. |
customDateFrom | string | — | Backward-compatible alias for dateFrom. If both are supplied, dateFrom takes precedence. |
customDateTo | string | — | Backward-compatible alias for dateTo. If both are supplied, dateTo takes precedence. |
sortOrder | string | "relevance" | Stable DOM relevance order, newest by visible machine timestamp, or most viewed by visible view count. Missing evidence retains original order. Allowed values: relevance, newest, most_viewed. |
sortBy | string | — | Compatibility alias for sortOrder. Supported compatibility values include relevance, recent, most_recent, and most_viewed; unsupported values fail validation rather than being guessed. |
dedupeMode | string | "canonical" | canonical suppresses overlapping entity rows by stable ID or normalized canonical URL (default). none preserves every observed non-status row; duplicate_records_suppressed in the manifest remains 0 because no suppression is performed. Allowed values: canonical, none. |
outputFields | array<string> | [] | Optional allowlist (up to 3,000 unique names) for compact rows, including bounded dotted paths through nested collections (for example answers.answer_text, comments.comment_text, feed.feed_text, or contributors.contributor_name where applicable). Child identity/URL fields and the actor identity, access/status, source, traversal, graph, filter, and parser-envelope fields remain protected. |
includeFieldEvidence | boolean | true | Retain field_sources and fields_present after projection; disabling removes only these evidence maps, not semantic or protected access fields. |
language | string | "auto" | Browser language hint. Short language codes and supported regional aliases are accepted and normalize to the short code; auto keeps the actor default. This does not translate content or prove access. Allowed values: auto, en, es, fr, de, pt, it, ja, ko, hi, id, nl, pl, tr, vi, zh, en-US, en-GB, de-DE, fr-FR, es-ES, pt-BR, it-IT, ja-JP, ko-KR, hi-IN, tr-TR, zh-CN, zh-TW. |
maxItems | integer | 0 | Maximum total records across all inputs. Use 0 for no total cap. |
maxItemsPerSource | integer | 0 | Maximum semantic records emitted from each individual query or direct URL. Status/access rows are preserved; 0 means unlimited. |
maxItemsPerQuery | integer | 0 | Compatibility alias for maxItemsPerSource. The canonical maxItemsPerSource takes precedence. |
requestDelay | number | 2 | Seconds to wait between search/page navigations. Set to 0 for no added delay; this controls pacing but cannot guarantee bypassing Quora access controls. |
maxNavigationRetries | integer | 2 | Maximum bounded retries after transient navigation failures (0–3). Access gates, 404/410 responses, and challenge pages are not retried. |
maxConcurrency | integer | 1 | Maximum independent indexed discovery queries processed concurrently in deterministic batches. Page extraction remains sequential for shared deduplication and incremental state correctness. |
maxAnswersPerQuestion | integer | 50 | Answer limit for each question URL, capped at maxResults. To scrape more than the current maxResults value, raise maxResults too. |
minAnswerUpvotes | integer | 0 | When positive, retain only answer records with explicit machine-readable upvote evidence at or above this threshold. Missing or ambiguous upvotes are never treated as zero and are reported in filter diagnostics. |
minUpvotes | integer | 0 | Compatibility alias for minAnswerUpvotes. The canonical minAnswerUpvotes value takes precedence when both are supplied. |
answerSortOrder | string | "observed" | Order retained answer records by rendered order or explicit typed upvotes, machine timestamps, or view counts. Rows missing the selected metric retain stable fallback order and are counted. Allowed values: observed, upvotes_desc, newest, views_desc. |
sortAnswersBy | string | "relevance" | Compatibility alias mapped to answerSortOrder: relevance=observed, recency=newest, upvotes=upvotes_desc. The canonical answerSortOrder takes precedence when both are supplied. Allowed values: relevance, recency, upvotes. |
filterAiAnswers | string | "include" | Include all answers, exclude only explicitly AI-labeled answers, or retain only explicitly AI-labeled answers. Unknown/missing AI labels are never classified and are counted in filter diagnostics. Legacy boolean true/false values are accepted by the API normalizer as exclude/include. Allowed values: include, exclude, only. |
topCommentsOnly | boolean | false | Retain only comments with an explicit visible Top/Most helpful/Popular comments label. Unknown comments are excluded and counted; votes and position are not used as top evidence. |
commentSortOrder | string | "observed" | Order comments by observed order or an explicitly typed metric/timestamp. Unknown selected metrics remain in stable fallback order and are counted. Allowed values: observed, upvotes_desc, reactions_desc, replies_desc, newest. |
maxScrollAttempts | integer | 20 | Bound the number of browser pagination/scroll attempts per question page. This is capped at 20 to control cost and runaway pages. |
scrapeDepth | string | "detail" | Overview extracts the initial visible entity without answer/comment/profile-activity expansion. Detail performs the configured bounded traversal; neither mode claims complete history. Allowed values: overview, detail. |
includeAnswers | boolean | true | Expand and emit answer records for question pages. |
includeMedia | boolean | true | Extract image/video URLs from visible question, answer, and post content. External links are controlled independently by includeOutboundLinks. |
includeOutboundLinks | boolean | true | Independently retain visible external anchors from question, answer, and post content; when omitted, follows includeMedia for backward compatibility. |
includeAnswerHtml | boolean | true | Include the rendered HTML of answer content. Disable to reduce dataset size. |
includeHtmlContent | boolean | true | Compatibility alias for includeAnswerHtml. It controls bounded answer HTML only; canonical includeAnswerHtml takes precedence. |
includeRelatedQuestions | boolean | true | Retain explicitly labeled related-question links from question pages. Disable to reduce question-context payload; this does not affect answer extraction. |
maxRelatedQuestions | integer | 0 | Bound the emitted related-question collection on question rows. Zero preserves all observed related links; positive values emit related_questions_scraped and related_questions_limit_reached diagnostics. |
includeQuestionTopics | boolean | true | Retain explicit topic links associated with question rows. Disable to reduce question-context payload; this does not affect answer or post traversal. |
maxQuestionTopics | integer | 0 | Bound retained explicit topic names/URLs on question rows. Zero preserves all observed topic links; positive values emit question_topics_scraped and question_topics_limit_reached diagnostics. |
topicFilter | array<string> | [] | For question and search-result rows, retain only rows with an exact explicitly observed topic label or topic URL matching one of these terms. Unknown topic evidence is excluded and counted; this is local payload filtering, not a server-side Quora filter. |
includeSpaceSettings | boolean | false | Opt in to explicitly labeled Space settings/content-permission evidence for Space rows; missing settings remain unknown and admin-only queues are not scraped. |
includeQuestionDetails | boolean | true | Retain rendered question description/body text on question rows. Disable to reduce payload while preserving question identity, metrics, answer/post traversal, and status envelopes. |
includeQuestionHtml | boolean | false | Retain bounded rendered HTML from the question scope after nested answer regions are removed. |
maxQuestionHtmlBytes | integer | 0 | Optional UTF-8 byte ceiling for question HTML; positive values emit question_html_bytes and question_html_truncated. |
maxAnswerHtmlBytes | integer | 0 | Optional UTF-8 byte ceiling for rendered answer HTML. Zero preserves uncapped backward-compatible HTML; positive values emit answer_html_bytes and answer_html_truncated. |
maxPostHtmlBytes | integer | 0 | Optional UTF-8 byte ceiling for rendered standalone Post HTML in unified Post rows. Zero disables HTML retention; positive values emit post_html_bytes and post_html_truncated. |
includeComments | boolean | false | Extract only comments rendered in visible answer comment-content nodes; hidden comments are not inferred. |
includePostComments | boolean | false | For standalone post URLs, retain only comments inside an explicitly labeled Top/Most helpful/Popular collection as top_comments. Ordinary comments, votes, and position never qualify. |
maxCommentsPerPost | integer | 50 | Bound the visible post-comment nodes inspected when includePostComments is enabled; this is not a complete comment-history guarantee. |
includeEngagementBreakdown | boolean | false | Opt in to explicitly labeled answer-scoped upvote/reaction items in nested answer rows. Comment controls and icons are never decomposed. |
includePostAuthorContext | boolean | false | Opt in to explicitly labeled About/Bio, Education, and Joined/Member since values inside the owning standalone-post author scope; no profile navigation or inference is performed. |
maxPostAuthorContextItems | integer | 10 | Bounds active Spaces and labeled public external links retained from the author scope; observed/emitted/cap diagnostics distinguish truncation. |
maxPostAuthorContextItems | integer | 10 | Maximum explicitly labeled active Space links retained per standalone-post author; observed, emitted, and cap diagnostics are preserved. |
maxCommentsPerAnswer | integer | 50 | Maximum visible comments extracted from each answer when comments are enabled. |
maxReplyExpandAttempts | integer | 0 | Bound clicks on clearly labeled visible reply/comment expansion controls per question page. |
maxCommentDepth | integer | 20 | Maximum explicit comment nesting depth retained from answer pages and standalone-post top-comment collections. Root comments are depth 0; deeper comments are omitted and marked with comment_depth_limit_reached. The maxDepth alias is accepted for compatibility. |
includeProfileActivity | boolean | false | Extract public activity links visibly rendered on profile pages. |
includeEntityGraph | boolean | false | Emit deterministic entity_edge relationship rows when explicit endpoint URLs or IDs are present. This is opt-in and does not invent relationships. |
maxGraphEdges | integer | 5000 | Independent cap for entity_edge rows. Graph rows may be added in addition to maxItems. |
profileActivityTypes | array<string> | ["answer","question","post"] | Activity categories to retain: answer, question, or post. |
maxProfileActivityItems | integer | 50 | Maximum visibly rendered public activity records per profile. |
onlyNewItems | boolean | false | Compare stable IDs and content hashes against a persistent state store and suppress unchanged rows. |
stateStoreId | string | — | Name of the Key-Value Store used for onlyNewItems state (auto-created on first use). Required when onlyNewItems is enabled. |
usernames | array<string> | [] | Quora profile usernames/slugs (with or without a leading @ or trailing slash). Normalized to https://www.quora.com/profile/{slug} and merged into the direct-URL queue; also seeds author-discovery and post-discovery when relevant. |
includeProfileSensitiveFields | boolean | true | When false, strip location, employment, education, and profile_joined_date from profile/author rows. |
includeSensitiveProfileFields | boolean | true | Deprecated alias for includeProfileSensitiveFields. The canonical field takes precedence whenever both are supplied. |
includeSocialLinks | boolean | true | When false, omit social_links/website_url from profile/author rows. |
activityTab | string | "all" | Click the given profile-activity tab before extracting profile_activity rows. Allowed values: all, answers, questions, posts. |
includeActivityText | boolean | true | Include bounded visible activity_text on profile_activity rows. |
includeActivityHtml | boolean | false | Include bounded, redacted rendered activity_html on profile_activity rows. |
maxItemsPerTab | integer | 0 | Independent cap on profile_activity rows per tab (answers/questions/posts), in addition to maxProfileActivityItems. |
maxDetailItems | integer | 0 | When scrapeDepth=detail, bounded number of profile_activity items opened individually for detail_text/detail_title extraction. |
includeUnchanged | boolean | true | When onlyNewItems is set, also retain change_type=unchanged rows instead of dropping them. |
maxAuthors | integer | 1000 | Max unique discovered author rows to emit, independent of maxItems. |
maxSourcePages | integer | 100 | Max author-discovery source pages visited (search/question/answer/topic/Space/post/profile). |
maxEvidencePerAuthor | integer | 10 | Max evidence_urls/evidence_question_urls/evidence_answer_urls/evidence_answer_previews entries retained per author. |
includeProfileSnapshot | boolean | false | When true and resultTypes includes author, open each discovered author's profile and merge a full profile snapshot onto the author row. |
includeRelatedContent | boolean | false | Opt in to an explicit 'related/attached content' region scoped to one answer or post card. |
maxRelatedContent | integer | 20 | Bound related_content entries retained per answer/post. |
includePostHtml | boolean | false | Legacy alias: true sets a 100000-byte maxPostHtmlBytes ceiling unless maxPostHtmlBytes is explicitly set. |
includeReactionBreakdown | boolean | false | Legacy alias for includeEngagementBreakdown. |
maxPostsPerSource | integer | 0 | Bound /post/... links discovered (then fully scraped as post rows) from one profile/Space/search source page. Distinct from maxItemsPerSource, which caps already-scraped rows rather than discovery candidates. Opt-in: 0 (default) disables post discovery from these source pages entirely (no post/post_discovery rows emitted); set a positive value to enable it. Direct standalone post URLs passed via directUrls/startUrls are always scraped regardless of this setting. |
includeGraphEdges | boolean | false | Compatibility alias for includeEntityGraph. |
includeTopicQuestions | boolean | true | Emit visible topic question links as topic_question rows. |
maxTopicQuestions | integer | 50 | Bound topic_question rows per topic page. |
includeTopicPosts | boolean | false | Emit explicitly linked standalone Posts from a topic page as topic_post rows. |
maxTopicPosts | integer | 50 | Bound topic_post rows per topic page. |
includeRelatedTopics | boolean | true | Emit visibly linked related topics as topic_related rows. |
maxRelatedTopics | integer | 50 | Bound topic_related rows per topic page. |
includeAnswerPreview | boolean | false | Opt in to bounded, already-rendered answer-preview fields on question/search_result cards, without navigating to the full answer page. |
maxAnswerPreviews | integer | 3 | Bound answer_preview_* entries per question/search_result card. |
minAnswerPreviewUpvotes | integer | 0 | Minimum typed upvote threshold an answer preview must meet to be retained (fail-closed; unknown upvotes are excluded). |
cookies | array<object> | — | Optional authenticated browser cookies. Use only for content you are authorized to view; cookies are stored as secret input and are never logged. |
cookieString | string | — | Optional browser-exported name=value; name2=value2 header. Parsed locally into authorized session cookies; the raw header is never emitted or logged. |
proxyConfiguration | object | Apify Residential | Proxy settings. Defaults to, and is internally enforced as, Apify Residential proxy for reliable access; explicit overrides (a different apifyProxyGroups or a custom proxyUrl) are respected as-is. |
Unified Quora question rows now preserve question_created_at alongside question_updated_at from the question-scoped extractor. These are conditional visible/structured dates, not run timestamps or completeness claims. |