ChatGPT Scraper – Answers, Citations & Sources by Country
Pricing
from $2.00 / 1,000 chatgpt searches
ChatGPT Scraper – Answers, Citations & Sources by Country
Send prompts to chatgpt.com and get every answer back as structured JSON, with the web sources it cited: url, title and publisher. Ask from 54 countries, run prompts concurrently, and see whether ChatGPT recommends you. No OpenAI API key, no ChatGPT account, no browser.
Pricing
from $2.00 / 1,000 chatgpt searches
Rating
0.0
(0)
Developer
R.L.
Maintained by CommunityActor stats
0
Bookmarked
13
Total users
7
Monthly active users
4 days ago
Last modified
Categories
Share
ChatGPT Scraper — Answers, Citations & Sources by Country
Ask ChatGPT anything at scale and get structured JSON back: the answer as Markdown text, plus the web sources it cited when it searched.
No OpenAI API key, no ChatGPT account, no browser. This Actor talks to chatgpt.com's own anonymous surface over plain HTTP requests, which is why a prompt is answered in seconds rather than in a browser session.
What you can use it for
- AI visibility / brand monitoring — track whether ChatGPT recommends you, and which sources it cites when it does.
- SEO and citation research — see which domains ChatGPT pulls from for a topic, with their titles and publishers.
- Market and competitor research — ask the same question from different countries and compare the answers.
- Content and product research — batch hundreds of prompts and get a clean dataset instead of copy-pasting from a chat window.
What inputs does it take?
| Field | Type | Description |
|---|---|---|
prompts | string[] | Required. One prompt per conversation, up to 4096 characters each. A longer prompt fails the run rather than being silently cut. |
country | string | Answer as if browsing from this country. Picked from a dropdown; it routes the proxy exit node, so it needs a proxy to have any effect. |
additionalPrompt | string | An optional second prompt asked after each prompt. Answered independently of the first, and its answer is the one saved. |
requireSources | boolean | Mark an answer that cites no web sources as failed. |
maxConcurrency | integer | Prompts answered at once. Default 5, max 20. |
maxRetries | integer | Retries per prompt, each from a fresh session and exit node. Default 3. |
proxyConfiguration | object | Residential proxy strongly recommended — see Why do I need a proxy? |
{"prompts": ["Top hotels in New York","What are the biggest business trends to watch in the next five years?"],"country": "us","additionalPrompt": "Which European capital has the largest old town?","maxConcurrency": 5,"proxyConfiguration": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"] }}
Every prompt in a run uses the same country and additionalPrompt. To compare countries, run
one task per country.
Which countries can I ask from?
54, plus "Any" for no geo-targeting. The code is passed to Apify Proxy as the exit country, which is the only thing ChatGPT actually sees:
| Region | Countries |
|---|---|
| North America | United States us, Canada ca, Mexico mx |
| South America | Brazil br, Argentina ar, Chile cl, Colombia co, Peru pe |
| British Isles & Oceania | United Kingdom gb, Ireland ie, Australia au, New Zealand nz |
| Western Europe | Germany de, France fr, Spain es, Italy it, Netherlands nl, Belgium be, Austria at, Switzerland ch, Portugal pt, Greece gr |
| Northern & Eastern Europe | Sweden se, Norway no, Denmark dk, Finland fi, Poland pl, Czechia cz, Romania ro, Lithuania lt, Latvia lv, Estonia ee, Ukraine ua, Russia ru |
| Middle East & Africa | Türkiye tr, Israel il, United Arab Emirates ae, Saudi Arabia sa, Egypt eg, South Africa za, Nigeria ng |
| Asia | India in, Pakistan pk, China cn, Hong Kong hk, Taiwan tw, Japan jp, South Korea kr, Singapore sg, Malaysia my, Indonesia id, Thailand th, Vietnam vn, Philippines ph |
What does a result row look like?
One dataset row per prompt. This is a real row, trimmed:
{"url": "https://chatgpt.com/c/6ab8ccd7-f160-83ea-9b34-a7581e05a71d","prompt": "What are the best noise cancelling headphones in 2026?","answer_text": "As of **September 2026**, the flagship noise-cancelling headphone field is led by …","answer_text_markdown": "As of **September 2026**, the flagship noise-cancelling headphone field is led by …","model": "gpt-5-6","web_search_triggered": true,"citations": [{"url": "https://www.techgearlab.com/topics/audio/best-noise-cancelling-headphones?utm_source=chatgpt.com","title": "Best Noise Cancelling Headphones | Lab Tested","source": "TechGearLab","snippet": "","published_at": null}],"search_sources": [ "… the same list as citations" ],"references": [],"links_attached": [],"country": "us","prompt_sent_at": "2026-09-27T07:59:15.162000+00:00","index": 0,"conversation_id": "6ab8ccd7-f160-83ea-9b34-a7581e05a71d"}
| Field | Description |
|---|---|
prompt | The prompt this row answers — the second prompt, if you set one. |
answer_text / answer_text_markdown | The answer. Both carry the same string: ChatGPT's own Markdown, including its headings, bold and tables. The inline citation markers ChatGPT embeds are stripped out, so the text reads cleanly. |
web_search_triggered | Whether ChatGPT went to the web for this answer. |
citations | One entry per cited webpage: {url, title, source, snippet, published_at}. source is the publisher's name; snippet and published_at are filled only when ChatGPT provides them. Note the utm_source=chatgpt.com ChatGPT appends to the URLs. |
search_sources | The same deduplicated source list as citations, under the name Bright Data's ChatGPT dataset uses, so a consumer written for either key works. |
references / links_attached | Present for schema compatibility and currently always empty on the anonymous surface. |
model | The model slug ChatGPT reported for the turn, e.g. gpt-5-6. The request asks for auto, so this is ChatGPT's choice, not yours, and it can be null if no slug came back. |
country | The country you asked from, or null for no geo-targeting. |
prompt_sent_at | UTC timestamp of the moment the prompt was submitted. |
index | Position of the prompt in your input. |
conversation_id / url | ChatGPT's id for the anonymous conversation, and the URL built from it. History and training are disabled for the turn, so treat it as an identifier rather than a page you can revisit. |
Prompts that fail after every retry are still saved, carrying error and error_code instead of
an answer, so a run never silently drops one.
Answers are pushed the moment they land, so with concurrency on, rows arrive out of order. Sort
by index to get your input order back.
How much does it cost?
$2 per 1000 answers ($0.002 each), charged per answer actually delivered. A prompt that ends
in an error — including one failed by requireSources — is written to the dataset and costs
nothing. If you set a maximum cost per run, the Actor stops starting new prompts once it is
reached instead of running past it.
Why do I need a proxy?
ChatGPT rate limits datacenter IPs hard: they start timing out and then get blocked by Cloudflare
after a few dozen requests. Run this on residential proxies, which is the default
configuration. country works by choosing the proxy's exit node, so it does nothing without one.
Running with no proxy at all is allowed, and the log says so at the start of the run.
How fast is it, and what happens when a prompt fails?
Each prompt gets its own session and its own proxy exit node, so maxConcurrency shortens a run
roughly proportionally — 5 at a time by default, up to 20. A single prompt answered in about 12
seconds in a check run on 2026-09-27; expect longer for long answers and for answers where
ChatGPT searches the web.
A failure is retried from a fresh session and a different exit node, backing off 5s, then 10s,
then 20s, capped at 60s, for up to maxRetries attempts. Three kinds get retried the same way: a
dead proxy exit, an HTTP error from ChatGPT, and ChatGPT deciding a particular session must sign
in. If every attempt fails, the row is saved with the error.
How does it work?
ChatGPT's anonymous surface (/backend-anon) answers a logged-out caller with JSON over
server-sent events. The Actor completes the site's sentinel handshake, including its proof-of-work
challenge, submits the prompt, and reads the streamed answer — the answer text, the model slug and
the cited sources all arrive in that stream. No browser is started, which is why runs are quick
and cheap compared with browser-driven scrapers.
FAQ
Do I need an OpenAI API key or a ChatGPT account?
No. The Actor uses the surface chatgpt.com serves to a logged-out visitor, so there is no key, no login and no per-token billing from OpenAI. You pay Apify per answer and nothing else.
Is this the same as calling the OpenAI API?
No, and the difference is the point. The API answers from the model alone; chatgpt.com decides for itself whether to search the web and attaches the pages it used. Those cited pages — which domains ChatGPT trusts for your query — are what this Actor is for. The trade-off is that you do not choose the model or the parameters.
How do I check whether ChatGPT recommends my brand?
Put the questions your buyers would type into prompts, set country, and read two things per
row: whether your brand appears in answer_text, and which domains appear in citations. Run it
on a schedule to see the answer drift over time — the same prompt does not always come back the
same way.
Can I compare answers across countries in one run?
Not in one run: a run has one country, because the country is the proxy exit node. Save one task
per country and run them side by side.
Does the second prompt remember the first answer?
No. Each prompt is submitted as its own turn with no shared thread, so the model never sees the first answer. A second prompt like "and which of those is cheapest?" cannot work — "those" refers to nothing. Write it so it stands alone.
Only the second prompt's answer is saved; the first answer is discarded. So additionalPrompt is
a way to ask a different question per run, not a way to have a conversation. If you want both
answers, put both questions in prompts instead and you get a row for each.
What happens to a prompt that keeps failing?
It is retried up to maxRetries times from a fresh session and exit node, and if it still fails
the row lands in the dataset with error and error_code filled in and no charge. A failing
prompt never takes the rest of the run down with it.
Can I throw away answers that cite nothing?
Set requireSources. An answer with no sources is then marked error: "no sources found" with
error_code: "RequireSourcesError" — kept in the dataset so you can see what happened, and not
charged.
Can I run it on a schedule?
Yes. A scheduled run over a fixed prompt list is the usual way to use it for visibility tracking.
The Actor keeps no state between runs, so every run is a fresh snapshot; keep prompt and
prompt_sent_at to line the snapshots up.
Limitations
- One country per run.
- Answers come from the logged-out ChatGPT, so there is no account history, no file uploads, no image generation, and no choice of model.
- ChatGPT is non-deterministic: the same prompt can return different answers and cite different sources on different runs.
- Answers can be wrong or out of date. Treat the cited sources as the ground truth, not the prose.
