ChatGPT Scraper – Answers, Citations & Sources by Country avatar

ChatGPT Scraper – Answers, Citations & Sources by Country

Pricing

from $2.00 / 1,000 chatgpt searches

Go to Apify Store
ChatGPT Scraper – Answers, Citations & Sources by Country

ChatGPT Scraper – Answers, Citations & Sources by Country

Send prompts to chatgpt.com and get every answer back as structured JSON, with the web sources it cited: url, title and publisher. Ask from 54 countries, run prompts concurrently, and see whether ChatGPT recommends you. No OpenAI API key, no ChatGPT account, no browser.

Pricing

from $2.00 / 1,000 chatgpt searches

Rating

0.0

(0)

Developer

R.L.

R.L.

Maintained by Community

Actor stats

0

Bookmarked

13

Total users

7

Monthly active users

4 days ago

Last modified

Categories

Share

ChatGPT Scraper — Answers, Citations & Sources by Country

Ask ChatGPT anything at scale and get structured JSON back: the answer as Markdown text, plus the web sources it cited when it searched.

No OpenAI API key, no ChatGPT account, no browser. This Actor talks to chatgpt.com's own anonymous surface over plain HTTP requests, which is why a prompt is answered in seconds rather than in a browser session.

What you can use it for

  • AI visibility / brand monitoring — track whether ChatGPT recommends you, and which sources it cites when it does.
  • SEO and citation research — see which domains ChatGPT pulls from for a topic, with their titles and publishers.
  • Market and competitor research — ask the same question from different countries and compare the answers.
  • Content and product research — batch hundreds of prompts and get a clean dataset instead of copy-pasting from a chat window.

What inputs does it take?

FieldTypeDescription
promptsstring[]Required. One prompt per conversation, up to 4096 characters each. A longer prompt fails the run rather than being silently cut.
countrystringAnswer as if browsing from this country. Picked from a dropdown; it routes the proxy exit node, so it needs a proxy to have any effect.
additionalPromptstringAn optional second prompt asked after each prompt. Answered independently of the first, and its answer is the one saved.
requireSourcesbooleanMark an answer that cites no web sources as failed.
maxConcurrencyintegerPrompts answered at once. Default 5, max 20.
maxRetriesintegerRetries per prompt, each from a fresh session and exit node. Default 3.
proxyConfigurationobjectResidential proxy strongly recommended — see Why do I need a proxy?
{
"prompts": [
"Top hotels in New York",
"What are the biggest business trends to watch in the next five years?"
],
"country": "us",
"additionalPrompt": "Which European capital has the largest old town?",
"maxConcurrency": 5,
"proxyConfiguration": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"] }
}

Every prompt in a run uses the same country and additionalPrompt. To compare countries, run one task per country.

Which countries can I ask from?

54, plus "Any" for no geo-targeting. The code is passed to Apify Proxy as the exit country, which is the only thing ChatGPT actually sees:

RegionCountries
North AmericaUnited States us, Canada ca, Mexico mx
South AmericaBrazil br, Argentina ar, Chile cl, Colombia co, Peru pe
British Isles & OceaniaUnited Kingdom gb, Ireland ie, Australia au, New Zealand nz
Western EuropeGermany de, France fr, Spain es, Italy it, Netherlands nl, Belgium be, Austria at, Switzerland ch, Portugal pt, Greece gr
Northern & Eastern EuropeSweden se, Norway no, Denmark dk, Finland fi, Poland pl, Czechia cz, Romania ro, Lithuania lt, Latvia lv, Estonia ee, Ukraine ua, Russia ru
Middle East & AfricaTürkiye tr, Israel il, United Arab Emirates ae, Saudi Arabia sa, Egypt eg, South Africa za, Nigeria ng
AsiaIndia in, Pakistan pk, China cn, Hong Kong hk, Taiwan tw, Japan jp, South Korea kr, Singapore sg, Malaysia my, Indonesia id, Thailand th, Vietnam vn, Philippines ph

What does a result row look like?

One dataset row per prompt. This is a real row, trimmed:

{
"url": "https://chatgpt.com/c/6ab8ccd7-f160-83ea-9b34-a7581e05a71d",
"prompt": "What are the best noise cancelling headphones in 2026?",
"answer_text": "As of **September 2026**, the flagship noise-cancelling headphone field is led by …",
"answer_text_markdown": "As of **September 2026**, the flagship noise-cancelling headphone field is led by …",
"model": "gpt-5-6",
"web_search_triggered": true,
"citations": [
{
"url": "https://www.techgearlab.com/topics/audio/best-noise-cancelling-headphones?utm_source=chatgpt.com",
"title": "Best Noise Cancelling Headphones | Lab Tested",
"source": "TechGearLab",
"snippet": "",
"published_at": null
}
],
"search_sources": [ "… the same list as citations" ],
"references": [],
"links_attached": [],
"country": "us",
"prompt_sent_at": "2026-09-27T07:59:15.162000+00:00",
"index": 0,
"conversation_id": "6ab8ccd7-f160-83ea-9b34-a7581e05a71d"
}
FieldDescription
promptThe prompt this row answers — the second prompt, if you set one.
answer_text / answer_text_markdownThe answer. Both carry the same string: ChatGPT's own Markdown, including its headings, bold and tables. The inline citation markers ChatGPT embeds are stripped out, so the text reads cleanly.
web_search_triggeredWhether ChatGPT went to the web for this answer.
citationsOne entry per cited webpage: {url, title, source, snippet, published_at}. source is the publisher's name; snippet and published_at are filled only when ChatGPT provides them. Note the utm_source=chatgpt.com ChatGPT appends to the URLs.
search_sourcesThe same deduplicated source list as citations, under the name Bright Data's ChatGPT dataset uses, so a consumer written for either key works.
references / links_attachedPresent for schema compatibility and currently always empty on the anonymous surface.
modelThe model slug ChatGPT reported for the turn, e.g. gpt-5-6. The request asks for auto, so this is ChatGPT's choice, not yours, and it can be null if no slug came back.
countryThe country you asked from, or null for no geo-targeting.
prompt_sent_atUTC timestamp of the moment the prompt was submitted.
indexPosition of the prompt in your input.
conversation_id / urlChatGPT's id for the anonymous conversation, and the URL built from it. History and training are disabled for the turn, so treat it as an identifier rather than a page you can revisit.

Prompts that fail after every retry are still saved, carrying error and error_code instead of an answer, so a run never silently drops one.

Answers are pushed the moment they land, so with concurrency on, rows arrive out of order. Sort by index to get your input order back.

How much does it cost?

$2 per 1000 answers ($0.002 each), charged per answer actually delivered. A prompt that ends in an error — including one failed by requireSources — is written to the dataset and costs nothing. If you set a maximum cost per run, the Actor stops starting new prompts once it is reached instead of running past it.

Why do I need a proxy?

ChatGPT rate limits datacenter IPs hard: they start timing out and then get blocked by Cloudflare after a few dozen requests. Run this on residential proxies, which is the default configuration. country works by choosing the proxy's exit node, so it does nothing without one. Running with no proxy at all is allowed, and the log says so at the start of the run.

How fast is it, and what happens when a prompt fails?

Each prompt gets its own session and its own proxy exit node, so maxConcurrency shortens a run roughly proportionally — 5 at a time by default, up to 20. A single prompt answered in about 12 seconds in a check run on 2026-09-27; expect longer for long answers and for answers where ChatGPT searches the web.

A failure is retried from a fresh session and a different exit node, backing off 5s, then 10s, then 20s, capped at 60s, for up to maxRetries attempts. Three kinds get retried the same way: a dead proxy exit, an HTTP error from ChatGPT, and ChatGPT deciding a particular session must sign in. If every attempt fails, the row is saved with the error.

How does it work?

ChatGPT's anonymous surface (/backend-anon) answers a logged-out caller with JSON over server-sent events. The Actor completes the site's sentinel handshake, including its proof-of-work challenge, submits the prompt, and reads the streamed answer — the answer text, the model slug and the cited sources all arrive in that stream. No browser is started, which is why runs are quick and cheap compared with browser-driven scrapers.

FAQ

Do I need an OpenAI API key or a ChatGPT account?

No. The Actor uses the surface chatgpt.com serves to a logged-out visitor, so there is no key, no login and no per-token billing from OpenAI. You pay Apify per answer and nothing else.

Is this the same as calling the OpenAI API?

No, and the difference is the point. The API answers from the model alone; chatgpt.com decides for itself whether to search the web and attaches the pages it used. Those cited pages — which domains ChatGPT trusts for your query — are what this Actor is for. The trade-off is that you do not choose the model or the parameters.

How do I check whether ChatGPT recommends my brand?

Put the questions your buyers would type into prompts, set country, and read two things per row: whether your brand appears in answer_text, and which domains appear in citations. Run it on a schedule to see the answer drift over time — the same prompt does not always come back the same way.

Can I compare answers across countries in one run?

Not in one run: a run has one country, because the country is the proxy exit node. Save one task per country and run them side by side.

Does the second prompt remember the first answer?

No. Each prompt is submitted as its own turn with no shared thread, so the model never sees the first answer. A second prompt like "and which of those is cheapest?" cannot work — "those" refers to nothing. Write it so it stands alone.

Only the second prompt's answer is saved; the first answer is discarded. So additionalPrompt is a way to ask a different question per run, not a way to have a conversation. If you want both answers, put both questions in prompts instead and you get a row for each.

What happens to a prompt that keeps failing?

It is retried up to maxRetries times from a fresh session and exit node, and if it still fails the row lands in the dataset with error and error_code filled in and no charge. A failing prompt never takes the rest of the run down with it.

Can I throw away answers that cite nothing?

Set requireSources. An answer with no sources is then marked error: "no sources found" with error_code: "RequireSourcesError" — kept in the dataset so you can see what happened, and not charged.

Can I run it on a schedule?

Yes. A scheduled run over a fixed prompt list is the usual way to use it for visibility tracking. The Actor keeps no state between runs, so every run is a fresh snapshot; keep prompt and prompt_sent_at to line the snapshots up.

Limitations

  • One country per run.
  • Answers come from the logged-out ChatGPT, so there is no account history, no file uploads, no image generation, and no choice of model.
  • ChatGPT is non-deterministic: the same prompt can return different answers and cite different sources on different runs.
  • Answers can be wrong or out of date. Treat the cited sources as the ground truth, not the prose.