# Keyword Research Scraper — Google, YouTube, Amazon (`eiv/keyword-research-scraper`) Actor

Turn one seed into hundreds of real keywords from Google, YouTube, Amazon, Bing, eBay and DuckDuckGo autocomplete at once. Grouped into questions and comparisons, with a count of how many engines agreed - so you can see which phrases carry buying intent. No API key.

- **URL**: https://apify.com/eiv/keyword-research-scraper.md
- **Developed by:** [Eimantas V](https://apify.com/eiv) (community)
- **Categories:** SEO tools, E-commerce, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.20 / 1,000 keyword founds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Keyword Research Scraper — Google, YouTube, Amazon & more

One seed in. Hundreds of real keywords out — from **six engines at once**, grouped into questions and comparisons, with a note of which engines agreed.

```
seed: "protein powder"     1,000 keywords in 5.6 seconds

6x  whey protein powder                [google youtube amazon bing ebay duckduckgo]  BUY
6x  protein powder chocolate           [google youtube amazon bing ebay duckduckgo]  BUY
5x  protein powder for weight loss     [google youtube amazon bing duckduckgo]       BUY

questions    267   what protein powder is best · how protein powder is made
comparisons   78   protein powder vs creatine · protein powder vs whey protein
long tail    805
```

No API key, no login, no browser.

***

### The number that matters is how many engines agreed

Every keyword tool gives you a list. This one tells you **which engines suggested each phrase**.

One engine offering a phrase means people type it. Two independent engines offering it means the demand is real rather than an artefact of one ranker. And when **Amazon** is one of them, the people typing it arrived somewhere they could buy — that is a different and more valuable fact than a search volume estimate.

Sort by `sourceCount`, filter on `hasBuyerIntent`, and the top of the list is where search demand and purchase intent overlap.

| Source | What a suggestion from it tells you |
|---|---|
| **Google** · **Bing** · **DuckDuckGo** | what people search for |
| **YouTube** | what people want to *watch* — different phrasing entirely |
| **Amazon** · **eBay** | what people are ready to **buy** |

***

### What you get

**Per keyword** — `keyword`, **`sourceCount`**, `sources`, **`hasBuyerIntent`**, `group`, `questionWord`, `bestPosition`, `positionsBySource`, `wordCount`, `country`, `language`

**Per seed** — keywords found, counts by question / comparison / preposition, **`multiSourceKeywords`**, `buyerIntentKeywords`, per-source coverage, median word count, long-tail count, and any sources that could not serve the country.

***

### Who this is for

- **SEO and content teams** — question keywords are article briefs that already have demand behind them.
- **Amazon and e-commerce sellers** — Amazon autocomplete is the only free window into what shoppers type into a buying box.
- **YouTube creators** — YouTube phrases its suggestions differently from Google. Titling from Google data leaves views on the table.
- **PPC** — comparisons name the competitors people weigh you against, which is where the cheap conquest terms live.

***

### Input

```json
{
  "keywords": ["protein powder", "creatine"],
  "sources": ["google", "youtube", "amazon"],
  "country": "US",
  "includeQuestions": true,
  "includeComparisons": true
}
```

| Option | Default | Notes |
|---|---|---|
| `sources` | all six | google, youtube, amazon, bing, ebay, duckduckgo |
| `country` | `US` | Sets the proxy exit too — see below |
| `language` | follows country | DE→de, JP→ja, BR→pt |
| `includeQuestions` | `true` | how, what, why, when, where, who, which, is, can, do… |
| `includeComparisons` | `true` | vs, versus, or, alternative |
| `includePrepositions` | `true` | for, with, without, near, like, best, cheap |
| `includeAlphabet` | `true` | a–z, the widest net |
| `includeDigits` | `false` | 0–9; useful for model numbers, noise otherwise |
| `maxKeywordsPerSeed` | `1000` | Seeds that hit it are flagged `truncated` |

***

### Five things worth knowing

Each was found by running against live data.

**Autocomplete is personalised on your IP, and the country parameter does not override it.** Asking Google for United States suggestions from a Lithuanian address returned `car insurance lithuania` and `protein powder kaina` — plausible-looking keywords for a market nobody asked about. This Actor therefore points the Apify Proxy at the country you are researching, automatically. If you pin a proxy country yourself, yours wins and the mismatch is called out in the log. **Run without a proxy and your keywords describe wherever the run happened to execute.**

**Amazon is addressed by marketplace hostname, not by marketplace id.** The id alone looks sufficient and is not: `completion.amazon.com` with a German id returns HTTP 200 and an empty list, which reads exactly like a term nobody searches for. `completion.amazon.de` with the same id returns `protein pulver` and `esn protein pulver`. Fifteen marketplaces are wired up; ask for one that Amazon does not operate in and the source is reported as unavailable rather than quietly answered with American keywords.

**The order of expansion is the truncation policy.** Questions, comparisons and prepositions run *before* the alphabet. One live seed reached its keyword cap after 187 of 336 lookups — with the alphabet leading it returned 65 questions, and with questions leading it returned **267 from the same budget**. If a seed truncates you lose long-tail letters, not the phrases you came for.

**An endpoint with nothing to offer hands your query straight back.** Ask Google for `how protein powder` and one of the suggestions is `how protein powder` — a phrase nobody types, arriving as a billable row. Ten of 55 expansion prefixes did this on one seed. Deleting every echo would be too blunt, because `protein powder alternatives` is both an echo of its own prefix and a real query. So a suggestion is kept when **some other prefix also surfaced it**, and dropped when the only way it ever appeared was by asking for it verbatim. Dropped rows are counted as `echoesDropped` and never charged. (`protein powder on` survives this test, and should: ON is Optimum Nutrition, and the `protein powder o` prefix found it independently.)

**Groups are read from the keyword, not from the prefix that found it.** `running shoes vs walking shoes` arrives from the letter "v" and is still a comparison. Matching is on whole words only, so `conversion rate` is not filed as a `vs` comparison.

***

### Output

```json
{
  "recordType": "keyword",
  "seed": "protein powder",
  "keyword": "whey protein powder",
  "sources": ["google", "youtube", "amazon", "bing", "ebay", "duckduckgo"],
  "sourceCount": 6,
  "hasBuyerIntent": true,
  "group": "alphabet",
  "bestPosition": 1,
  "positionsBySource": { "google": 2, "amazon": 1, "bing": 3 },
  "wordCount": 3,
  "country": "US", "language": "en"
}
```

Rows arrive sorted: most engines first, then best rank. Three ready-made views — **Keywords**, **Questions & comparisons**, **Seed summary**. Set `flattenOutput: true` for CSV.

***

### Honest limits

- **No search volume, CPC or difficulty.** Autocomplete says what people type, not how many. `sourceCount` is a demand signal, not a volume estimate, and this Actor will not invent one.
- **Suggestions are live and personalised.** Two runs a week apart will differ, and that is the data being current rather than the Actor being unstable.
- **Ten suggestions per lookup** is the engines' limit. Breadth comes from asking many prefixes, which is exactly what the expansions do.
- **Amazon covers fifteen marketplaces.** Elsewhere it is reported in `sourcesUnavailable`.
- **Seeds must be short phrases.** A pasted URL is rejected rather than expanded into a hundred pointless lookups.
- **Seeds no source can serve are never charged**, nor are query echoes, nor duplicates across sources.

***

### Pricing

| Event | Price | When |
|---|---|---|
| Actor start | $0.005 | Once per run |
| Seed researched | $0.01 | Per seed, covering every lookup across every source |
| Keyword found | $0.0002 | Per unique keyword |

**$0.20 per 1,000 keywords.** A seed returning 1,000 keywords costs about **$0.21**. AnswerThePublic and Keyword Tool start at $89–99 a month.

***

### Tips

- **Sort by `sourceCount`, then filter `hasBuyerIntent`.** That is your commercial shortlist, and it takes one sort in the dataset view.
- **Feed the output back in as seeds.** Point `sourceDatasetId` at a finished run to go a second level deep on the phrases that scored highest.
- **Run the same seed for several countries** to size a market before translating anything. The German and US lists for one seed rarely map onto each other.
- **Use YouTube-only for titles and Google-only for articles.** They phrase the same intent differently, and the gap between the two lists is content nobody has written yet.
- **Compare the question list against your site.** Every question with no matching page is a brief.

# Actor input Schema

## `keywords` (type: `array`):

One per line. Short phrases, not URLs — "protein powder", "crm software", "running shoes". Each seed is expanded into hundreds of real queries.

## `sourceDatasetId` (type: `string`):

Read seeds from an existing dataset instead of typing them — for example the keywords found by a previous run.

## `sourceDatasetField` (type: `string`):

Which field on the source dataset holds the seed.

## `sources` (type: `array`):

Leave empty for all six. Valid values: google, youtube, amazon, bing, ebay, duckduckgo. Google and Bing show what people search; YouTube shows what they want to watch; Amazon and eBay show what they are ready to buy.

## `country` (type: `string`):

ISO code such as US, GB, DE, JP. Autocomplete is personalised on the address the request comes from, so the Apify Proxy is pointed at this country automatically — otherwise a US search run from Europe returns European keywords.

## `language` (type: `string`):

Two-letter code. Left empty it follows the country: DE gives German, JP Japanese, BR Portuguese.

## `includeQuestions` (type: `boolean`):

Prefix the seed with question words. These are the keywords that map directly onto articles and videos people are already looking for.

## `includeComparisons` (type: `boolean`):

Surfaces the competitors and substitutes people weigh against your term.

## `includePrepositions` (type: `boolean`):

Opens up qualified long-tail intent.

## `includeAlphabet` (type: `boolean`):

Appends each letter. The widest net: one seed's alphabet pass returned 258 distinct keywords where the bare seed returned ten.

## `includeDigits` (type: `boolean`):

Appends each digit. Useful for model numbers, sizes and years; noise for most other seeds, so it is off by default.

## `maxKeywordsPerSeed` (type: `integer`):

Stops a broad seed from running away. Seeds that hit it are flagged truncated.

## `includeSeedSummary` (type: `boolean`):

Add one rollup record per seed: counts by question, comparison and preposition, cross-source agreement, buyer-intent totals and per-source coverage. Not billed as a keyword.

## `maxConcurrency` (type: `integer`):

Lookups in flight at once. These endpoints exist to answer a search box on every keystroke, so they tolerate this comfortably.

## `requestTimeoutSecs` (type: `integer`):

Per-lookup timeout. Responses are a few hundred bytes, so this rarely matters.

## `maxRetries` (type: `integer`):

Retries for connection resets and 5xx responses.

## `flattenOutput` (type: `boolean`):

Emit flat dot-notation columns. Mainly affects the sources list and the per-source position map.

## `proxyConfiguration` (type: `object`):

Strongly recommended. Autocomplete is personalised on the exit address, so without a proxy your keywords describe wherever the run happens to execute rather than the country you asked for. The country is set for you to match.

## Actor input object example

```json
{
  "keywords": [
    "running shoes"
  ],
  "sourceDatasetField": "keyword",
  "sources": [
    "google",
    "youtube",
    "amazon"
  ],
  "country": "US",
  "includeQuestions": true,
  "includeComparisons": true,
  "includePrepositions": true,
  "includeAlphabet": true,
  "includeDigits": false,
  "maxKeywordsPerSeed": 1000,
  "includeSeedSummary": true,
  "maxConcurrency": 6,
  "requestTimeoutSecs": 30,
  "maxRetries": 2,
  "flattenOutput": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `results` (type: `string`):

Keyword records carry recordType 'keyword'; rollups carry 'seed-summary'.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": [
        "running shoes"
    ],
    "sources": [
        "google",
        "youtube",
        "amazon"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("eiv/keyword-research-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keywords": ["running shoes"],
    "sources": [
        "google",
        "youtube",
        "amazon",
    ],
}

# Run the Actor and wait for it to finish
run = client.actor("eiv/keyword-research-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": [
    "running shoes"
  ],
  "sources": [
    "google",
    "youtube",
    "amazon"
  ]
}' |
apify call eiv/keyword-research-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,eiv/keyword-research-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/GrVJhJZIp5U1vxlVG/builds/U2T299ByLm1YUeVmX/openapi.json
