# Google Patents Scraper (`confidential_gnat/google-patents-scraper`) Actor

Scrapes patent data from Google Patents (patents.google.com) by keyword or URL, including title, abstract, inventors, assignee, filing and publication dates, citations, figures and PDF links.

- **URL**: https://apify.com/confidential\_gnat/google-patents-scraper.md
- **Developed by:** [ActorFlow](https://apify.com/confidential_gnat) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Google Patents Scraper

**Scrape patent data from Google Patents (patents.google.com)** by keyword or URL. This **Google Patents scraper** extracts the title, abstract, inventors, assignee, filing and publication dates, application number, cited patents, figure images and PDF links for every matching patent, and exports them to JSON, CSV or Excel — or straight into your code through the **Apify API**. Paste a search URL or type a keyword, press Start, and the results land in a structured dataset.

**Target website:** [patents.google.com](https://patents.google.com)

### ✨ Features of this Google Patents scraper

- **Patent data extraction** — title, abstract, inventors, assignee, filing date, publication date, application number, citations, figures and the patent PDF
- **Keyword and URL input** — search by keyword, or paste any Google Patents search URL and keep its filters (country, date range, status, inventor, assignee)
- **Pagination support** — walks through result pages automatically until your item limit is reached
- **Proxy support** — optional, and switched off by default
- **Language selection** — fetch the machine-translated patent page in any of 10 languages
- **Repeat-run caching** — name a cache project and later runs skip patents already scraped
- **No browser required** — runs on plain HTTP requests, which makes it fast and cheap

### 🚀 How to scrape Google Patents in 5 steps

1. [Sign up](https://apify.com/sign-up) for a free Apify account — includes **$5 monthly credit**.
2. Open the actor page and click **Try for free**.
3. Fill in the **Input** fields — a search keyword or a start URL is required.
4. Click **Start** and wait for the run to complete.
5. Download results from the **Output** tab in JSON, CSV, or Excel format.

You can also run this actor via the [Apify API](https://docs.apify.com/api/v2) or integrate it directly into your workflows using [Zapier](https://zapier.com/apps/apify), [Make](https://www.make.com/), or [n8n](https://n8n.io/).

### 💰 Pricing

This actor uses **pay-per-result** billing based on the compute units a run consumes.

- New Apify accounts include **$5 of free monthly credit**.
- It runs on plain HTTP requests rather than a headless browser, so it costs significantly less to run than browser-based patent scrapers.
- Proxies are disabled by default, which keeps runs at their cheapest.

### 🔧 Input configuration

| Field                | Type    | Required | Default                    | Description                                                                                                                         |
| -------------------- | ------- | -------- | -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
| `searchQueries`      | array   | —        | `["solid state battery"]`  | Keywords to search for, one search per entry. Each keyword is searched independently and the item limit applies to each separately. |
| `startUrls`          | array   | —        | —                          | Google Patents URLs. Accepts search URLs and single patent URLs; the page type is detected automatically.                           |
| `maxItems`           | integer | —        | `5`                        | Maximum patents to scrape **per keyword or start URL**. Set to `0` for no limit.                                                    |
| `language`           | string  | —        | `en`                       | Language of the patent page to fetch. One of `en`, `de`, `fr`, `es`, `it`, `ja`, `ko`, `zh`, `pt`, `ru`.                            |
| `cacheProjectName`   | string  | —        | —                          | Name a project to remember scraped patents across runs so repeat runs skip them.                                                    |
| `proxyConfiguration` | object  | —        | `{"useApifyProxy": false}` | Proxy settings. Off by default.                                                                                                     |

**Supported URL types:**

- Search results — `https://patents.google.com/?q=battery`
- Search with filters — `https://patents.google.com/?q=battery&country=US&after=priority:20200101`
- Single patent — `https://patents.google.com/patent/US9739567B2/en`

### 📦 Google Patents scraper output data

Each result is a JSON object with the keys `url`, `publicationNumber`, `title`, `abstract`, `inventors`, `assignee`, `filingDate`, `publicationDate`, `applicationNumber`, `citations`, `pdfUrl`, `thumbnail`, `figures` and `searchSnippet`. The dataset ships with two views: **Overview**, a compact table of publication number, title, assignee, publication date and URL; and **Full patent details**, which adds the abstract, inventors, application number, citations and figure images.

**Sample output:**

```json
[
    {
        "url": "https://patents.google.com/patent/US11942620B2/en",
        "publicationNumber": "US11942620B2",
        "title": "Solid state battery with uniformly distributed electrolyte, and methods of fabrication relating thereto",
        "abstract": "The present disclosure relates to a solid-state electrochemical cell having a uniformly distributed solid-state electrolyte and methods of fabrication relating thereto. The method may include forming a plurality of apertures within the one or more solid-state electrodes; impregnating the one or more solid-state electrodes with a solid-state electrolyte precursor solution so as to fill the plurality of apertures and any other void or pores within the one or more electrodes with the solid-state electrolyte precursor solution; and heating the one or more electrodes so as to solidify the solid-state electrolyte precursor solution and to form the distributed solid-state electrolyte.",
        "inventors": ["Yong Lu", "Zhe Li", "Xiaochao Que", "Haijing Liu", "Meiyuan Wu"],
        "assignee": "GM Global Technology Operations LLC",
        "filingDate": "2021-12-06",
        "publicationDate": "2024-03-26",
        "applicationNumber": "US:17/543,160",
        "citations": ["JP:H1021963:A", "JP:2000090979:A"],
        "pdfUrl": "https://patentimages.storage.googleapis.com/f8/02/22/a9e3950cab3340/US11942620.pdf",
        "thumbnail": "https://patentimages.storage.googleapis.com/90/b8/0e/ed6c106ab89ebf/US11942620-20240326-D00000.png",
        "figures": [
            "https://patentimages.storage.googleapis.com/23/4f/5d/37847c3cb9dfb3/US11942620-20240326-D00000.png",
            "https://patentimages.storage.googleapis.com/78/9e/7c/171f5b5c46f975/US11942620-20240326-D00001.png"
        ],
        "searchSnippet": "In the instances of solid-state batteries, which include solid-state electrolyte layers disposed between solid-state electrodes, the solid-state electrolyte layer physically separates the solid-state electrodes so that a distinct separator is not required. Solid-state batteries have advantages over \u2026"
    },
    {
        "url": "https://patents.google.com/patent/US11239459B2/en",
        "publicationNumber": "US11239459B2",
        "title": "Low-expansion composite electrodes for all-solid-state batteries",
        "abstract": "A composite electrode for use in an all-solid-state electrochemical cell that cycles lithium ions is provided. The composite electrode comprises a solid-state electroactive material that undergoes volumetric expansion and contraction during cycling of the electrochemical cell and a solid-state electrolyte. The solid-state electroactive material is in the form of a plurality of particles and each particle has a plurality of internal pores formed therewithin. Each particle has an average porosity ranging from about 10% to about 75%, and the composite electrode has an interparticle porosity between the solid-state electroactive material and solid-state electrolyte particles ranging from about 5% to about 40%. The intraparticle pores and the interparticle porosity accommodate the volumetric expansion and contraction of the solid-state electroactive material so to minimize outward expansion of the electroactive particles, micro-cracking of the solid-state electrolyte, and delamination within the electrochemical cell.",
        "inventors": ["Thomas A. Yersak", "Mei Cai"],
        "assignee": "GM Global Technology Operations LLC",
        "filingDate": "2018-10-18",
        "publicationDate": "2022-02-01",
        "applicationNumber": "US:16/164,525",
        "citations": ["US:20160017266:A1", "US:20120100438:A1"],
        "pdfUrl": "https://patentimages.storage.googleapis.com/9b/93/8b/6a8616b3676c91/US11239459.pdf",
        "thumbnail": "https://patentimages.storage.googleapis.com/31/2c/35/2f94d4ccf0b302/US11239459-20220201-D00000.png",
        "figures": [
            "https://patentimages.storage.googleapis.com/63/96/43/46838a8256d58f/US11239459-20220201-D00000.png",
            "https://patentimages.storage.googleapis.com/79/cb/33/17d1ca13183ae1/US11239459-20220201-D00001.png"
        ],
        "searchSnippet": "The drawings described herein are for illustrative purposes only of selected embodiments and not all possible implementations, and are not intended to limit the scope of the present disclosure. FIG. 1 is a schematic of an example of an all-solid-state electrochemical battery cell; FIG. 2A is an \u2026"
    }
]
```

### 🐍 How to scrape Google Patents with Python, JavaScript or the API

Run the actor programmatically with the official Apify clients. Replace `<YOUR_API_TOKEN>` with the token from your [Apify Console](https://console.apify.com/account/integrations).

**Python** (`pip install apify-client`):

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_API_TOKEN>")

run = client.actor("<username>/google-patents-scraper").call(run_input={
    "searchQueries": ["solid state battery"],
    "maxItems": 5,
    "language": "en",
})

for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)
```

**JavaScript** (`npm install apify-client`):

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: '<YOUR_API_TOKEN>' });

const run = await client.actor('<username>/google-patents-scraper').call({
    searchQueries: ['solid state battery'],
    maxItems: 5,
    language: 'en',
});

const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);
```

**cURL** — start a run and wait for the dataset:

```bash
curl -X POST "https://api.apify.com/v2/acts/<username>~google-patents-scraper/run-sync-get-dataset-items?token=<YOUR_API_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"searchQueries": ["solid state battery"], "maxItems": 5, "language": "en"}'
```

### 💡 What you can use Google Patents data for

- **Prior-art searches** — find existing patents before filing to assess novelty
- **Competitive intelligence** — track what a named assignee is filing and when
- **Technology landscaping** — map how a field evolves by filing and publication dates
- **Freedom-to-operate analysis** — collect the patents that might read on a product
- **Citation network analysis** — follow which patents cite which to find foundational work
- **R\&D trend monitoring** — watch filing volume in a technology area over time

Patent attorneys, IP analysts, R\&D teams, technology-transfer offices and competitive-intelligence researchers use this data across pharmaceuticals, electronics, automotive, materials science and software.

### ⚠️ Limitations & known issues

- **Machine-translated text** — patents filed in other languages are served through Google's own translation; the original wording may differ.
- **Abstract length** — the abstract comes from the page's metadata and may be truncated for very long patents.
- **Rate limiting** — very high-volume runs may be throttled; enable proxies if you hit limits.
- **Full claims and description** — this actor collects the abstract and bibliographic data, not the complete claims text.

### ❓ Frequently asked questions

#### Can I scrape Google Patents legally?

This actor only collects data that is already publicly visible on Google Patents — no login, paywall, or private content is accessed. Scraping publicly available data is generally considered lawful (see *hiQ Labs v. LinkedIn* as precedent), and patent documents are public records by design. You remain responsible for complying with Google's Terms of Service and any applicable laws in your jurisdiction.

#### Does this scraper get the full patent claims text?

No. It extracts the bibliographic data and the abstract — title, inventors, assignee, dates, application number, citations, figures and the PDF link. The full claims and description are available in the linked PDF.

#### How many patents can I scrape from one search?

As many as the search returns. `maxItems` limits results **per keyword or start URL**, so three keywords with `maxItems: 100` gives up to 300 patents. Set it to `0` to remove the limit entirely.

#### Do I need a proxy to scrape patents.google.com?

No. Google Patents responds reliably without one, which is why proxies are off by default. Enable them only if you run at high volume and start seeing throttling.

#### How do I scrape Google Patents with Python?

Install `apify-client`, then call the actor with your search keywords and iterate the dataset — see the Python example above. The run returns structured JSON you can load straight into pandas.

#### Can I run this Google Patents scraper on a schedule?

Yes. Use [Apify Schedules](https://docs.apify.com/platform/schedules) to run it hourly, daily or weekly. Set `cacheProjectName` so each scheduled run skips patents already collected and only returns newly published ones.

#### What output formats are supported?

JSON, CSV, Excel, XML and RSS, either from the **Output** tab or through the Apify API.

### 🔗 Other actors you may find useful

- 📰 [Google News AI Scraper](https://apify.com/confidential_gnat/google-news-ai-scraper) — scrape Google News articles and summarize them with AI
- 🏠 [Uniacco Property Scraper](https://apify.com/confidential_gnat/uniacco-property-scraper) — extract student accommodation listings from Uniacco
- 🏘️ [HousingAnywhere Property Scraper](https://apify.com/confidential_gnat/housinganywhere-property-scraper) — extract rental listings from HousingAnywhere
- 🛒 [Woolworths Products Scraper](https://apify.com/confidential_gnat/woolworths-products-scraper) — scrape product listings and prices from Woolworths
- 🛏️ [Spotahome Property Scraper](https://apify.com/confidential_gnat/spotahome-property-scraper) — extract mid- to long-term rental listings from Spotahome

### 💬 Support & contact

If you encounter any issues or have questions, please [open an issue](https://github.com/) or reach out via the [Apify Community Forum](https://community.apify.com/).

You can also find more of our actors on the [Apify Store](https://apify.com/store).

# Actor input Schema

## `searchQueries` (type: `array`):

Keywords to search Google Patents for, one search per entry. Each keyword is searched independently and the Max items limit applies to each one separately.

## `startUrls` (type: `array`):

Google Patents URLs to scrape. Accepts search URLs (https://patents.google.com/?q=battery) and single patent URLs (https://patents.google.com/patent/US9739567B2/en). Page type is detected automatically.

## `maxItems` (type: `integer`):

Maximum number of patents to scrape per search keyword or start URL. Set to 0 for no limit.

## `language` (type: `string`):

Language version of the patent page to fetch. Google Patents machine-translates non-English patents into this language.

## `cacheProjectName` (type: `string`):

Optional. Name a project to remember which patents have already been scraped across runs, so repeated runs skip them. Leave empty to scrape everything every time.

## `proxyConfiguration` (type: `object`):

Proxy settings. Google Patents is reachable without a proxy, so proxies are disabled by default to keep runs cheap. Enable them if you hit rate limits at high volume.

## Actor input object example

```json
{
  "searchQueries": [
    "solid state battery"
  ],
  "startUrls": [],
  "maxItems": 5,
  "language": "en",
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQueries": [
        "solid state battery"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("confidential_gnat/google-patents-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "searchQueries": ["solid state battery"] }

# Run the Actor and wait for it to finish
run = client.actor("confidential_gnat/google-patents-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQueries": [
    "solid state battery"
  ]
}' |
apify call confidential_gnat/google-patents-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,confidential_gnat/google-patents-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/FA1CJPrxh28VNWUZ7/builds/gx0SoTDES3CBR6FNF/openapi.json
