# Google Patents Scraper — Claims & Citations (`zenomastro/google-patents-intelligence`) Actor

Google Patents scraper for prior-art and IP intelligence. Search or look up patents; filter assignee, inventor, CPC, country, date and status; extract claims, citations, legal events and PDFs; build family/citation graphs and free landscape summaries.

- **URL**: https://apify.com/zenomastro/google-patents-intelligence.md
- **Developed by:** [Rosario Vitale](https://apify.com/zenomastro) (community)
- **Categories:** Business, AI, Automation
- **Stats:** 1 total users, 0 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 patent results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Google Patents Scraper API — Claims, Landscape & Graphs

### Why use this Actor?

Google Patents scraper for prior-art and IP intelligence. Search or look up patents; filter assignee, inventor, CPC, country, date and status; extract claims, citations, legal events and PDFs; build family/citation graphs and free landscape summaries.

### Features

- **Google Patents search expressions** — Search terms or Google Patents query expressions.
- **Maximum results per query** — Maximum results per query. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Maximum result pages per query** — Maximum result pages per query. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Optional country filter** — Optional country filter. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Optional language filter** — Optional language filter. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Enrich patent detail pages** — Adds abstract, detail title, inventors/assignees, application/patent numbers, PDF and references when exposed.
- **Maximum detail pages per query** — Maximum detail pages per query. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Detail request concurrency** — Detail request concurrency. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Persistent monitor key** — Persistent monitor key. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Emit only newly observed publications** — Emit only newly observed publications. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Request timeout** — Request timeout. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Retries** — Retries. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

### Use cases

- Prior-art research.
- Competitor ip monitoring.
- Assignee and inventor research.
- Patent intelligence datasets.

### Example input

```json
{
  "maxResultsPerQuery": 250,
  "maxPagesPerQuery": 10,
  "includeDetails": false,
  "detailsLimitPerQuery": 50,
  "detailConcurrency": 5,
  "onlyNew": false
}
```

### Pricing & cost control

Use the bounded input limits and filters to keep runs predictable. Pay-per-result Actors only charge primary result rows; summary, status and monitoring metadata are designed to add context without inflating result volume.

### FAQ

**What is this Actor for?**\
It is designed for prior-art research, competitor IP monitoring, assignee and inventor research.

**Can I run it on a schedule?**\
Yes. You can schedule Actor runs on Apify and send the resulting dataset into automations, webhooks, storage, or downstream APIs.

**How do I control cost and run size?**\
Use the input limits and filters shown in the Actor input form. The Actor applies bounded defaults and hard caps so large jobs remain predictable.

### Search keywords

google patents scraper, google patent scraper github, google patents scraping, what is google patents, how to use google patents, patent search api, patent search api key, patent office api, uspto patent search api, google patent search api, espacenet patent search api, wipo patent search api, patent public search api, epo patent search api

Search Google Patents through the same public XHR result surface used by the web application and export structured research datasets without a browser, account, API key or paid patent provider.

### Search output

Each patent result can include:

- Publication number and Google Patents URL.
- Clean title and search-result snippet.
- Priority, filing, grant and publication dates.
- Inventor and assignee.
- Language.
- Direct PDF and thumbnail URLs when Google exposes them.
- Figure metadata.
- Family country/status metadata.
- Search rank.
- Optional persistent `isNew` state.

### Optional deep detail enrichment

Enable `includeDetails` to fetch a bounded number of patent detail pages concurrently. The Actor extracts public meta fields including abstract/description, detail title, inventors, assignees, filing/issue dates, application and patent numbers, PDF URL, and references/citations exposed in the page metadata.

Unlike products that force a full expensive detail crawl, search-only mode stays lightweight. Deep enrichment is explicitly bounded with `detailsLimitPerQuery` and `detailConcurrency`.

### Monitoring

Use a stable `monitorKey` in scheduled runs to keep a bounded set of publication numbers already observed. `onlyNew=true` then turns a saved query into a patent-watch feed.

### Example

```json
{
  "queries":["artificial intelligence controller","solid state battery"],
  "maxResultsPerQuery":250,
  "includeDetails":true,
  "detailsLimitPerQuery":50,
  "monitorKey":"research-radar"
}
```

### Reliability

The Actor uses the public Google Patents XHR JSON endpoint for result pages instead of scraping visual search-result markup. Duplicate publications across queries are removed. Detail-page failures do not destroy search results. Inputs, pages, detail requests, concurrency, retries and timeouts are all bounded.

Google Patents is a public research interface rather than a guaranteed commercial API, so response-shape changes are possible. Deployment includes a live cloud smoke test against the current XHR endpoint.

Patent metadata is informational. Verify critical legal or bibliographic facts against official patent-office records.

### Extended capabilities

- Search or directly look up patent publications with assignee, inventor, CPC, date, status, and locale filters.
- Enrich bounded results with claims, classifications, citations, legal events, PDFs, figures, and family metadata.
- Optionally emit citation and patent-family graphs with bounded node counts.

# Actor input Schema

## `queries` (type: `array`):

Search terms or Google Patents query expressions.

## `maxResultsPerQuery` (type: `integer`):

Maximum results per query. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `maxPagesPerQuery` (type: `integer`):

Maximum result pages per query. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `country` (type: `string`):

Optional country filter. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `language` (type: `string`):

Optional language filter. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `includeDetails` (type: `boolean`):

Adds abstract, detail title, inventors/assignees, application/patent numbers, PDF and references when exposed.

## `detailsLimitPerQuery` (type: `integer`):

Maximum detail pages per query. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `detailConcurrency` (type: `integer`):

Detail request concurrency. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `monitorKey` (type: `string`):

Persistent monitor key. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `onlyNew` (type: `boolean`):

Emit only newly observed publications. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `requestTimeoutSecs` (type: `integer`):

Request timeout. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `retries` (type: `integer`):

Retries. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `publicationNumbers` (type: `array`):

Optional direct patent lookups such as US20240256345A1 or a patents.google.com/patent/... URL.

## `includeClaims` (type: `boolean`):

Extract bounded public claim text from patent detail pages. Enable only when needed because claims can make records large.

## `claimsLimit` (type: `integer`):

Maximum claim blocks kept per enriched or direct patent detail.

## `assignee` (type: `string`):

Optional Google Patents assignee filter, e.g. Google LLC.

## `inventor` (type: `string`):

Optional inventor name filter.

## `cpc` (type: `string`):

Optional CPC code filter such as G06F.

## `after` (type: `string`):

Optional Google Patents date lower bound, e.g. 2024-01-01.

## `before` (type: `string`):

Optional Google Patents date upper bound, e.g. 2026-01-01.

## `legalStatus` (type: `string`):

Optional Google Patents legal-status filter when supported by the current search surface.

## `sort` (type: `string`):

Requested Google Patents ordering.

## `includeGraph` (type: `boolean`):

Emit a free patent\_graph row per search query with patent nodes, directed citation edges and family links derived from enriched metadata.

## `graphMaxNodes` (type: `integer`):

Bound graph size for predictable memory and output volume.

## `includeLandscapeSummary` (type: `boolean`):

Emit a free per-query intelligence row with top assignees, inventors, CPC codes, publication-year mix, legal-status distribution and citation counts.

## `landscapeTopN` (type: `integer`):

Maximum ranked assignee, inventor, CPC and legal-status values included in each landscape summary.

## Actor input object example

```json
{
  "queries": [],
  "maxResultsPerQuery": 250,
  "maxPagesPerQuery": 10,
  "includeDetails": false,
  "detailsLimitPerQuery": 50,
  "detailConcurrency": 5,
  "onlyNew": false,
  "requestTimeoutSecs": 30,
  "retries": 2,
  "publicationNumbers": [],
  "includeClaims": false,
  "claimsLimit": 20,
  "sort": "relevance",
  "includeGraph": false,
  "graphMaxNodes": 500,
  "includeLandscapeSummary": true,
  "landscapeTopN": 20
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("zenomastro/google-patents-intelligence").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("zenomastro/google-patents-intelligence").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call zenomastro/google-patents-intelligence --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,zenomastro/google-patents-intelligence"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/sZLvpHdDC0cLI5zRT/builds/dGXDIvX1mWQXBwgfc/openapi.json
