# Patents Scraper — Google Patents Search, Assignees, Citations (`chorelet/patents-scraper`) Actor

Search patents worldwide by text, assignee, inventor, country, status and date, and export one row per patent: title, abstract, assignee, inventor, all four dates, PDF link and publication number — with an option to open each patent for its full abstract, CPC classes, legal status and citations.

- **URL**: https://apify.com/chorelet/patents-scraper.md
- **Developed by:** [Chorelet](https://apify.com/chorelet) (community)
- **Categories:** Business, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.70 / 1,000 patents

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Patents Scraper — Google Patents Search, Assignees, Citations

Search patents worldwide the way the website does — full text, assignee, inventor, patent office, granted or application, utility or design, language, and a date range on the priority, filing or publication date — and get one row per patent: publication number, title, snippet, assignee, inventor, all four dates, family size, drawing thumbnail and a direct PDF link.

Turn on **Open each patent page** and every row gains the parts a search result leaves out: the full abstract, every inventor, CPC classifications, legal status, the patents it cites and how many patents cite it.

No API key and no account. One search request returns 100 patents, so a 1,000-patent sweep is ten requests and about twenty seconds.

### What you get

- **A portfolio in one run**: put a company in *Assignees* and a date in *From date* and you have everything it has filed since, sorted newest first
- **Prior-art lists** with the snippet that matched, the PDF link and the family size next to each hit
- **Citations in both directions** with the detail pages on: what a patent cites, and how many later patents cite it — the usual proxy for how important it turned out to be
- **CPC classifications** as a clean list of full subgroups (`A47K3/022`), not the hierarchy above them
- **Legal status** as the registry words it — *Active*, *Expired - Fee Related*, *Withdrawn*
- Filters that actually narrow: patent office, granted versus application, design patents, language, and which of the three dates the range applies to

Google serves at most 1,000 results for a single query, so a wide sweep is best split by year, country or assignee — the run says so in the log when it hits that ceiling.

# Actor input Schema

## `queries` (type: `array`):

Full-text queries, one run per line: `solid state battery`, `mRNA vaccine delivery`. Operators the website understands work here too, for example `(battery AND anode) NOT lithium`.

## `assignees` (type: `array`):

The company or institution that owns the patent: `Toyota`, `Samsung Electronics`, `MIT`.

## `inventors` (type: `array`):

Inventor names as they appear on the patent.

## `countries` (type: `array`):

Two-letter codes: `US`, `EP`, `WO`, `CN`, `JP`, `DE`, `KR`. Empty = everywhere.

## `status` (type: `string`):

Granted patents or published applications.

## `type` (type: `string`):

Utility patents or design patents.

## `language` (type: `string`):

`ENGLISH`, `GERMAN`, `CHINESE`… Empty = any.

## `dateField` (type: `string`):

The priority date is when the invention was first claimed, the filing date when the application was submitted, the publication date when it became public.

## `after` (type: `string`):

`2024-01-01`. Leave empty for no lower bound.

## `before` (type: `string`):

`2026-12-31`. Leave empty for no upper bound.

## `sortBy` (type: `string`):

Relevance is Google's own ranking; newest first sorts by publication date.

## `includeDetails` (type: `boolean`):

Adds the full abstract, every inventor, CPC classifications, legal status, the patents it cites and how many cite it. One page request per patent.

## `concurrency` (type: `integer`):

How many patent pages to open at once when detail pages are on.

## `maxResults` (type: `integer`):

Google serves at most 1,000 results for one query — narrow by date, country or assignee to go deeper.

## Actor input object example

```json
{
  "queries": [
    "solid state battery"
  ],
  "assignees": [],
  "inventors": [],
  "countries": [],
  "status": "",
  "type": "",
  "dateField": "priority",
  "sortBy": "relevance",
  "includeDetails": false,
  "concurrency": 4,
  "maxResults": 100
}
```

# Actor output Schema

## `patents` (type: `string`):

Everything scraped — items of the default dataset. Use ?format=csv or xlsx on this URL for spreadsheets.

## `summary` (type: `string`):

Patents per search, how many matched, detail pages opened and errors.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "solid state battery"
    ],
    "assignees": [],
    "inventors": [],
    "countries": [],
    "status": "",
    "type": "",
    "language": "",
    "dateField": "priority",
    "after": "",
    "before": "",
    "sortBy": "relevance",
    "includeDetails": false,
    "concurrency": 4,
    "maxResults": 100
};

// Run the Actor and wait for it to finish
const run = await client.actor("chorelet/patents-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": ["solid state battery"],
    "assignees": [],
    "inventors": [],
    "countries": [],
    "status": "",
    "type": "",
    "language": "",
    "dateField": "priority",
    "after": "",
    "before": "",
    "sortBy": "relevance",
    "includeDetails": False,
    "concurrency": 4,
    "maxResults": 100,
}

# Run the Actor and wait for it to finish
run = client.actor("chorelet/patents-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "solid state battery"
  ],
  "assignees": [],
  "inventors": [],
  "countries": [],
  "status": "",
  "type": "",
  "language": "",
  "dateField": "priority",
  "after": "",
  "before": "",
  "sortBy": "relevance",
  "includeDetails": false,
  "concurrency": 4,
  "maxResults": 100
}' |
apify call chorelet/patents-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,chorelet/patents-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/4OMVphWtuBPHTTbXk/builds/MTR05pSYJlrafaiwW/openapi.json
