# Terms of Service Clause Finder - Scraping, AI & Arbitration (`neverempty/terms-of-service-clause-finder`) Actor

Before an AI agent or scraper uses a website: finds its terms of service and quotes, word for word with section numbers, the clauses on scraping and bots, circumventing limits, AI training and agents, resale, liquidated damages and arbitration. Amazon: 17 clauses, all verbatim (2026-09-25).

- **URL**: https://apify.com/neverempty/terms-of-service-clause-finder.md
- **Developed by:** [NeverEmpty](https://apify.com/neverempty) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $8.40 / 1,000 terms analyzeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Terms of Service Clause Finder - Scraping, AI & Arbitration

**For AI agents and scrapers that need to know what a website's terms say before they use it:** give a list of domains, get back each site's terms of service page and the clauses on **scraping and bots**, **circumventing rate limits or security**, **AI training and AI agents**, **building databases, resale and commercial use**, **liquidated damages**, and **arbitration and governing law** - quoted word for word, with the section number, the heading above it, the page URL and the terms' last-updated date.

> **This is not legal advice.** The Actor finds and quotes clauses by their wording; it never says whether something is allowed or forbidden. A category with no quoted clause does **not** mean the use is permitted. Always read the full terms.

### Input and output in one look

Input:

```json
{
  "domains": ["amazon.com", "www.importyeti.com", "https://themen.kleinanzeigen.de/nutzungsbedingungen/"],
  "categories": ["scraping-and-automated-access", "ai-training-and-ai-agents"]
}
```

One row per domain (real row from amazon.com, 2026-09-25, two of its 17 clauses shown):

```json
{
  "status": "ok",
  "domain": "amazon.com",
  "termsUrl": "https://www.amazon.com/gp/help/customer/display.html?nodeId=GLSBYFE9MGKKQXXM&ref_=footer_cou",
  "termsFoundVia": "homepage-link",
  "termsPageTitle": "Conditions of Use - Amazon Customer Service",
  "termsLastUpdated": "2026-08-14",
  "termsLastUpdatedText": "Last updated: August 14, 2026",
  "clauseCount": 17,
  "categoriesFound": ["scraping-and-automated-access", "circumventing-access-controls-or-rate-limits", "ai-training-and-ai-agents", "database-building-redistribution-commercial-use", "arbitration-and-governing-law"],
  "clauses": [
    {
      "categories": ["ai-training-and-ai-agents"],
      "section": null,
      "heading": "Agents",
      "text": "Transparency and Consent. No Agent may access, use, or interact with Amazon Services unless, at all times, it identifies itself and operates in strict accordance with the requirements in section 3 of these Agent Terms. ...",
      "matchedPhrases": { "ai-training-and-ai-agents": ["No Agent may", "Agent Terms"] }
    },
    {
      "categories": ["scraping-and-automated-access", "database-building-redistribution-commercial-use"],
      "heading": "LICENSE AND ACCESS",
      "text": "... This license does not include any resale or commercial use of any Amazon Service, or its contents; ... or any use of data mining, robots, or similar data gathering and extraction tools. ...",
      "matchedPhrases": { "scraping-and-automated-access": ["data mining", "robots"], "database-building-redistribution-commercial-use": ["commercial use"] }
    }
  ],
  "note": "Quoted clauses are matched by phrases; this is not legal advice and says nothing about what is allowed. Read the full terms."
}
```

### What it looks for

| Category (`categories`) | Examples of the wording it matches |
|---|---|
| `scraping-and-automated-access` | scrape, crawl, spider, robot, bots, data mining, harvest, "automated means"; Scraping, Crawler, automatisierte Abfragen; aspiration, extraction automatisée; arañas, medios automatizados; スクレイピング, クローラ, ロボット, 自動化された手段 |
| `circumventing-access-controls-or-rate-limits` | circumvent / bypass / avoid security or technical measures, rate limits, CAPTCHAs, IP blocks, unreasonable load; umgehen, übermäßige Last; contourner; eludir; アクセス制御の回避, 過度な負荷 |
| `ai-training-and-ai-agents` | training AI or machine learning models, training data, text and data mining (incl. § 44b UrhG), AI agents, "Agent" terms; 人工知能の学習, 情報解析 |
| `database-building-redistribution-commercial-use` | create a database or compilation, resell, redistribute, republish, frame or mirror, commercial use; Datenbank, gewerblich; base de données, fins commerciales; base de datos, fines comerciales; データベース, 商用, 転載, 再配布 |
| `liquidated-damages-and-penalties` | liquidated damages, a fixed amount per violation or per page (for example "$15,000 USD per 1,000,000 posts"), Vertragsstrafe, clause pénale, cláusula penal, 違約金 |
| `arbitration-and-governing-law` | arbitration, class action waiver, governing law, exclusive jurisdiction or venue; Gerichtsstand, anwendbares Recht; droit applicable; ley aplicable, jurisdicción; 準拠法, 専属的合意管轄 |

A paragraph under a heading such as "Governing Law", "Arbitration", "Agents", "Gerichtsstand" or "準拠法" is also returned in that category even if its own sentence does not repeat the word (`matchedPhrases` then says `heading: …`).

Wording that looks similar but is about something else is **not** returned. These were found in real terms and are checked by the Actor's tests: "automated text messages" and autodialer consent, "automated decision-making", "eBay's automated systems scan messages" (the site's own systems), "we use automated tools to translate", "commercially reasonable efforts", "under penalty of perjury", a site describing its own chatbot ("Due to the nature of Generative AI…"), a company describing itself ("an AI safety and research company"), and force majeure ("回避することのできない災厄").

### How it finds the terms page

1. **You gave a full URL with a path** - that page is read as the terms (`termsFoundVia: input-url`).
2. **You gave a domain** - the home page is read and its links are scored by their text and address in English, German, French, Spanish, Italian, Portuguese, Dutch and Japanese ("Terms of Service", "Conditions of Use", "User Agreement", "AGB", "Nutzungsbedingungen", "CGU", "Conditions générales", "Términos y condiciones", "Aviso legal", "利用規約" …). Links for other audiences (API, sellers, advertisers, merchants) score lower and are listed in `otherTermsLinks` instead.
3. If the best link is a **legal hub page** ("Legal", "Policies", a footer link to /policies), the Actor follows one more step to the terms on that page (`legal-page-link`).
4. If nothing is linked, it tries the **usual addresses** - /terms, /terms-of-service, /terms-of-use, /legal/terms, /tos, /legal … and, by country or page language, /agb and /nutzungsbedingungen (de), /cgu (fr), /aviso-legal (es), /kiyaku and /rule (ja) (`common-path`). If the home page redirected to another host (spotify.com → open.spotify.com), the original domain is tried too.
5. **Terms that live in a OneTrust notice** (the page is an empty shell and the text is in a JSON file on privacyportal-cdn.onetrust.com) are read from that JSON (`termsFoundVia` ends with `+onetrust`).

Each page is accepted as terms only if its title, headings or address name terms, conditions or an agreement (in any of the languages above), it has at least 1,200 characters of text and reads like a contract. A privacy or cookie policy is **not** treated as terms, even when you give its URL directly. A domain whose terms are not found is **not charged**.

### What it will not do

- **Robots.txt is respected for every page it reads**, including the terms page itself (RFC 9309: a robots.txt that cannot be read because of a server error is treated as "do not read"). A disallowed page comes back as a free `robots-disallowed` row. Sites whose robots.txt disallows everything for all bots (for example x.com, linkedin.com, reddit.com) cannot be read this way - give the terms to your agent another way.
- **No check pages are solved or bypassed**, no proxies are used to get around a refusal, and nothing that needs a sign-in is read. Those domains come back as free `blocked` or `login-required` rows.
- **No JavaScript is run.** When the terms page is found but its text is loaded later by JavaScript, the row says `terms-text-not-readable` (free) - it does **not** claim the terms have no such clauses.
- It never guesses: a date that is not on the page is `null`, a section number that is not printed is `null`.

Real result (production run on 2026-09-25, 40 well-known domains in 8 languages, 256 MB, about 12 to 100 seconds): terms read for **20 domains with 218 quoted clauses**; 14 domains refused this reader or showed a check page (for example etsy.com, booking.com, openai.com, nytimes.com); 5 disallow bots in robots.txt; 1 resolved to a private address. Browser check of the quoted text: GitHub 12/12, Amazon 17/17, Airbnb 25/25, Spotify 14/14, Kleinanzeigen 17/17 and Mercari 2/2 quotes found word for word on the live page, with the same last-updated date.

### Output fields

| Field | Meaning |
|---|---|
| `status` | `ok` (terms read and returned; charged), or a free reason: `terms-not-found`, `terms-text-not-readable`, `robots-disallowed`, `robots-unreachable`, `blocked`, `login-required`, `not-found`, `http-error`, `unreachable`, `unreadable`, `not-html`, `bad-input`, `budget-reached`, `no-change` |
| `input`, `domain`, `position` | What you gave and its place in the list |
| `termsUrl` | The terms page that was read (after redirects) |
| `termsFoundVia` | `input-url`, `input-url-link`, `homepage-link`, `legal-page-link` or `common-path` (+`+onetrust`) |
| `termsPageTitle`, `termsLanguage` | Page title and the `<html lang>` value |
| `termsLastUpdated`, `termsLastUpdatedText` | ISO date and the sentence it came from ("Last updated", "Effective date", "Stand", "gelten ab", "Dernière mise à jour", "Última actualización", "最終更新日", "改定" …); null when the page states no date |
| `termsTextLength`, `termsTruncated` | Characters of text read; true if the page was larger than 3 MB and read only up to that size |
| `clauseCount`, `categoriesFound`, `clauseCountByCategory` | How many clauses were quoted, in which categories (up to 12 per category) |
| `clauses` | Array of `{ categories, section, heading, text, matchedPhrases, truncated }`. `text` is the paragraph word for word; paragraphs longer than 1,200 characters are cut to the matching sentences with "…" and `truncated: true` |
| `otherTermsLinks` | Other legal links found on the way (for example API terms, consumer vs commercial terms) as `{ text, url }` - pass one of them as a URL to read it |
| `changeType`, `changedClauses`, `changedCategories`, `previousCheckedAt`, `termsUrlChanged`, `previousTermsUrl`, `watchName` | Monitor mode (below) |
| `httpStatus`, `note`, `fetchedAt` | Last HTTP status, a plain-English note, time of reading |

### Monitor mode: only terms that changed

Set `watchName` (and `onlyChanges: true` to receive only changed domains). The Actor remembers each domain's terms text paragraph by paragraph. On the next run it returns only domains whose text changed, with `changedClauses` as `{ change: "added" | "removed" | "modified", categories, before, after }`, so an agent can see at once whether a scraping, AI or arbitration clause was added.

- **A new "last updated" date or copyright year alone is not a change** - dates are masked before comparing.
- Menus and lists of links are not compared.
- A change is read a second time a moment later and reported only if both reads show exactly the same changes; otherwise nothing is reported or remembered that time, and the next check compares again.
- If the terms are found on a different page than last time, that check only records the new page; changes are reported from the next check.
- The first run of a watch returns every domain as `first-check`. A run where nothing changed returns one free `no-change` row and charges only the run start fee.

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| `domains` | array of strings | example `github.com` (read free of charge) | Domains or full terms URLs, 1 to 500 per run. If empty, the example domain github.com is read and nothing is charged |
| `categories` | array | all six | Which clause categories to return |
| `onlyChanges` | boolean | false | Monitor mode: return only domains whose terms changed |
| `watchName` | string | - | Name of the remembered terms (letters, digits, `.`, `-`, `_`; up to 40) |
| `resetMonitoringState` | boolean | false | Forget what this watch remembered |
| `maxConcurrency` | integer | 4 | Domains read in parallel (1 to 8) |
| `requestTimeoutSecs` | integer | 20 | Time limit per request (5 to 60). HTTP 429 and 5xx are asked again up to two more times |

### Pricing

Pay per event: **one charge per domain whose terms were read and returned** (`terms-analyzed`) plus **one small start fee per run** (`actor-start`), charged only in runs that return at least one domain (in monitor mode: runs that read and compared terms, even if nothing changed). Rows for domains whose terms could not be found or read are free. If the maximum total charge you set for a run has no room for the start fee plus one domain, nothing is requested and nothing is charged.

### Tips for agents

- Put the domain, not a deep page, unless you already know the terms URL.
- Use `categories` to keep the output small; the price is the same.
- Read `termsLastUpdated` and store `termsUrl`; schedule monitor mode weekly to learn when a site adds an AI or scraping clause.

### Support

Open an issue in the Issues tab with the domain and what you expected. Domains that are blocked, disallowed or JavaScript-only are reported as such on purpose.

# Actor input Schema

## `domains` (type: `array`):

Websites whose terms of service you want checked, one per line (1 to 500 per run). Give a domain (example.com) and the Actor finds the terms page from the home page links, a legal page it links to, or the usual addresses (/terms, /tos, /legal, /agb, /cgu, /aviso-legal …). Give a full URL with a path (https://example.com/legal/terms) and that page is read as the terms. The same domain given twice is read and charged once. If this is empty, the example domain github.com is read and nothing is charged.

## `categories` (type: `array`):

Which kinds of clauses to return. Leave empty for all six. The price is per domain whose terms were read, whatever you select.

## `onlyChanges` (type: `boolean`):

Return only domains whose terms text changed since the last run with the same watch name (plus domains new to the watch), with each changed paragraph before and after. The first run returns every domain as the starting point. Dates alone (a new "last updated" date, a copyright year) are not a change. A change is read twice a moment apart and reported only if both reads agree. A run in which nothing changed returns a free row saying so and charges only the run start fee.

## `watchName` (type: `string`):

Name of the remembered terms used to compare runs (letters, digits, dot, dash, underscore; up to 40). Setting it (or turning on monitor mode) fills changeType and changedClauses. Use a different name for each list of domains you track on its own schedule. With monitor mode on and no name, the name "default" is used.

## `resetMonitoringState` (type: `boolean`):

Start this watch over: forget the remembered terms before this run, so every domain is returned as a first check.

## `maxConcurrency` (type: `integer`):

How many domains are read in parallel (1 to 8).

## `requestTimeoutSecs` (type: `integer`):

How long to wait for one page to answer (5 to 60 seconds). HTTP 429, 500, 502, 503 and 504 are asked again up to two more times; a page that does not answer in time is asked once more.

## Actor input object example

```json
{
  "domains": [
    "github.com",
    "www.importyeti.com",
    "https://themen.kleinanzeigen.de/nutzungsbedingungen/"
  ],
  "onlyChanges": false,
  "resetMonitoringState": false,
  "maxConcurrency": 4,
  "requestTimeoutSecs": 20
}
```

# Actor output Schema

## `results` (type: `string`):

One row per domain: the terms of service URL, how it was found, last updated date, language, and the matching clauses (scraping and bots, circumventing limits, AI training and agents, databases and resale, liquidated damages, arbitration and governing law) quoted with section number and heading. In monitor mode, the changed paragraphs before and after. A domain whose terms could not be found or read comes back as a free row that says why.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "github.com",
        "www.importyeti.com",
        "https://themen.kleinanzeigen.de/nutzungsbedingungen/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("neverempty/terms-of-service-clause-finder").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "domains": [
        "github.com",
        "www.importyeti.com",
        "https://themen.kleinanzeigen.de/nutzungsbedingungen/",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("neverempty/terms-of-service-clause-finder").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "github.com",
    "www.importyeti.com",
    "https://themen.kleinanzeigen.de/nutzungsbedingungen/"
  ]
}' |
apify call neverempty/terms-of-service-clause-finder --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,neverempty/terms-of-service-clause-finder"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/WQyfqUlKSHkdDMWfw/builds/YLNXQRjQM9FCDYdKe/openapi.json
