Terms of Service Clause Finder - Scraping, AI & Arbitration avatar

Terms of Service Clause Finder - Scraping, AI & Arbitration

Pricing

from $8.40 / 1,000 terms analyzeds

Go to Apify Store
Terms of Service Clause Finder - Scraping, AI & Arbitration

Terms of Service Clause Finder - Scraping, AI & Arbitration

Before an AI agent or scraper uses a website: finds its terms of service and quotes, word for word with section numbers, the clauses on scraping and bots, circumventing limits, AI training and agents, resale, liquidated damages and arbitration. Amazon: 17 clauses, all verbatim (2026-09-25).

Pricing

from $8.40 / 1,000 terms analyzeds

Rating

0.0

(0)

Developer

NeverEmpty

NeverEmpty

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

For AI agents and scrapers that need to know what a website's terms say before they use it: give a list of domains, get back each site's terms of service page and the clauses on scraping and bots, circumventing rate limits or security, AI training and AI agents, building databases, resale and commercial use, liquidated damages, and arbitration and governing law - quoted word for word, with the section number, the heading above it, the page URL and the terms' last-updated date.

This is not legal advice. The Actor finds and quotes clauses by their wording; it never says whether something is allowed or forbidden. A category with no quoted clause does not mean the use is permitted. Always read the full terms.

Input and output in one look

Input:

{
"domains": ["amazon.com", "www.importyeti.com", "https://themen.kleinanzeigen.de/nutzungsbedingungen/"],
"categories": ["scraping-and-automated-access", "ai-training-and-ai-agents"]
}

One row per domain (real row from amazon.com, 2026-09-25, two of its 17 clauses shown):

{
"status": "ok",
"domain": "amazon.com",
"termsUrl": "https://www.amazon.com/gp/help/customer/display.html?nodeId=GLSBYFE9MGKKQXXM&ref_=footer_cou",
"termsFoundVia": "homepage-link",
"termsPageTitle": "Conditions of Use - Amazon Customer Service",
"termsLastUpdated": "2026-08-14",
"termsLastUpdatedText": "Last updated: August 14, 2026",
"clauseCount": 17,
"categoriesFound": ["scraping-and-automated-access", "circumventing-access-controls-or-rate-limits", "ai-training-and-ai-agents", "database-building-redistribution-commercial-use", "arbitration-and-governing-law"],
"clauses": [
{
"categories": ["ai-training-and-ai-agents"],
"section": null,
"heading": "Agents",
"text": "Transparency and Consent. No Agent may access, use, or interact with Amazon Services unless, at all times, it identifies itself and operates in strict accordance with the requirements in section 3 of these Agent Terms. ...",
"matchedPhrases": { "ai-training-and-ai-agents": ["No Agent may", "Agent Terms"] }
},
{
"categories": ["scraping-and-automated-access", "database-building-redistribution-commercial-use"],
"heading": "LICENSE AND ACCESS",
"text": "... This license does not include any resale or commercial use of any Amazon Service, or its contents; ... or any use of data mining, robots, or similar data gathering and extraction tools. ...",
"matchedPhrases": { "scraping-and-automated-access": ["data mining", "robots"], "database-building-redistribution-commercial-use": ["commercial use"] }
}
],
"note": "Quoted clauses are matched by phrases; this is not legal advice and says nothing about what is allowed. Read the full terms."
}

What it looks for

Category (categories)Examples of the wording it matches
scraping-and-automated-accessscrape, crawl, spider, robot, bots, data mining, harvest, "automated means"; Scraping, Crawler, automatisierte Abfragen; aspiration, extraction automatisée; arañas, medios automatizados; スクレイピング, クローラ, ロボット, 自動化された手段
circumventing-access-controls-or-rate-limitscircumvent / bypass / avoid security or technical measures, rate limits, CAPTCHAs, IP blocks, unreasonable load; umgehen, übermäßige Last; contourner; eludir; アクセス制御の回避, 過度な負荷
ai-training-and-ai-agentstraining AI or machine learning models, training data, text and data mining (incl. § 44b UrhG), AI agents, "Agent" terms; 人工知能の学習, 情報解析
database-building-redistribution-commercial-usecreate a database or compilation, resell, redistribute, republish, frame or mirror, commercial use; Datenbank, gewerblich; base de données, fins commerciales; base de datos, fines comerciales; データベース, 商用, 転載, 再配布
liquidated-damages-and-penaltiesliquidated damages, a fixed amount per violation or per page (for example "$15,000 USD per 1,000,000 posts"), Vertragsstrafe, clause pénale, cláusula penal, 違約金
arbitration-and-governing-lawarbitration, class action waiver, governing law, exclusive jurisdiction or venue; Gerichtsstand, anwendbares Recht; droit applicable; ley aplicable, jurisdicción; 準拠法, 専属的合意管轄

A paragraph under a heading such as "Governing Law", "Arbitration", "Agents", "Gerichtsstand" or "準拠法" is also returned in that category even if its own sentence does not repeat the word (matchedPhrases then says heading: …).

Wording that looks similar but is about something else is not returned. These were found in real terms and are checked by the Actor's tests: "automated text messages" and autodialer consent, "automated decision-making", "eBay's automated systems scan messages" (the site's own systems), "we use automated tools to translate", "commercially reasonable efforts", "under penalty of perjury", a site describing its own chatbot ("Due to the nature of Generative AI…"), a company describing itself ("an AI safety and research company"), and force majeure ("回避することのできない災厄").

How it finds the terms page

  1. You gave a full URL with a path - that page is read as the terms (termsFoundVia: input-url).
  2. You gave a domain - the home page is read and its links are scored by their text and address in English, German, French, Spanish, Italian, Portuguese, Dutch and Japanese ("Terms of Service", "Conditions of Use", "User Agreement", "AGB", "Nutzungsbedingungen", "CGU", "Conditions générales", "Términos y condiciones", "Aviso legal", "利用規約" …). Links for other audiences (API, sellers, advertisers, merchants) score lower and are listed in otherTermsLinks instead.
  3. If the best link is a legal hub page ("Legal", "Policies", a footer link to /policies), the Actor follows one more step to the terms on that page (legal-page-link).
  4. If nothing is linked, it tries the usual addresses - /terms, /terms-of-service, /terms-of-use, /legal/terms, /tos, /legal … and, by country or page language, /agb and /nutzungsbedingungen (de), /cgu (fr), /aviso-legal (es), /kiyaku and /rule (ja) (common-path). If the home page redirected to another host (spotify.com → open.spotify.com), the original domain is tried too.
  5. Terms that live in a OneTrust notice (the page is an empty shell and the text is in a JSON file on privacyportal-cdn.onetrust.com) are read from that JSON (termsFoundVia ends with +onetrust).

Each page is accepted as terms only if its title, headings or address name terms, conditions or an agreement (in any of the languages above), it has at least 1,200 characters of text and reads like a contract. A privacy or cookie policy is not treated as terms, even when you give its URL directly. A domain whose terms are not found is not charged.

What it will not do

  • Robots.txt is respected for every page it reads, including the terms page itself (RFC 9309: a robots.txt that cannot be read because of a server error is treated as "do not read"). A disallowed page comes back as a free robots-disallowed row. Sites whose robots.txt disallows everything for all bots (for example x.com, linkedin.com, reddit.com) cannot be read this way - give the terms to your agent another way.
  • No check pages are solved or bypassed, no proxies are used to get around a refusal, and nothing that needs a sign-in is read. Those domains come back as free blocked or login-required rows.
  • No JavaScript is run. When the terms page is found but its text is loaded later by JavaScript, the row says terms-text-not-readable (free) - it does not claim the terms have no such clauses.
  • It never guesses: a date that is not on the page is null, a section number that is not printed is null.

Real result (production run on 2026-09-25, 40 well-known domains in 8 languages, 256 MB, about 12 to 100 seconds): terms read for 20 domains with 218 quoted clauses; 14 domains refused this reader or showed a check page (for example etsy.com, booking.com, openai.com, nytimes.com); 5 disallow bots in robots.txt; 1 resolved to a private address. Browser check of the quoted text: GitHub 12/12, Amazon 17/17, Airbnb 25/25, Spotify 14/14, Kleinanzeigen 17/17 and Mercari 2/2 quotes found word for word on the live page, with the same last-updated date.

Output fields

FieldMeaning
statusok (terms read and returned; charged), or a free reason: terms-not-found, terms-text-not-readable, robots-disallowed, robots-unreachable, blocked, login-required, not-found, http-error, unreachable, unreadable, not-html, bad-input, budget-reached, no-change
input, domain, positionWhat you gave and its place in the list
termsUrlThe terms page that was read (after redirects)
termsFoundViainput-url, input-url-link, homepage-link, legal-page-link or common-path (++onetrust)
termsPageTitle, termsLanguagePage title and the <html lang> value
termsLastUpdated, termsLastUpdatedTextISO date and the sentence it came from ("Last updated", "Effective date", "Stand", "gelten ab", "Dernière mise à jour", "Última actualización", "最終更新日", "改定" …); null when the page states no date
termsTextLength, termsTruncatedCharacters of text read; true if the page was larger than 3 MB and read only up to that size
clauseCount, categoriesFound, clauseCountByCategoryHow many clauses were quoted, in which categories (up to 12 per category)
clausesArray of { categories, section, heading, text, matchedPhrases, truncated }. text is the paragraph word for word; paragraphs longer than 1,200 characters are cut to the matching sentences with "…" and truncated: true
otherTermsLinksOther legal links found on the way (for example API terms, consumer vs commercial terms) as { text, url } - pass one of them as a URL to read it
changeType, changedClauses, changedCategories, previousCheckedAt, termsUrlChanged, previousTermsUrl, watchNameMonitor mode (below)
httpStatus, note, fetchedAtLast HTTP status, a plain-English note, time of reading

Monitor mode: only terms that changed

Set watchName (and onlyChanges: true to receive only changed domains). The Actor remembers each domain's terms text paragraph by paragraph. On the next run it returns only domains whose text changed, with changedClauses as { change: "added" | "removed" | "modified", categories, before, after }, so an agent can see at once whether a scraping, AI or arbitration clause was added.

  • A new "last updated" date or copyright year alone is not a change - dates are masked before comparing.
  • Menus and lists of links are not compared.
  • A change is read a second time a moment later and reported only if both reads show exactly the same changes; otherwise nothing is reported or remembered that time, and the next check compares again.
  • If the terms are found on a different page than last time, that check only records the new page; changes are reported from the next check.
  • The first run of a watch returns every domain as first-check. A run where nothing changed returns one free no-change row and charges only the run start fee.

Input

FieldTypeDefaultDescription
domainsarray of stringsexample github.com (read free of charge)Domains or full terms URLs, 1 to 500 per run. If empty, the example domain github.com is read and nothing is charged
categoriesarrayall sixWhich clause categories to return
onlyChangesbooleanfalseMonitor mode: return only domains whose terms changed
watchNamestring-Name of the remembered terms (letters, digits, ., -, _; up to 40)
resetMonitoringStatebooleanfalseForget what this watch remembered
maxConcurrencyinteger4Domains read in parallel (1 to 8)
requestTimeoutSecsinteger20Time limit per request (5 to 60). HTTP 429 and 5xx are asked again up to two more times

Pricing

Pay per event: one charge per domain whose terms were read and returned (terms-analyzed) plus one small start fee per run (actor-start), charged only in runs that return at least one domain (in monitor mode: runs that read and compared terms, even if nothing changed). Rows for domains whose terms could not be found or read are free. If the maximum total charge you set for a run has no room for the start fee plus one domain, nothing is requested and nothing is charged.

Tips for agents

  • Put the domain, not a deep page, unless you already know the terms URL.
  • Use categories to keep the output small; the price is the same.
  • Read termsLastUpdated and store termsUrl; schedule monitor mode weekly to learn when a site adds an AI or scraping clause.

Support

Open an issue in the Issues tab with the domain and what you expected. Domains that are blocked, disallowed or JavaScript-only are reported as such on purpose.