Website JSON-LD Structured Data Scraper avatar

Website JSON-LD Structured Data Scraper

Pricing

from $1.80 / 1,000 page result saveds

Go to Apify Store
Website JSON-LD Structured Data Scraper

Website JSON-LD Structured Data Scraper

Extract JSON-LD objects from public HTML URLs. Keep schema types, original values, script indexes and graph paths. Get clear invalid JSON and page error results.

Pricing

from $1.80 / 1,000 page result saveds

Rating

0.0

(0)

Developer

Cliqto Media

Cliqto Media

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

6 hours ago

Last modified

Share

Read JSON-LD structured data from public HTML pages. Get the original objects, their schema types, and the script and graph path for each object. Use this Actor for SEO checks and for moving structured web data into your own tools.

Start with one public URL in startUrls. The price is $0.0018 per saved page result + $0.001 per run, plus Apify platform usage. Empty pages and page errors also return a saved result. Check status before using the objects.

Website JSON-LD Structured Data Scraper: public input, structured results and supported export

Table of contents

What it does

The Actor reads only the URLs you add. It finds all HTML scripts with the application/ld+json type. It reads single objects, arrays, and @graph nodes. Each saved page has its own result, even when there are no objects or the page has an error.

A graph wrapper with only @context and @graph becomes its graph nodes. A wrapper with its own other fields is kept as an object too. Normal nested fields stay inside the original object. Repeated objects in different scripts stay separate, so you can check each source path.

It uses HTTP and does not run page JavaScript. It does not follow page links, load remote JSON-LD contexts, expand JSON-LD, or test Google rich result rules. See the JSON-LD standard for these terms.

Data

Get one Dataset row per processed URL. objects keeps each original object in data, with scriptIndex, jsonPath, types, and the local or inherited context. jsonLd gives the same original objects without the path fields. schemaTypes lists the types found before filtering.

issues shows invalid JSON and extraction limits. An invalid block does not hide valid blocks from the same page. Error messages do not contain the invalid block text.

Quick start

  1. Open Input.
  2. Add https://apify.com to Page URLs (startUrls).
  3. Keep the default settings. To keep only a website object, set typeFilters to ["WebSite"].
  4. Click Start.
  5. Open Pages and JSON-LD. Check status, schemaTypes, and objects.
  6. In the key-value store, open RUN_SUMMARY to check unprocessed pages.
{"startUrls":[{"url":"https://apify.com"}],"typeFilters":["WebSite"]}

The checked page returned a WebSite object named Apify. A site can change its HTML later. The input editor prefill has apify.com and example.com; replace it with your own list.

Input

FieldTypeRequiredDefault / prefillRangeEffectRecommendation
startUrlsarray of { "url": string }YesNo runtime default; editor prefill: Apify and Example Domain1–100 URLs; each URL up to 2048 charactersPages to read; only public HTTP(S), standard ports, no URL loginStart with 1–3 pages. Use public URLs without secrets.
typeFiltersstring arrayNo[]Up to 20 values, each 1–200 charactersKeeps objects whose direct @type matches any value exactlyUse Product, Recipe, or another exact source value. Empty keeps all.
maxObjectsPerPageintegerNo10001–1000Limits objects before the type filterKeep the default for a full check. A hit marks the page partial.
requestTimeoutSecsintegerNo201–20 secondsTotal time for one page, including retries and redirectsLower it for slow sources.
maxRunSecsintegerNo1801–180 secondsTotal work time; stops new pages and cancels active workKeep 180. Check the summary if pages are left.

An empty input or unknown key fails validation. URLs are normalized and exact duplicates are fetched once. Fragments do not make separate URLs. Query values still affect the request, but are removed from saved url and finalUrl fields. Use inputIndex to map a result back to your input. Do not put credentials in URLs.

Settings together

The object limit comes before typeFilters. A smaller object limit may hide objects that would match your filter. A filter changes output only; it does not reduce requests or page charges. schemaTypes and objectCount describe the extracted objects before the filter. An empty match set on a page with objects is still complete if there are no issues.

Input examples

Keep all objects on two pages

{"startUrls":[{"url":"https://apify.com"},{"url":"https://example.com"}]}

At the check date, the first page returned Organization, WebSite, and SoftwareApplication objects. The second returned not_found. Both have saved page results and page charges.

Keep one type

{"startUrls":[{"url":"https://apify.com"}],"typeFilters":["WebSite"]}

Returns only the WebSite object. The page charge is the same with or without the filter.

Small extraction limit

{"startUrls":[{"url":"https://apify.com"}],"maxObjectsPerPage":1}

The checked page returned one Organization object with status: "partial" and an OBJECT_LIMIT issue. Increase the limit to read more objects.

Output example

This is a shortened result for the WebSite input above. All fields are listed below. Values come from a checked public page; the scrape time is left out here.

{
"url": "https://apify.com/",
"finalUrl": "https://apify.com/",
"status": "complete",
"schemaTypes": ["Organization", "SoftwareApplication", "WebSite"],
"objectCount": 3,
"matchedObjectCount": 1,
"objects": [{
"scriptIndex": 1,
"jsonPath": "",
"types": ["WebSite"],
"context": "https://schema.org",
"data": {
"@context": "https://schema.org",
"@type": "WebSite",
"@id": "https://apify.com/#website",
"name": "Apify",
"url": "https://apify.com/",
"publisher": {"@id": "https://apify.com/#organization"}
}
}],
"issues": [],
"errorCode": null
}

jsonPath: "" means the script root. A path such as /0/@graph/1 means the second graph node in the first array item. jsonLd contains the same values as objects[].data. JSON values stay as published; the Actor does not add missing source fields or change HTML entities inside scripts.

Output fields

Page row

FieldTypeCan be null?ExampleMeaning
inputIndexintegerNo0Zero-based position in the original URL input
urlstringNohttps://apify.com/Input URL without credentials, query, or fragment; invalid URLs use [invalid URL]
finalUrlstringYeshttps://apify.com/Final URL after a successful HTML fetch; null on fetch errors
httpStatusintegerYes200HTTP response code; null when no response code is available
statusstringNocompletecomplete, partial, not_found, or error
schemaTypesstring arrayNo["WebSite"]Sorted unique direct types before the type filter; empty if none
jsonLdobject arrayNo[{"@type":"WebSite"}]Original retained objects; empty if no retained objects
objectsobject arrayNoSee exampleOriginal retained objects and their source paths
issuesobject arrayNo[{"scriptIndex":2,"jsonPath":"","code":"INVALID_JSON"}]Parsing and extraction issues; empty means none
scriptCountintegerNo3Number of JSON-LD scripts found in fetched HTML
objectCountintegerNo3Objects extracted before filtering and output size limits; may be limited by extraction caps
matchedObjectCountintegerNo1Objects kept in this saved row after filters and size limits
errorCodestringYesTIMEOUTSafe fetch error code; null when the HTML was read successfully
requestsintegerNo1Transport attempts, including retries and redirects; a blocked DNS check can count as an attempt without a connection
retriesintegerNo0Temporary error retries used
responseBytesintegerNo492185Bytes of a successful HTML body; zero when the fetch failed
scrapeDateISO date-time stringNo2026-10-05T06:25:16.009ZTime processing this page started, in UTC

Each object in objects

FieldTypeCan be null?ExampleMeaning
scriptIndexintegerNo1Zero-based index among JSON-LD scripts, not all HTML scripts
jsonPathstringNo/0/@graph/1JSON Pointer path inside that script; empty string is root
typesstring arrayNo["Organization","LocalBusiness"]This object's direct string @type values; empty if not provided
contextany JSON valueYeshttps://schema.orgOwn or inherited context, copied as data; never fetched
dataobjectNo{"@type":"WebSite"}Original parsed object, including its nested fields

Each issue in issues

FieldTypeCan be null?ExampleMeaning
scriptIndexintegerNo2Script index; -1 for a page output-size limit
jsonPathstringNo/@graph/1Location when known; empty for a script-wide issue
codestringNoINVALID_JSONINVALID_JSON, INVALID_JSONLD_SHAPE, OBJECT_LIMIT, or DEPTH_LIMIT

Key-value store records

OUTPUT is { "pages": [page rows], "summary": {run summary} }. The Dataset has the same page rows. RUN_SUMMARY is the same run summary without the pages. On rejected runtime input, the summary has status: "error", errorCode: "INVALID_INPUT", and processedUrlCount: 0.

Summary fieldType / nullExampleMeaning
statusstring; not nullpartialOverall result: complete, partial, not_found, or error
inputCountinteger; not null4Submitted URL entries
uniqueUrlCountinteger; not null3Entries after exact URL deduplication
duplicateUrlCountinteger; not null1Duplicate entries not fetched
processedUrlCountinteger; not null3Saved page rows
unprocessedUrlCountinteger; not null0Unique entries without a saved page row
objectCountinteger; not null3Sum of saved pages' extracted object counts
matchedObjectCountinteger; not null1Sum of saved pages' kept object counts
requestsinteger; not null3All source transport attempts, including a fetched page stopped by the total output limit
retriesinteger; not null0All source retries
pageStatusesobject of integer counts; not null{"complete":1,"partial":0,"not_found":2,"error":0}Saved page counts by status
stopReasonstring or nulloutput_limitnull on normal finish; otherwise run_time_limit, cancelled, charge_limit, output_limit, or storage_error

Pricing

  • $0.0018 per saved page result (page-result). The price covers all retained JSON-LD objects and diagnostics on that page.
  • $0.001 per run (apify-actor-start) at the supported 256–512 MB memory settings.
  • Apify platform usage is extra and paid by you. Compute, storage, and traffic depend on your plan and run. Check the run's usage in Console.

A not_found page, a page with no filter matches, a partial page, and an error page each cost one page event when their result is saved. A duplicate input costs no extra page event. Results that are not saved are not charged as page events. A run that is cancelled can still have the start charge and charges for saved pages. A hard stop can leave Dataset rows that did not reach OUTPUT; use the Dataset and run log in that case.

Saved page rowsRun countEvent price, before extra usage
11$0.0028
101$0.019
1001$0.181
1000 across ten runs of 10010$1.81

These are price calculations, not speed promises. One run accepts up to 100 URLs. The 100-URL input limit was checked with local test pages, not a load test on real sites. Total output limits and slow pages can stop a run earlier.

The minimum run cost limit is $0.0028. Use Maximum cost per run to limit event charges. Extra platform usage is separate. The Actor stops when the page-event limit is reached; saved rows remain available. Read the Apify pricing guide for platform billing details.

Use cases

TaskInputFields to use
Check JSON-LD coveragePublic page URLs; no filterstatus, schemaTypes, issues
Collect product or recipe objectsExact source type in typeFiltersobjects[].data, jsonLd
Find a broken scriptPublic page URLsissues[].scriptIndex, issues[].code
Track where a graph node came fromPages with @graphurl, inputIndex, scriptIndex, jsonPath

Scheduling and monitoring

You can save the input as an Apify task and run it on a platform schedule. Each run reads the current page HTML and uses the normal price. There is no built-in change tracking or discount for unchanged pages. Compare stored objects in your own tool, for example by @id and source URL. Do not compare scrapeDate as a content change.

API and integrations

Send the same JSON input to POST /v2/acts/cliqtomedia~website-jsonld-structured-data-scraper/runs. Authenticate with an Apify token in the request header. Use the run's defaultDatasetId to read GET /v2/datasets/{datasetId}/items. Read status before passing jsonLd or objects[].data to the next step.

Export page rows as JSON or CSV from the Dataset. JSON is the useful choice for nested JSON-LD objects. For another tool, pass url, schemaTypes, and objects[].data to your mapping. Apify MCP can run the Actor with the same input. There is no browser setup or website API key.

Limits and partial results

  • Only public HTTP(S) pages on standard ports. Private network IPs, URL login, and unsafe redirects are blocked.
  • HTML only. UTF-8 and ASCII are supported; missing charset is read as UTF-8. A source that ignores Accept-Encoding: identity is rejected rather than read as broken text.
  • 1 MB HTML body, measured while reading. Larger pages give BODY_LIMIT; they are not reported as empty pages.
  • Up to 3 redirects, 2 retries, and 6 transport attempts per URL. 429 and 5xx responses can be retried within the same page deadline.
  • Up to 1000 extracted objects per page. Any JSON nesting deeper than 64 levels or more than 10,000 JSON values per script gives DEPTH_LIMIT. That script is skipped; other scripts can still return objects. The graph walk also stops after 10,000 visited entries per page.
  • 2 MB per saved page row and 8 MB total page output. Whole objects can be omitted to meet a row limit; the row becomes partial. Up to 200 issues are saved per page. If there are more, the issue list is shortened and a page-level OBJECT_LIMIT is added. A total output limit stops new saved pages. Check unprocessedUrlCount.
  • Limits are separate: 100 large pages with 1000 objects each may hit the output or time limit first.
  • JavaScript-only JSON-LD is not visible to this HTTP version. Linked contexts are kept as values; they are never loaded.
  • No recursive crawl, login, cookies, browser, Microdata, RDFa, Open Graph, or Google eligibility check.
  • complete means the extracted objects passed this parser, not that the page has valid SEO markup. With no valid objects and malformed JSON, the page is partial.
  • Mixed page errors give an overall partial. All page errors give error. All valid pages without objects give not_found.
  • A platform run can finish with SUCCEEDED and still have page errors. Check RUN_SUMMARY and page status.
  • Storage and billing are separate actions. A storage error stops the run. Restarting or resurrecting a run that already has Dataset rows is rejected with RESUME_NOT_SUPPORTED; start a new run.

Troubleshooting

No objects: Check status. not_found means the fetched HTML has no valid objects. The source may add them with JavaScript. If objectCount is positive but matchedObjectCount is zero, remove the type filter.

INVALID_JSON: A JSON-LD script contains invalid JSON. Check its scriptIndex in the public page source. Other valid scripts can still have objects.

partial: Read issues and stopReason. Raise maxObjectsPerPage if its limit is too low. Split the URL list if the output or run time limit was reached.

HTTP_ERROR, NETWORK_ERROR, or TIMEOUT: Check that the public page opens and try a smaller URL list. A 403 or 429 can mean the site blocks this HTTP request. This version has no browser or proxy fallback.

UNSAFE_ADDRESS, UNSAFE_HOST, or UNSAFE_URL: Use a public website URL without login or a custom port. Private and local services are not supported.

BODY_LIMIT, NON_HTML, UNSUPPORTED_ENCODING, or UNSUPPORTED_CHARSET: Use a smaller UTF-8 HTML page. These failures do not mean no JSON-LD exists.

RESUME_NOT_SUPPORTED: Start a new run. Export saved Dataset rows from the earlier run before repeating the input.

FAQ

Do I need a website account?

No. Add public pages that can be read without login.

Does this change the JSON-LD objects?

No. It parses JSON and keeps original object values. It adds source fields outside data. It does not expand contexts or repair invalid JSON.

Do I pay if there are no matching objects?

Yes. A saved empty-page result or a filter with no matches costs one page event. The start charge and platform usage also apply.

Can this replace a scraper that also returns meta tags?

Only the core url, jsonLd, schemaTypes, and scrapeDate fields overlap. This Actor adds source paths and statuses. Meta and social tag fields need a separate tool.

Support

Open an issue on this Actor's Apify page. Include the run ID, input field names, and the safe error code. Keep tokens, cookies, and private page data out of the message.