Embedded Page Data Extractor - JSON-LD, __NEXT_DATA__ & Nuxt
Pricing
from $2.80 / 1,000 page with embedded data returneds
Embedded Page Data Extractor - JSON-LD, __NEXT_DATA__ & Nuxt
For AI agents and scrapers: every JSON a page ships in its HTML (__NEXT_DATA__, Next.js RSC, Nuxt, Apollo, __INITIAL_STATE__, JSON-LD, microdata), with IDs, numeric prices and stock the page never shows, a map of paths, or only the values at paths you name. Monitor mode returns only changes.
Pricing
from $2.80 / 1,000 page with embedded data returneds
Rating
0.0
(0)
Developer
NeverEmpty
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Embedded Page Data Extractor - JSON-LD, NEXT_DATA & Nuxt
For AI agents and scrapers: give it page URLs, get back every JSON the page already ships inside its HTML, as clean JSON, sorted by kind, with a map of the paths inside it. Name a path such as props.pageProps.product.price and you get just that value.
Many modern sites put the data you want into the HTML before any JavaScript runs: Next.js writes it to __NEXT_DATA__ or to its React Server Components stream (self.__next_f), Nuxt to __NUXT_DATA__, Apollo to window.__APOLLO_STATE__, Redux and friends to window.__INITIAL_STATE__ or __PRELOADED_STATE__, YouTube to ytInitialData, and almost every shop, recipe, event and article page adds schema.org JSON-LD or microdata. Reading that JSON is faster and more exact than parsing the visible HTML, and it holds fields the page never shows (IDs, prices in numbers, stock, coordinates, timestamps). This Actor finds all of it in one pass so you (or your agent) do not have to dig through <script> tags by hand.
Input and output in one look:
{"urls": ["https://www.nike.com/t/air-force-1-07-mens-shoes-jBrhbr/CW2288-111", "https://www.goodreads.com/book/show/5907.The_Hobbit"],"paths": ["next-data:props.pageProps.selectedProduct.prices.currentPrice", "json-ld:aggregateRating.ratingValue", "json-ld:[0].name"],"includeData": false}
{"url": "https://www.nike.com/t/air-force-1-07-mens-shoes-jBrhbr/CW2288-111","status": "ok","kindsFound": ["next-data", "json-ld"],"values": {"next-data:props.pageProps.selectedProduct.prices.currentPrice": 115,"json-ld:aggregateRating.ratingValue": null,"json-ld:[0].name": "Nike Air Force 1 '07 Men's Shoes"},"missingPaths": ["json-ld:aggregateRating.ratingValue"]}
(Real values from a run on 2026-09-25; the full row is shown under Output.)
What it finds
kind | Where it lives in the HTML | Notes |
|---|---|---|
next-data | <script id="__NEXT_DATA__" type="application/json"> | Next.js pages router. The page's data is usually under props.pageProps. |
next-flight | self.__next_f.push([1,"..."]) scripts | Next.js app router (React Server Components). The chunks are joined and split into rows { id, tag, value }; JSON rows are parsed, T text rows are kept as text. |
nuxt | <script id="__NUXT_DATA__"> / window.__NUXT__ | Nuxt 3 payloads are stored in the compact devalue format; they are decoded back into normal JSON (format: "devalue"). A Nuxt 2 window.__NUXT__=(function(a,b){...})(...) is code, not data, and is reported as unparsed (see below). |
apollo | window.__APOLLO_STATE__ and similar | Apollo Client cache (ROOT_QUERY and normalized objects). |
initial-state | window.__INITIAL_STATE__, __PRELOADED_STATE__, __INITIAL_DATA__, __SERVER_DATA__, __REACT_QUERY_STATE__, __remixContext, __staticRouterHydrationData, ytInitialData, ytInitialPlayerResponse, Inertia.js data-page, ... | Also window.X = JSON.parse("..."). |
json-ld | <script type="application/ld+json"> | One block per script. schemaTypes lists the schema.org @type values (also inside @graph). |
microdata | itemscope / itemprop attributes | Returned as { items: [{ type, id, properties }] }, the same shape as the WHATWG microdata-to-JSON algorithm. |
script-json | any other <script type="application/json">, application/*+json, text/x-magento-init | For example GitHub's react-app.embeddedData or Rotten Tomatoes' score blocks. |
script-variable | any other window.X = {...} / var longName = {...} | Only values that read as JSON are returned. |
data-attribute | data-*='{...}' attributes | Off by default (includeDataAttributes), because on most sites these are analytics tags. |
JSON that is almost JSON is repaired and read: undefined, !0/!1, raw line breaks inside strings, unquoted keys, single quotes, trailing commas (the block's format says json-repaired or js-object), and HTML-entity-encoded JSON ({"...).
Input
| Field | Type | Default | What it does |
|---|---|---|---|
urls | array of strings | (empty: the example https://www.bbc.com/news is used) | Pages to read, 1 to 1,000 per run. A URL without https:// is read as https://. The same URL given twice is read and charged once. |
paths | array of strings | (none) | Up to 50 dot-paths whose values go to values. Prefix with a kind (next-data:, json-ld:, nuxt:, apollo:, initial-state:, microdata:, script-json:, ...) or blockN: to pick the block; without a prefix the first block where the path exists is used. [0] or .0 for an array item, [*] for every item, ["key.with.dots"] for unusual keys. Not found = null in values and listed in missingPaths. |
kinds | array of strings | (all except data-attribute) | Only look for these kinds. |
includeDataAttributes | boolean | false | Also return JSON in data-* attributes. |
includeData | boolean | true | Return each block's full JSON in blocks[].data. Turn off to get only the block list, valueMap and values (small rows, good for agents). |
maxDataBytes | integer | 1000000 | Upper limit (0 to 5,000,000) for the JSON returned per page. Blocks that do not fit come back with dataOmitted: true and the reason; valueMap and values still use them. |
valueMapLimit | integer | 150 | How many paths to list in valueMap (0 to 5,000). |
onlyChanges | boolean | false | Monitor mode: return only pages whose watched values changed since the last run with the same watchName, plus pages new to the watch. |
watchName | string | (none) | Name of the remembered state (letters, digits, ., -, _; up to 40). Setting it fills changeType and changes. |
ignorePaths | array of strings | (none) | Paths that monitor mode should not compare (a path also ignores everything under it). |
resetMonitoringState | boolean | false | Forget what this watch remembered. |
maxConcurrency | integer | 4 | Pages read in parallel, 1 to 8. |
requestTimeoutSecs | integer | 20 | Time limit for one request, 5 to 60 seconds. HTTP 429, 500, 502, 503 and 504 are asked again up to two more times. |
Output
One row per URL. The Nike row from the example above (the valueMap shortened to 3 of its 150 entries; run vqEyOPQSgs1f0I0oN):
{"status": "ok","position": 0,"url": "https://www.nike.com/t/air-force-1-07-mens-shoes-jBrhbr/CW2288-111","finalUrl": "https://www.nike.com/t/air-force-1-07-mens-shoes-DZejrQoC/CW2288-111","httpStatus": 200,"contentType": "text/html; charset=UTF-8","redirects": [{ "url": "https://www.nike.com/t/air-force-1-07-mens-shoes-jBrhbr/CW2288-111", "status": 302 }],"blockCount": 2,"kindsFound": ["next-data", "json-ld"],"kindCounts": { "next-data": 1, "json-ld": 1 },"blocks": [{"index": 0, "kind": "next-data", "source": "script#__NEXT_DATA__[type=application/json]","scriptId": "__NEXT_DATA__", "scriptType": "application/json", "variable": null, "attribute": null,"bytes": 115152, "format": "json", "parsed": true, "parseError": null, "cutByPageLimit": false,"topLevelType": "object", "topLevelKeys": ["props", "page", "query", "buildId", "assetPrefix", "isFallback", "isExperimentalCompile", "gssp", "scriptLoader"],"itemCount": null, "schemaTypes": null, "dataOmitted": true, "dataOmittedReason": "includeData is off", "data": null, "rawText": null},{"index": 1, "kind": "json-ld", "source": "script[type=application/ld+json]","scriptId": null, "scriptType": "application/ld+json", "variable": null, "attribute": null,"bytes": 19605, "format": "json", "parsed": true, "parseError": null, "cutByPageLimit": false,"topLevelType": "array", "topLevelKeys": null, "itemCount": 1, "schemaTypes": ["ProductGroup"],"dataOmitted": true, "dataOmittedReason": "includeData is off", "data": null, "rawText": null}],"valueMap": [{ "block": 0, "kind": "next-data", "path": "props.pageProps.styleColor", "type": "string", "count": 1, "example": "CW2288-111" },{ "block": 0, "kind": "next-data", "path": "props.pageProps.groupKey", "type": "string", "count": 1, "example": "DZejrQoC" },{ "block": 1, "kind": "json-ld", "path": "[*].@type", "type": "string", "count": 1, "example": "ProductGroup" }],"valueMapTotalPaths": 1002,"values": {"next-data:props.pageProps.selectedProduct.prices.currentPrice": 115,"json-ld:aggregateRating.ratingValue": null,"json-ld:[0].name": "Nike Air Force 1 '07 Men's Shoes"},"valueSources": { "next-data:props.pageProps.selectedProduct.prices.currentPrice": 0, "json-ld:[0].name": 1 },"missingPaths": ["json-ld:aggregateRating.ratingValue"],"likelyNeedsJavaScript": false,"jsEvidence": null,"visibleTextChars": 5969,"pageTitle": "Nike Air Force 1 '07 Men's Shoes. Nike.com","htmlBytesRead": 774793,"htmlTruncated": false,"dataBytes": 0,"dataOmittedBlocks": null,"changeType": null,"changes": null,"previousCheckedAt": null,"watchName": null,"note": null,"fetchedAt": "2026-09-24T22:08:28.801Z"}
With includeData on (the default), each parsed block also carries its full JSON in data.
Useful columns:
valueMap: paths inside the JSON (arrays shown as[*]), the value type, how many values share the path, and an example value. Paths underprops.pagePropsof__NEXT_DATA__and shallow paths come first; analytics, config, translation and ad branches come last. Use it to pickpathsfor the next run.values/valueSources/missingPaths: the values at yourpaths, which block each came from, and which were not found.likelyNeedsJavaScript/jsEvidence: the page's body has almost no text (an empty#root/#appcontainer, a "please enable JavaScript" message, or under 300 characters next to scripts), so some of its data is loaded by JavaScript after it opens and is not in the HTML.blocks[].parseError/rawText: a block that was found but could not be read as JSON, with the first 2,000 characters.
Rows that are not charged
status | Meaning |
|---|---|
no-embedded-data | The page was read but has no embedded data. When the page is an empty JavaScript shell the note says so: the data is loaded later by JavaScript, which this Actor does not run (this is not "the page has no data"). Pages that only carry small script settings (a script-json, initial-state, script-variable or data-attribute block under 1KB) are also here, with those settings listed in blocks. |
no-parsable-data | Embedded data was found, but none of it reads as JSON (for example a Nuxt 2 function). The blocks are listed with parseError and rawText. |
robots-disallowed / robots-unreachable | The site's robots.txt does not allow this address, or could not be read (RFC 9309: treated as disallowed). The page was not requested. |
login-required | The page needs a sign-in (401 or a redirect to a sign-in page). |
blocked | The site refused the request or showed a check page (Cloudflare, DataDome, PerimeterX, AWS WAF, Akamai, Kasada, ...), including check pages sent with HTTP 200 or 202. |
not-found / http-error / unreachable / unreadable / not-html / bad-input | As named; the note says why. |
paths-not-found | Monitor mode with paths: none of the paths exist in the page's embedded data, so there is nothing to compare (the run start fee is not charged for this page). |
budget-reached / no-change | The run's maximum charge was reached / monitor mode found no change. |
Monitor mode
Set watchName (and onlyChanges to get only changed pages). The first run returns every page as first-check. Later runs compare:
- with
paths: only those values; - without
paths: every value in the embedded data (measured 2026-09-25: on pages with a live feed or recommendations, such as a news home page or a YouTube watch page, this reports the feed itself changing between runs; to watch only what matters, namepathsor list the noisy branches inignorePaths), except keys that change on every load (token,nonce,csrf,session,requestId,traceId,correlationId,buildId,timestamp,etag,hashand similar) and yourignorePaths.
Arrays whose items are all small (up to 200 characters of JSON each, up to 500 items) are compared as one value regardless of order, so a list that is only shuffled is not a change. A change is re-read 1.5 seconds later and reported only when the second read shows the same new value. Without paths, values that differ between two reads a moment apart are ignored for that page from then on; with paths, such a value is not reported that run and is compared again in the next run. Without paths, a path that appears or disappears (for example an FAQ entry or a recommendation slot) is reported only when two runs in a row see it; the first time it is held back (listed in the note, not charged), and if it flips back in the next run, that path is ignored for that page from then on. Changed values are reported in the same run. Changed rows carry changes: [{ path, change, before, after }]. Without paths, each path starts with the block it belongs to, named by kind and identity so that a new block elsewhere on the page does not shift the others: json-ld(Product):offers.price, next-data(__NEXT_DATA__):props.pageProps.title, apollo(__APOLLO_STATE__):ROOT_QUERY... (before is kept for values up to 200 characters; longer values are compared by a hash). A page read only up to the 5 MB limit never reports values as removed. If you change paths, each page starts over as new.
What it can and cannot see
- It reads the HTML the server sends, up to 5 MB per page, with an ordinary HTTP client. It does not run JavaScript, so data that a page fetches after it opens (XHR/fetch calls) is not included;
likelyNeedsJavaScriptflags those pages. - Measured on 2026-09-25 from Apify's servers on 42 commonly used pages: 21 pages returned embedded data (Next.js, Nuxt, Apollo, YouTube, JSON-LD, microdata), 3 had none, 11 refused automated requests or showed a check page, 2 were disallowed by robots.txt and 5 did not answer or could not be used. A person with a browser may still see pages this Actor cannot read.
- It respects each site's robots.txt (RFC 9309, as
NeverEmptyPageData; if a site has no group for that name,User-agent: *applies) for every URL and every redirect. It never signs in, never solves or bypasses check pages and does not switch to proxies to get around a refusal. - It only reads public web addresses (not internal network addresses).
Pricing
Pay per event:
- Page with embedded data returned: charged once per URL that comes back with
status: "ok"(at least one block read as JSON). Rows with any other status are free. - Run start: charged once per run that returns at least one page, or, in monitor mode, once per run that read and compared pages (even if nothing changed). A run that returns nothing is not charged at all.
If a run's maximum total charge has no room for the start fee plus one page, the Actor requests nothing and charges nothing.
Calling it from code
curl -X POST "https://api.apify.com/v2/acts/neverempty~embedded-page-data-extractor/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \-H "Content-Type: application/json" \-d '{"urls":["https://www.goodreads.com/book/show/5907.The_Hobbit"],"paths":["json-ld:aggregateRating.ratingValue"],"includeData":false}'
AI agents can call it through the Apify MCP server (https://mcp.apify.com) as neverempty/embedded-page-data-extractor.
Limits
- Up to 1,000 URLs and 50 paths per run; up to 5 MB of HTML per page; up to 5 MB of returned JSON per page (
maxDataBytes). - Next.js flight rows are returned as React Server Components rows, not as a rebuilt component tree.
- RDFa and Open Graph
<meta>tags are not extracted (they are not embedded JSON). - In monitor mode, avoid two overlapping schedules with the same
watchName.
Support
Something missing or wrong? Open an issue in the Issues tab with the URL and what you expected. This Actor is not affiliated with Next.js, Vercel, Nuxt, Apollo or schema.org.