Webpage to Verified JSON avatar

Webpage to Verified JSON

Pricing

from $0.04 / verified json created

Go to Apify Store
Webpage to Verified JSON

Webpage to Verified JSON

Extract source-backed JSON from one public static HTTPS page. Returns flat values, excerpts, and missing fields. Read the default dataset. Set maxTotalChargeUsd above $0.04.

Pricing

from $0.04 / verified json created

Rating

0.0

(0)

Developer

Dravara LLC

Dravara LLC

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

8 days ago

Last modified

Share

Turn one static public webpage into source-backed structured JSON.

Provide a URL and a small flat schema. The Actor returns only values it can tie to the downloaded page, along with missing fields and safe failure details. Here, "verified" means backed by a normalized, bounded excerpt from the downloaded page; it is not an independent guarantee that the publisher's statement is true.

This Actor is self-service. Normal runs return one structured dataset item; handled failures return one safe failure item without requiring a Dravara account, operator, or support fulfillment.

Pricing and charge controls

The live Apify Store pricing section is authoritative. This Pay per event + usage product can charge the customer for (1) one Dravara apify-actor-start event per Actor run, (2) one Dravara product result event only after a chargeable success, and (3) customer-paid Apify platform usage. The underlying Apify platform usage remains separate. Dravara does not reduce its event charge to absorb the underlying Apify platform usage. Before running, set Apify's maxTotalChargeUsd option to a limit you accept.

The custom result event is charged only for a success result. Approved partial and failed paths do not create that custom event. Each Actor run can still record the separate one-time apify-actor-start creator event, and the underlying Apify platform usage can still apply. The apify-default-dataset-item event is configured at $0 for this release.

Input

url is one public HTTPS page. schema is shorthand that maps each desired field name to its primitive type:

{
"url": "https://example.com/product",
"schema": {
"title": {"type": "string"},
"price": {"type": "number"},
"available": {"type": "boolean"},
"tier": {"type": "enum", "enum_values": ["free", "pro"]}
},
"include_evidence": true,
"idempotency_key": "550e8400-e29b-41d4-a716-446655440000"
}

Supported types are text, string, number, boolean, date, url, and enum. A schema can contain 1-20 fields. Nested objects and arrays are not part of this version. The URL is limited to 4,096 characters; each field name is limited to 100 characters; and an enum accepts 1-20 values. include_evidence remains accepted for compatibility, but evidence can no longer be disabled: both true and false return it.

idempotency_key is optional. It accepts 1-128 ASCII letters, digits, periods, underscores, tildes, or hyphens. Dravara HMAC-scopes it to the authenticated Apify customer and this product. While the replay record remains retained, reusing the same key with the same canonical input can return the exact durable result in a later Actor run; reusing it with changed input is rejected. Other customers and products have separate namespaces. A confirmed terminal replay does not create another Dravara result event or result artifact, but the new Actor run can still incur its separate Actor-start event and Apify platform usage. Dravara does not store or return the raw key. Use a unique opaque job ID, such as a UUID—not a secret or personal identifier.

Use cases

  • Convert a product page into a consistent flat record for a catalog or spreadsheet.
  • Extract selected plan, feature, or metadata fields for a downstream API workflow.
  • Check that required public-page fields exist while retaining source excerpts for review.

Output and evidence

The default dataset receives one status-bearing item. A complete extraction is success; verified values plus missing fields are partial; zero extracted values or a fetch failure is failed. The item contains data, per-field evidence, missing_fields, confidence, retrieval provenance, bounded usage telemetry, warnings, and structured errors. The confidence field is the completed-field coverage ratio, not an accuracy probability. Evidence includes a mandatory, non-empty whitespace-normalized source excerpt for every value, bounded to 4,000 characters, plus a selector when available and its source URL. The excerpt is not a byte-for-byte HTML quote: whitespace is normalized, and long elements are truncated to that limit. Unsupported fields stay in missing_fields; they are never returned as values without evidence.

An Apify run can complete while this item has status: failed. Integrations must inspect the dataset item's status and errors, not the Actor run status alone.

Failures and retries

Read the dataset item's status, errors, and each error's retryable value. The same idempotency key and canonical input replay the same terminal failure; changed input with that key is rejected. Inspect the original dataset and Charged events before starting another run. Correct nonretryable input first. For a retryable fetch or capacity failure, wait for recovery and use a new unique key only when billing is clear. That new run can incur its Actor-start event, Apify platform usage, and a result event after a chargeable success.

If the Actor itself fails after a dataset item appears, do not start another run yet. Compare that run's default dataset with its Charged events first.

Complete output example

Actual identifiers, timestamps, timing, and page values vary.

{
"run_id": "f82b6a45-7b8d-45a8-8439-de74fa0b2a88",
"status": "success",
"next_action": null,
"warnings": [
"This result does not include the final Apify platform-usage amount. Check this run's Usage and Charged events for final billing.",
"Evidence excerpts normalize whitespace and are bounded to 4,000 characters; use the returned source URL for verbatim review."
],
"errors": [],
"retrieval": {
"requested_url": "https://example.com/product",
"final_url": "https://example.com/product",
"retrieved_at": "2026-07-27T12:00:00Z",
"method": "static_https",
"http_status": 200,
"redirect_count": 0,
"content_type": "text/html",
"bytes_received": 286
},
"usage": {
"duration_ms": 18,
"bytes_processed": 286,
"extraction_attempts": 4
},
"cost": {
"currency": "USD",
"total_cost_usd": null,
"billable_events": []
},
"url": "https://example.com/product",
"data": {
"title": "Starter plan",
"price": 12,
"available": true,
"tier": "pro"
},
"evidence": {
"title": [{
"text": "Starter plan",
"selector": "script[type=\"application/ld+json\"]",
"source_url": "https://example.com/product"
}],
"price": [{
"text": "12",
"selector": "script[type=\"application/ld+json\"]",
"source_url": "https://example.com/product"
}],
"available": [{
"text": "True",
"selector": "script[type=\"application/ld+json\"]",
"source_url": "https://example.com/product"
}],
"tier": [{
"text": "pro",
"selector": "script[type=\"application/ld+json\"]",
"source_url": "https://example.com/product"
}]
},
"missing_fields": [],
"confidence": 1.0
}

Expected runtime

A responsive static page usually finishes in seconds. Remote-server speed and shared capacity can vary. One fetch uses an 8-second connect timeout and a 12-second read timeout; the Actor has a five-minute default run timeout.

Durable replay and retention

Paid launch fails closed unless Dravara's durable result store is enabled. For a paid run, the validated result JSON is stored as a compressed canonical response envelope in a private R2 bucket and becomes eligible for deletion after 365 days. PostgreSQL stores bounded run, settlement, artifact-reference, quota, retention-claim, cost/revenue, and replay-identity metadata; it does not store the full fetched page or raw request body. This persistence supports exact result replay; it is not page monitoring. When supplied, idempotency_key is stored by Dravara only through the namespaced HMAC replay identity and the request digest, not as raw text.

Apify separately retains normal run input and dataset output under the runner's storage settings. Deleting that Apify storage does not itself delete Dravara's durable records. After a terminal run is at least 365 days old, Dravara's retention worker can irreversibly minimize its platform identity, replay identity, input digest, and free-form run metadata only when its artifacts have been deleted, reconciliation is final, and no hold or failed retention claim remains. Run status, timestamps, result-integrity hash, economic records, and settlement evidence remain under a separate accounting and dispute-retention policy; not all PostgreSQL data expires after 365 days. Customer-key replay is no longer available after those replay identities and result artifacts are minimized. See Dravara Privacy for the complete boundary.

Version 1 limits

  • URL input is static public HTTPS only, on port 443.
  • A fetched page is limited to 2 MB of decoded HTML and up to three redirects; every redirect destination is revalidated.
  • There is no login support, JavaScript execution, browser rendering, proxy support, CAPTCHA bypass, access-control bypass, or form submission.
  • One page is processed per run. There is no crawling, screenshot capture, model-based guessing, or recursive schema support.
  • There is no product-level monitoring, scheduling, or webhook delivery.

Do not put passwords, session tokens, API keys, or private customer data in a URL or idempotency_key. Credentials and URL fragments are rejected. Page content is processed by Dravara, LLC's private gateway using deterministic extraction rules. The destination site receives a request from that gateway and may retain it under the site's own logging and privacy practices.

Support: support@dravarahq.com