# Page Publish Date Finder - Article Date & Last Modified (`neverempty/page-publish-date-finder`) Actor

For AI agents citing web pages: when a page was published and last updated, with the evidence (JSON-LD, meta tags, <time>, text, URL, Last-Modified) and a high/medium/low confidence. On 28 pages with a known date it answered 27, all the right day. Optional Internet Archive first-seen date.

- **URL**: https://apify.com/neverempty/page-publish-date-finder.md
- **Developed by:** [NeverEmpty](https://apify.com/neverempty) (community)
- **Categories:** AI, News, SEO tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.80 / 1,000 page dateds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Page Publish Date Finder - Article Date & Last Modified

**For AI agents that cite web pages:** give it a URL, get back **when the page was published and when it was last updated**, each with a **confidence (high / medium / low)**, **where the date was read from**, and **why that one was chosen**, plus every other date found on the page. When a page has no date, it says so instead of guessing.

An answer that quotes a page without its date can present a five-year-old fact as today's news. The date is usually on the page, but in a dozen different places (JSON-LD, Open Graph meta tags, `<time>` elements, a "Published:" or "更新日" label, the URL, the HTTP header), and the places often disagree. This Actor reads all of them in one pass and tells you which one to trust.

Input and output in one look:

```json
{ "urls": ["https://mag.executive.itmedia.co.jp/executive/article/2609/25/2000001737/"] }
```

```json
{
  "status": "ok",
  "url": "https://mag.executive.itmedia.co.jp/executive/article/2609/25/2000001737/",
  "publishedDate": "2026-09-25T10:53:03+09:00",
  "publishedConfidence": "medium",
  "publishedSource": "meta tag",
  "publishedReason": "from meta tag name=\"cxenseparse:recs:publishtime\"; agrees with <time> element, date element, visible text; differs from JSON-LD (2026-09-25T10:53:03Z); JSON-LD gives the same clock time with a different time zone (2026-09-25T10:53:03Z); this one is used because another source's time supports it",
  "publishedTimezoneKnown": true,
  "daysSincePublished": 0,
  "modifiedDate": "2026-09-25T10:53:03Z",
  "modifiedConfidence": "low",
  "sourcesFound": ["meta tag", "date element", "JSON-LD", "<time> element", "visible text", "HTTP Last-Modified"],
  "pageLanguage": "ja"
}
```

(Real values from a run on 2026-09-25. This Japanese news page writes Japan time with a `Z` in its JSON-LD; the meta tag and the server's Last-Modified show the real time zone, so the meta tag wins and the conflict is spelled out.)

### What it reads

Every date found is listed in `evidence` with its source, field and raw text. Sources are grouped into families; when two different families point to the same day, the date is more trustworthy.

| Source (`publishedSource` / `modifiedSource`) | Where on the page | Notes |
|---|---|---|
| `JSON-LD` | `<script type="application/ld+json">`: `datePublished`, `dateCreated`, `uploadDate`, `dateModified` | Only the nodes that describe the page itself (top level, `@graph`, `mainEntity`). Dates of other articles listed on the page (`itemListElement` and similar) are not counted. The node whose `url`/`@id` equals the page wins. |
| `meta tag` | `article:published_time`, `article:modified_time`, `og:updated_time`, `citation_publication_date`, `DC.date.issued`, `dcterms.modified`, `parsely-pub-date`, `sailthru.date`, `pubdate`, `date` and 30 more; also any meta name ending in `published`, `first_published`, `last_updated`, `modified` | Names that do not say which date they are (`date`, `DC.date`) are used last. |
| `microdata` | `itemprop="datePublished"` / `dateModified` / `uploadDate` | |
| `<time> element` | `<time datetime="...">` | Labelled as published or updated by `pubdate`, `itemprop` or class names (`published`, `updated`, ...). Times inside comments, related-article lists, sidebars and footers are skipped. An unlabelled `<time>` that equals another source's update time is treated as the update date. |
| `date element` | Elements whose class or id says `pubdate`, `post-date`, `entry-date`, `published`, `updated`, `date`, ... and whose text starts with a date | For example `<div class="pubdate">2026年9月25日</div>`. |
| `visible text` | Labels followed by a date: `Published`, `Posted on`, `Updated`, `Last updated`, `Last modified`, `Last reviewed or updated`, `This page was last edited on`, `Released`, `公開日`, `公開`, `掲載日`, `投稿日`, `更新日`, `最終更新日`, `作成日`, `発表日` | Text inside `<code>` and `<pre>` is ignored (so the example `Last-Modified: Wed, 21 Oct 2015 ...` on a documentation page is not read as the page's date). |
| `URL` | `/2026/09/24/`, `/2026/sep/24/`, `2026-09-24`, `/20260925-...`; `/2026/09/` gives the month only | Weak on its own; strong as agreement. |
| `HTTP Last-Modified` | The response header | **Ignored when it equals the time the server answered (within 5 minutes), or is within 1 hour of it with no time on the page to support it** (what many dynamic sites and CDN caches send); `httpLastModifiedIsServeTime` says so. |
| `sitemap lastmod` | `<lastmod>` and `<news:publication_date>` for the page in the sitemaps listed in robots.txt | Only with `checkSitemap`. |
| Internet Archive | First successful capture in the Wayback Machine | Only with `includeWayback`. Returned as `firstArchivedAt` and `publishedNoLaterThan`, never as the publish date. |

Date formats read: ISO 8601 (with or without time and zone), RFC 2822 (`Thu, 24 Sep 2026 18:10:00 -0400`), `September 24, 2026`, `24 Sep 2026`, `28-Jun-2026`, `2026/09/24 15:42`, `2026年9月25日 10時53分`, `令和8年9月24日`, and US zone names (EDT, PST, ...). Dates later than now and years before 1991 are ignored.

### How the date is chosen

1. For each family, one date is taken (the one that describes this page).
2. The publish date comes from the first family in this order: JSON-LD, meta tag, microdata, `<time>`, date element, visible text, URL. The last-modified date: JSON-LD, meta tag, microdata, visible text, `<time>`, date element, sitemap, HTTP Last-Modified.
3. If that date is contradicted by two other strong sources that agree with each other and supported by none, the majority wins (and the reason says so).
4. If two sources give the same clock time with different time zones, the one whose instant another source supports is used.
5. **Confidence:**
   - `high`: another family agrees and no strong source disagrees. "Agrees" means within 2 hours when both have a time and zone, otherwise the same date as written on the page (a URL, sitemap or HTTP date may be one day off, because those are often in UTC).
   - `medium`: only one strong source, or more sources agree than disagree.
   - `low`: only weak sources (URL, unlabelled `<time>`, sitemap, HTTP header), only a month, or the sources disagree.
6. A last-modified date that equals the time the server answered (within 2 minutes) is treated as generated on each load and not used.
7. Contradictions are listed in `conflicts` (for example, an update date earlier than the publish date, or an Internet Archive capture older than the stated publish date).

**Time zones are kept as the page wrote them.** `2026-08-08T01:40:17+09:00` stays in Japan time (converting it to UTC would make it look like August 7). A time with no zone is returned without a `Z` and `publishedTimezoneKnown` is `false`. A date without a time has `publishedTimezoneKnown: null`.

### Measured on real pages

49 URLs collected on 2026-09-25 (news: BBC, Ars Technica, The Verge, TechCrunch, Engadget, Al Jazeera, Wired, NHK, Asahi, ITmedia, Impress, GIGAZINE; blogs: Cloudflare, GitHub, Apify, dev.to, Smashing Magazine, Martin Fowler, Simon Willison, Qiita, Zenn, Publickey; government: GOV.UK, NASA, CDC, IRS, 厚生労働省, デジタル庁; shops: IKEA, Allbirds, Gymshark, Patagonia, Apple, Steam, 価格.com, 楽天ブックス; reference: Wikipedia, MDN, arXiv, YouTube).

- **Right day, every time it answered:** 28 of the pages have an independent answer (the publish date in the site's own RSS/Atom feed, or a known record such as the arXiv submission date). A publish date was returned for 27 of the 28, and **all 27 are the right day**. The 28th (a 厚生労働省 press release) has no date anywhere in its HTML, and the row says so instead of guessing.
- **Per source** (pages where the source had a publish date / of those, the right day): JSON-LD 18/18, meta tag 14/14, date element 12/12, visible text 7/7, microdata 1/1, `<time>` 15/14, URL 7/6 (the miss is a month-only URL, `/2026/09/`).
- **What a reader sees:** in a browser, the chosen publish date appears as written on the page for 22 of 29 dated pages. Of the other 7: 2 show only a relative time ("2 hours ago", "16 years ago"), 1 showed the browser a consent wall, 3 publish dates exist only in the page's metadata (Wikipedia's article creation date, CDC's first-published date), and 1 (Qiita) shows a posted date one day later than its JSON-LD, which is why that row is `low` with both dates in the reason.
- **Run on Apify** (one run, 49 URLs, 70 seconds, 256MB): 40 pages read, 33 with a date, 7 with none (mostly shop product pages, which rarely carry a date), 9 not read (2 check pages or refusals, 1 disallowed by robots.txt, 2 robots.txt that did not answer, 2 no longer existing, 2 no answer).
- **Sitemaps** added a date for only 2 of 15 pages in a separate test (both already had one), which is why `checkSitemap` is off by default.

### Input

| Field | Type | Default | What it does |
|---|---|---|---|
| `urls` | array of strings | (empty: the example `https://blog.apify.com/fixing-actors-invisible-to-agents/` is used) | Pages to date, 1 to 1,000 per run. A URL without `https://` is read as `https://`. The same URL given twice is read and charged once. |
| `includeWayback` | boolean | `false` | Also ask the Internet Archive when it first saved each URL (`firstArchivedAt`, `publishedNoLaterThan`). Slower (1 to 60 seconds per URL) and charged separately, only when a capture is found. |
| `checkSitemap` | boolean | `false` | Also look for the page's `<lastmod>` in the sitemaps listed in robots.txt (at most 4 sitemap files per page). |
| `includeEvidence` | boolean | `true` | Return every date found in `evidence`. Turn off for smaller rows. |
| `onlyChanges` | boolean | `false` | Monitor mode: return only pages whose chosen dates changed since the last run with the same watch name. |
| `watchName` | string | (none) | Name of the remembered state for monitor mode (letters, digits, `.`, `-`, `_`; up to 40). |
| `resetMonitoringState` | boolean | `false` | Forget what this watch remembered before this run. |
| `maxConcurrency` | integer | `4` | Pages read in parallel (1 to 8). |
| `requestTimeoutSecs` | integer | `20` | Time limit per request (5 to 60 seconds). |

### Output

One row per URL. Main columns:

| Column | Meaning |
|---|---|
| `status` | `ok`, or why there is no date: `no-date-found`, `robots-disallowed`, `robots-unreachable`, `login-required`, `blocked`, `not-found`, `http-error`, `unreachable`, `unreadable`, `bad-input`, `budget-reached`, `no-change` |
| `publishedDate`, `modifiedDate` | The chosen dates (ISO 8601, time zone as written on the page), or `null` |
| `publishedConfidence`, `modifiedConfidence` | `high`, `medium` or `low` |
| `publishedSource`, `modifiedSource` | Where the chosen date was read (`JSON-LD`, `meta tag`, `microdata`, `<time> element`, `date element`, `visible text`, `URL`, `sitemap lastmod`, `HTTP Last-Modified`) |
| `publishedReason`, `modifiedReason` | Which field, what agrees, what disagrees, time zone notes |
| `publishedTimezoneKnown`, `modifiedTimezoneKnown` | `true`, `false` (time without zone) or `null` (date only) |
| `daysSincePublished`, `daysSinceModified` | Age in days at the time of the check |
| `publishedNoLaterThan`, `firstArchivedAt`, `firstArchiveUrl`, `waybackStatus` | Internet Archive first capture (`found`, `no-capture`, `archive-unavailable`, `not-requested`) |
| `sitemapLastmod`, `sitemapUrl` | With `checkSitemap` |
| `httpLastModified`, `httpLastModifiedIsServeTime` | The header and whether it is only the time the server answered |
| `sourcesFound` | Families that had at least one usable date |
| `evidence` | Every date found: `source`, `field`, `dateType` (`published`, `modified`, `unknown`), `value`, `raw`, `timezoneKnown`, `precision`, `usedFor`, `ignored`, `detail` |
| `conflicts` | Contradictions and limits (for example "The Internet Archive saved this address on 2023-03-22, before the publish date the page gives") |
| `pageTitle`, `pageLanguage`, `canonicalUrl`, `finalUrl`, `redirects`, `httpStatus`, `contentType`, `htmlTruncated` | About the page that was read |
| `changeType`, `previousPublishedDate`, `previousModifiedDate`, `previousCheckedAt`, `watchName` | Monitor mode (`first-check`, `new`, `changed`, `unchanged`, `unconfirmed`) |
| `note`, `checkedAt` | Why a row has no date; when the page was read |

### What it will not do

- **It does not invent a date.** A page without any date comes back as `no-date-found` (free). An unlabelled date that is simply today's date or the current time (a clock or a "today" display in the page header) is not used unless a labelled date on the page agrees. An HTTP Last-Modified header that is only the time the page was built or cached for the request is not used.
- **It respects robots.txt** for every URL (RFC 9309; its own name is `NeverEmptyPageDate`, otherwise the `*` group). Disallowed URLs are not requested.
- It does not sign in, does not solve or bypass check pages (Cloudflare, DataDome, AWS WAF, ...) and does not switch to proxies to get around a refusal. Such pages come back as a free row that says why.
- It does not run page JavaScript. Dates that only appear after scripts run cannot be seen.
- It reads up to 5MB of HTML per page. Non-HTML addresses (PDF, images) are checked for a URL date and the Last-Modified header only.

### Monitor mode

With `onlyChanges: true` and a `watchName`, the first run returns every page as the starting point; later runs return only pages whose chosen publish date or last-modified date changed (a change is re-read after 1.5 seconds and reported only if both reads agree). A run in which nothing changed returns one free row and charges only the run start fee.

With a `watchName` but without `onlyChanges`, every page is returned as usual and `changeType` tells you what happened since the last run; a page whose dates changed but differed again on the second read is returned with `changeType: "unconfirmed"` and is not remembered, so the next run checks it again.

### Pricing

Pay per event:

- **Run start** (`actor-start`): once per run that returns at least one dated page (in monitor mode: once per run that could read and compare).
- **Date found** (`date-found`): per URL for which a publish or last-modified date was returned.
- **Internet Archive capture** (`wayback-first-capture`): per URL for which you asked the Internet Archive and it had a capture.

Rows without a date (no date on the page, robots.txt, sign-in, check page, 404, unreachable, bad input, budget) are free. If the maximum total charge you set has no room for the start fee plus one page, nothing is requested and nothing is charged.

### Calling it from code

```bash
curl -X POST "https://api.apify.com/v2/acts/neverempty~page-publish-date-finder/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"urls":["https://arstechnica.com/health/2026/09/cdc-opens-state-ordering-for-covid-19-vaccines-after-unexplained-delay/"],"includeEvidence":false}'
```

### Support

Found a page where the date is wrong? Open an issue in the Issues tab with the URL and the date you expected. This is an unofficial tool that reads only what the page publicly shows.

# Actor input Schema

## `urls` (type: `array`):

Pages to date, one per line (1 to 1,000 per run). A URL without http:// or https:// is read as https://. The same URL given twice is read and charged once. If this is empty, the example page https://blog.apify.com/fixing-actors-invisible-to-agents/ is used.

## `includeWayback` (type: `boolean`):

Look up the first successful capture of each URL in the Internet Archive Wayback Machine (official CDX API). It tells you the page existed no later than that date (publishedNoLaterThan), which helps when the page itself has no date, and it flags pages whose stated publish date is later than their first capture. Slower (the archive takes 1 to 60 seconds per URL) and charged separately, only when a capture is found.

## `checkSitemap` (type: `boolean`):

Look for the page's <lastmod> (and Google News publication date) in the sitemaps listed in robots.txt, opening at most 4 sitemap files per page. Off by default: in a test of 15 pages the sitemap added a date for only 2, and both already had a date on the page, while large sites can have sitemaps of tens of megabytes.

## `includeEvidence` (type: `boolean`):

Return the list of every date found on the page with where it came from (JSON-LD field, meta tag name, <time> element, visible label, URL pattern, HTTP header). Turn off for smaller rows with only the chosen dates.

## `onlyChanges` (type: `boolean`):

Return only pages whose chosen publish date or last-modified date changed since the last run with the same watch name (plus pages new to the watch). The first run returns every page as the starting point. A change is re-read after 1.5 seconds and reported only if the second read shows the same dates. A run in which nothing changed returns a free row saying so and charges only the run start fee.

## `watchName` (type: `string`):

Name of the remembered state used to compare runs (letters, digits, dot, dash, underscore; up to 40). Setting it (or turning on monitor mode) fills changeType and the previous dates. Use a different name for each list of pages you track on its own schedule. With monitor mode on and no name, the name "default" is used.

## `resetMonitoringState` (type: `boolean`):

Start this watch over: forget the remembered dates before this run, so every page is returned as a first check.

## `maxConcurrency` (type: `integer`):

How many pages are read in parallel (1 to 8).

## `requestTimeoutSecs` (type: `integer`):

How long to wait for one page to answer (5 to 60 seconds). HTTP 429, 500, 502, 503 and 504 are asked again up to two more times; a page that does not answer in time is asked once more.

## Actor input object example

```json
{
  "urls": [
    "https://blog.apify.com/fixing-actors-invisible-to-agents/"
  ],
  "includeWayback": false,
  "checkSitemap": false,
  "includeEvidence": true,
  "onlyChanges": false,
  "resetMonitoringState": false,
  "maxConcurrency": 4,
  "requestTimeoutSecs": 20
}
```

# Actor output Schema

## `results` (type: `string`):

One row per URL: the most likely publish date and last-modified date, each with a confidence (high, medium, low), where it was read from and why it was chosen, plus every date found on the page (evidence) and any conflicts. A page with no date, a page that robots.txt does not allow, that needs a sign-in or that shows a check page comes back as a free row that says why.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://blog.apify.com/fixing-actors-invisible-to-agents/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("neverempty/page-publish-date-finder").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://blog.apify.com/fixing-actors-invisible-to-agents/"] }

# Run the Actor and wait for it to finish
run = client.actor("neverempty/page-publish-date-finder").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://blog.apify.com/fixing-actors-invisible-to-agents/"
  ]
}' |
apify call neverempty/page-publish-date-finder --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,neverempty/page-publish-date-finder"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/a6kg2CcY5yAZQ4AIQ/builds/83NguLAv8JaI3mUNx/openapi.json
