# New York Times Scraper: Articles, Feeds and Historical Archive (`getascraper/nytimes-scraper`) Actor

Extract full New York Times articles by section feed, historical date archive back to 1970, or specific URL. Get headline, byline, authors, tags, images, word count and full text in one dataset. Export to Google Sheets, Slack or your API. Skip juggling two separate scrapers. $0.0132 per article.

- **URL**: https://apify.com/getascraper/nytimes-scraper.md
- **Developed by:** [GetAScraper](https://apify.com/getascraper) (community)
- **Categories:** News, Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $9.90 / 1,000 articles

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## 🗞️ New York Times Scraper: Articles, Feeds and Historical Archive

<table width="100%">
<tr>
<td style="padding:24px 28px;background:#F1F5F9;border:1px solid #CBD5E1;border-top:4px solid #1F2937;border-radius:12px">
<span style="font-size:23px;font-weight:800;color:#1C1917;line-height:1.3">Every New York Times article, past or present, in one Actor</span><br>
<span style="font-size:15px;color:#57534E;line-height:1.6">Pull today's coverage from live section feeds or walk the full archive back to 1970. Real article text, not a summary.</span>
</td>
</tr>
</table>

<table width="100%">
<tr>
<td style="padding:14px 12px;width:25%;background:#FFFFFF;border:1px solid #CBD5E1;border-radius:10px 0 0 10px;vertical-align:top">
<span style="font-size:15px;font-weight:800;color:#1F2937">🔀 Real-time and archive, one Actor</span><br>
<span style="font-size:12px;color:#57534E">Today's news and any date back to 1970, without running two separate tools.</span>
</td>
<td style="padding:14px 12px;width:25%;background:#FFFFFF;border:1px solid #CBD5E1;border-left:none;vertical-align:top">
<span style="font-size:15px;font-weight:800;color:#1F2937">📄 The real article, not a snippet</span><br>
<span style="font-size:12px;color:#57534E">Full text, paragraph by paragraph, not a truncated preview.</span>
</td>
<td style="padding:14px 12px;width:25%;background:#FFFFFF;border:1px solid #CBD5E1;border-left:none;vertical-align:top">
<span style="font-size:15px;font-weight:800;color:#1F2937">📊 19 fields per article</span><br>
<span style="font-size:12px;color:#57534E">Word count, paragraph count, and a flag for thin extractions, on top of the usual headline and byline.</span>
</td>
<td style="padding:14px 12px;width:25%;background:#FFFFFF;border:1px solid #CBD5E1;border-left:none;border-radius:0 10px 10px 0;vertical-align:top">
<span style="font-size:15px;font-weight:800;color:#1F2937">🏷️ Sub-brand aware</span><br>
<span style="font-size:12px;color:#57534E">Tags whether a story ran under the main paper, The Athletic, Wirecutter, or Cooking.</span>
</td>
</tr>
</table>

Extract [New York Times](https://www.nytimes.com) articles by section, date archive, or specific URL. Export to JSON, CSV or Excel, or connect straight into Google Sheets and your own pipeline via the API. Run it on demand or on a schedule. No coding required.

### ✨ Why use this Actor

**Built for anyone who needs New York Times coverage without stitching two scrapers together.**

- 📚 **Media researchers and journalists**: track how the Times covered a story over the years, not just what ran this morning. The same Actor pulls live coverage and the historical archive.
- 📣 **PR, SEO and brand monitoring teams**: catch every mention of your brand or client across NYT's sections, business, tech, sports, culture, the moment it publishes.
- 🤖 **NLP and data teams**: get the real article text your model needs to learn from, not a two-sentence summary, with accurate word and paragraph counts.

**One Actor instead of two.** Other NYT scrapers split real-time scraping and the historical archive into separate Actors, or skip the archive entirely. This Actor covers both in one run, plus 19 output fields including a sub-brand tag (NYT main, The Athletic, Wirecutter, Cooking) and a thin-extraction quality flag that no competing Actor exposes.

### ⚙️ How it works

<table width="100%">
<tr>
<td style="padding:16px 14px;width:33%;background:#F1F5F9;border:1px solid #CBD5E1;border-radius:10px 0 0 10px;vertical-align:top">
<span style="font-size:12px;font-weight:800;color:#1F2937;letter-spacing:1px">STEP 1</span><br>
<span style="font-size:14px;font-weight:700;color:#1C1917">Pick your source</span><br>
<span style="font-size:12px;color:#57534E">Choose a section to watch, a date range to walk, or paste specific article URLs.</span>
</td>
<td style="padding:16px 14px;width:33%;background:#F1F5F9;border:1px solid #CBD5E1;border-left:none;vertical-align:top">
<span style="font-size:12px;font-weight:800;color:#1F2937;letter-spacing:1px">STEP 2</span><br>
<span style="font-size:14px;font-weight:700;color:#1C1917">Run the Actor</span><br>
<span style="font-size:12px;color:#57534E">It pulls each article's real text and metadata as the Times itself publishes it.</span>
</td>
<td style="padding:16px 14px;width:33%;background:#F1F5F9;border:1px solid #CBD5E1;border-left:none;border-radius:0 10px 10px 0;vertical-align:top">
<span style="font-size:12px;font-weight:800;color:#1F2937;letter-spacing:1px">STEP 3</span><br>
<span style="font-size:14px;font-weight:700;color:#1C1917">Get your data</span><br>
<span style="font-size:12px;color:#57534E">Headline, byline, dates, tags, images and full text land in your dataset, ready to export or pipe downstream.</span>
</td>
</tr>
</table>

### 📥 Input

| Field | Type | Required | Description |
|---|---|---|---|
| `mode` | enum | No | "section" for live feeds, "archive" for a historical date range, or "urls" for specific articles. |
| `sections` | array of strings | No | NYT section feed names to watch, used when mode is "section". Defaults to the homepage feed. |
| `urls` | array of URLs | No | Specific article URLs to scrape, used when mode is "urls". |
| `dateFrom` | string | No | Start date (YYYY-MM-DD), used when mode is "archive". |
| `dateTo` | string | No | End date (YYYY-MM-DD), used when mode is "archive". |
| `fetchFullText` | boolean | No | Fetch each article's full body, authors, tags and images. On by default. Turn off in archive mode for a fast, free index-only sweep. |
| `maxItems` | integer | No | Stop after collecting this many articles. |
| `proxyConfiguration` | object | No | Proxy settings. Only used when fetching full text; discovery alone needs no proxy. |

### 📤 Output

Every result is one row in the dataset. A typical article looks like this:

```json
{
  "title": "Wendell Berry, Writer Who Extolled America's Agrarian Past, Dies at 92",
  "url": "https://www.nytimes.com/2026/08/31/us/wendell-berry-dead.html",
  "byline": "By Robert D. McFadden",
  "authors": ["Robert D. McFadden"],
  "publishedAt": "2026-09-01T00:07:13.000Z",
  "section": "U.S.",
  "tags": ["Agriculture and Farming", "Writing and Writers", "Poetry and Poets"],
  "wordCount": 1680,
  "paragraphCount": 34,
  "bodyUnusuallyShort": false,
  "subBrand": "The New York Times",
  "imageUrls": ["https://static01.nyt.com/images/2026/08/31/multimedia/31berry-wendell-tpjh/31berry-wendell-tpjh-videoSixteenByNineJumbo1600.jpg"]
}
```

Download the dataset in JSON, CSV, Excel, HTML or XML from the Apify Console, or pull it through the API.

### 📊 Data table

| Field | Type | Description |
|---|---|---|
| `title` / `url` | string | Headline and the article's own link. |
| `byline` / `authors` | string / array | Writer credit, as a string and as a list of names. |
| `publishedAt` / `modifiedAt` | string | ISO publish and last-update timestamps. |
| `section` / `subsection` | string | Where the article sits on the Times. |
| `tags` | array | The Times' own keyword tags for the story. |
| `summary` | string | The article's own dek or description. |
| `imageUrls` | array | Photo URLs published with the article. |
| `wordCount` / `paragraphCount` | number | Real length, computed from the actual article body. |
| `articleBody` | string | Full article text. |
| `bodyUnusuallyShort` | boolean | Flags a result whose body came back far shorter than a normal article, usually a sign of a thin extraction. |
| `archiveListingDate` | string | The archive day this article was listed under, when found via archive mode. |
| `subBrand` | string | Which masthead the story ran under: the main paper, The Athletic, Wirecutter, or Cooking. |

The Output tab also ships three pre-built views: an articles overview, full text, and media & tags.

### 💰 Pricing

This Actor is pay per result: you only pay for the articles you actually collect, and a run that returns nothing costs nothing. There is no subscription and no minimum spend.

### ⭐ Enjoying New York Times Scraper?

<table width="100%" style="display:table;width:100%">
<tr>
<td style="padding:20px 24px 14px;background:#F1F5F9;border:1px solid #CBD5E1;border-left:5px solid #1F2937;border-radius:10px 10px 0 0">
<span style="font-size:20px;letter-spacing:4px">⭐ ⭐ ⭐ ⭐ ⭐</span><br>
<span style="font-size:17px;font-weight:800;color:#1C1917">Saved you from stitching together two scrapers?</span><br>
<span style="font-size:14px;color:#57534E">A 5-star rating takes 10 seconds and helps other researchers and brand teams find it. Your feedback also tells us what to build next.</span>
</td>
</tr>
<tr>
<td style="padding:0;background:#1F2937;border:1px solid #CBD5E1;border-top:none;border-radius:0 0 10px 10px;text-align:center">
<a href="https://apify.com/getascraper/nytimes-scraper/reviews" style="display:block;padding:13px 16px;color:#FFFFFF;text-decoration:none;font-weight:800;font-size:15px;letter-spacing:0.3px">★&nbsp;&nbsp;Rate this Actor on Apify</a>
</td>
</tr>
</table>

### 🛠️ Tips for better runs

- Use `mode: section` with a schedule to monitor new coverage on a topic or brand as it publishes.
- Use `mode: archive` with `fetchFullText` off for a fast, free index of everything published in a date range, then re-run specific URLs with full text once you know which ones you need.
- Combine sections from different mastheads (`Sports`, `Automobiles`, `DiningandWine`) in one run to cover NYT's full range in a single Actor call.

### ❓ FAQ

**Does this get around the New York Times paywall?**
No. This Actor extracts the same article content the Times itself serves on its public pages. It does not unlock subscriber-only features or bypass any access control beyond what a public visitor already sees.

**Is it legal to scrape nytimes.com?**
This Actor only collects data from New York Times pages. You are responsible for how you use the data and for complying with the New York Times' terms of service.

**Why do some articles have no image or tags?**
Not every article publishes every field. This Actor never invents one: if the Times doesn't publish it, the field is left out rather than filled with a placeholder.

**Can I get notified of new articles automatically?**
Yes. Schedule this Actor to run daily or hourly from the Apify Console and pipe new results into Google Sheets, Slack, or your own database with no code.

Found a bug or need a custom version of this Actor? Open an issue from the Actor's Issues tab and it'll be looked at directly.

### 🔗 Other actors

- [BookMyShow Scraper: Movies, Showtimes & Concerts](https://apify.com/getascraper/bookmyshow-scraper) ↗ - India's major ticketing platform, movies and live events.
- [District Event Scraper: Venue, Dates & Ticket Pricing](https://apify.com/getascraper/district-event-scraper) ↗ - India event listings with full ticket-tier pricing.
- [CryptoPanic News Scraper: No Login, No API Key](https://apify.com/getascraper/cryptopanic-news-scraper) ↗ - crypto news aggregation, another news-monitoring Actor.
- [Cashify Scraper: Used Phone Resale Prices India](https://apify.com/getascraper/cashify-scraper) ↗ - India consumer marketplace pricing data.

# Actor input Schema

## `mode` (type: `string`):

"section": pull the newest articles from one or more NYT section feeds. "archive": walk NYT's own date archive over a date range. "urls": scrape a specific list of article URLs.

## `sections` (type: `array`):

NYT section feed names, used when mode is "section". Confirmed live values include: HomePage, World, US, Politics, Business, Technology, Science, Health, Sports, Arts, Movies, Music, Travel, and 45+ more (see nytimes.com/rss for the full current list). Case-sensitive, exact feed name.

## `urls` (type: `array`):

Specific nytimes.com article URLs to scrape, used when mode is "urls".

## `dateFrom` (type: `string`):

Start date (YYYY-MM-DD), used when mode is "archive". NYT's own archive covers back to 1970.

## `dateTo` (type: `string`):

End date (YYYY-MM-DD), used when mode is "archive".

## `fetchFullText` (type: `boolean`):

Fetch each article's full body text, authors, tags, images and word count, instead of just headline and URL. Fetching full text requires Apify's UNBLOCKER proxy, which has a real per-request cost - discovery alone (RSS feeds, the date archive index) is free.

## `maxItems` (type: `integer`):

Stop after collecting this many articles. Keep this low for a quick, cheap run; raise it once you know how many you need.

## `proxyConfiguration` (type: `object`):

Proxy settings. Only used when fetching full article text - discovery alone needs no proxy. Defaults to Apify's UNBLOCKER proxy group, confirmed live as the only mechanism that clears NYT's anti-bot protection on article pages.

## Actor input object example

```json
{
  "mode": "section",
  "sections": [
    "HomePage"
  ],
  "urls": [],
  "dateFrom": "",
  "dateTo": "",
  "fetchFullText": true,
  "maxItems": 20,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "UNBLOCKER"
    ]
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mode": "section",
    "sections": [
        "HomePage"
    ],
    "urls": [],
    "dateFrom": "",
    "dateTo": "",
    "fetchFullText": true,
    "maxItems": 20,
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "UNBLOCKER"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("getascraper/nytimes-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "mode": "section",
    "sections": ["HomePage"],
    "urls": [],
    "dateFrom": "",
    "dateTo": "",
    "fetchFullText": True,
    "maxItems": 20,
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["UNBLOCKER"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("getascraper/nytimes-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mode": "section",
  "sections": [
    "HomePage"
  ],
  "urls": [],
  "dateFrom": "",
  "dateTo": "",
  "fetchFullText": true,
  "maxItems": 20,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "UNBLOCKER"
    ]
  }
}' |
apify call getascraper/nytimes-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,getascraper/nytimes-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/pYfoxgj3q4zuJOyg5/builds/Ea1pFztMrSl2nxuae/openapi.json
