# Techmeme News Headlines Scraper (`automation-lab/techmeme-news-headlines-source-links`) Actor

Export current official Techmeme RSS headlines, timestamps, Techmeme story permalinks, and attributed primary publisher article URLs.

- **URL**: https://apify.com/automation-lab/techmeme-news-headlines-source-links.md
- **Developed by:** [Automation Lab](https://apify.com/automation-lab) (community)
- **Categories:** News
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.48 / 1,000 item extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Techmeme News Headlines Scraper

Export **techmeme news headlines** from Techmeme's official current RSS feed. Each dataset row pairs a story's title and publication time with its Techmeme permalink and the primary publisher article URL when Techmeme attributes one. This is a lightweight feed export for daily technology-news monitoring, not a full-site crawl or an article-body extractor.

### Who is it for?

- Editorial teams compile a morning scan of technology stories with original reporting links.
- Researchers compare fresh headline snapshots in their own spreadsheet or warehouse.
- Communications analysts filter the current feed for a topic before routing matches to their own alerts.

### Why use this Actor?

The output exposes both the Techmeme story permalink and its credited primary publisher URL in separate fields. It uses the official RSS feed instead of rendering pages, so there is no browser, login, or proxy configuration. A scheduled Actor run gives you a fresh snapshot; compare successive datasets in your own workflow for additions.

### What data comes back?

| Field | Meaning |
| --- | --- |
| `title` | Headline as supplied in the RSS item. |
| `techmemeUrl` | Canonical Techmeme story permalink from RSS. |
| `publisherUrl` | Primary publisher article URL from the item description's lead-image link; `null` if absent. |
| `publishedAt` | RSS `pubDate` converted to UTC ISO 8601. |
| `fetchedAt` | UTC timestamp of this retrieval. |

Only current RSS items are returned. Related story clusters, article text, images, and historical archive search are outside scope.

### Get started

1. Open the Actor and accept the default `maxItems: 20` for a current snapshot.
2. Optionally set `keyword` to a substring of the headline (case-insensitive).
3. Optionally supply `publishedAfter` as a UTC ISO timestamp to discard older current-feed entries.
4. Run, then download the default dataset as JSON, CSV, or Excel. Schedule repeated runs in Apify if you need periodic snapshots.

### Input parameters

| Parameter | Default | Description |
| --- | --- | --- |
| `maxItems` | `20` | Maximum matching rows from the feed, between 1 and 100. The feed itself may contain fewer. |
| `keyword` | none | Case-insensitive headline substring; no match returns an empty dataset. |
| `publishedAfter` | none | Inclusive UTC ISO cutoff (e.g. `2026-01-01T00:00:00Z`). This is not an archive query. |

Example input for topical monitoring:

```json
{"keyword":"AI","maxItems":10}
```

### Output example

A representative shape of a current-feed record (URLs and headline shortened for illustration):

```json
{
  "title": "PitchBook: VCs have invested $4B+ in quantum computing companies YTD (Financial Times)",
  "techmemeUrl": "https://www.techmeme.com/260926/p7#a260926p7",
  "publisherUrl": "https://www.ft.com/content/ae9eedd2-4530-47e4-be4b-242d0e2a6253",
  "publishedAt": "2026-09-26T10:20:01.000Z",
  "fetchedAt": "2026-09-26T14:04:19.891Z"
}
```

The live publisher link may contain query parameters supplied by the feed. Some articles require a separate publisher subscription; this Actor does not bypass it.

### How much does it cost to export Techmeme headlines?

This Actor uses pay-per-event pricing: one `$0.005` start event per run and one item event for each emitted record. The item unit price depends on your **Apify Store monthly spend tier**, not the number of records in this run:

| Spend tier | Item price |
| --- | ---: |
| FREE | $0.010497 |
| BRONZE | $0.0091276 |
| SILVER | $0.0071195 |
| GOLD | $0.0054766 |
| PLATINUM | $0.0054766 |
| DIAMOND | $0.0054766 |

At BRONZE the estimated charge is $0.050638 for 5 headlines, $0.096276 for 10, or $0.187552 for 20 (start fee included); at GOLD the same examples cost $0.032383, $0.059766, or $0.114532. The current Pricing tab is authoritative. Empty filtered runs do not charge for items, but the start event still applies. These are usage estimates, not guaranteed invoices: refunds, fraud, disputes, taxes, corrections and clawbacks can change final payout or billing. Apify platform execution costs are separate from the event count.

### Integrations and scheduled workflows

Use an Apify Schedule for a daily snapshot, then connect the default dataset to Google Sheets, Make, Zapier, or your own ETL. Store `techmemeUrl` as a stable story key to deduplicate across scheduled runs. Compare `publishedAt` rather than `fetchedAt` when ordering stories; the latter identifies the snapshot retrieval time. The Actor does not maintain an internal history, send notifications, or deliver an alert itself.

### API access

Start an Actor run with your Apify token (replace the token securely, do not embed it in shared files):

```bash
curl -X POST 'https://api.apify.com/v2/acts/automation-lab~techmeme-news-headlines-source-links/runs?token=YOUR_TOKEN' \
  -H 'Content-Type: application/json' -d '{"maxItems":20}'
```

JavaScript:

```js
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('automation-lab/techmeme-news-headlines-source-links').call({ maxItems: 20 });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);
```

Python:

```python
from apify_client import ApifyClient
import os
client = ApifyClient(os.environ['APIFY_TOKEN'])
run = client.actor('automation-lab/techmeme-news-headlines-source-links').call(run_input={'maxItems': 20})
items = client.dataset(run['defaultDatasetId']).list_items().items
print(items)
```

### MCP usage

To expose this Actor to Claude Code via Apify MCP:

```bash
claude mcp add --transport http apify \
  'https://mcp.apify.com?tools=automation-lab/techmeme-news-headlines-source-links'
```

For Claude Desktop, Cursor, and VS Code MCP setup, configure the equivalent remote server (each client's remote HTTP-server configuration UI may differ):

```json
{"mcpServers":{"apify":{"url":"https://mcp.apify.com?tools=automation-lab/techmeme-news-headlines-source-links"}}}
```

Example prompt: “Export the current Techmeme headlines and publisher URLs; show only headlines containing AI.” The assistant still sees only the current RSS window.

### Limits and reliability

The feed is controlled by Techmeme and may contain fewer than your `maxItems`. Feed entries can change or disappear between runs. The headline substring filter is literal, not semantic or full-text search; a no-match filter succeeds with zero records. Malformed upstream RSS, HTTP failures after bounded retries, and non-XML responses fail visibly rather than returning a misleading empty success. The Actor does not crawl linked publisher sites.

### Legality and responsible use

This Actor does not use AI or send feed entries to an AI provider. It fetches only public RSS and writes the selected metadata into the run's default dataset; Apify controls dataset retention and deletion through your account settings. No user-provided credentials are accepted or logged. For questions about a run, open the Actor's Issues tab with a run link (do not paste private tokens).

Techmeme and the publishers own their content. Check their terms and your use case before redistributing headlines or links. This Actor exports public RSS metadata, not paywalled article bodies or personal account information. Publisher links may be gated by subscriptions or expire if they include share tokens.

### Data quality checks

Confirm `techmemeUrl` is a Techmeme permalink and use it for deduplication. A missing publisher link is an explicit `null`, not the Techmeme permalink copied into the publisher column. Filtered feeds may contain fewer rows than the requested cap; inspect your input cutoff before treating a small dataset as a failure.

### Frequently asked questions

#### Why did I get no records?

Remove `keyword` and `publishedAfter` to see the full current feed, then retry with a narrower cutoff. This feed does not provide historical search.

#### Why is a publisher URL null or inaccessible?

A feed item may omit the lead publisher link. A non-null link can still be behind a publisher paywall, expire, or change independently of this Actor.

#### Can I get all related stories in a Techmeme cluster?

No. This product deliberately exports the primary attributed article URL from each headline item, not cluster coverage or article bodies.

### Related automation-lab Actors

For other news sources, see [Reuters Latest News Feed Scraper](https://apify.com/automation-lab/reuters-latest-news-feed-scraper) and [Ground News Bias & Coverage Scraper](https://apify.com/automation-lab/ground-news-bias-coverage-scraper). Those are separate source workflows, not extensions of Techmeme's RSS feed.

# Changelog

This Actor's version history is a separate document: https://apify.com/automation-lab/techmeme-news-headlines-source-links/changelog.md

# Actor input Schema

## `maxItems` (type: `integer`):

Maximum matching headlines to export from the current feed (up to 100; the feed may contain fewer).

## `keyword` (type: `string`):

Case-insensitive substring in the headline. Only current-feed items are searched; no matching items produces an empty dataset.

## `publishedAfter` (type: `string`):

UTC ISO timestamp, inclusive, such as 2026-09-26T00:00:00Z. Only current RSS entries are available.

## Actor input object example

```json
{
  "maxItems": 20
}
```

# Actor output Schema

## `overview` (type: `string`):

Default dataset view with headline, Techmeme permalink, publisher link and UTC timestamps.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxItems": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("automation-lab/techmeme-news-headlines-source-links").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "maxItems": 20 }

# Run the Actor and wait for it to finish
run = client.actor("automation-lab/techmeme-news-headlines-source-links").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxItems": 20
}' |
apify call automation-lab/techmeme-news-headlines-source-links --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,automation-lab/techmeme-news-headlines-source-links"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Q39CPM0rLzy1dsP5N/builds/npakRLPw68FMzhZwr/openapi.json
