# Article Extractor: Full Text, Author & Date from Any News URL (`hydrafetch/article-extractor`) Actor

Extract the full text, author, publish date and site details from any article or blog post URL.

- **URL**: https://apify.com/hydrafetch/article-extractor.md
- **Developed by:** [Hydrafetch](https://apify.com/hydrafetch) (community)
- **Categories:** News, AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.70 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

<div align="center">

<img src="https://hydrafetch.com/brand/mark-512.png" alt="Hydrafetch" height="64" />

## Article Extractor: Full Text, Author & Date from Any News URL

**Extract the full text, author, publish date and site details from any article or blog post URL.**

`Pay per result` · `$3 per 1,000 articles` · `No API key` · Powered by [Hydrafetch](https://hydrafetch.com?utm_source=apify\&utm_medium=readme\&utm_content=article-extractor)

</div>

***

### Article Extractor at a glance

- **Input:** A list of URLs.
- **Output:** one row per article with 12 fields, including title, author, publishedTime, wordCount.
- **Price:** $3 per 1,000 articles, charged only for results.
- **You need:** nothing else. No cookies, no logins, no proxies and no API keys.
- **Export:** JSON, CSV, Excel or XML, or straight into your own tools with the Apify API and MCP.

### What is Article Extractor?

Extracts clean, readable articles from any news site or blog. Each result has the full text as markdown, the title, author and publish date when the page states them, the site name, language, description, lead image and word count. Menus, related links, ads and comment sections are removed. Pages that are not articles are skipped and not charged, and a missing author or date comes back as null rather than a guess. Built for news monitoring, media tracking, AI summarisation and research datasets.

- **Full article text.** The body of the article as clean markdown, without menus, related links, ads or comment sections.
- **Author and date.** Every author and the publish time, left empty when the page does not state them, so your data stays clean.
- **Publication details.** Site name, language, description and lead image.
- **Word count.** For filtering short pieces or budgeting tokens.

### What data can I extract with Article Extractor?

Every result is one row per article, with these fields:

| Field | What it holds |
| --- | --- |
| `url` | The URL you asked for |
| `finalUrl` | Where the URL ended up after redirects |
| `title` | Article headline |
| `author` | Every author, joined into one line |
| `authors` | Every author as a list, in byline order |
| `publishedTime` | Publish date, when the page states it |
| `siteName` | Publication name |
| `language` | Detected language |
| `description` | Summary line from the page |
| `image` | Lead image URL |
| `wordCount` | Words in the article |
| `text` | The full article as clean markdown |

The table view in Apify shows URL, Title, Authors, Published, Words. The full record is in the JSON, CSV, Excel and XML exports, and over the API.

### How to use Article Extractor

1. Open Article Extractor in the Apify Console. A free Apify account is enough to try it.
2. Paste your URLs, one per line, or upload a list.
3. Click **Start**. Each URL is processed on its own, so one bad entry never loses the batch.
4. Download the results as JSON, CSV or Excel, or send them to your own tools with an integration.

### Input

One field: **URLs**. Use full URLs, including `https://`.

Every line is checked before the run starts. Each must be a web address like https://stripe.com/pricing; a run with any other line is refused with that line named, so a typo never costs you anything.

```json
{
  "urls": [
    "https://stripe.com/blog/idempotency",
    "https://github.blog/engineering/architecture-optimization/how-we-improved-push-processing-on-github/"
  ]
}
```

### Output

A article we cannot resolve is skipped rather than returned empty, and you are not charged for it. The run log names every one that was skipped, so a short result is never a mystery.

```json
{
  "url": "https://stripe.com/blog/idempotency",
  "title": "Designing robust and predictable APIs with idempotency",
  "author": "Brandur Leach",
  "authors": [
    "Brandur Leach"
  ],
  "publishedTime": "2017-02-22",
  "siteName": "Stripe",
  "language": "en",
  "wordCount": 1415,
  "text": "**Authors:** Brandur Leach\n\n**Published:** 2017-02-22\n\n# Designing robust and predictable APIs with idempotency\n\nNetworks are unreliable. ..."
}
```

### How much does Article Extractor cost?

Article Extractor costs **$3 per 1,000 articles** returned, which is $0.003 each, with no separate compute charge. 10,000 articles cost $30.

- **You pay only for results.** A article that cannot be resolved is skipped and free.
- **Try it on the free plan.** Each free Apify account can run up to 50 articles through this Actor. On a paid plan, $5 covers about 1,666 articles.
- **Volume discounts.** Scale plans pay 5% less per result and Business plans 10% less.
- **Cap any run.** Set a maximum cost per run in the run options and the Actor stops cleanly when it is reached.

### Common use cases

- **News monitoring.** Turn a list of article links into readable text with dates you can sort by.
- **Media and PR tracking.** Collect coverage of a brand or topic with author and publication.
- **AI summarisation.** Feed clean article text to a model instead of cluttered HTML.
- **Research datasets.** Build a corpus of articles from many publishers in one run.

### Integrate Article Extractor with other apps

Article Extractor works with the integrations on the Apify platform, including [Make](https://apify.com/integrations/make), [Zapier](https://apify.com/integrations/zapier), [n8n](https://apify.com/integrations/n8n), [Google Sheets](https://apify.com/integrations/google-sheets), [Slack](https://apify.com/integrations/slack), [Airbyte](https://apify.com/integrations/airbyte), [LangChain](https://apify.com/integrations/langchain), [LlamaIndex](https://apify.com/integrations/llamaindex). Results can also go anywhere with a [webhook](https://docs.apify.com/platform/integrations/webhooks) when a run finishes.

To keep a list current, save your input as a task and put it on a [schedule](https://docs.apify.com/platform/schedules).

### Article Extractor API

Run Article Extractor from your own code with the Apify API clients.

**JavaScript**

```js
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: 'YOUR_APIFY_TOKEN' });
const run = await client.actor('hydrafetch/article-extractor').call({ urls: ["https://stripe.com/blog/idempotency","https://github.blog/engineering/architecture-optimization/how-we-improved-push-processing-on-github/"] });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);
```

**Python**

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("hydrafetch/article-extractor").call(run_input={"urls": ["https://stripe.com/blog/idempotency","https://github.blog/engineering/architecture-optimization/how-we-improved-push-processing-on-github/"]})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)
```

AI agents can call it too. Connect Claude, Cursor or any MCP client to the [Apify MCP server](https://mcp.apify.com) and add `hydrafetch/article-extractor` as a tool.

### FAQ

#### What if a URL is not an article?

Pages without a real article body, such as homepages or short listings, are skipped and not charged.

#### Why is the author or date sometimes null?

Because the page does not state it. Fields stay empty rather than filled with a guess.

#### What if an article has several authors?

All of them come back, in the order the byline lists them: as a list in authors, and joined into one line in author. That covers posts that credit authors only through their author links, not just ones that declare them in metadata.

#### Which sites does it work on?

Any news site or blog. There is no per-site setup, so a new publisher works the day you add it.

#### Do I need cookies, a login or proxies?

No. Article Extractor works without cookies, accounts, proxies or API keys. Paste your URLs and run it; everything else is handled for you.

#### Is it legal to use Article Extractor?

Article Extractor collects only publicly available article text and metadata. As with any data collection, you are responsible for how you use the results: follow the terms of the websites you work with and data protection laws such as GDPR wherever personal data is involved. If you are unsure about your use case, check with a lawyer.

#### Does Article Extractor have an API?

Yes. Every run is available through the Apify API, with the JavaScript and Python examples above, and through the Apify MCP server for AI agents.

### Other Actors by Hydrafetch

- [Company Enrichment API](https://apify.com/hydrafetch/company-enrichment-api): turn a list of domains into full company records: name, description, logo, brand colors, fonts, socials and industry.
- [Bulk Company Logo Finder](https://apify.com/hydrafetch/bulk-company-logo-finder): give it a list of domains and get a direct image URL for each company logo, with dimensions and dominant color.
- [Website Color Palette Extractor](https://apify.com/hydrafetch/website-color-palette-extractor): read a site's real design system: colors by role with contrast ratios, the type scale, corner radius and button styling.
- [AI SEO & GEO Audit](https://apify.com/hydrafetch/ai-seo-geo-audit): check whether an LLM or AI agent can actually read your pages, and get the specific reasons when it cannot.
- [Website to Markdown](https://apify.com/hydrafetch/website-to-markdown): turn a list of URLs into clean markdown for LLMs, RAG and AI agents, with navigation, banners and boilerplate removed.
- [Company Social Links Finder](https://apify.com/hydrafetch/company-social-links-finder): find the LinkedIn company page and every social profile a company links from its own website, from just the domain.
- [LinkedIn Company Scraper](https://apify.com/hydrafetch/linkedin-company-scraper): get industry, company size, headcount, followers, headquarters and more from LinkedIn company pages, by URL or by domain.
- [Tech Stack Detector](https://apify.com/hydrafetch/website-tech-stack-detector): find the CMS, ecommerce platform, analytics, frameworks, hosting and payment tools any website runs, from just the domain.
- [PDF to Markdown](https://apify.com/hydrafetch/pdf-to-markdown): convert PDF links into clean markdown with headings and tables intact, ready for LLMs, RAG and search.
- [Website Crawler to Markdown](https://apify.com/hydrafetch/website-crawler-to-markdown): crawl a whole website or docs section into clean markdown for LLMs, RAG and AI agents, one row per page.
- [Company Jobs Scraper](https://apify.com/hydrafetch/company-jobs-scraper): get every open job at a company from just its domain: titles, teams, locations, pay and full descriptions, from the hiring platform it uses.
- [YouTube Transcript Scraper](https://apify.com/hydrafetch/youtube-transcript-scraper): get the full transcript of any YouTube video with timestamps, title and channel, ready for AI summaries, RAG and search.
- [Google Ads Library Scraper](https://apify.com/hydrafetch/google-ads-library-scraper): get every Google ad a company runs from just its domain: creatives, formats and first and last shown dates, from the public Ads Transparency Center.
- [Company Website Finder](https://apify.com/hydrafetch/company-website-finder): turn a list of company names into their official websites and domains, each checked against the company homepage so directories never slip through.
- [Website Screenshot API](https://apify.com/hydrafetch/website-screenshot-api): capture screenshots of any web pages as hosted PNG links, full page or first screen, rendered in a real browser and priced per screenshot.

### Terms of use

**Your results are yours.** We claim no rights in the inputs you submit or the data you get back. Full terms: [hydrafetch.com/terms](https://hydrafetch.com/terms?utm_source=apify\&utm_medium=readme\&utm_content=article-extractor)

### Your feedback

Something not working, or a field you need? Open an issue on the Issues tab with the input you used, and we read every one.

***

<div align="center">
Built by <a href="https://hydrafetch.com?utm_source=apify&utm_medium=readme&utm_content=article-extractor">Hydrafetch</a>. Clean web data for developers and agents.
</div>

# Actor input Schema

## `urls` (type: `array`):

The URLs to process, one per line. Full URLs including https://.

## Actor input object example

```json
{
  "urls": [
    "https://stripe.com/blog/idempotency",
    "https://github.blog/engineering/architecture-optimization/how-we-improved-push-processing-on-github/"
  ]
}
```

# Actor output Schema

## `results` (type: `string`):

Article Extractor: Full Text, Author & Date from Any News URL records (the run's default dataset).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://stripe.com/blog/idempotency",
        "https://github.blog/engineering/architecture-optimization/how-we-improved-push-processing-on-github/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("hydrafetch/article-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://stripe.com/blog/idempotency",
        "https://github.blog/engineering/architecture-optimization/how-we-improved-push-processing-on-github/",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("hydrafetch/article-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://stripe.com/blog/idempotency",
    "https://github.blog/engineering/architecture-optimization/how-we-improved-push-processing-on-github/"
  ]
}' |
apify call hydrafetch/article-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,hydrafetch/article-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/mke71rd97mU0VoDLW/builds/5KZa85HmVE7gTFa97/openapi.json
