# Article Text Extractor — Clean Content, Authors & Dates (`dallel-data/article-text-metadata-extractor`) Actor

Extract readable article text, titles, authors, dates and publisher metadata from public URLs. Batch news and blog pages, filter by keyword, and export JSON or CSV.

- **URL**: https://apify.com/dallel-data/article-text-metadata-extractor.md
- **Developed by:** [GUIR Dallel](https://apify.com/dallel-data) (community)
- **Categories:** News
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.70 / 1,000 delivered results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Article Text Extractor — Clean Content, Authors & Dates

Extract readable article text, titles, authors, dates and publisher metadata from public URLs. Batch news and blog pages, filter by keyword, and export JSON or CSV.

### Quick start

Paste your targets into the input form and run. The example below fetches a small sample:

```json
{
  "targets": [
    "https://blog.apify.com/web-scraping-python/"
  ],
  "maxResults": 10,
  "maxItemsPerTarget": 10,
  "maxPages": 3
}
```

The dataset supports JSON, CSV and Excel export. Use the API, schedules, Make, n8n or other Apify integrations to repeat the same collection. No external API key, browser cookies or login is required by this Actor.

### Input and limits

| Field | Meaning |
|---|---|
| `targets` | 1–50 source URLs. See README for supported URL formats. |
| `maxResults` | Stops the run after this many unique matching results. Always set a sensible cap. |
| `maxItemsPerTarget` | Maximum delivered results per input target, after filters. |
| `maxPages` | Maximum pages per target. Contact extraction follows same-host contact/about/support pages. Shopify, Lever and SmartRecruiters paginate. Ignored for single-page article/product extraction and single-response Ashby/Greenhouse boards. |
| `proxyConfiguration` | Optional Apify proxy settings. Default requests connect directly. Proxy costs are included in listed per-result pricing; using expensive custom proxies can reduce run efficiency. |
| `keyword` | Optional case-insensitive substring. Matches title/description for jobs and feeds, title/vendor/tags for products, URL for sitemaps. Applied before billing. |

Duplicate targets are processed once. Filters apply before result billing. Runs are bounded snapshots; limits can stop collection before the source is exhausted.

### Output

Each result includes `source`, `id`, `title`, `url` and `fetchedAt` where applicable, plus source-specific fields. The Results view highlights `title`, `publisher`, `author`, `publishedAt`, `wordCount`, `url`, `text`. JSON export retains every field, including nested arrays. Missing source values are null rather than estimated.

The free `SUMMARY` record in the default key-value store reports each target's result count and errors, HTTP requests, and whether the overall result or spending limit was reached. Errors are not inserted into the paid dataset. A run in which every source fails is marked failed; mixed success is reported explicitly in the summary.

### Pricing

$2 per 1,000 delivered results ($0.002 each). No Actor start fee. Filtered rows, duplicates within the same context and source errors are not charged. Results are deduplicated within each run. Google Search preserves a URL appearing in distinct queries. Source requests and the default proxy costs are included in paid per-result pricing. Set Apify's maximum charge per run as an additional budget control.

### Source coverage and limitations

Fetches the supplied URLs only; no link discovery or Google News redirect decoding. Uses Trafilatura heuristic text extraction, not AI. No JavaScript rendering, paywall bypass, login or PDF extraction. Layouts vary and extracted text may retain boilerplate or omit sections. Pages yielding fewer than 40 words are omitted. Metadata dates are heuristic and may be missing or inaccurate; verify against the linked source. Comments are excluded.

Public sources can change, throttle requests or return no matches. The Actor retries transient network failures up to twice, follows a bounded number of redirects, and rejects private-network targets. It does not bypass login or CAPTCHA challenges. Use data within the source's applicable terms and your intended permissions.

### Support

Open an issue on this Actor's Issues tab with the failing public target, input, and run link. Never include passwords, tokens or personal account cookies.

# Actor input Schema

## `targets` (type: `array`):

1–50 source URLs. See README for supported URL formats.

## `maxResults` (type: `integer`):

Stops the run after this many unique matching results. Always set a sensible cap.

## `maxItemsPerTarget` (type: `integer`):

Maximum delivered results per input target, after filters.

## `maxPages` (type: `integer`):

Maximum pages per target. Contact extraction follows same-host contact/about/support pages. Shopify, Lever and SmartRecruiters paginate. Ignored for single-page article/product extraction and single-response Ashby/Greenhouse boards.

## `proxyConfiguration` (type: `object`):

Optional Apify proxy settings. Default requests connect directly. Proxy costs are included in listed per-result pricing; using expensive custom proxies can reduce run efficiency.

## `keyword` (type: `string`):

Optional case-insensitive substring. Matches title/description for jobs and feeds, title/vendor/tags for products, URL for sitemaps. Applied before billing.

## Actor input object example

```json
{
  "targets": [
    "https://blog.apify.com/web-scraping-python/"
  ],
  "maxResults": 25,
  "maxItemsPerTarget": 25,
  "maxPages": 3
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "targets": [
        "https://blog.apify.com/web-scraping-python/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("dallel-data/article-text-metadata-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "targets": ["https://blog.apify.com/web-scraping-python/"] }

# Run the Actor and wait for it to finish
run = client.actor("dallel-data/article-text-metadata-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "targets": [
    "https://blog.apify.com/web-scraping-python/"
  ]
}' |
apify call dallel-data/article-text-metadata-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,dallel-data/article-text-metadata-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/eZFaWsX1m6HsbtTd4/builds/OcbA30aUNgZahgXYg/openapi.json
