# CoinDesk Scraper (`publicmoney/coindesk-scraper`) Actor

Extract CoinDesk policy, regulation and market-structure coverage as structured news: headline, summary, author, categories, an ISO 8601 publication time and the assets each story mentions. Export data, run via API, schedule and monitor runs, or integrate with other tools.

- **URL**: https://apify.com/publicmoney/coindesk-scraper.md
- **Developed by:** [Public Money](https://apify.com/publicmoney) (Apify)
- **Categories:** Business
- **Stats:** 4 total users, 3 monthly users, 100.0% runs succeeded, 1 bookmarks
- **User rating**: 5.00 out of 5 stars

## Pricing

from $1.00 / 1,000 records

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

CoinDesk is the original crypto trade paper and the desk institutional readers still watch on policy, regulation and market structure. This Actor turns its feed into structured records: headline, summary, author, categories and a real ISO 8601 timestamp, each tagged with the assets the story actually names so you can route a headline to the position it affects.

### What it does

- Covers the **institutional end of crypto coverage**: policy, regulation, ETFs and market structure, which is where CoinDesk is strongest.
- Tags **which assets each story is about** in `mentionedTickers`, resolved against a live CoinGecko coin list and the SEC company ticker file rather than a hardcoded table, with a guard so an ordinary word like "optimism" does not ship `OP` as a mention.
- Converts the feed's RFC 822 date to a real **ISO 8601 `datePublished`**, so stories sort and filter correctly instead of arriving as "Sat, 06 Sep 2026 05:30:00 GMT" strings.
- Returns the **summary body** in `articleBody`, decoded from the feed's HTML entities, so a headline arrives with enough context to classify without a second fetch.
- Carries the feed's own `guid`, so a scheduled run can drop stories it has already seen instead of reprocessing the whole feed.

### Use cases

| You need to | How this Actor does it |
| --- | --- |
| Route a headline to a position | Filter on `mentionedTickers` and alert only on the assets you hold |
| Watch regulation and policy | Filter `categories` on the policy tags and schedule the run hourly |
| Build a news timeline for an asset | Accumulate runs and group by `mentionedTickers` and `datePublished` |
| Feed a research agent | Call the Actor over MCP and let the model read the day's stories |
| Deduplicate a feed | Key your store on `guid` and only process what is new |
| Join news to prices | Run the CoinGecko or CoinMarketCap Actor on the same schedule and match on ticker |

### Quick start

1. Click **Try for free**.
2. Leave **Feeds** on `all`. CoinDesk publishes a single feed covering the whole newsroom.
3. Set **Articles per feed** to control how many stories come back.
4. Click **Start**. Rows appear within seconds.
5. Export as JSON, CSV, Excel or XML, or read the dataset over the API.

### Input

| Field | Type | Default | What it controls |
| --- | --- | --- | --- |
| `feeds` | array | `all` | Which CoinDesk feed to read. CoinDesk publishes one, so `all` is the whole newsroom |
| `articlesPerFeed` | integer | `50` | How many stories each feed returns, counted from the newest |
| `maxItems` | integer | `0` | Caps how many records are written. `0` writes them all |

```json
{
    "articlesPerFeed": 50,
    "maxItems": 0
}
```

### Output

One dataset item per story. Fields the feed does not carry for a story are dropped rather than returned as `null`, so a story with no byline carries no `author`.

| Field group | Fields |
| --- | --- |
| Story | `status`, `headline`, `articleBody`, `url`, `guid` |
| Attribution | `publisher`, `author`, `categories`, `feed` |
| Assets | `mentionedTickers` |
| Timing | `datePublished`, `scrapedAt` |

```json
{
    "status": "ok",
    "headline": "Bitcoin ETF inflows top $1bn as institutional demand returns",
    "articleBody": "Spot bitcoin ETFs recorded their strongest week of net inflows since March, with issuers reporting sustained institutional allocation.",
    "publisher": "CoinDesk",
    "author": "Helene Braun",
    "categories": [
        "Markets",
        "Policy"
    ],
    "feed": "all",
    "mentionedTickers": [
        "BTC",
        "ETH"
    ],
    "datePublished": "2026-09-06T05:30:00.000Z",
    "scrapedAt": "2026-09-06T06:02:11.418Z",
    "url": "https://www.coindesk.com/markets/2026/09/06/bitcoin-etf-inflows-top-1bn"
}
```

### Integrations

Run it over the API and get the rows back in one call:

```bash
curl -X POST "https://api.apify.com/v2/acts/publicmoney~coindesk-scraper/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"articlesPerFeed": 50, "maxItems": 0}'
```

From Python:

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_TOKEN")
run = client.actor("publicmoney/coindesk-scraper").call(run_input={"articlesPerFeed": 50, "maxItems": 0})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item["headline"])
```

Give an AI agent the Actor over MCP:

```json
{
    "mcpServers": {
        "apify": {
            "url": "https://mcp.apify.com/?actors=publicmoney/coindesk-scraper"
        }
    }
}
```

Schedules run it on any cron, webhooks fire when a run finishes, and platform integrations push the
dataset to Google Sheets, Slack, Airtable, Zapier or your own endpoint.

### Cost

Pay per event, so you pay for records rather than compute time.

| Event | Free tier | Top volume tier |
| --- | --- | --- |
| News article | $0.002 | $0.0007 |
| Actor start | $0.00005 per GB | Same |

A record that returned no data is published as a failure row and is **never charged**. Six volume tiers apply, so the per-record price falls with monthly volume.

### Troubleshooting

| Issue | Solution |
| --- | --- |
| Every story returns `failed` | CoinDesk's feed was unreachable on that run. It is a public endpoint, so this is usually transient. Check the log and rerun. |
| `mentionedTickers` is empty on a story that clearly names an asset | The ticker index is resolved live and a story that names a company only in prose, with no cashtag and no bracketed symbol, will not match. An index that fails to load leaves the field empty rather than failing the run, so check the log. |
| A scheduled run returns stories I already have | The feed carries a rolling window, so each run sees recent stories again. Deduplicate on `guid`, which is stable per story. |
| `articleBody` is a summary, not the full article | That is what the feed publishes. This Actor does not fetch the article page, so you get the headline, the summary and the link. |

### FAQ

#### Does CoinDesk have an API?

CoinDesk sells data products and its indices through a commercial API, and there is no free tier for the newsroom. This Actor reads the public syndication feed instead and is charged per article.

#### How far back does the feed go?

As far as CoinDesk's feed window, which is the recent stories rather than the archive. To build a history, schedule the Actor and accumulate the dataset, deduplicating on `guid`.

#### Why is `datePublished` different from the site?

It is the feed's own publication time converted to ISO 8601 in UTC. The site renders it in your local timezone.

#### How does it know which assets a story is about?

It resolves cashtags and bracketed symbols in the headline and body against a live CoinGecko coin list and the SEC company ticker file. Both are fetched per run rather than hardcoded, so new listings resolve without an Actor update, and there is a guard against common words that happen to be tickers.

#### Does it return the full article text?

No. It returns what the feed publishes: headline, summary body, author, categories, link and date. Fetching and republishing full articles is a copyright question, so the Actor gives you the metadata and the URL.

#### Do I need a CoinDesk API key?

No. CoinDesk publishes its feed openly and this Actor reads it, so there is no CoinDesk credential anywhere. You need an Apify token to call the Actor over the API.

#### Can I get this data in Python?

Yes, with the `apify-client` package as shown above. It returns parsed JSON, so there is no HTML or response handling on your side.

#### Can I get the data into Excel or Google Sheets?

Yes. Export the dataset as XLSX or CSV, or connect the Google Sheets integration so each run appends to a sheet.

#### Can an AI agent call this Actor?

Yes. Add it to an MCP client with the config above and the model can request what it needs on its own. Every record is flat JSON with named fields, so no post-processing is needed.

#### Is it legal to scrape CoinDesk?

This Actor reads CoinDesk's own public feed, which the publisher publishes for syndication, and returns headline, summary, byline and link rather than full article text. Republishing the text of an article is a copyright question and is your responsibility, so take your own legal advice for how you use it.

### Changelog

- **0.0.2** Added mentioned-ticker resolution against the live coin and SEC ticker indexes.
- **0.0.1** First release. CoinDesk articles as structured records.

### Feedback

Found a field CoinDesk publishes that this Actor misses, or an input it rejects? Open an issue on the Issues tab with the input and what you expected. A daily test runs every Actor in the fleet against live sources, so parser fixes ship fast.

# Actor input Schema

## `feeds` (type: `array`):

Which CoinDesk feed to read. CoinDesk publishes a single feed covering the whole newsroom, so 'all' is everything it puts out. Leave the field empty for the same result. Examples: 'all'. Default is 'all'.

## `articlesPerFeed` (type: `integer`):

How many stories to return per feed, newest first. The feed carries a rolling window of recent stories rather than an archive, so a value above what it holds returns everything available. Examples: 10, 50, 100. Default is 50.

## `maxItems` (type: `integer`):

Maximum number of stories to read, counted from the top of the list. Use it to cap spend on a long list without editing the list itself. Examples: 10, 50, 200. Default is 0, which reads every story given.

## Actor input object example

```json
{
  "feeds": [
    "all"
  ],
  "articlesPerFeed": 50,
  "maxItems": 0
}
```

# Actor output Schema

## `results` (type: `string`):

One item per requested input, in the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "feeds": [
        "all"
    ],
    "articlesPerFeed": 50,
    "maxItems": 0
};

// Run the Actor and wait for it to finish
const run = await client.actor("publicmoney/coindesk-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "feeds": ["all"],
    "articlesPerFeed": 50,
    "maxItems": 0,
}

# Run the Actor and wait for it to finish
run = client.actor("publicmoney/coindesk-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "feeds": [
    "all"
  ],
  "articlesPerFeed": 50,
  "maxItems": 0
}' |
apify call publicmoney/coindesk-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,publicmoney/coindesk-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/HUhDuPF66tgMmEWdM/builds/9Dy9OJv1dysNNBGTL/openapi.json
