# ATP & WTA Match Dataset - Sackmann CSV Format (`crawlplant/tennis-match-dataset`) Actor

Whole seasons of ATP and WTA singles results in the columns of Jeff Sackmann's tennis\_atp / tennis\_wta CSV files, from 1990 to this week: winner and loser, rank, seed, entry, hand, height, age, score, round, minutes, serve statistics, plus pre-match Elo. Updated every hour. Export as CSV.

- **URL**: https://apify.com/crawlplant/tennis-match-dataset.md
- **Developed by:** [CrawlPlant](https://apify.com/crawlplant) (community)
- **Categories:** Sports, AI
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## ATP & WTA Match Dataset - Sackmann CSV Format

*Independent tool, not affiliated with, endorsed by or connected to Jeff Sackmann, Tennis Abstract, Flashscore, the ATP,
the WTA or the ITF. It reads the public match data that flashscore.com shows to every visitor and the official ATP and
WTA ranking lists and player profiles.*

**Whole seasons of ATP and WTA singles results in the columns of Jeff Sackmann's `atp_matches_YYYY.csv` and
`wta_matches_YYYY.csv`**, from 1990 to this week, updated every hour. Code and notebooks that read his files by column name
keep working: winner and loser with **rank and ranking points of that week, seed, entry (Q, WC, LL, PR), hand, height,
country and age**, the score (`7-6(5) 6-4`, `RET`, `W/O`), best of, round, minutes and the **serve statistics**
(`w_ace` ... `l_bpFaced`), plus each player's **Elo before the match**. Export the dataset as CSV.

### Why this one

- **Current.** This week's matches are in the dataset within the hour: no waiting for a yearly update.
- **Same columns, same order.** `tourney_id`, `tourney_name`, `surface`, `draw_size`, `tourney_level`, `tourney_date`,
  `match_num`, `winner_*`, `loser_*`, `score`, `best_of`, `round`, `minutes`, `w_*` / `l_*` serve columns,
  `winner_rank`, `winner_rank_points`, `loser_rank`, `loser_rank_points`.
- **Seeds and entries from the official draws**, hand and height from the official ATP and WTA player profiles, ranks
  from the official weekly lists.
- **Elo included.** Four extra columns close each row: `match_id`, `tour`, `winner_elo` and `loser_elo` (pre-match Elo,
  overall).
- **ATP, WTA, Challenger and ITF.** `tours: ["challenger-men"]` gives his qual\_chall files, `["itf-men"]` his futures.
- **Any range.** Whole seasons, or the last 7 days to keep your own copy current.
- **See it live.** The size of the database, the ranking weeks behind the rank columns and the Elo leaders of both
  tours: [crawlplant.com/tennis-match-dataset](https://crawlplant.com/tennis-match-dataset/).

### What can you use it for?

- **Prediction and betting models** trained on the same columns as the best-known open tennis dataset, kept current.
- **Research and journalism**: season-by-season results with rankings, seeds and serve statistics.
- **Replacing a stale copy**: refresh your atp\_matches / wta\_matches tables weekly with one scheduled task.
- **AI agents**: structured match history as JSON or CSV.

### Quick start

1. Click **Try for free** with the default input (the first 1,000 matches of the 2025 ATP season).
2. Set **Years** and **Tours**, and **Max results** above the season size (ATP about 4,300 matches, WTA about 3,800).
3. Export the dataset as **CSV**: the columns come in his order.

#### Copy to your AI assistant

Paste this into ChatGPT, Claude or any agent so it knows how to use the Actor:

```
crawlplant/tennis-match-dataset on Apify: ATP / WTA / Challenger / ITF singles results in the columns of Jeff Sackmann's
atp_matches_YYYY.csv (tourney_id, tourney_name, surface, draw_size, tourney_level, tourney_date, match_num, winner_id,
winner_seed, winner_entry, winner_name, winner_hand, winner_ht, winner_ioc, winner_age, loser_*, score, best_of, round,
minutes, w_ace ... l_bpFaced, winner_rank, winner_rank_points, loser_rank, loser_rank_points) + match_id, tour,
winner_elo, loser_elo. From 1990, updated hourly. Input: years (e.g. [2024, 2025]), tours (atp, wta, challenger-men,
challenger-women, itf-men, itf-women), dateFrom / dateTo (when years is empty), maxItems (default 1000).
Price: $1.00 per 1,000 matches on the Free plan.
```

### Ready-to-use examples

**1. A full ATP season**

```json
{ "years": [2025], "tours": ["atp"], "maxItems": 5000 }
```

**2. Three WTA seasons**

```json
{ "years": [2023, 2024, 2025], "tours": ["wta"], "maxItems": 15000 }
```

**3. This season so far, both tours**

```json
{ "years": [], "tours": ["atp", "wta"], "maxItems": 10000 }
```

**4. The last 7 days, to keep your copy current**

```json
{ "years": [], "dateFrom": "-7", "dateTo": "today", "tours": ["atp", "wta"], "maxItems": 5000 }
```

**5. Challenger season (his qual\_chall file)**

```json
{ "years": [2025], "tours": ["challenger-men"], "maxItems": 20000 }
```

**6. The 1990s, ATP**

```json
{ "years": [1990, 1991, 1992, 1993, 1994, 1995, 1996, 1997, 1998, 1999], "tours": ["atp"], "maxItems": 40000 }
```

### How to…

#### Replace Jeff Sackmann's tennis\_atp and tennis\_wta CSV files

Run one season per tour (example 1, then `"tours": ["wta"]`) and export as CSV. Code that reads his files by column name
works unchanged; keep the rows whose `round` doesn't start with "Q" for his main-draw-only files.

#### Keep a match dataset up to date every week

Save example 4 as a task and schedule it weekly: each run returns the last 7 days, and `match_id` lets you upsert
without duplicates.

#### Get ATP and WTA results with rankings for a model

Every row carries both players' rank and ranking points from the official list of that week, the seeds and entries of
the draw, and pre-match Elo (`winner_elo`, `loser_elo`).

### Input options

| Option | Default | Description |
|---|---|---|
| `years` | `[2025]` | Seasons, from 1990; empty = the From / To days |
| `tours` | `atp` | `atp`, `wta`, `challenger-men`, `challenger-women`, `itf-men`, `itf-women`; empty = ATP and WTA |
| `dateFrom` / `dateTo` | this season / today | When `years` is empty: `2026-01-01`, `-7`, `today` |
| `maxItems` | `1000` | Maximum matches in total |

### Example output

The 2025 Hong Kong final, a real row:

```json
{
  "tourney_id": "2024-C8wqzbIc",
  "tourney_name": "Hong Kong",
  "surface": "Hard",
  "draw_size": 16,
  "tourney_level": "A",
  "tourney_date": 20241230,
  "match_num": 15,
  "winner_id": "nLLdS9q0",
  "winner_seed": null,
  "winner_entry": null,
  "winner_name": "Alexandre Muller",
  "winner_hand": "R",
  "winner_ht": 183,
  "winner_ioc": "FRA",
  "winner_age": 27.9,
  "loser_id": "YNkOjUIB",
  "loser_seed": null,
  "loser_entry": "WC",
  "loser_name": "Kei Nishikori",
  "loser_hand": "R",
  "loser_ht": 178,
  "loser_ioc": "JPN",
  "loser_age": 35,
  "score": "2-6 6-1 6-3",
  "best_of": 3,
  "round": "F",
  "minutes": 104,
  "w_ace": 2, "w_df": 1, "w_svpt": 77, "w_1stIn": 60, "w_1stWon": 44, "w_2ndWon": 7, "w_SvGms": 12, "w_bpSaved": 4, "w_bpFaced": 6,
  "l_ace": 4, "l_df": 1, "l_svpt": 75, "l_1stIn": 48, "l_1stWon": 28, "l_2ndWon": 14, "l_SvGms": 12, "l_bpSaved": 4, "l_bpFaced": 8,
  "winner_rank": 67,
  "winner_rank_points": 778,
  "loser_rank": 106,
  "loser_rank_points": 578,
  "match_id": "C032DjD6",
  "tour": "atp",
  "winner_elo": 2110,
  "loser_elo": 2259
}
```

### Output fields

His columns, with these conventions:

- Player ids (`winner_id`, `loser_id`) and `tourney_id` use Flashscore's ids (strings); `match_id` is Flashscore's match id.
- `tourney_date` is the Monday of the main-draw week (YYYYMMDD, as in his files); `tourney_level`: G Grand Slam,
  M Masters 1000, F tour finals, D Davis Cup / Billie Jean King Cup, A other tour level (WTA 1000 included), C Challenger,
  S ITF.
- Qualifying rounds are in the same rows (`round` Q1-Q3); `draw_size` is counted from the main-draw matches played.
- Serve columns: ATP and WTA matches from 2012; earlier seasons have results, ranks, seeds, hand, height and age.
- Extra columns: `match_id`, `tour`, `winner_elo`, `loser_elo`, then `scrapedAt`, `source`, `store`.

### Pricing

Pay per result: platform usage is included.

| Event | No discount (Free plan) | Bronze (Starter) | Silver (Scale) | Gold (Business) |
|---|---|---|---|---|
| Match (per 1,000) | $1.00 | $0.90 | $0.80 | $0.70 |
| Actor start (per run) | $0.00005 | $0.00005 | $0.00005 | $0.00005 |

| Example on the Free plan | Matches | Cost |
|---|---|---|
| Default run: 1,000 matches of the 2025 ATP season | 1,000 | ~$1.00 |
| The full 2025 ATP season with qualifying | 4,285 | ~$4.30 |
| The 1990 ATP season | 3,580 | ~$3.58 |
| Last 7 days, ATP and WTA (week to 2026-10-02) | 193 | ~$0.19 |

Set a **maximum cost per run** in the run options to stop a large run at your budget.

### Reliability

- Rows come from our tennis database (more than 900,000 matches, updated every hour, ranks and player profiles from the
  official ATP and WTA sources): a season comes back in seconds, nothing is scraped during your run.
- Each run writes a summary to the key-value store (`OUTPUT`) with the rows saved and any warnings.

### Run it through the API

JavaScript:

```js
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('crawlplant/tennis-match-dataset').call({ years: [2025], tours: ['atp'], maxItems: 5000 });
// the dataset as CSV, columns in Sackmann's order
const csv = await client.dataset(run.defaultDatasetId).downloadItems('csv');
```

Python:

```python
import pandas as pd
from apify_client import ApifyClient
client = ApifyClient("YOUR_TOKEN")
run = client.actor("crawlplant/tennis-match-dataset").call(run_input={"years": [2025], "tours": ["wta"], "maxItems": 5000})
df = pd.DataFrame(client.dataset(run["defaultDatasetId"]).list_items().items)
```

### Use with AI agents

The Actor works as a tool in the [Apify MCP server](https://mcp.apify.com) for Claude, ChatGPT, Cursor and others:

```
https://mcp.apify.com?tools=crawlplant/tennis-match-dataset
```

Try *"which players won the most tiebreaks on clay this season?"*.

### More tennis data

- [Flashscore Tennis Scraper](https://apify.com/crawlplant/flashscore-tennis): odds from about 20 bookmakers, head-to-head,
  form, live scores, rankings, player statistics and Elo lists.
- [Tennis Match History](https://apify.com/crawlplant/tennis-match-history): every match of one player's career.
- [Tennis Point-by-Point Data](https://apify.com/crawlplant/tennis-point-by-point): every game and point of any match.

### FAQ

#### Is this Jeff Sackmann's data?

No. It's an independent dataset in the same column layout, built from Flashscore results and the official ATP and WTA
ranking lists, draws and player profiles. His repositories are a great resource; this Actor gives you the same shape,
current to the hour.

#### How far back does it go?

ATP and WTA from 1990, Challenger from 2008, ITF from 2011. Serve statistics from 2012.

#### Is it legal?

The data is public sporting results and official rankings, read without logging in. How you use and publish them is your
responsibility, so check the sources' terms for your use case.

### Privacy

The data is about professional players' public sporting results and profiles (names, countries, ages, hand, height).
Each run sends the developer anonymous feature-usage statistics (the options used); your Apify account id is replaced by
a one-way hash.

# Actor input Schema

## `years` (type: `array`):

The seasons, e.g. \[2024, 2025] (from 1990). Empty = the From / To days below, or this season so far.

## `tours` (type: `array`):

atp = his atp\_matches files, wta = wta\_matches, challenger-men = qual\_chall, itf-men = futures. Empty = ATP and WTA.

## `dateFrom` (type: `string`):

Used when Years is empty: the first day, e.g. 2026-01-01 or -7 (seven days ago). Empty Years and From = this season so far.

## `dateTo` (type: `string`):

Used when Years is empty: the last day, inclusive. Empty = today.

## `maxItems` (type: `integer`):

Maximum matches to return. A full ATP season is about 4,300 matches with qualifying, WTA about 3,800.

## Actor input object example

```json
{
  "years": [
    2025
  ],
  "tours": [
    "atp"
  ],
  "maxItems": 1000
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "years": [
        2025
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("crawlplant/tennis-match-dataset").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "years": [2025] }

# Run the Actor and wait for it to finish
run = client.actor("crawlplant/tennis-match-dataset").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "years": [
    2025
  ]
}' |
apify call crawlplant/tennis-match-dataset --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,crawlplant/tennis-match-dataset"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/yfsGxzjjQX5LSdwga/builds/M1nov5iH1qWZxvvtq/openapi.json
