LMArena LLM Leaderboard Scraper avatar

LMArena LLM Leaderboard Scraper

Pricing

from $0.50 / 1,000 record scrapeds

Go to Apify Store
LMArena LLM Leaderboard Scraper

LMArena LLM Leaderboard Scraper

Scrape the LMArena (Chatbot Arena) ELO leaderboard — ranks, ratings, vote counts, and confidence intervals for the text, code, agent, text-to-image, and text-to-video arena categories the leaderboard currently publishes. Returns one row per model per leaderboard variant.

Pricing

from $0.50 / 1,000 record scrapeds

Rating

0.0

(0)

Developer

BowTiedRaccoon

BowTiedRaccoon

Maintained by Community

Actor stats

0

Bookmarked

15

Total users

5

Monthly active users

12 days ago

Last modified

Share

Scrape the LMArena Chatbot Arena ELO leaderboard — the most widely cited blind-evaluation ranking for large language models. Every major model launch references its Chatbot Arena ELO position. This actor returns one row per model per leaderboard variant, covering every arena category the leaderboard page currently publishes in a single run.

What You Get

Each record contains:

  • Leaderboard variant — arena slug (text, code, agent, agent-code, agent-work, text-to-image, text-to-video, plus any further categories the site adds later)
  • Rank — current position with 95% confidence interval bounds (rank_upper, rank_lower)
  • Model name and organization
  • ELO score with confidence interval low/high
  • Vote count — number of blind battles used to compute this rating
  • License — Proprietary, Apache-2.0, etc.
  • Context window tokens
  • Input/output pricing (per million tokens, where publicly available)
  • Profile URL — link to official model announcement
  • Scraped at — ISO 8601 timestamp

Usage

Basic run (all arenas, default 10 records)

No configuration needed — just run with default settings to get a sample of the leaderboard.

Full leaderboard

Set maxItems to a high number (e.g., 1000) to get every entry across all arenas currently on the leaderboard page.

Filter by arena

Use the arenas input to limit results to specific variants:

{
"maxItems": 500,
"arenas": ["text", "code"]
}

Available arena slugs (the leaderboard page decides which categories exist; this list reflects what it currently publishes):

  • text — overall text/conversation benchmark (the main leaderboard)
  • code — coding tasks (WebDev models)
  • agent — best overall agents
  • agent-code — best coding agents
  • agent-work — best agents for work tasks
  • text-to-image — image generation
  • text-to-video — video generation

Input Schema

FieldTypeDefaultDescription
maxItemsinteger10Maximum total records to return
arenasarray[] (all)Arena slugs to include. Empty = return all arenas

Output Sample

{
"leaderboard_variant": "text",
"rank": 1,
"rank_upper": 1,
"rank_lower": 4,
"model_name": "claude-opus-4-6-thinking",
"organization": "Anthropic",
"elo_score": 1502.17,
"elo_confidence_low": 1497.91,
"elo_confidence_high": 1506.43,
"vote_count": 34186,
"license": "Proprietary",
"context_window_tokens": 1000000,
"input_price_per_million": 5,
"output_price_per_million": 25,
"profile_url": "https://www.anthropic.com/news/claude-opus-4-6",
"scraped_at": "2026-06-01T15:05:33.265Z"
}

Technical Notes

  • Single HTTP request — all arena data comes from a single page. No pagination, no API calls.
  • Fast — typical run completes in under 30 seconds.
  • The leaderboard updates daily as new arena battles are processed.

Use Cases

  • Track ELO rankings over time for specific models or organizations
  • Compare models across different task categories (code vs. vision vs. overall)
  • Feed into model-routing or evaluation pipelines
  • Monitor competitive positioning for newly launched models
  • Research and academic benchmarking datasets