LLM Model Comparison (340+ models)
Pricing
Pay per event
LLM Model Comparison (340+ models)
Compare LLMs (Claude, GPT, Gemini, Grok, Llama, DeepSeek, Mistral, Qwen) on price, context, modality, capabilities and live benchmarks. OpenRouter-powered. No key.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Dev D
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
LLM Model Comparison ⚖️🤖
Compare large language models — Claude, GPT, Gemini, Grok, Llama, DeepSeek, Mistral, Qwen and more — head-to-head on price, context window, modality, capabilities and live benchmark scores. One call, no API key.
Choosing between models means juggling pricing pages, context limits, and scattered benchmark leaderboards. This Actor pulls the whole hosted-LLM catalog from OpenRouter's public model API and normalizes every model into one clean, comparison-ready row — so you can line up the models you care about and sort them by whatever matters: intelligence, price, context, or Elo.
Covers every major provider: Anthropic · OpenAI · Google · xAI · Meta · DeepSeek · Mistral · Qwen · Cohere · and dozens more (340+ models).
Three modes
- Compare (default) — name the models you want (
["claude", "gpt-5", "gemini", "grok"]) and get them lined up and ranked. Broad terms match every variant; exact ids (anthropic/claude-sonnet-5) pin one. - Search the catalog — filter all 340+ models by provider, price ceiling, minimum context, modality, reasoning/tools support.
- Market digest — a single summary row: price ranges, provider mix, and top models by intelligence and by arena Elo.
What you get per model
| Field | Example |
|---|---|
price_input_per_m / price_output_per_m | $2.00 / $10.00 per 1M tokens |
context_length / max_output_tokens | 1,000,000 / 128,000 |
modality / input_modalities | text+image+file->text |
supports_reasoning / supports_tools | true / true |
intelligence_index · coding_index · agentic_index | 53.4 · 71.2 · 48.0 (Artificial Analysis) |
arena_avg_elo · arena_best_rank | 1278 · 3 (Design Arena) |
knowledge_cutoff · created_at · url | … |
Example — Claude vs GPT vs Gemini vs Grok
{"mode": "compare","models": ["claude", "gpt-5", "gemini", "grok"],"sortBy": "intelligence"}
Returns each matching model ranked by intelligence index, with price, context, capabilities and benchmark scores side by side.
Example — cheapest capable models
{"mode": "models","maxOutputPricePerM": 5,"minContext": 200000,"reasoningOnly": true,"sortBy": "price_output"}
Every reasoning model with a 200k+ context under $5 / 1M output tokens, cheapest first.
Run it on a schedule
New models and price changes land constantly. Schedule a daily run to keep a live LLM price/benchmark table in a sheet, dashboard, or your own comparison page — or diff the digest over time to catch new releases and price cuts.
Notes & limitations
- No API key — the OpenRouter model catalog is a public endpoint.
- Pricing is OpenRouter-listed (USD per 1M tokens). It mirrors provider pricing closely; some models have long-context surcharge tiers (flagged via
has_long_context_surcharge). Treat it as OpenRouter's rate, not an official provider quote. - Benchmark scores are partial — Artificial Analysis (intelligence/coding/agentic) and Design Arena (Elo/rank) are present for many but not all models (roughly half); they're surfaced when available and
nullotherwise, never estimated. - Only models offered via OpenRouter appear; a model not on OpenRouter won't be listed.
- Data source: OpenRouter
/api/v1/models.
Keywords
LLM comparison, AI model comparison, compare LLMs, Claude vs GPT, GPT vs Gemini, Grok vs Claude, LLM pricing, LLM price comparison, model context window, LLM benchmarks, intelligence index, arena Elo, OpenRouter, Anthropic Claude, OpenAI GPT, Google Gemini, xAI Grok, Llama, DeepSeek, Mistral, Qwen, model catalog, AI model registry, tokens per dollar, reasoning models, multimodal models.