JailbreakBench LLM Robustness Leaderboard Scraper avatar

JailbreakBench LLM Robustness Leaderboard Scraper

Pricing

Pay per event

Go to Apify Store
JailbreakBench LLM Robustness Leaderboard Scraper

JailbreakBench LLM Robustness Leaderboard Scraper

Scrapes the public JailbreakBench leaderboard tracking attack-success-rate for jailbreak techniques (PAIR, GCG, AIM, and more) against open- and closed-source LLMs, with and without defenses (SmoothLLM, perplexity filter, etc). Snapshot each run to track technique-vs-model ASR movement over time.

Pricing

Pay per event

Rating

0.0

(0)

Developer

BowTiedRaccoon

BowTiedRaccoon

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

9 days ago

Last modified

Share

Scrape the JailbreakBench leaderboard — the open academic robustness benchmark tracking jailbreak attack success rate (ASR) against open- and closed-source LLMs, with and without defenses. This actor returns one row per (model, attack technique, defense) combination, stamped with a per-run snapshot date so repeated runs build a time-series of ASR movement.

What You Get

Each record contains:

  • Leaderboard — always jailbreakbench (the source benchmark)
  • Leaderboard groupopen_source or closed_source, depending on which JailbreakBench table the row came from
  • Model and model provider — target LLM (e.g. Vicuna-13B, GPT-4-0125-Preview) and its organization, inferred from a static prefix map (Meta, LMSYS, OpenAI, Anthropic, Google, Mistral AI, Alibaba, Cohere); left empty when the model isn't recognized rather than guessed
  • Attack method and attack method (full name) — short code (e.g. PAIR, GCG, AIM) plus the full published technique name
  • Attack success rate — ASR percentage reported by JailbreakBench for this row
  • Defense — defense mechanism applied during the evaluation (e.g. SmoothLLM, Perplexity filter, None)
  • Dataset — always JBB-Behaviors, the behavior dataset JailbreakBench evaluates against
  • Metric type — always attack (JailbreakBench evaluates defenses as attack outcomes against a defended model, not as a separate metric)
  • Threat model — access assumption the attack makes (e.g. Black-box access)
  • Average queries — average number of queries the attack used to succeed, as published (- when not reported for that technique)
  • Submission date — date the result was published/updated on the leaderboard
  • Paper URL and jailbreak artifacts URL — links to the technique's paper and published attack transcripts, when available
  • Snapshot date — this run's capture date (YYYY-MM-DD), the time-series anchor for tracking ASR movement across runs
  • Source URL and scraped at — exact page fetched and per-record scrape timestamp

refusal_rate and benign_refusal_rate are included in the schema for parity with the broader jailbreak-benchmark data model but are always empty for JailbreakBench — the source only reports a single ASR figure per row, not a separately-measured refusal rate.

Scope

This actor scrapes JailbreakBench only. JailbreakBench's own leaderboard page names it as the reference source; adjacent benchmarks (HarmBench, TrustLLM) are not included — HarmBench in particular renders its leaderboard client-side from a heavier SPA shell and would need a materially different (browser-based) build.

Usage

Basic run (both leaderboards, default 10 records)

No configuration needed — just run with default settings.

Full leaderboard

Set maxItems to a high number (e.g., 200) to get every row across both tables (~96 rows total as of this actor's build).

Filter by leaderboard

Use the leaderboardType input to scrape just one table:

{
"leaderboardType": "open_source",
"maxItems": 100
}

Input Schema

FieldTypeDefaultDescription
leaderboardTypeenumbothWhich leaderboard table(s) to scrape: both, open_source, or closed_source
maxItemsinteger10Maximum total records to return

Output Sample

{
"leaderboard": "jailbreakbench",
"leaderboard_group": "open_source",
"model": "Vicuna-13B",
"model_provider": "LMSYS",
"attack_method": "PAIR",
"attack_method_full": "Prompt Automatic Iterative Refinement (PAIR)",
"attack_success_rate": 69,
"refusal_rate": "",
"defense": "None",
"benign_refusal_rate": "",
"dataset": "JBB-Behaviors",
"metric_type": "attack",
"threat_model": "Black-box access",
"average_queries": "34",
"submission_date": "12 Oct 2023",
"paper_url": "https://arxiv.org/abs/2310.08419",
"jailbreak_artifacts_url": "https://github.com/JailbreakBench/artifacts/blob/main/attack-artifacts/PAIR/black_box/vicuna-13b-v1.5.json",
"snapshot_date": "2026-07-13",
"source_url": "https://jailbreakbench.github.io/tables/table_open_source.html",
"scraped_at": "2026-07-13T07:57:28.824Z"
}

Technical Notes

  • Two-URL crawl, no pagination. JailbreakBench's homepage doesn't carry the table data itself — the leaderboard sections load client-side via jQuery from two static HTML fragment pages, which this actor fetches directly: tables/table_open_source.html and tables/table_closed_source.html.
  • No proxy required — GitHub Pages static hosting with no anti-bot measures.
  • Fast — a full run across both tables completes in a few seconds.
  • Time-series by design. Run this actor periodically (e.g. weekly) to build a history of ASR movement per technique/model/defense combination — each run's snapshot_date is the time-series anchor. This is a small, always-complete-in-one-run crawl; there's no large-crawl resume workflow to manage.

Use Cases

  • Track jailbreak/robustness leaderboard movement over time for specific models, attacks, or defenses
  • Compare open-source vs. closed-source model resilience to published jailbreak techniques
  • Feed AI-safety or guardrail-vendor evaluation pipelines with structured ASR data
  • Research and academic benchmarking of LLM robustness trends