Cybench leaderboard: which LLMs score highest

Cybench (cybersecurity capture-the-flag tasks) currently has scores for 16 models. The top score is 43.0% by grok-4-0709; the median model scores 17.5% and the lowest scores 5.0%.

Cybench leaderboard: top 16 of 16 models (highest-effort row per model)
#ModelOrganisationCybench scoreCheapest input $/MSourceCaptured
1grok-4-0709xAI43.0%—Epoch AI Benchmarking Hub2026-10-05
2claude-opus-4-1-20250805Anthropic42.0%—Epoch AI Benchmarking Hub2026-10-05
3grok-4-1xAI39.0%—Epoch AI Benchmarking Hub2026-10-05
4Claude Opus 4Anthropic38.0%$15.00Epoch AI Benchmarking Hub2026-10-05
5claude-sonnet-4-20250514Anthropic35.0%—Epoch AI Benchmarking Hub2026-10-05
6grok-4-fastxAI30.0%—Epoch AI Benchmarking Hub2026-10-05
7claude-3-7-sonnet-20250219Anthropic20.0%—Epoch AI Benchmarking Hub2026-10-05
8Claude 3.5 Sonnet (Jun 2024)Anthropic17.5%$3.00Epoch AI Benchmarking Hub2026-10-05
9GPT-4.5 Preview (Feb 2025)OpenAI17.5%$75.00Epoch AI Benchmarking Hub2026-10-05
10GPT-4o (Nov 2024)OpenAI12.5%$2.50Epoch AI Benchmarking Hub2026-10-05
11claude-3-opus-20240229Anthropic10.0%$15.00Epoch AI Benchmarking Hub2026-10-05
12o1-previewOpenAI10.0%$15.00Epoch AI Benchmarking Hub2026-10-05
13gemini-1.5-pro-001-feb24Google DeepMind7.5%—Epoch AI Benchmarking Hub2026-10-05
14Llama-3.1-405B-InstructMeta AI7.5%$3.50Epoch AI Benchmarking Hub2026-10-05
15Mixtral-8x22B-Instruct-v0.1Mistral AI7.5%$0.600Epoch AI Benchmarking Hub2026-10-05
16Meta-Llama-3-70B-InstructMeta AI5.0%$0.120Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to GPT-4o (Nov 2024) at $2.50 per million input tokens (score 12.5%).

What Cybench measures

Cybench measures cybersecurity capture-the-flag tasks.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks