Fiction.liveBench leaderboard: which LLMs score highest

Fiction.liveBench (long-context story comprehension (16k)) currently has scores for 47 models. The top score is 94.4% by grok-4-0709; the median model scores 61.1% and the lowest scores 25.0%.

Fiction.liveBench leaderboard: top 47 of 47 models (highest-effort row per model)
#ModelOrganisationFiction.liveBench scoreCheapest input $/MSourceCaptured
1grok-4-0709xAI94.4%—Epoch AI Benchmarking Hub2026-10-05
2grok-4-fastxAI94.4%—Epoch AI Benchmarking Hub2026-10-05
3Gemini 2.5 Pro Preview (Jun 2025)Google DeepMind91.7%$1.25Epoch AI Benchmarking Hub2026-10-05
4Kimi K2.5Moonshot86.1%$0.450Epoch AI Benchmarking Hub2026-10-05
5DeepSeek-V3.2-Exp (high)DeepSeek83.3%$0.270Epoch AI Benchmarking Hub2026-10-05
6QwQ-32BAlibaba83.3%$0.150Epoch AI Benchmarking Hub2026-10-05
7gemini-2.5-flash-preview-05-20Google DeepMind77.8%—Epoch AI Benchmarking Hub2026-10-05
8DeepSeek-R1 (May 2025)DeepSeek75.0%$0.400Epoch AI Benchmarking Hub2026-10-05
9Qwen3-235B-A22B-Thinking-2507Alibaba75.0%$0.230Epoch AI Benchmarking Hub2026-10-05
10Qwen3-32BAlibaba74.2%$0.080Epoch AI Benchmarking Hub2026-10-05
11DeepSeek-R1DeepSeek69.4%$0.400Epoch AI Benchmarking Hub2026-10-05
12MiniMax-M1-80kMiniMax69.4%$0.550Epoch AI Benchmarking Hub2026-10-05
13qwen3-235b-a22bAlibaba67.7%$0.180Epoch AI Benchmarking Hub2026-10-05
14chatgpt-4o-01-29-2025OpenAI66.7%—Epoch AI Benchmarking Hub2026-10-05
15Gemini 2.5 Pro Exp (Mar 2025)Google DeepMind66.7%—Epoch AI Benchmarking Hub2026-10-05
16Gemini 2.5 Pro Preview (Mar 2025)Google DeepMind66.7%$1.25Epoch AI Benchmarking Hub2026-10-05
17Kimi-K2-Instruct-0905Moonshot66.7%$0.500Epoch AI Benchmarking Hub2026-10-05
18qwen-max-2025-01-25Alibaba66.7%—Epoch AI Benchmarking Hub2026-10-05
19Qwen3-Max-InstructAlibaba66.7%—Epoch AI Benchmarking Hub2026-10-05
20GPT-4.1OpenAI63.9%$2.00Epoch AI Benchmarking Hub2026-10-05
21GPT-4.5 Preview (Feb 2025)OpenAI63.9%$75.00Epoch AI Benchmarking Hub2026-10-05
22qwen3-14bAlibaba62.5%$0.080Epoch AI Benchmarking Hub2026-10-05
23qwen3-8bAlibaba62.1%$0.040Epoch AI Benchmarking Hub2026-10-05
24Claude Opus 4Anthropic61.1%$15.00Epoch AI Benchmarking Hub2026-10-05
25Gemini 2.0 Flash (Feb 2025)Google DeepMind,Google61.1%$0.150Epoch AI Benchmarking Hub2026-10-05
26Kimi K2 InstructMoonshot61.1%$0.500Epoch AI Benchmarking Hub2026-10-05
27GLM-4.5-FP8Z.ai58.3%—Epoch AI Benchmarking Hub2026-10-05
28grok-3-betaxAI58.3%—Epoch AI Benchmarking Hub2026-10-05
29Qwen3-Next-80B-A3B-InstructAlibaba55.6%$0.090Epoch AI Benchmarking Hub2026-10-05
30parasail-qwen3-235b-a22b-instruct-2507Alibaba52.9%—Epoch AI Benchmarking Hub2026-10-05
31DeepSeek-V3.1DeepSeek52.8%$0.250Epoch AI Benchmarking Hub2026-10-05
32Gemini 2.0 Flash Thinking ExpGoogle DeepMind,Google52.8%—Epoch AI Benchmarking Hub2026-10-05
33claude-3-7-sonnet-20250219Anthropic50.0%—Epoch AI Benchmarking Hub2026-10-05
34DeepSeek-V3 (Mar 2025)DeepSeek50.0%$0.200Epoch AI Benchmarking Hub2026-10-05
35Gemini 2.5 Flash Preview (Apr 2025)Google DeepMind47.2%$0.300Epoch AI Benchmarking Hub2026-10-05
36gemini-2.5-flash-lite-preview-06-17 (Default thinking length)Google DeepMind47.2%—Epoch AI Benchmarking Hub2026-10-05
37claude-sonnet-4-20250514Anthropic46.9%—Epoch AI Benchmarking Hub2026-10-05
38Llama-4-Maverick-17B-128E-InstructMeta AI46.2%$0.630Epoch AI Benchmarking Hub2026-10-05
39GPT-4.1 miniOpenAI44.4%$0.400Epoch AI Benchmarking Hub2026-10-05
40gpt-oss-120b (high)OpenAI44.4%$0.030Epoch AI Benchmarking Hub2026-10-05
41Gemini 2.0 Pro Exp (Feb 2025)Google DeepMind41.7%—Epoch AI Benchmarking Hub2026-10-05
42Qwen3-30B-A3BAlibaba40.6%$0.090Epoch AI Benchmarking Hub2026-10-05
43Llama-4-Scout-17B-16E-InstructMeta AI36.0%$0.050Epoch AI Benchmarking Hub2026-10-05
44gemma-3-27b-itGoogle DeepMind33.3%$0.080Epoch AI Benchmarking Hub2026-10-05
45Llama-3.3-70B-InstructMeta AI33.3%$0.120Epoch AI Benchmarking Hub2026-10-05
46GPT-4.1 nanoOpenAI25.0%$0.100Epoch AI Benchmarking Hub2026-10-05
47nvidia.nemotron-nano-9b-v2NVIDIA25.0%$0.040Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to Qwen3-32B at $0.080 per million input tokens (score 74.2%).

What Fiction.liveBench measures

Fiction.liveBench measures long-context story comprehension (16k).

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks