Furniture Assembly leaderboard: which LLMs score highest

Furniture Assembly (planning a physical assembly) currently has scores for 25 models. The top score is 83.3% by Claude Opus 5.5 (max); the median model scores 40.0% and the lowest scores 20.0%.

Furniture Assembly leaderboard: top 25 of 25 models (highest-effort row per model)
#ModelOrganisationFurniture Assembly scoreCheapest input $/MSourceCaptured
1Claude Opus 5.5 (max)Anthropic83.3%$4.00Epoch AI Benchmarking Hub2026-10-05
2GPT-6 Astra (max)OpenAI80.0%$10.00Epoch AI Benchmarking Hub2026-10-05
3GPT-6.1 Sol (max)OpenAI80.0%$2.00Epoch AI Benchmarking Hub2026-10-05
4Claude Sonnet 5.5 (max)Anthropic75.0%$2.00Epoch AI Benchmarking Hub2026-10-05
5Claude Fable 5.1 (max)Anthropic70.0%$10.00Epoch AI Benchmarking Hub2026-10-05
6Claude Opus 5 (max)Anthropic60.8%$5.00Epoch AI Benchmarking Hub2026-10-05
7GPT-6 Sol (max)OpenAI58.3%$2.00Epoch AI Benchmarking Hub2026-10-05
8GPT-5.6 Sol (max)OpenAI56.7%$2.00Epoch AI Benchmarking Hub2026-10-05
9GPT-5.6 Terra (max)OpenAI54.2%$2.00Epoch AI Benchmarking Hub2026-10-05
10GPT-5.5 (xhigh)OpenAI44.2%$5.00Epoch AI Benchmarking Hub2026-10-05
11GPT-6 Luna (max)OpenAI44.2%$0.100Epoch AI Benchmarking Hub2026-10-05
12GPT-5.6 Luna (max)OpenAI42.5%$0.200Epoch AI Benchmarking Hub2026-10-05
13Grok 4.6 (xhigh)xAI40.0%$1.25Epoch AI Benchmarking Hub2026-10-05
14GPT-5.2 (xhigh)OpenAI38.3%$1.75Epoch AI Benchmarking Hub2026-10-05
15GPT-5.4 (xhigh)OpenAI37.5%$2.50Epoch AI Benchmarking Hub2026-10-05
16Claude Fable 5 (max)Anthropic35.8%$10.00Epoch AI Benchmarking Hub2026-10-05
17Kimi K3 (max)Moonshot34.2%$1.29Epoch AI Benchmarking Hub2026-10-05
18Claude Opus 4.7 (max)Anthropic33.3%$5.00Epoch AI Benchmarking Hub2026-10-05
19Gemini 3.8 Flash (high)Google DeepMind31.7%$0.750Epoch AI Benchmarking Hub2026-10-05
20Claude Opus 4.6 (max)Anthropic28.3%$5.00Epoch AI Benchmarking Hub2026-10-05
21Gemini 3.7 Flash (high)Google DeepMind26.7%$0.750Epoch AI Benchmarking Hub2026-10-05
22Grok 4.5 (high)xAI22.5%$2.00Epoch AI Benchmarking Hub2026-10-05
23Kimi K2.6Moonshot21.7%$0.650Epoch AI Benchmarking Hub2026-10-05
24Grok 4.7 (xhigh)xAI20.8%$2.00Epoch AI Benchmarking Hub2026-10-05
25Qwen3.8 Max (0902) (xhigh)Alibaba20.0%$1.65Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to GPT-6.1 Sol (max) at $2.00 per million input tokens (score 80.0%).

What Furniture Assembly measures

Furniture Assembly measures planning a physical assembly.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks