Anthropic
·Released on Sep 1, 2026

Claude Fable 5.1 API Benchmarks, Pricing & Provider Data

Compare Claude Fable 5.1 with another model

Choose a model to open its comparison page.

Share on X
LLM

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

Quality
#4of 101
79.0
LMSpeed score

Category Performance

Observed-capability estimates with uncertainty reported separately from independent benchmark families.

Coverage
4 / 8
Methodology
V3.0
Category PerformanceObserved-capability estimates with uncertainty reported separately from independent benchmark families.Agents64.4Coding67.9Reasoning61.3Knowledge68.7Math-Multilingual-Multimodal-Instruction following-
#1Knowledge68.7Provisional1/4 Measured dimensions
80% interval: 54.782.7
broad knowledge68.7
aa_omniscience · aa_omniscience_index / benchlm_category_knowledge · benchlm_category_knowledge
professional knowledgePrior only
factualityPrior only
retrieval open bookPrior only
#2Coding67.9RatedGlobal rank #14/4 Measured dimensions
80% interval: 61.074.8
code generation65.5
scicode · scicode
repository engineering71.8
swe_pro · swe_pro
debugging testing74.8
swe_multilingual · benchlm_coding_swe_multilingual
tooling quality59.6
terminalbench · aa_terminal_bench21
#3Agents64.4RatedGlobal rank #43/4 Measured dimensions
80% interval: 55.873.0
planningPrior only
tool use62.7
tau · aa_tau3_banking
environment execution66.5
gdpval_aa · benchlm_agentic_gdpval_aa / osworld · os_world2
recovery reliability64
briefcase · aa_briefcase_elo / harvey_lab · aa_harvey_lab
#4Reasoning61.3RatedGlobal rank #13/4 Measured dimensions
80% interval: 52.670.0
abstract logic59.1
arc_agi · arc_agi2
scientific causal66.7
critpt · critpt / gpqa · gpqa / hle · hle / hle · hle_no_tools
multistep constraintsPrior only
evidence verification58.1
lcr · lcr
No data:MathMultilingualMultimodalInstruction following

Specifications

Input and output token limits for this model, plus how it ranks on long-context understanding.

INPUT
1Mtokens
1.2K pages of text
OUTPUT
128Ktokens
8K128K1M4M
1M

Features

Technical Details

Input
Output
Released
Sep 2026
Documentation
Tokenizer
Claude
Architecture
text+image+file->text
Moderated
Yes
Supported parameters
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstoolsverbosity

Rankings

Excels at

Falls behind in

Detailed scores

Updated: Sep 2, 2026

Overall

undefined metric} other undefined metrics}}

Overall score79.0#4 / 101DeepSWE67.4#5 / 17
Overall score
LMSpeed rank#4
Score79.0
UpdatedSep 2, 2026
Confidence2
DeepSWE
LMSpeed rank#5
Score67.4
UpdatedSep 2, 2026
Confidence2

Speed & latency

undefined metric} other undefined metrics}}

Output speed61.4 tok/s#54 / 78Time to first token16.07 s#69 / 78
Output speed
LMSpeed rank#54
Score61.4 tok/s
UpdatedSep 2, 2026
Confidence4
Time to first token
LMSpeed rank#69
Score16.07 s
UpdatedSep 2, 2026
Confidence4

Pricing

undefined metric} other undefined metrics}}

Input price$10.00/M#172 / 182Output price$50.00/M#173 / 182
Input price
LMSpeed rank#172
Score$10.00/M
UpdatedSep 2, 2026
Confidence4
Output price
LMSpeed rank#173
Score$50.00/M
UpdatedSep 2, 2026
Confidence4

Agents

V3.0

undefined metric} other undefined metrics}} · Rated

Score64.4#480% interval55.8–73.03/4 Measured dimensionsAgentic score85.0#7 / 70Terminal-Bench 4.055.8#1 / 1
Dimensions and evidence
Planning & decompositionPrior only
Tool use62.7
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
tau · aa_tau3_banking · z 1.29 · q 1.00
Environment & long-horizon execution66.5
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
gdpval_aa · benchlm_agentic_gdpval_aa · z 1.88 · q 1.00
osworld · os_world2 · z 0.81 · q 1.00
Recovery & completion reliability64
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
briefcase · aa_briefcase_elo · z 0.91 · q 1.00
harvey_lab · aa_harvey_lab · z 0.55 · q 1.00
Agentic score
LMSpeed rank#7
Score85.0
UpdatedSep 2, 2026
Confidence2
Terminal-Bench 4.0
LMSpeed rank#1
Score55.8
UpdatedSep 2, 2026
Confidence2
Terminal-Bench-Science 0.1
LMSpeed rank#1
Score52.6
UpdatedSep 2, 2026
Confidence2
OSWorld 2.0V3 evidence
LMSpeed rank#6
Score41.7
UpdatedSep 2, 2026
Confidence2
AutomationBench
LMSpeed rank#2
Score31.4
UpdatedSep 2, 2026
Confidence2
Toolathlon-Verified
LMSpeed rank#2
Score77.8
UpdatedSep 2, 2026
Confidence2
Toolathlon Verified Pass@3
LMSpeed rank#2
Score81.5
UpdatedSep 2, 2026
Confidence2
Toolathlon Verified Pass³
LMSpeed rank#1
Score73.1
UpdatedSep 2, 2026
Confidence2
Toolathlon Verified avg. turns
LMSpeed rank#1
Score23.7
UpdatedSep 2, 2026
Confidence2
AA Agentic Index
LMSpeed rank#1
Score61.3
UpdatedSep 2, 2026
Confidence2
GDPval-AA
LMSpeed rank#1
Score67.7
UpdatedSep 2, 2026
Confidence2
GDPval-AAV3 evidence
LMSpeed rank#2
Score1853.0
UpdatedSep 2, 2026
Confidence2
AA BriefcaseV3 evidence
LMSpeed rank#2
Score1694.0
UpdatedSep 2, 2026
Confidence2
AA Harvey LABV3 evidence
LMSpeed rank#5
Score93.0
UpdatedSep 2, 2026
Confidence2
AA Tau3 BankingV3 evidence
LMSpeed rank#2
Score47.2
UpdatedSep 2, 2026
Confidence2

Coding

V3.0

undefined metric} other undefined metrics}} · Rated

Score67.9#180% interval61.0–74.84/4 Measured dimensionsSciCode57.6%#5 / 207Coding score80.3#4 / 82
Dimensions and evidence
Code generation65.5
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
scicode · scicode · z 1.94 · q 1.00
Repository engineering71.8
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
swe_pro · swe_pro · z 3.00 · q 1.00
Debugging & testing74.8
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
swe_multilingual · benchlm_coding_swe_multilingual · z 3.00 · q 1.00
Tool-assisted development & quality59.6
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
terminalbench · aa_terminal_bench21 · z 1.10 · q 1.00
SciCodeV3 evidence
LMSpeed rank#5
Score57.6%
UpdatedSep 2, 2026
Confidence4
Coding score
LMSpeed rank#4
Score80.3
UpdatedSep 2, 2026
Confidence2
AA Terminal-Bench 2.1V3 evidence
LMSpeed rank#1
Score91.4
UpdatedSep 2, 2026
Confidence2
SWE-bench ProV3 evidence
LMSpeed rank#1
Score81.2
UpdatedSep 2, 2026
Confidence2
SWE MultilingualV3 evidence
LMSpeed rank#2
Score89.1
UpdatedSep 2, 2026
Confidence2
SWE Multimodal
LMSpeed rank#2
Score54.7
UpdatedSep 2, 2026
Confidence2
ProgramBench
LMSpeed rank#2
Score87.6
UpdatedSep 2, 2026
Confidence2
CursorBench v3.2
LMSpeed rank#1
Score73.4
UpdatedSep 2, 2026
Confidence2
AA Coding Index
LMSpeed rank#1
Score81.6
UpdatedSep 2, 2026
Confidence2
AA-SciCodeV3 evidence
LMSpeed rank#1
Score62.0
UpdatedSep 2, 2026
Confidence2

Reasoning

V3.0

undefined metric} other undefined metrics}} · Rated

Score61.3#180% interval52.6–70.03/4 Measured dimensionsGPQA90.6%#26 / 214HLE55.9%#3 / 211
Dimensions and evidence
Abstract logic59.1
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
arc_agi · arc_agi2 · z 1.07 · q 1.00
Scientific & causal reasoning66.7
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
critpt · critpt · z 1.08 · q 1.00
gpqa · gpqa · z 1.36 · q 1.00
hle · hle · z 1.43 · q 1.00
hle · hle_no_tools · z 1.90 · q 1.00
Multi-step constraintsPrior only
Evidence integration & verification58.1
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
lcr · lcr · z 0.71 · q 1.00
GPQAV3 evidence
LMSpeed rank#26
Score90.6%
UpdatedSep 2, 2026
Confidence4
HLEV3 evidence
LMSpeed rank#3
Score55.9%
UpdatedSep 2, 2026
Confidence4
Reasoning score
LMSpeed rank#4
Score83.1
UpdatedSep 2, 2026
Confidence2
ARC-AGI-1
LMSpeed rank#1
Score97.5
UpdatedSep 2, 2026
Confidence2
ARC-AGI-2V3 evidence
LMSpeed rank#3
Score90.0
UpdatedSep 2, 2026
Confidence2
AA-LCRV3 evidence
LMSpeed rank#5
Score80.0
UpdatedSep 2, 2026
Confidence2
CritPtV3 evidence
LMSpeed rank#5
Score29.7
UpdatedSep 2, 2026
Confidence2

Knowledge

V3.0

undefined metric} other undefined metrics}} · Provisional

Score68.780% interval54.7–82.71/4 Measured dimensionsKnowledge score97.0#1 / 70HLE w/o tools60.9#1 / 22
Dimensions and evidence
Broad knowledge68.7
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
aa_omniscience · aa_omniscience_index · z 1.60 · q 1.00
benchlm_category_knowledge · benchlm_category_knowledge · z 1.83 · q 1.00
Professional knowledgePrior only
FactualityPrior only
Retrieval & open-book usePrior only
Knowledge scoreV3 evidence
LMSpeed rank#1
Score97.0
UpdatedSep 2, 2026
Confidence2
HLE w/o tools
LMSpeed rank#1
Score60.9
UpdatedSep 2, 2026
Confidence2
Artificial Analysis Intelligence IndexV3 evidence
LMSpeed rank#1
Score65.7
UpdatedSep 2, 2026
Confidence2
AA-GPQA Diamond
LMSpeed rank#5
Score93.7
UpdatedSep 2, 2026
Confidence2
AA-HLE
LMSpeed rank#1
Score59.1
UpdatedSep 2, 2026
Confidence2
AA-Omniscience IndexV3 evidence
LMSpeed rank#1
Score43.5
UpdatedSep 2, 2026
Confidence2
AA-Omniscience Accuracy
LMSpeed rank#1
Score67.2
UpdatedSep 2, 2026
Confidence2
AA-Omniscience Hallucination Rate
LMSpeed rank#54
Score72.6
UpdatedSep 2, 2026
Confidence2

Math

V3.0

undefined metric} other undefined metrics}} · No data

No data80% interval30.8–69.20/4 Measured dimensions
Dimensions and evidence
Foundational mathPrior only
Competition mathPrior only
Advanced proofsPrior only
Applied & tool-assisted mathPrior only

Multilingual

V3.0

undefined metric} other undefined metrics}} · No data

No data80% interval30.8–69.20/4 Measured dimensions
Dimensions and evidence
Cross-language understandingPrior only
Multilingual generationPrior only
Reasoning transferPrior only
Low-resource robustnessPrior only

Multimodal

V3.0

undefined metric} other undefined metrics}} · No data

No data80% interval30.8–69.20/4 Measured dimensions
Dimensions and evidence
Perception & OCRPrior only
Document & spatial understandingPrior only
Visual reasoningPrior only
Video & grounded actionPrior only

Instruction following

V3.0

undefined metric} other undefined metrics}} · No data

No data80% interval30.8–69.20/4 Measured dimensions
Dimensions and evidence
Constraint followingPrior only
Structured outputPrior only
Novel-instruction generalizationPrior only
Multi-turn & long instructionsPrior only

Alternatives & Similar Models

Also known as

anthropic/claude-fable-5.1

Rankings are based on community-submitted tests and periodic health probes. Advisory only, not official data.Standard benchmark data may include BenchLM and other public sources.

Build with LMSpeed Data

Free

Free public API for LLM pricing, benchmarks & provider data

  • Real-time pricing & availability
  • Speed & latency benchmarks
  • 300+ models, 600+ providers
  • 1,000 requests/day
View API Documentation