Meta
·Released on Aug 5, 2026

Muse Spark 1.2 API Benchmarks, Pricing & Provider Data

Compare Muse Spark 1.2 with another model

Choose a model to open its comparison page.

Share on X
LLM

The page also shows measured API speed and first-token latency.

Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context...

Quality
#20of 112
69.0
LMSpeed score
Speed
#17of 80
69char/s
22.53 s

Category Performance

Observed-capability estimates with uncertainty reported separately from independent benchmark families.

Coverage
4 / 8
Methodology
V3.0
Category PerformanceObserved-capability estimates with uncertainty reported separately from independent benchmark families.Agents56.3Coding55.7Reasoning58.5Knowledge58.8Math-Multilingual-Multimodal-Instruction following-
#1Knowledge58.8Provisional1/4 Measured dimensions
80% interval: 44.872.8
broad knowledge58.8
aa_omniscience · aa_omniscience_index / benchlm_category_knowledge · benchlm_category_knowledge
professional knowledgePrior only
factualityPrior only
retrieval open bookPrior only
#2Reasoning58.5Estimated2/4 Measured dimensions
80% interval: 47.769.2
abstract logicPrior only
scientific causal61.4
critpt · critpt / gpqa · gpqa / hle · hle
multistep constraintsPrior only
evidence verification55.5
lcr · lcr
#3Agents56.3RatedGlobal rank #203/4 Measured dimensions
80% interval: 47.465.3
planningPrior only
tool use50.7
tau · aa_tau3_banking
environment execution62.9
enterprise_ops_gym · aa_enterprise_ops_gym / gdpval_aa · benchlm_agentic_gdpval_aa
recovery reliability55.4
briefcase · aa_briefcase_elo
#4Coding55.7Estimated2/4 Measured dimensions
80% interval: 43.867.6
code generation61.7
scicode · scicode
repository engineeringPrior only
debugging testingPrior only
tooling quality49.7
terminalbench · aa_terminal_bench21
No data:MathMultilingualMultimodalInstruction following

Specifications

Input and output token limits for this model, plus how it ranks on long-context understanding.

INPUT
1.0Mtokens
1.3K pages of text
OUTPUT
943.7Ktokens
8K128K1M4M
1.0M

Features

Technical Details

Input
Output
Released
Aug 2026
Documentation
Tokenizer
Other
Architecture
text+image+file+audio+video->text
Moderated
Yes
Supported parameters
include_reasoningmax_tokensreasoningreasoning_effortrepetition_penaltyresponse_formatstructured_outputstemperaturetool_choicetoolstop_ktop_p

Rankings

Excels at

Falls behind in

Detailed scores

Updated: Sep 14, 2026

Overall

undefined metric} other undefined metrics}}

Overall score69.0#20 / 112DeepSWE59.3#14 / 21
Overall score
LMSpeed rank#20
Score69.0
UpdatedSep 14, 2026
Confidence2
DeepSWE
LMSpeed rank#14
Score59.3
UpdatedSep 10, 2026
Confidence2

Speed & latency

undefined metric} other undefined metrics}}

Output speed211.3 tok/s#17 / 80Time to first token12.79 s#68 / 80
Output speed
LMSpeed rank#17
Score211.3 tok/s
UpdatedSep 14, 2026
Confidence4
Time to first token
LMSpeed rank#68
Score12.79 s
UpdatedSep 14, 2026
Confidence4

Pricing

undefined metric} other undefined metrics}}

Input price$1.25/M#123 / 186Output price$4.25/M#116 / 186
Input price
LMSpeed rank#123
Score$1.25/M
UpdatedSep 14, 2026
Confidence4
Output price
LMSpeed rank#116
Score$4.25/M
UpdatedSep 14, 2026
Confidence4

Agents

V3.0

undefined metric} other undefined metrics}} · Rated

Score56.3#2080% interval47.4–65.33/4 Measured dimensionsTerminal-Bench 2.182.9#7 / 10GDPval-AA1631.0#9 / 72
Dimensions and evidence
Planning & decompositionPrior only
Tool use50.7
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
tau · aa_tau3_banking · z -0.43 · q 1.00
Environment & long-horizon execution62.9
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
enterprise_ops_gym · aa_enterprise_ops_gym · z 0.45 · q 1.00
gdpval_aa · benchlm_agentic_gdpval_aa · z 1.17 · q 1.00
Recovery & completion reliability55.4
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
briefcase · aa_briefcase_elo · z 0.07 · q 1.00
Terminal-Bench 2.1
LMSpeed rank#7
Score82.9
UpdatedSep 14, 2026
Confidence2
GDPval-AAV3 evidence
LMSpeed rank#9
Score1631.0
UpdatedSep 14, 2026
Confidence2
Agentic score
LMSpeed rank#14
Score81.2
UpdatedSep 14, 2026
Confidence2
AA Agentic Index
LMSpeed rank#15
Score44.0
UpdatedSep 14, 2026
Confidence2
GDPval-AA
LMSpeed rank#12
Score51.2
UpdatedSep 14, 2026
Confidence2
AA EnterpriseOps-GymV3 evidence
LMSpeed rank#6
Score47.3
UpdatedSep 14, 2026
Confidence2
Terminal-Bench 2.1 (Vals)
LMSpeed rank#18
Score69.7
UpdatedSep 14, 2026
Confidence2
AA BriefcaseV3 evidence
LMSpeed rank#10
Score1361.0
UpdatedSep 3, 2026
Confidence1
AA Tau3 BankingV3 evidence
LMSpeed rank#14
Score34.8
UpdatedSep 3, 2026
Confidence1

Coding

V3.0

undefined metric} other undefined metrics}} · Estimated

Score55.780% interval43.8–67.62/4 Measured dimensionsSciCode57.4%#9 / 89Terminal-Bench 2.182.9#7 / 11
Dimensions and evidence
Code generation61.7
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
scicode · scicode · z 1.39 · q 1.00
Repository engineeringPrior only
Debugging & testingPrior only
Tool-assisted development & quality49.7
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
terminalbench · aa_terminal_bench21 · z -0.54 · q 1.00
SciCodeV3 evidence
LMSpeed rank#9
Score57.4%
UpdatedSep 14, 2026
Confidence4
Terminal-Bench 2.1
LMSpeed rank#7
Score82.9
UpdatedSep 14, 2026
Confidence2
Coding score
LMSpeed rank#25
Score61.0
UpdatedSep 14, 2026
Confidence2
AA Coding Index
LMSpeed rank#16
Score72.2
UpdatedSep 14, 2026
Confidence2
AA-SciCodeV3 evidence
LMSpeed rank#7
Score57.4
UpdatedSep 14, 2026
Confidence2
FrontierSWE v2
LMSpeed rank#10
Score12.0
UpdatedSep 14, 2026
Confidence2
SWE-bench (Vals)
LMSpeed rank#11
Score86.6
UpdatedSep 14, 2026
Confidence2
VulcanBench v3
LMSpeed rank#3
Score87.0
UpdatedSep 14, 2026
Confidence2
DeepSWE
LMSpeed rank#13
Score59.3
UpdatedSep 14, 2026
Confidence2
AA Terminal-Bench 2.1V3 evidence
LMSpeed rank#14
Score80.1
UpdatedSep 3, 2026
Confidence1

Reasoning

V3.0

undefined metric} other undefined metrics}} · Estimated

Score58.580% interval47.7–69.22/4 Measured dimensionsGPQA90.4%#32 / 218HLE45.5%#17 / 216
Dimensions and evidence
Abstract logicPrior only
Scientific & causal reasoning61.4
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
critpt · critpt · z 0.73 · q 1.00
gpqa · gpqa · z 0.92 · q 1.00
hle · hle · z 1.01 · q 1.00
Multi-step constraintsPrior only
Evidence integration & verification55.5
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
lcr · lcr · z 0.38 · q 1.00
GPQAV3 evidence
LMSpeed rank#32
Score90.4%
UpdatedSep 14, 2026
Confidence4
HLEV3 evidence
LMSpeed rank#17
Score45.5%
UpdatedSep 14, 2026
Confidence4
AA-LCRV3 evidence
LMSpeed rank#29
Score79.0
UpdatedSep 14, 2026
Confidence2
CritPtV3 evidence
LMSpeed rank#19
Score17.7
UpdatedSep 14, 2026
Confidence2
Reasoning score
LMSpeed rank#21
Score75.3
UpdatedSep 14, 2026
Confidence2

Knowledge

V3.0

undefined metric} other undefined metrics}} · Provisional

Score58.880% interval44.8–72.81/4 Measured dimensionsKnowledge score76.0#26 / 83Artificial Analysis Intelligence Index39.8#25 / 117
Dimensions and evidence
Broad knowledge58.8
undefined family} other undefined families}} · undefined metric} other undefined metrics}}
aa_omniscience · aa_omniscience_index · z 1.03 · q 1.00
benchlm_category_knowledge · benchlm_category_knowledge · z 0.38 · q 1.00
Professional knowledgePrior only
FactualityPrior only
Retrieval & open-book usePrior only
Knowledge scoreV3 evidence
LMSpeed rank#26
Score76.0
UpdatedSep 14, 2026
Confidence2
Artificial Analysis Intelligence IndexV3 evidence
LMSpeed rank#25
Score39.8
UpdatedSep 14, 2026
Confidence2
AA-GPQA Diamond
LMSpeed rank#31
Score90.4
UpdatedSep 14, 2026
Confidence2
AA-HLE
LMSpeed rank#14
Score45.5
UpdatedSep 14, 2026
Confidence2
AA-Omniscience IndexV3 evidence
LMSpeed rank#11
Score27.2
UpdatedSep 14, 2026
Confidence2
AA-Omniscience Accuracy
LMSpeed rank#25
Score45.4
UpdatedSep 14, 2026
Confidence2
AA-Omniscience Hallucination Rate
LMSpeed rank#95
Score33.3
UpdatedSep 14, 2026
Confidence2
MMLU-Pro (Vals)
LMSpeed rank#17
Score88.3
UpdatedSep 14, 2026
Confidence2

Math

V3.0

undefined metric} other undefined metrics}} · No data

No data80% interval30.8–69.20/4 Measured dimensions
Dimensions and evidence
Foundational mathPrior only
Competition mathPrior only
Advanced proofsPrior only
Applied & tool-assisted mathPrior only

Multilingual

V3.0

undefined metric} other undefined metrics}} · No data

No data80% interval30.8–69.20/4 Measured dimensions
Dimensions and evidence
Cross-language understandingPrior only
Multilingual generationPrior only
Reasoning transferPrior only
Low-resource robustnessPrior only

Multimodal

V3.0

undefined metric} other undefined metrics}} · No data

80% interval30.8–69.20/4 Measured dimensionsDesign Arena Website1322.0#3 / 78
Dimensions and evidence
Perception & OCRPrior only
Document & spatial understandingPrior only
Visual reasoningPrior only
Video & grounded actionPrior only
Design Arena Website
LMSpeed rank#3
Score1322.0
UpdatedSep 14, 2026
Confidence2

Instruction following

V3.0

undefined metric} other undefined metrics}} · No data

No data80% interval30.8–69.20/4 Measured dimensions
Dimensions and evidence
Constraint followingPrior only
Structured outputPrior only
Novel-instruction generalizationPrior only
Multi-turn & long instructionsPrior only

OpenRouter endpoints

1 endpoints

Third-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.

Provider endpointInputOutput1d uptime30m latency30m throughputContext / output
Meta
meta
$1.25/M$4.25/M100%undefined tokens / undefined tokens

Alternatives & Similar Models

Also known as

meta/muse-spark-1.2muse-spark-1.2

Rankings are based on community-submitted tests and periodic health probes. Advisory only, not official data.Standard benchmark data may include BenchLM and other public sources.

Build with LMSpeed Data

Free

Free public API for LLM pricing, benchmarks & provider data

  • Real-time pricing & availability
  • Speed & latency benchmarks
  • 300+ models, 600+ providers
  • 1,000 requests/day
View API Documentation