Muse Spark 1.1 is the best available model for producing finished legal work product, but only a mid-pack legal researcher
Muse Spark 1.1 is the best available model for producing finished legal work product, and a mid-pack choice for legal research and legal knowledge. Muse Spark 1.1 ranks first of 27 on Vals AI's run of Harvey's held-out Legal Agent Benchmark at 20.00% task pass, ahead of Grok 4.5 at 12.92%, according to the Vals AI Harvey results reviewed July 29, 2026. Muse Spark 1.1 costs $0.80 per test on that benchmark, the cheapest cost per test among the top seven systems, according to the Vals AI Harvey results reviewed July 29, 2026.
That combination makes Muse Spark 1.1 the strongest first choice for turning matter files, instructions, and tools into a completed legal deliverable. Muse Spark 1.1 ranks tenth of 27 at 37.98% all-pass on Vals AI's Legal Research Bench, according to the Vals AI Legal Research Bench reviewed July 29, 2026. Muse Spark 1.1 ranks nineteenth of 129 at 84.98% on Vals AI's LegalBench, according to the Vals AI LegalBench results reviewed July 29, 2026. Choose Muse Spark 1.1 for execution, but supply strong sources and close review when research drives the work.
Muse Spark 1.1 leads the held-out agentic legal benchmark while costing less per test than every other top-seven model
Muse Spark 1.1 leads Vals AI's run of Harvey's held-out Legal Agent Benchmark with 20.00% task pass, followed by Grok 4.5 at 12.92% and Claude Fable 5 at 11.25%, according to the Vals AI Harvey results reviewed July 29, 2026. Muse Spark 1.1 costs $0.80 per test, the lowest cost per test in the top seven, according to the Vals AI Harvey results reviewed July 29, 2026. Muse Spark 1.1 therefore combines the board's clearest finished-work-product result with unusually low test cost.
Vals AI gives each agent six file and shell tools plus skills for docx, pptx, and xlsx work, according to the Vals AI Harvey benchmark reviewed July 29, 2026. The benchmark asks an agent to carry a legal assignment through a file and office-document workflow, according to the Vals AI Harvey benchmark reviewed July 29, 2026. Harvey grades a task as resolved only when every criterion passes, according to the Vals AI Harvey benchmark reviewed July 29, 2026. Muse Spark 1.1 satisfies 92.86% of individual criteria, the highest criteria pass rate on the board, according to the Vals AI Harvey results reviewed July 29, 2026.
A deliverable can satisfy most criteria and still fail because one required detail is wrong. Muse Spark 1.1 completes the entire rubric more often than every other tested system, according to the Vals AI Harvey results reviewed July 29, 2026. The result is comparative evidence for execution, not permission to skip review.
Muse Spark 1.1 is the best executor without being the best lawyer
Muse Spark 1.1 ranks tenth of 27 on Vals AI's Legal Research Bench at 37.98% all-pass, according to the Vals AI Legal Research Bench reviewed July 29, 2026. Muse Spark 1.1 ranks nineteenth of 129 on Vals AI's LegalBench at 84.98%, according to the Vals AI LegalBench results reviewed July 29, 2026. Those mid-pack placements are the necessary counterweight to Muse Spark 1.1's agentic lead.
Muse Spark 1.1 is nonetheless both the cheapest and fastest system on the Legal Research Bench board at $0.38 per test and 312.99 seconds, and Vals AI tags Muse Spark 1.1 Best Budget and Best Speed, according to the Vals AI Legal Research Bench reviewed July 29, 2026. The benchmark asks agents to search case law, the web, and documents before producing supported answers, according to the Vals AI Legal Research Bench reviewed July 29, 2026. Speed and cost make Muse Spark 1.1 appealing for repeated research runs, but they do not turn a tenth-place answer into the leading answer.
The LegalBench spread from first to nineteenth is only 3.58 points, so Muse Spark 1.1's rank sounds more dramatic than the score gap, according to the Vals AI LegalBench results reviewed July 29, 2026. Rank and score belong together: Muse Spark 1.1 is not a LegalBench leader, but Muse Spark 1.1 remains competitive on the displayed percentage.
The benchmark Muse Spark 1.1 wins measures whether an agent can carry a tool-using assignment through to a finished deliverable, according to the Vals AI Harvey benchmark reviewed July 29, 2026. The benchmarks where Muse Spark 1.1 sits mid-pack measure legal research and knowledge, according to the Vals AI Legal Research Bench reviewed July 29, 2026 and Vals AI LegalBench results reviewed July 29, 2026. A model can be the best executor without being the best lawyer. The Claude Opus 5 report shows the opposite split. Choose based on the workflow's bottleneck: finding law or producing finished work.
Muse Spark 1.1's broader benchmark record confirms exceptional execution rather than across-the-board legal leadership
| Benchmark | Muse Spark 1.1 result | What the result suggests |
|---|---|---|
| Artificial Analysis Harvey LAB-AA | Muse Spark 1.1 ranks third at a 93.1% criterion pass rate, behind Kimi K3 at 94.6% and Claude Fable 5 with an Opus 4.8 fallback at 93.6%, according to the Artificial Analysis Harvey LAB-AA results reviewed July 29, 2026. | Muse Spark 1.1 also performs strongly in a different agentic legal run. |
| Vals AI LegalBench | Muse Spark 1.1 ranks nineteenth of 129 at 84.98%, while Claude Fable 5 leads at 88.56%, according to the Vals AI LegalBench results reviewed July 29, 2026. | Muse Spark 1.1 is competitive but not leading on broad legal knowledge and reasoning. |
| Artificial Analysis Intelligence Index | Muse Spark 1.1 scores 51 against a median of 33 among comparable models, according to the Artificial Analysis model page reviewed July 29, 2026. | Muse Spark 1.1 has strong general capability, but the index is non-legal evidence. |
Artificial Analysis labels Muse Spark 1.1's 93.1% Harvey LAB-AA result as the criterion pass rate, meaning the share of individual rubric criteria satisfied, according to the Artificial Analysis Harvey LAB-AA results reviewed July 29, 2026. Artificial Analysis's criterion pass rate is not Vals AI's task pass rate and is not an all-pass rate, according to the Artificial Analysis Harvey LAB-AA results reviewed July 29, 2026. Artificial Analysis runs the Stirrup harness without Harvey's custom document-generation skill scripts, so the results are not directly comparable, according to the Artificial Analysis Harvey LAB-AA results reviewed July 29, 2026.
Artificial Analysis's Intelligence Index is a general, non-legal measure, according to the Artificial Analysis model page reviewed July 29, 2026. The legal verdict should rest on legal benchmarks.
Muse Spark 1.1 is a closed-weight, API-only model with low token pricing and a million-token context window
Muse Spark 1.1 is a closed-weight, API-only model from Meta Superintelligence Labs that was released on July 9, 2026, according to Meta's July 9, 2026 announcement. Meta opened the Meta Model API in public preview with the release, according to Meta's July 9, 2026 announcement.
Muse Spark 1.1 is available through the Meta Model API public preview, the Meta AI app in Thinking mode, and meta.ai, according to Meta's July 9, 2026 announcement. Muse Spark 1.1 accepts text, images, video, audio, and PDF inputs and produces text-only output, according to Meta's July 9, 2026 announcement. Muse Spark 1.1 has a 1 million-token context window with active context management, according to Meta's July 9, 2026 announcement.
Muse Spark 1.1 costs $1.25 per million input tokens and $4.25 per million output tokens, according to the Artificial Analysis model page reviewed July 29, 2026. Hosted access and long context make Muse Spark 1.1 straightforward to test on large matters. Closed weights constrain teams that require self-hosting.
Muse Spark 1.1 still needs current legal knowledge and careful professional review
Muse Spark 1.1's strength at executing work product makes the legal knowledge supplied to Muse Spark 1.1 more important, not less important. A strong executor working from wrong or stale law produces confident, well-formatted, wrong output. Legal teams should supply current authorities, matter-specific instructions, and approved templates before asking Muse Spark 1.1 to complete a legal assignment.
Muse Spark 1.1's benchmark split argues for testing the whole workflow. Evaluate Muse Spark 1.1 with the research services, source documents, tools, and controls that Muse Spark 1.1 would receive in practice. Inspect citations, adverse authority, jurisdictional fit, calculations, formatting, and every task-specific requirement.
For current legal inputs, start with OpenAgreements' most-used practice guides, then turn recurring work into a more reliable workflow with legal checklists. Better legal inputs cannot replace professional judgment, but better legal inputs can help Muse Spark 1.1 execute from current law instead of generic memory.
OpenAgreements is independent and is not affiliated with Meta.