# Is Claude Opus 5 good at legal work?[^about]

Claude Opus 5 is the best available model for legal research, combining the leading Vals AI result with the lowest task cost among the top three, but Claude Opus 5 is a top-tier rather than leading choice for agentic legal work.

## Claude Opus 5 is the best available model for legal research, but not for legal work overall {#claude-opus-5-legal-verdict}

Claude Opus 5 is the best available model for legal research, and a top-tier but not leading choice for agentic legal work. Claude Opus 5 ranks first on Vals AI's Legal Research Bench at 55.29% all-pass and costs $6.76 per test, the lowest cost among the top three systems, according to the [Vals AI Legal Research Bench reviewed July 28, 2026](https://www.vals.ai/benchmarks/legal_research).

Claude Opus 5 does not lead legal work overall. Claude Fable 5 beats Claude Opus 5 on Vals AI's LegalBench and Vals AI's independent run of Harvey's held-out Legal Agent Benchmark, according to the [Vals AI LegalBench results reviewed July 28, 2026](https://www.vals.ai/benchmarks/legal_bench) and [Vals AI Harvey results reviewed July 28, 2026](https://www.vals.ai/benchmarks/hlab). Choose Claude Opus 5 first for demanding legal research, but compare Claude Opus 5 with other leaders for long, tool-using legal assignments.

## Claude Opus 5 leads legal research while costing less per task than the nearest leaders {#claude-opus-5-leads-legal-research}

Claude Opus 5 leads all 27 systems on Vals AI's proprietary Legal Research Bench at 55.29% all-pass, a 5.77-point lead over the prior leader, according to the [Vals AI Legal Research Bench reviewed July 28, 2026](https://www.vals.ai/benchmarks/legal_research). Claude Opus 5 costs $6.76 per test, compared with $9.79 for second-place Claude Fable 5 and $21.61 for third-place GPT-5.6 Sol, according to the [Vals AI Legal Research Bench reviewed July 28, 2026](https://www.vals.ai/benchmarks/legal_research).

Vals AI asks agents to use case-law search, web search, and document retrieval to answer legal questions across eight practice areas, according to the [Vals AI Legal Research Bench reviewed July 28, 2026](https://www.vals.ai/benchmarks/legal_research). Claude Opus 5 reaches 90.58% weighted pass with partial credit, compared with 55.29% when every rubric check must pass, according to the [Vals AI Legal Research Bench reviewed July 28, 2026](https://www.vals.ai/benchmarks/legal_research). Every model loses 6 to 17 points on questions requiring synthesis across jurisdictions, courts, or legal regimes, according to the [Vals AI Legal Research Bench reviewed July 28, 2026](https://www.vals.ai/benchmarks/legal_research). Claude Opus 5's lead matters for research-heavy practice, but the synthesis drop remains a reason to inspect authorities and reasoning.

## Claude Opus 5's agentic legal record is top-tier but splits sharply across Harvey runs {#claude-opus-5-agentic-legal-gap}

Harvey reports Claude Opus 5 at 11.7% all-pass on Harvey's own Legal Agent Benchmark and calls Claude Opus 5 a meaningful step up over prior Opus models, according to the [Harvey Opus 5 report reviewed July 28, 2026](https://www.harvey.ai/blog/opus-5-in-harvey). Harvey sells legal AI and deploys Claude Opus 5 in Harvey's own product, according to the [Harvey Opus 5 report reviewed July 28, 2026](https://www.harvey.ai/blog/opus-5-in-harvey). Harvey's figure is vendor evidence from a company with direct product experience and a commercial stake in the model's reception.

Vals AI's independent held-out run puts Claude Opus 5 seventh at 6.67% task pass, below Claude Opus 4.8 at 9.58% and Claude Fable 5 at 11.25%, according to the [Vals AI Harvey results reviewed July 28, 2026](https://www.vals.ai/benchmarks/hlab). Claude Opus 5 costs $23.67 per test, more than twice Claude Opus 4.8's $10.22, and trails Claude Opus 4.8 on criteria pass, 87.7% to 87.86%, according to the [Vals AI Harvey results reviewed July 28, 2026](https://www.vals.ai/benchmarks/hlab). Claude Opus 5's independent result therefore sits below Claude Opus 4.8 on the same held-out set, which is the opposite of the direction Harvey reports.

Harvey's 11.7% all-pass figure and Vals AI's 6.67% task-pass figure are not directly comparable: Harvey reports all-pass from Harvey's own run, while Vals AI reports task pass from an independent run of Harvey's held-out set, according to the [Harvey Opus 5 report reviewed July 28, 2026](https://www.harvey.ai/blog/opus-5-in-harvey) and [Vals AI Harvey results reviewed July 28, 2026](https://www.vals.ai/benchmarks/hlab). Vals AI gives the agent six file and shell tools plus skills for `docx`, `pptx`, and `xlsx` work, according to the [Vals AI Harvey benchmark reviewed July 28, 2026](https://www.vals.ai/benchmarks/hlab). Harvey reports that Claude Opus 5 can match Claude Opus 4.8 at maximum reasoning while using lower reasoning levels and 26% fewer tokens on average, according to the [Harvey Opus 5 report reviewed July 28, 2026](https://www.harvey.ai/blog/opus-5-in-harvey).

Our hypothesis is that run configuration, including reasoning settings and the generation harness, contributes to the divergence. Different configurations can change the work product before grading, while different metrics change how success is summarized. The public evidence does not isolate one definitive cause, so neither figure should silently stand in for the other.

Claude Opus 5 has no Artificial Analysis Harvey LAB-AA result, according to the [Artificial Analysis Harvey LAB-AA results reviewed July 28, 2026](https://artificialanalysis.ai/evaluations/harvey-lab-aa). A Claude Opus 5 Harvey LAB-AA result is the most useful next evidence because it would add another independent measurement of finished legal deliverables.

## Claude Opus 5 is strong but not first on broader legal reasoning {#claude-opus-5-other-legal-benchmarks}

Claude Opus 5's other results reinforce the narrow verdict without making Claude Opus 5 the winner of every benchmark.

| Benchmark | Claude Opus 5 result | What the result suggests |
| --- | --- | --- |
| Vals AI LegalBench | Claude Opus 5 ranks fourth among 129 systems at 86.97%, tied with GPT-5.6 Sol on the displayed score and behind Claude Fable 5 at 88.56%, according to the [Vals AI LegalBench results reviewed July 28, 2026](https://www.vals.ai/benchmarks/legal_bench). | Claude Opus 5 is near the top on broad legal reasoning, but Claude Opus 5 does not lead. |
| Artificial Analysis Intelligence Index | Claude Opus 5 scores 61 against a median of 33, according to the [Artificial Analysis model page reviewed July 28, 2026](https://artificialanalysis.ai/models/claude-opus-5). | Claude Opus 5 has leading general capability, but this index is not legal evidence. |

Claude Opus 5 sits only 1.59 points behind Claude Fable 5 on LegalBench, so fourth place sounds more dramatic than the score gap, according to the [Vals AI LegalBench results reviewed July 28, 2026](https://www.vals.ai/benchmarks/legal_bench). Claude Fable 5's higher score still prevents a broader first-place claim.

Artificial Analysis's general, non-legal Intelligence Index combines nine evaluations, none of which is legal, according to the [Artificial Analysis model page reviewed July 28, 2026](https://artificialanalysis.ai/models/claude-opus-5). The general score supports Claude Opus 5's frontier status, but the legal verdict should rest on legal benchmarks rather than a non-legal index.

## Claude Opus 5 is closed-weight and broadly available through hosted services {#claude-opus-5-availability}

Claude Opus 5 is a closed-weight model available through the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, and Microsoft Foundry, according to [Anthropic's July 24, 2026 announcement](https://www.anthropic.com/news/claude-opus-5). Claude Opus 5 is also available through claude.ai, Claude Code, and Claude Cowork, according to [Anthropic's July 24, 2026 announcement](https://www.anthropic.com/news/claude-opus-5). Claude Opus 5 uses the Claude API model ID `claude-opus-5`, the AWS Bedrock ID `anthropic.claude-opus-5`, and the Google Cloud ID `claude-opus-5`, according to the [Anthropic model overview reviewed July 28, 2026](https://platform.claude.com/docs/en/docs/about-claude/models/overview).

Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens, according to [Anthropic's July 24, 2026 announcement](https://www.anthropic.com/news/claude-opus-5). Claude Opus 5 has a 1 million-token context window and a 128,000-token maximum output, according to the [Anthropic model overview reviewed July 28, 2026](https://platform.claude.com/docs/en/docs/about-claude/models/overview). Claude Opus 5 supports adaptive thinking, and Claude Opus 5 defaults to `high` effort in the Claude API and Claude Code unless the effort parameter is set explicitly, according to the [Anthropic model overview reviewed July 28, 2026](https://platform.claude.com/docs/en/docs/about-claude/models/overview).

Claude Opus 5's hosted reach makes Claude Opus 5 straightforward to test across several enterprise clouds and Anthropic products. Claude Opus 5's closed weights remain a practical constraint for legal teams that require self-hosting, independent inspection of the complete model, or control over the full inference environment.

## Claude Opus 5 still needs current legal knowledge and careful review {#claude-opus-5-needs-legal-knowledge}

Claude Opus 5's legal-research lead does not make current legal knowledge, matter-specific instructions, or professional review optional. Claude Opus 5's reliable knowledge cutoff and training-data cutoff are both May 2026, according to the [Anthropic model overview reviewed July 28, 2026](https://platform.claude.com/docs/en/docs/about-claude/models/overview). A research-strong model still works from a fixed training date, so legal teams should supply current authorities and verify that the final work reflects the law and facts that govern the matter.

Claude Opus 5's split agentic results also show why legal teams should evaluate the complete workflow. Test Claude Opus 5 with the documents, research tools, office-file skills, instructions, and review controls that Claude Opus 5 would receive in practice. Inspect citations, adverse authority, cross-jurisdiction synthesis, document formatting, and every task-specific requirement before relying on a finished work product.

For current legal inputs, start with OpenAgreements' most-used [practice guides](/practice-guides), then turn recurring work into a more reliable workflow with [legal checklists](/checklists). Better legal inputs cannot replace professional judgment, but better legal inputs can help Claude Opus 5 work from current law instead of generic memory.

OpenAgreements is independent and is not affiliated with Anthropic.


[^about]: By Steven Obiajulu, J.D. Published by [openagreements.org](https://openagreements.org). Last reviewed 2026-07-28. License: CC BY 4.0. Steven Obiajulu, J.D. wrote this essay. It states the author's views, synthesizes public sources, and is not legal advice. This article is for informational purposes only and does not create an attorney-client relationship. CC BY 4.0. Cite as Steven Obiajulu, *Is Claude Opus 5 good at legal work?*, OpenAgreements (last updated July 28, 2026), https://openagreements.org/models/claude-opus-5.
