On this pageKimi K3 is one of the best models for agentic legal work
Model Report

Is Kimi K3 good at legal work?

Kimi K3 ranks first on the Artificial Analysis Harvey legal-agent benchmark, fourth on the Vals AI implementation, and third on legal research. Our view: Kimi K3 is one of the best, and arguably the best, models for agentic legal work when raw model capability matters most.

More details about this document
Editor
, OpenAgreements editor
License
CC BY 4.0

Kimi K3 ranks first on Artificial Analysis's Harvey legal-agent benchmark, fourth on Vals AI's implementation, and third on Vals AI's legal-research benchmark; our view is that Kimi K3 is one of the best, and arguably the best, models for agentic legal work when raw model capability matters most. The ranking spread comes from Artificial Analysis's Harvey LAB-AA results reviewed July 19, 2026, Vals AI's Harvey results reviewed July 19, 2026, and Vals AI's Legal Research Bench reviewed July 19, 2026. Kimi K3's first-place result is a serious signal because Harvey LAB tests hard, realistic assignments rather than short legal questions, according to the Artificial Analysis methodology reviewed July 19, 2026.

Kimi K3 is a frontier-class, 2.8-trillion-parameter model from Moonshot AI that launched on July 16, 2026, according to Axios's July 16, 2026 release report and VentureBeat's July 16, 2026 model report. Kimi K3 also ranks fourth of 187 models with a score of 57 on Artificial Analysis's general, non-legal Intelligence Index, according to the Artificial Analysis model page reviewed July 19, 2026.

Kimi K3 leads a genuinely hard legal-agent benchmark

Kimi K3 ranks first on Artificial Analysis's Harvey LAB-AA with a 26.7% all-pass rate, ahead of Claude Fable 5 at 14.2% and Grok 4.5 at 13.3%, according to the Artificial Analysis results reviewed July 19, 2026. Harvey LAB-AA's all-pass metric requires every rubric criterion for a task to pass, so Kimi K3's 26.7% score gives no partial credit for an almost-complete deliverable, according to the Artificial Analysis methodology reviewed July 19, 2026.

Harvey LAB-AA uses Harvey's 120 private tasks across 24 legal practice areas, with assignments such as memos, disclosure schedules, deposition summaries, and redlines, according to the Artificial Analysis methodology reviewed July 19, 2026. Harvey LAB-AA asks an agent to read matter documents, work inside a sandbox, and produce finished legal deliverables, according to the Artificial Analysis benchmark page reviewed July 19, 2026. Harvey LAB-AA's demanding scope is why Kimi K3's first-place finish matters: a large lead on realistic, long-horizon legal work is stronger evidence than a narrow win on legal trivia or isolated classification.

Artificial Analysis grades each criterion with a single judge, Gemini 3.1 Pro, according to the Artificial Analysis methodology reviewed July 19, 2026. Gemini 3.1 Pro ranks below Kimi K3 on Harvey LAB-AA, according to the Artificial Analysis results reviewed July 19, 2026. Using a competitor as the judge does not eliminate evaluator bias, but using a competitor makes a simple same-family advantage an unconvincing explanation for Kimi K3's lead.

Kimi K3's split Harvey rankings probably reflect the harness more than the judge

Kimi K3 ranks fourth in Vals AI's Harvey implementation with a 10.8% task-pass rate and about a 90.8% criteria-pass rate, according to the Vals AI results reviewed July 19, 2026. Vals AI uses two judges, GPT-5.5 and Claude Sonnet 4.6, while Artificial Analysis uses one judge, Gemini 3.1 Pro, according to the Vals AI methodology reviewed July 19, 2026 and Artificial Analysis methodology reviewed July 19, 2026.

Kimi K3's move from first place at Artificial Analysis to fourth place at Vals AI might come from the judges, or Kimi K3's move might come from the tools and skill files available during generation. Our hypothesis is that the generation harness is the more likely driver. A judge swap can, in principle, be tested cheaply by rescoring the same saved outputs, while a harness change alters how a model reads files, edits documents, and creates the deliverables that later reach the judges.

Artificial Analysis runs Kimi K3 through its open-source Stirrup harness with sandboxed code execution and document-processing software, but Artificial Analysis does not provide Harvey's custom document-generation skill scripts, according to the Artificial Analysis methodology reviewed July 19, 2026. Vals AI follows Harvey's environment with six file-system tools plus docx, xlsx, and pptx skills, according to the Vals AI methodology reviewed July 19, 2026. Vals AI's skill files contain instructions and scripts for producing office documents, according to the Vals AI methodology reviewed July 19, 2026. Artificial Analysis describes the Stirrup score as raw model capability because Kimi K3 must solve the document workflow without Harvey's custom skills, according to the Artificial Analysis methodology reviewed July 19, 2026.

Artificial Analysis also requires exact output filenames, grades only text extracted from submitted files, and states that Harvey LAB-AA is not directly comparable with Harvey's published results, according to the Artificial Analysis methodology reviewed July 19, 2026. Artificial Analysis's stricter filename matching and text-only grading could move scores too, so the public evidence does not isolate one definitive cause for Kimi K3's ranking gap.

Artificial Analysis's raw-capability setup may be more representative for a reader who would not give an AI a polished set of Harvey-style office-document skills. Moonshot reportedly focused Kimi K3's training on office file formats such as .docx, .xlsx, and .pptx, an emphasis that adds useful context to Kimi K3's performance when document helpers are withheld, according to VentureBeat's July 16, 2026 report. Kimi K3's ability to handle office formats with less scaffolding would be a practical advantage for legal teams without custom document helpers.

Kimi K3's other legal results support a top-tier verdict without making Kimi K3 the winner of every kind of legal task.

BenchmarkKimi K3 resultWhat the result suggests
Vals AI Legal Research BenchKimi K3 ranks third at 44.2% all-pass, behind Claude Fable 5 at 49.5% and GPT-5.6 Sol at 48.1%, according to Vals AI results reviewed July 19, 2026.Kimi K3 is a leading choice for open-ended, multi-source legal research.
Vals AI LegalBenchKimi K3 ranks ninth at 86.0%, while Claude Fable 5 leads at 88.6%, according to Vals AI results reviewed July 19, 2026.Kimi K3 is strong on broad legal reasoning tasks, but Kimi K3 is less distinctive on this less-agentic eval.
Artificial Analysis Intelligence IndexKimi K3 ranks fourth of 187 with a score of 57, according to the Artificial Analysis model page reviewed July 19, 2026.Kimi K3's legal performance sits on top of frontier-level general capability.

Kimi K3's legal-research result deserves particular weight for research-heavy practice because Vals AI requires agents to find and synthesize statutes, cases, and other authorities across realistic questions, according to the Vals AI Legal Research Bench reviewed July 19, 2026. Kimi K3's LegalBench result is still strong, but the narrow spread between Kimi K3's 86.0% and the leading 88.6% makes ninth place sound more dramatic than the underlying score difference, according to the Vals AI LegalBench results reviewed July 19, 2026.

Kimi K3 is API-only until the full weights arrive

Kimi K3 is available through hosted access today, but Kimi K3 is not yet downloadable or self-hostable, according to Axios's July 16, 2026 report. Moonshot AI plans to release the full Kimi K3 weights on July 27, 2026, according to Axios's July 16, 2026 report and VentureBeat's July 16, 2026 report.

Kimi K3's API-only period matters for legal teams that require on-premises deployment, independent inspection, or control over the complete inference environment. Kimi K3's planned open-weight status may become an important advantage after July 27, but Kimi K3 remains unavailable for self-hosting as of July 19, 2026, according to Axios's July 16, 2026 report.

Kimi K3's Harvey results suggest that legal-AI performance depends on more than the model name. A strong model can look materially different when one harness supplies document skills and another harness asks the model to build the workflow itself. Legal teams should therefore evaluate both the model and the legal knowledge, instructions, tools, and quality controls surrounding the model.

OpenAgreements offers a free source of up-to-date legal knowledge for areas of law that change quickly. Start with the most-used practice guides, then turn recurring work into a more reliable workflow with legal checklists. Better legal inputs cannot replace professional review, but better legal inputs can help a capable model work from current law instead of generic memory.

OpenAgreements is independent and is not affiliated with Moonshot AI.