Model reports
For each frontier model, one page answers a single question: is it good at legal work? Each report synthesizes dated, sourced measurement and stays current as new results land. See also our legal AI evaluations.
Model reportIs Kimi K3 good at legal work?Kimi K3 ranks first on the Artificial Analysis Harvey legal-agent benchmark, fourth on the Vals AI implementation, and third on legal research. Our view: Kimi K3 is one of the best, and arguably the best, models for agentic legal work when raw model capability matters most.
Model reportIs Gemini 3.6 Flash good at legal work?Gemini 3.6 Flash pairs first-ranked output speed with lower pricing, making it a strong candidate for high-volume legal workflows; the model is new enough that legal-specific benchmarks have not yet published scores.
Model reportIs Claude Opus 5 good at legal work?Claude Opus 5 is the best available model for legal research, combining the leading Vals AI result with the lowest task cost among the top three, but Claude Opus 5 is a top-tier rather than leading choice for agentic legal work.
Model reportIs Muse Spark 1.1 good at legal work?Muse Spark 1.1 is the best available model for producing finished legal work product, and a mid-pack choice for legal research and legal knowledge, pairing a first-place agentic result with unusually low benchmark costs.- Model reportIs Grok 4.6 good at legal work?Grok 4.6 is one of the best models for legal work: third on a strict end-to-end legal-agent benchmark and fourth on legal research, but most benchmark assignments still contain at least one miss.