Research

OpenAgreements maintains a 50-state legal knowledge base of practice guides, contract templates, and contract review checklists. In the course of maintaining it, we study how AI models handle legal work, and we publish each study with its method and data.

The evaluation papers replicate and measure concrete failure modes we observe in ordinary maintenance work. The principal one we document is what we call institutional-knowledge leakage: some frontier models carry internal-facing analysis from the knowledge base into external-facing deliverables.

The dataset paper describes the structured expert-correction data we capture as part of our Git-based maintenance workflow, including diffs, rationales, and the authorities considered and not applied. These corrections often involve reconciling disparate and potentially conflicting authority, a task type that an independent leaderboard has found to be among the most difficult.

The single research page previously published at /for-labs has been separated into the evals and dataset write-ups above.