On this pageAbstract
Evals

What three Gemini models miss about current non-compete law

Gemini 3.5 Flash-Lite, 3.8 Flash and 3.1 Pro Preview, asked without web access, compared on settled state non-compete doctrine versus 2026 earnings thresholds and statutes effective since 2024.

About this essay
Editor
  • Lawyer
  • Harvard Law '18 (J.D.)
  • MIT '13 (S.B.)
  • Former Ropes & Gray (6 yrs)
  • Admitted in NY
License
CC BY 4.0

Abstract

If an assistant gives you a 2026 non-compete earnings threshold, verify the figure against the state source before using it. In our no-web test, Gemini 3.5 Flash-Lite, Gemini 3.8 Flash, and Gemini 3.1 Pro Preview each matched the keyed threshold in 1 of 7 states, though they did not match the same state. On recent legislation, they said no qualifying statute had been enacted in 16, 15, and 13 of 19 selected states, respectively, where the key recorded one. The two larger models answered most questions about settled rules correctly. The gap was in what changed recently, so check current figures and recent statutes against a dated source even when an assistant handles the settled rules well.

The answer key came from the same OpenAgreements guides shown in the oa condition. An oa score measures extraction of a keyed fact from that guide, not independent legal accuracy. The next-state-guide control tests whether the answer depends on receiving the correct state's guide. For an external check, agents spot-checked a seeded sample of 25 key items against primary sources: 23 were confirmed, 2 were partly correct, and 0 were contradicted. We dropped both partly correct items. This was an agent primary-source spot-check, not lawyer adjudication; two confirmed items relied on secondary reproductions of official text.

Setup

We asked the models the verbatim questions in longtail.py, with the instruction to answer as of September 2026 for an agreement signed today. Each item used a single temperature-0 generateContent run on 2026-09-23, in JSON response mode, without web access or tools. Gemini 3.5 Flash-Lite ran in three conditions: cold received the question without context; oa received the target state's pinned practice guide; control received the next state's guide alphabetically. Gemini 3.8 Flash and Gemini 3.1 Pro Preview ran cold.

The families asked about earnings thresholds, statutes effective since 2024-01-01, consideration for an existing employee's covenant, statutory duration, choice-of-law restrictions, and advance notice. We scored answers against the frozen key, with a threshold match allowed within ±1% of the keyed dollars. The guide version is pinned to commit 30b1281ef738e7df8351ac09bd310ceb19e83610.

Answer key and audit

Labelling agents proposed answers with guide excerpts. A gate retained an item when its quoted support was an exact substring of the pinned guide and its numbers or dates appeared in that excerpt. Human review dropped labels where a defensible answer could have scored wrong. We then sampled across question families, including cases where cold and oa agreed, for the agent source check described above. The partly correct Connecticut consideration item and Maryland recency item were removed after the audit. The Connecticut opinion qualifies the treatment of continued at-will employment; Maryland's chapter law distinguishes its effective date from its later applicability date.

Results

Correct answers on the retained items:

Question familyFlash-Lite coldFlash coldPro Preview coldFlash-Lite oaFlash-Lite control
2026 earnings threshold1/71/71/77/70/7
Recent statute effective date2/194/193/1919/190/19
Consideration14/3028/3029/3030/308/30
Statutory duration6/109/109/1010/100/10
Choice of law7/87/87/88/81/8
Advance notice5/55/55/55/50/5

Current earnings thresholds

The model matches occurred in different states: Flash-Lite matched Washington, while Flash and Pro Preview matched Illinois. The table shows the keyed figure and each cold answer in dollars. None means the model returned no dollar figure. The Colorado key is supported by the state's 2026 PAY CALC rule; Tennessee's enacted bill status supports its threshold and effective date.

StateKeyFlash-LiteFlashPro Preview
Colorado$130,014$123,833$123,750$123,750
District of Columbia$162,164$156,426$150,000$150,000
Illinois$75,000$90,780$75,000$75,000
Oregon$119,541$121,944$100,533$100,000
Tennessee$70,000nonenonenone
Virginia$78,364.52$81,120$73,320$151,164
Washington$126,858.83$125,744$120,560none

Recent statutes

Each retained recency state had a qualifying statute in the key. The prompt asked whether a statute changing non-compete enforceability for any class of workers had taken effect on or after 2024-01-01 and, if so, for its effective date. A false-enactment response means the model answered that no such statute existed. That happened in 16/19 Flash-Lite answers, 15/19 Flash answers, and 13/19 Pro Preview answers. We report those separately from effective-date recall: Flash-Lite matched a keyed date in 2/19 states, Flash in 4/19, and Pro Preview in 3/19. Pro Preview said a statute existed but gave no matching date in 3 states; Flash-Lite did so in 1. For scale, a comparator that always answered enacted and supplied 2025-07-01 matched 4/19. This positive-case design does not measure how well the models distinguish enacted changes from states without a qualifying change.

Settled doctrine and citations

The two larger models answered most questions about settled rules correctly: consideration was 28/30 for Flash and 29/30 for Pro Preview; statutory duration was 9/10 for each; choice of law was 7/8 for each; and notice was 5/5 for each. Flash-Lite's 14/30 consideration result is a small-model observation in this run, rather than a claim about the larger models. In a separate short-form test asking for controlling statutes, the citation column showed no lift from the guide: 22/26 correct cold versus 21/26 with oa.

Excluded question family

We excluded court narrowing from the published aggregates. A binary narrow-or-void answer does not fit strict blue-pencil rules and near-ban states cleanly enough to make those scores interpretable. The underlying items remain in the evaluation materials.

Limits and reproducibility

This is a single-vendor snapshot with one temperature-0 run per item, not a measure of repeat-run reliability. The tested recency states were selected positive cases. Guide-assisted scores share their answer source with the key, and the external audit is a sample checked by agents rather than a complete lawyer review. The cold errors do not establish why the models answered as they did.

The prompts, gated key, exclusions, audit, scorer, and raw runs are in evals/non-compete-currency in github.com/open-agreements/open-agreements. The guide commit is 30b1281ef738e7df8351ac09bd310ceb19e83610; the frozen results.json SHA-256 is 2f385abcd96cd31b80292e770eac44e0ad89564a98fbaf500b22f9e434b0536d. The generated RESULTS.md displays the scored cells and exclusions.

Changelog

  • 2026-09-25: Drafted from the frozen results and retained-item audit. The Connecticut consideration and Maryland recency labels were excluded following partly correct audit verdicts.

Agents can load the current practice guides through openagreements.org/mcp.