# What three Gemini models miss about current non-compete law[^about]

Gemini 3.5 Flash-Lite, 3.8 Flash and 3.1 Pro Preview, asked without web access, compared on settled state non-compete doctrine versus 2026 earnings thresholds and statutes effective since 2024.

## Abstract {#abstract}

If an assistant gives you a 2026 non-compete earnings threshold, verify the figure against the state source before using it. In our no-web test, Gemini 3.5 Flash-Lite, Gemini 3.8 Flash, and Gemini 3.1 Pro Preview each matched the keyed threshold in 1 of 7 states, though they did not match the same state. On recent legislation, they said no qualifying statute had been enacted in 16, 15, and 13 of 19 selected states, respectively, where the key recorded one. The two larger models answered most questions about settled rules correctly. The gap was in what changed recently, so check current figures and recent statutes against a dated source even when an assistant handles the settled rules well.

The answer key came from the same OpenAgreements guides shown in the `oa` condition. An `oa` score measures extraction of a keyed fact from that guide, not independent legal accuracy. The next-state-guide control tests whether the answer depends on receiving the correct state's guide. For an external check, agents spot-checked a seeded sample of 25 key items against primary sources: 23 were confirmed, 2 were partly correct, and 0 were contradicted. We dropped both partly correct items. This was an agent primary-source spot-check, not lawyer adjudication; two confirmed items relied on secondary reproductions of official text.

## Setup {#setup}

We asked the models the verbatim questions in [`longtail.py`](https://github.com/open-agreements/open-agreements/blob/main/evals/non-compete-currency/longtail.py), with the instruction to answer as of September 2026 for an agreement signed today. Each item used a single temperature-0 `generateContent` run on 2026-09-23, in JSON response mode, without web access or tools. Gemini 3.5 Flash-Lite ran in three conditions: `cold` received the question without context; `oa` received the target state's pinned practice guide; `control` received the next state's guide alphabetically. Gemini 3.8 Flash and Gemini 3.1 Pro Preview ran `cold`.

The families asked about earnings thresholds, statutes effective since 2024-01-01, consideration for an existing employee's covenant, statutory duration, choice-of-law restrictions, and advance notice. We scored answers against the frozen key, with a threshold match allowed within ±1% of the keyed dollars. The guide version is pinned to commit `30b1281ef738e7df8351ac09bd310ceb19e83610`.

## Answer key and audit {#answer-key}

Labelling agents proposed answers with guide excerpts. A gate retained an item when its quoted support was an exact substring of the pinned guide and its numbers or dates appeared in that excerpt. Human review dropped labels where a defensible answer could have scored wrong. We then sampled across question families, including cases where `cold` and `oa` agreed, for the agent source check described above. The partly correct Connecticut consideration item and Maryland recency item were removed after the audit. The [Connecticut opinion](https://www.jud.ct.gov/external/supapp/Cases/AROcr/CR349/CR349.25.pdf) qualifies the treatment of continued at-will employment; [Maryland's chapter law](https://mgaleg.maryland.gov/2024RS/Chapters_noln/CH_378_hb1388e.pdf) distinguishes its effective date from its later applicability date.

## Results {#results}

Correct answers on the retained items:

| Question family | Flash-Lite cold | Flash cold | Pro Preview cold | Flash-Lite `oa` | Flash-Lite control |
| --- | --- | --- | --- | --- | --- |
| 2026 earnings threshold | 1/7 | 1/7 | 1/7 | 7/7 | 0/7 |
| Recent statute effective date | 2/19 | 4/19 | 3/19 | 19/19 | 0/19 |
| Consideration | 14/30 | 28/30 | 29/30 | 30/30 | 8/30 |
| Statutory duration | 6/10 | 9/10 | 9/10 | 10/10 | 0/10 |
| Choice of law | 7/8 | 7/8 | 7/8 | 8/8 | 1/8 |
| Advance notice | 5/5 | 5/5 | 5/5 | 5/5 | 0/5 |

### Current earnings thresholds {#thresholds}

The model matches occurred in different states: Flash-Lite matched Washington, while Flash and Pro Preview matched Illinois. The table shows the keyed figure and each `cold` answer in dollars. None means the model returned no dollar figure. The Colorado key is supported by the [state's 2026 PAY CALC rule](https://www.sos.state.co.us/CCR/GenerateRulePdf.do?ruleVersionId=12310&fileName=7+CCR+1103-14); Tennessee's enacted [bill status](https://wapp.capitol.tn.gov/apps/BillInfo/Default.aspx?BillNumber=HB1034&GA=114) supports its threshold and effective date.

| State | Key | Flash-Lite | Flash | Pro Preview |
| --- | --- | --- | --- | --- |
| Colorado | $130,014 | $123,833 | $123,750 | $123,750 |
| District of Columbia | $162,164 | $156,426 | $150,000 | $150,000 |
| Illinois | $75,000 | $90,780 | $75,000 | $75,000 |
| Oregon | $119,541 | $121,944 | $100,533 | $100,000 |
| Tennessee | $70,000 | none | none | none |
| Virginia | $78,364.52 | $81,120 | $73,320 | $151,164 |
| Washington | $126,858.83 | $125,744 | $120,560 | none |

### Recent statutes {#recency}

Each retained recency state had a qualifying statute in the key. The prompt asked whether a statute changing non-compete enforceability for any class of workers had taken effect on or after 2024-01-01 and, if so, for its effective date. A false-enactment response means the model answered that no such statute existed. That happened in 16/19 Flash-Lite answers, 15/19 Flash answers, and 13/19 Pro Preview answers. We report those separately from effective-date recall: Flash-Lite matched a keyed date in 2/19 states, Flash in 4/19, and Pro Preview in 3/19. Pro Preview said a statute existed but gave no matching date in 3 states; Flash-Lite did so in 1. For scale, a comparator that always answered enacted and supplied 2025-07-01 matched 4/19. This positive-case design does not measure how well the models distinguish enacted changes from states without a qualifying change.

### Settled doctrine and citations {#doctrine}

The two larger models answered most questions about settled rules correctly: consideration was 28/30 for Flash and 29/30 for Pro Preview; statutory duration was 9/10 for each; choice of law was 7/8 for each; and notice was 5/5 for each. Flash-Lite's 14/30 consideration result is a small-model observation in this run, rather than a claim about the larger models. In a separate short-form test asking for controlling statutes, the citation column showed no lift from the guide: 22/26 correct `cold` versus 21/26 with `oa`.

### Excluded question family {#narrowing}

We excluded court narrowing from the published aggregates. A binary narrow-or-void answer does not fit strict blue-pencil rules and near-ban states cleanly enough to make those scores interpretable. The underlying items remain in the evaluation materials.

## Limits and reproducibility {#limits}

This is a single-vendor snapshot with one temperature-0 run per item, not a measure of repeat-run reliability. The tested recency states were selected positive cases. Guide-assisted scores share their answer source with the key, and the external audit is a sample checked by agents rather than a complete lawyer review. The cold errors do not establish why the models answered as they did.

The prompts, gated key, exclusions, audit, scorer, and raw runs are in [`evals/non-compete-currency`](https://github.com/open-agreements/open-agreements/tree/main/evals/non-compete-currency) in `github.com/open-agreements/open-agreements`. The guide commit is `30b1281ef738e7df8351ac09bd310ceb19e83610`; the frozen `results.json` SHA-256 is `2f385abcd96cd31b80292e770eac44e0ad89564a98fbaf500b22f9e434b0536d`. The generated [`RESULTS.md`](https://github.com/open-agreements/open-agreements/blob/main/evals/non-compete-currency/RESULTS.md) displays the scored cells and exclusions.

## Changelog {#changelog}

- 2026-09-25: Drafted from the frozen results and retained-item audit. The Connecticut consideration and Maryland recency labels were excluded following partly correct audit verdicts.

Agents can load the current practice guides through [openagreements.org/mcp](https://openagreements.org/mcp).

<!-- NUMBER LEDGER
Frontmatter and changelog: 2026-09-25 is the publication review date and audit date; audit/audit_a_verdicts.json and audit/audit_b_verdicts.json verdict records; results.json exclusions.
Title, abstract, setup, and thresholds heading/table: model identifiers, 2026 scope, run date 2026-09-23, temperature 0, guide commit, threshold tolerance, keyed values, cold answers, and 1/7 matches: results.json models.*.cells.threshold, models.*.threshold_cold_answers, models.*.threshold_cold_passed, threshold_key, threshold_tolerance, run_date, temperature, guide_ref; longtail.py AS_OF and Q.threshold.
Abstract and answer-key section: audit sample 25, 23 confirmed, 2 partly, 0 contradicted, 2 secondary reproductions, and dropped item identities: audit/audit_a_verdicts.json and audit/audit_b_verdicts.json verdict and source_tier fields; results.json exclusions.consideration:connecticut and exclusions.recency:maryland.
Setup: 2024-01-01 recency boundary from longtail.py Q.recency; model conditions and run method from README.md and results.json api, models, run_date, temperature, guide_ref.
Results table: every cell from results.json models.<model>.cells.<family>.<condition>; no narrowing cells included.
Thresholds table: results.json threshold_key.<state> and models.<model>.threshold_cold_answers.<state>; Colorado and Tennessee legal-source links from audit/audit_b_verdicts.json id threshold:colorado and audit/audit_evidence.json tennessee.source.
Recency paragraph: 19-state denominator and date matches from results.json models.<model>.cells.recency.cold; false-enactment and unmatched-date counts from models.<model>.recency_breakdown.cold; 2025-07-01 and 4/19 from recency_always_true_comparator; 2024-01-01 from longtail.py Q.recency.
Doctrine paragraph: results.json models.<model>.cells.consideration, duration, choice_of_law, notice; short-form 22/26 and 21/26 from python3 -B score_bench.py citation column.
Reproducibility: results.json guide_ref and RESULTS.md header for results.json SHA-256 2f385abcd96cd31b80292e770eac44e0ad89564a98fbaf500b22f9e434b0536d.
-->


[^about]: By Steven Obiajulu, J.D. Published by [openagreements.org](https://openagreements.org). Last reviewed 2026-09-25. License: CC BY 4.0. Steven Obiajulu, J.D. wrote this essay. It states the author's views, synthesizes public sources, and is not legal advice. This article is for informational purposes only and does not create an attorney-client relationship. Source excerpts and linked materials belong to their owners. CC BY 4.0. Cite as Steven Obiajulu, *What three Gemini models miss about current non-compete law*, OpenAgreements (last updated September 25, 2026), https://openagreements.org/research/current-law-gap.
