Abstract
Correction — 29 July 2026. Two defects were found in this study after publication. The original elicitation asked each model for its fairness score before any reasoning, and a later item-by-item review found 24 of the 35 clause archetypes defective — 18 carried a closing sentence that gave the answer away, and 6 were unrealistic or confounded as drafted. One descriptive claim is withdrawn (see the changelog). The results in Results are retained as the historical output of that instrument, not as standalone corrected evidence. A second elicitation on revised items under corrected prompt ordering is reported in Replication; because it changes both the item set and the field order, differences between the two runs do not isolate either change. This correction also applies the pre-registered multiplicity control to the verdicts themselves, which it had never gated.
When a language model rates the fairness of a contract provision, does the rating depend on what the clause does, or on who it does it to? We constructed 105 clause groups across three contracting relationships (employment, residential lease, B2B SaaS) in which the two one-sided variants of each clause are the same text with the party role tokens swapped (schematically, 'The Company may terminate this Agreement without notice' versus 'The Employee may terminate this Agreement without notice'), so the two variants are textually identical apart from the party labels. Four frontier models (GPT-5.1, Claude Sonnet 5, Gemini 3.1 Pro Preview, Gemini 3.1 Flash Lite) rated each variant's fairness under the original role labels and under counterbalanced, opaque Party A / Party B labels, over 10,000 determinations in all.
Three of the four models rated the identical text as significantly less fair when it favored the stronger party (employer or landlord): group-level mean asymmetry +0.86 points on a 10-point scale for GPT-5.1 (95% CI 0.51 to 1.31), +0.50 for Claude Sonnet 5, +0.21 for Gemini 3.1 Pro Preview, with Gemini 3.1 Flash Lite showing no reliable tilt (+0.09, n.s.). Under the opaque labels the asymmetry collapsed to between 0.00 and 0.06 in all four models (90 percent confidence intervals, the equivalence-test convention, all inside ±0.13). The tilt is carried by the party role labels.
We do not label the effect bias. A model that treats an employer terminating without notice differently from an employee quitting without notice may be correctly calibrated to asymmetric legal and economic reality even though the sentences match. What this evaluation establishes is narrower: the differential treatment rides on the role-label presentation. It survives when the words are identical, and it vanishes when the role labels, together with the agreement framing that names them, are replaced with opaque ones.
The text-level null — no fairness tilt once the role labels are replaced with opaque ones — holds in all four models in both elicitations, and is the most robust result in this work.
Results — original v2 elicitation
Labeled role names produced a fairness asymmetry in three of four models; opaque Party A / Party B labels collapsed it in all four. Per-model verdicts on the five hypotheses, as produced by the original instrument:
| Hypothesis | Flash Lite | Pro Preview | GPT-5.1 | Sonnet 5 |
|---|---|---|---|---|
| H1 — labeled fairness tilt > 0 | no (+0.09) | yes (+0.21) | yes (+0.86) | yes (+0.50) |
| H2 — labels carry part of the tilt | no | yes (+0.19) | yes (+0.80) | yes (+0.51) |
| H3 — the weak-party label specifically | no | no (−0.01) | yes (+0.24) | no |
| H4 — no text-level tilt under Party A/B (TOST ±0.5) | yes (0.00) | yes (+0.03) | yes (+0.06) | yes (0.00) |
| H5 — advocacy-condition tilt survives anonymization | no | no | no | no |
Decision rule and scope, fixed before data collection: employment and landlord–tenant groups pooled (up to 70 complete-case group pairs per cell; four of the Claude Sonnet 5 cells have n=69 after missing-data exclusions), with the entity-vs-entity B2B family as a descriptive comparator. A hypothesis passes on a group-level mean whose seeded BCa bootstrap 95% CI excludes zero upward and whose Holm-adjusted bootstrap p-value falls below 0.05, Holm applied across the five hypotheses within each model; the no-text-tilt hypothesis used a two-one-sided-tests (TOST) equivalence criterion at ±0.5 points, reported with the conventional 90 percent intervals.
The multiplicity control did not gate the verdicts until this correction. The pre-registration committed to Holm correction across the five hypotheses within each model, and the adjusted p-values were computed and reported from the beginning — but the pass/fail decision was taken on the unadjusted interval alone, so Holm never gated anything. Applying the registered rule changes no cell in this table; it changes two in the replication, both reported there as they now fall.
Two descriptive patterns worth reporting alongside the confirmatory cells. First, the tilt is larger in the individual-versus-entity families than in B2B for three of four models (GPT-5.1: employment 1.29, landlord–tenant 0.43, B2B 0.34). Second, direction identification is essentially intact without labels: models still read which party a clause favors at 97 to 100 percent accuracy under Party A/B; what changes is how unfair they say it is.
A third pattern reported here in the original version of this page — that the tilt concentrated in procedural and dignity clauses — has been withdrawn. See the changelog.
Methods and materials
Mirror-by-construction items. Each of 35 clause archetypes (termination, cure periods, fee shifting, venue, indemnification, consent standards, and so on) is a template with two role slots. The strong-party variant and the weak-party variant are rendered from the same template with the slots filled in opposite ways, so the variants are byte-identical up to role tokens — a validator re-renders every committed item from its template and fails on any byte difference. Each group also carries a mutual base variant that applies the clause to both parties symmetrically, for 315 rated items across the 105 groups. This removes by construction the confound that dominated our pilot: hand-written variant pairs where one side's edit was simply harsher than the other's.
Counterbalanced opaque labels. In the opaque-label conditions, which party becomes Party A versus Party B is assigned per group by seeded randomization, balanced 18/17 within each family, so a generic preference for the first-named party cannot masquerade as a role-label effect.
Analysis plan fixed in advance. The design document — hypotheses H1 through H5 as algebraic group-level estimands, the confirmatory scope, the equivalence margin, the multiplicity correction, and the sampling policy — was committed to version control alone, before any item text existed. The items, generator, validator, and analysis code were committed before the first API call. Three post-data amendments are content-neutral: two JSON parse-repair rungs for one model's malformed output (recovering 82 of 97 unparseable responses from cached bytes without re-calling any model), and a third that fixed a display-layer bug in the descriptive report and added the per-archetype descriptive table the design promised but the pre-data commit omitted — each recorded in the design file with its rationale, and none touching any estimand, decision rule, or scope. Missing data after repairs: 15 of 10,080 responses (0.15 percent), never resampled.
A fourth amendment is not content-neutral, and every figure above was produced under it. The prompts asked for the scalar before the prose — the model emitted its fairness score and its favored-party label before writing a word of explanation. This is the same field-order artifact reported against a third-party judge schema in harveyai/harvey-labs#106. In a controlled replay there, one borderline criterion with one judge model at temperature 0 returned fail in three of three runs under verdict-first ordering and pass in three of three runs under reasoning-first ordering — a demonstrated effect on a single criterion, not a benchmark-wide result. The same ordering was present in this study's own instrument and went unnoticed until 2026-07-28, after publication. The re-elicitation under corrected ordering is in Replication.
What the machinery does not decide. One human-judgment layer remains: each archetype carries a one-line declaration of which slot-holder the one-sided clause favors. These 35 declarations are published with the item set for audit. All 35 have since been re-read individually and the declared direction confirmed on each — see Replication. That re-reading was performed by this study's author, an M&A lawyer; it is an unblinded review of the author's own instrument, not independent adjudication, and is a weaker check than an outside labeler would provide. The central interpretive question — whether label-conditioned ratings are miscalibration or correct legal judgment — is left open by design; it requires practitioner ground truth from outside this work, which is the next phase.
Limits
Text symmetry is not consequence symmetry: identical sentences can describe different real-world events depending on who acts, and a fairness rater is entitled to know that. Results are scoped to this fixed 35-archetype benchmark; templates are deliberately family-generic (no salaries, rents, or service-specific terms), which buys the mirror guarantee at some cost in realism. Ratings come from one sample per prompt: the Gemini models were called at temperature 0 with JSON output mode, while the GPT-5.1 and Claude Sonnet 5 calls set no temperature (their current APIs reject the parameter for these models), so no within-prompt variance is measured anywhere. The opaque-label condition removes the role labels and incidental framing (agreement type), so the label effect is measured as a presentation bundle. This evaluation measures reading — rating existing provisions — not drafting; whether generation shows the same label-conditioned tilt is an open question we consider harder and more consequential.
Replication
Two defects were found in the study above after publication: the elicitation ordering described in Methods, and item quality. All 35 archetypes were re-read individually and the declared favored-party direction confirmed on each, but the items themselves were flagged extensively — 18 carried a closing sentence that gave the answer away, making the test too easy, and 6 were unrealistic or confounded as drafted. That re-reading was performed by this study's author, an M&A lawyer; it is an unblinded review of the author's own instrument rather than independent adjudication.
The study was re-run as v3: prompts reordered to elicit reasoning before the scalar, the giveaway sentences deleted, three clauses rewritten, and the six flagged archetypes dropped. That leaves 29 archetypes, 87 mirror groups, 261 items, 8,352 determinations. The v3 instrument and analysis plan were committed before its first API call, but this was not an independent pre-registration: v2 had already been seen, the revised items derive from v2, and the estimands and decision rules were inherited rather than chosen blind to the earlier results. v3 data is never pooled with v2 — the two are separate elicitations of a changed instrument on a changed item set, reported side by side.
| Hypothesis | Flash Lite | Pro Preview | GPT-5.1 | Sonnet 5 |
|---|---|---|---|---|
| H1 — labeled fairness tilt > 0 | yes (+0.67) | no (+0.10) | yes (+0.95) | yes (+0.33) |
| H2 — labels carry part of the tilt | yes (+0.57) | no (+0.26) | yes (+0.91) | no (+0.31) |
| H3 — the weak-party label specifically | no (+0.22) | no (+0.10) | no (+0.29) | no (+0.21) |
| H4 — no text-level tilt under Party A/B (TOST ±0.5) | yes (+0.10) | yes (−0.16) | yes (+0.05) | yes (+0.02) |
| H5 — advocacy-condition tilt survives anonymization | no (−0.31) | no (−0.04) | no (−0.26) | no (+0.34) |
Complete-case n per cell ranges from 50 to 58 — 29 archetypes × 2 confirmatory families gives 58 group pairs, less missing-data exclusions detailed below. Verdicts apply the pre-registered rule stated in Results, including the Holm adjustment the original publication computed but did not gate on.
The central claim survives. The text-level null (H4) holds in all four models in both elicitations — it is the most robust result in this work, with Holm-adjusted p below 0.001 in every v3 cell. H1 replicates in GPT-5.1 and Sonnet 5 and is newly supported in Flash Lite; H2 is supported in Flash Lite and GPT-5.1.
The unchanged-item comparison did not detect a shift from the field ordering, but it cannot establish that the ordering had no effect. Ten archetypes were left byte-identical between the runs, so their only instrument difference is the reordering. On those items the pooled labeled asymmetry is 0.412 in each run, and the paired change across the 20 item-family units is 0.000, with a 95% BCa interval of −0.175 to +0.138 — so an ordering contribution of up to roughly 0.18 points, about 40 percent of the 0.412 estimate, is not excluded. Per-model paired changes span −0.150 to +0.150 and cancel in the pooled mean, so the exact equality of the two pooled figures is partly arithmetic. In the anonymized control the same paired contrast is −0.100, 95% interval −0.225 to −0.050, which excludes zero: the reordering did measurably move that condition, and since H2 is the labeled-minus-anonymized contrast, that shift if anything widens H2 in v3. The runs are also a week apart on hosted models. This rules out a large pooled shift in the labeled condition; it does not isolate the ordering effect and it does not establish a null.
Six verdicts changed between the runs, reported as they fell. Flash Lite gained both H1 (+0.09 → +0.67) and H2 (+0.09 → +0.57). Pro Preview lost both (H1 +0.21 → +0.10, H2 +0.19 → +0.26). Sonnet 5 lost H2 (+0.51 → +0.31). GPT-5.1 lost H3 (+0.24 → +0.29), the one cell that had supported weak-party-label localization, so that hypothesis is now supported nowhere. Note that two of the four losses have larger point estimates in v3 than in v2: they fall to the smaller n and the multiplicity correction, not to a shrinking effect, and should not be read as the effect disappearing. Per-model attribution is unstable across these two runs.
Direction identification remained high after item revision. Among 694 parsed role-labeled variant determinations in v3 there were no wrong favored-party calls. The v2 run produced ten, every one naming the stronger party for a clause that favored the weaker (exact sign test p = 0.002). This is not a like-for-like improvement estimate: six of those ten errors came from two archetypes dropped before v3, which the v3 design recorded in advance as a reason to expect fewer errors by construction.
Missing data, by model. After parsing, 99 of 8,352 v3 responses were unusable: 83 Gemini Pro, 12 GPT-5.1, 4 Claude Sonnet 5, none from Flash Lite. Confirmatory analyses are complete-case per hypothesis. Gemini Pro's failures fall entirely in the represented-advocacy conditions and none in the neutral conditions, so its H1 through H4 cells rest on the full 58 group pairs and its lost H1 and H2 verdicts are not a missing-data artifact; it is why its H5 cell rests on 50. Those failures are also uneven across conditions — 25 in rep_weak_anon against 8 in rep_strong_anon — so the Pro H5 result should be treated cautiously. Seven of the GPT-5.1 failures would yield valid, shape-complete JSON under an exact terminal-brace repair, two of them H4 responses; that repair is deliberately not applied here. Whether it is admissible under the same cached-byte standard as the v2 parse-repair amendments is a decision on its own merits, and folding it into a correction pass would be the wrong way to make it.
A reusable methods finding. Reordering the fields moved the missing-data burden onto the measured quantity, and changed which model carries it. The aggregate rate of unparseable responses is similar across the runs — 0.96 percent of v2 calls, 1.22 percent of v3 calls — but the composition changed completely. In v2 the truncation was Claude Sonnet 5's, 96 of 2,520 calls, and because the field that cut off was prose, 82 were recovered from cached bytes. In v3 it is Gemini Pro, which truncated on none of its 2,520 v2 calls and on 83 of 2,088 v3 calls, and not one is recoverable: the last field is now the integer under measurement, so completing the brace could silently turn a truncated 10 into an 8. Unrecoverable missingness went from 0.15 to 1.19 percent. If you reorder a judge schema to put reasoning first, put a throwaway prose field last.
Changelog
- 2026-07-29 (second pass, after external review) — Applied the pre-registered Holm correction to the confirmatory verdicts themselves; it had been computed and reported since publication but never gated the pass/fail decision, contrary to the design document. Two v3 cells lose support as a result (Pro Preview H2, Sonnet 5 H2); no v2 cell changes. The count of verdicts differing between the runs therefore rises from three to six. Corrected the description of harveyai/harvey-labs#106: the controlled replay there covered one borderline criterion with one judge model and three runs per arm, not ten verdicts, and that scope caveat should not have been dropped. Withdrew the claim that the field ordering had not moved the estimate — the paired analysis on byte-identical items gives 0.000 with a 95% interval of −0.175 to +0.138, and the anonymized control shifted by −0.100 with an interval excluding zero. Qualified v3 as not an independent pre-registration. Recast the direction-identification heading from improved to remained high, since six of v2's ten errors came from archetypes dropped before v3. Added per-model missing-data counts, the n range, and the note that Gemini Pro's truncation is confined to the represented conditions. Moved the correction above the fold and relabelled the original results section. Softened the signature of an item confound to consistent with an item-specific confound, and rewrote the ground-truth attribution to name it as unblinded author review.
- 2026-07-29 — Withdrew the per-archetype concentration claim (that labeled tilt concentrated in procedural and dignity clauses, led by one-way written-reasons-on-termination, jury-trial waiver, and records-retention clauses). All three archetypes driving it were judged defective on re-reading — unrealistic or confounded — and were dropped from the v3 item set, so the claim is withdrawn rather than retested. The largest of the three also carried this benchmark's largest residual under opaque labels, which is consistent with an item-specific confound rather than a pure label effect. Disclosed the score-before-reasoning elicitation ordering (Methods) under which all v2 figures were produced. Added Replication reporting the v3 re-run.
- 2026-07-23 — Original publication.
Relation to our other work
Like the institutional-knowledge leakage evaluation, this study reports a narrow, verifiable failure-mode measurement rather than a leaderboard: one mechanism, isolated with controls, negative results included. The expert-correction dataset described in our data paper supplies the practitioner-labeling pipeline that the open calibration question above requires.