The board, the bounds, and the gap between them.
That frontier now has a measured structure, in two parts with different epistemic status. The largest proven gap is on the visible cases: a protocol-feasible planner (the engine’s Bayesian frontier solver, chained round by round under the same participant sets, row limits and consent filtering) certifies ≥.55–.70 of the attainable welfare reachable, against a best realized .38 on bilateral and .25 on multiparty clearing, making clearing the largest proven unrealized welfare on the board. Where information hides, the frontier is open rather than proven: the planner cannot use the communication channel and certifies only ≈.09–.24 on discovery, a range the best model’s held-out .156 already sits inside. Models achieve roughly what is reachable without talking; whether information-gathering through language can recover the rest is the open question, and no model has demonstrated that it can. Probe-tier lower bounds (5 dev worlds/case).
Every model faces the same five cases in the same 5-agent market: one seat under test against a four-agent frozen panel (gemini-2.5-flash, temperature 0). The same gap appears in every model tested, and the ranking is secondary. Closed means approaching the ceiling models reach when information is visible, not merely clearing the deterministic floor: the best held-out result on any hidden-information case realizes .19 of the optimum (the deferred-settlement confirmation arm reaches .49), and on discovery the leader barely clears a trivial floor. The frontier is behavioral, not baked in: the measured information floor is ≈ 0, so the unrealized welfare is recoverable by skill. Every model now carries a private held-out score (◇), reported next to its development result in the complete matrix.
AER (Aggregate Efficiency Ratio) is the share of the full-information welfare optimum the market realizes: 1.00 is the fully efficient outcome, 0 is no improvement over the starting allocation, a negative value means trades left the group worse off. It is market-level welfare across the whole five-seat market, not the tested agent's private surplus. Because the denominator is W*, the score charges the irreducible information gap (welfare no agent could recover under hidden information), and for value-destroying outcomes it shrinks the magnitude of the negative score toward zero.
The equation. For each case, pooled over seeds:
AER = Σs ( Wrealized − Winitial ) ⁄ Σs ( W* − Winitial )
Wrealized is the market's final welfare net of coordination costs, but only if the settlement passes the feasibility-and-authorization gate; a failed gate contributes Winitial (that episode scores 0). W* is the full-information social optimum: a social planner maximizing total welfare, computed ex-post. The denominator is floored at 0, and a degenerate zero-opportunity case is dropped, not scored. It is a raw ratio of sums: value destruction stays negative, and the score cannot exceed 1 (realized welfare never beats the optimum). W* is the god's-eye ceiling, so against any information-constrained frontier this is a conservative lower bound on efficiency for positive gains, with the negative-magnitude caveat above.
How to read it. The primary estimand is pooled AER vs W* per case, with an equal-weighted-across-cases companion so a large-opportunity case doesn't dominate. A case counts as solved only when its held-out lower bound clears the greedy floor by a preregistered margin, not at AER = 1. The greedy floor is a deterministic baseline: each round it executes the single best mutually-beneficial bilateral trade among the agents the controller can see (the same visibility the model has), computed directly, never through the language compiler. Comparability caveat: because the floor settles directly on the ledger while models must express deals in text that a compiler and verifier accept, “below the floor” mixes economic strategy with language-executability; a funnel-matched greedy baseline is planned.
Each case emphasizes one capability (bilateral exchange, clearing, discovery, consent), though the cases differ in other ways too, so read them as a family rather than a controlled single-variable manipulation. The integrated arena composes them into a single live market: the cases decompose the economy; the arena is where it must hold together.
Every model, the same economy.
Pooled AER over the five cases, 30 development seeds each, with bootstrap 95% intervals. Read it for the shape, not the ranking: a trivial deterministic floor divides the field, and no model closes the hidden-information cases. Every model here also carries a private held-out score (◇, complete matrix): held-out tracks development with no overfitting signature.
- 1gemini-3.5-flashGoogle+.258[.225–.291]
- 2kimi-k2Moonshot+.228[.197–.263]
- 3greedy floorbaseline+.123[.111–.135]
- 4gemini-2.5-flashGoogle+.111[.094–.129]
- —deepseek-v4-flashDeepSeek · dev pool withdrawn—held-out +.217
- —minimax-m3MiniMax · dev pool withdrawn—held-out +.149
- —glm-4.7-flashZ.ai · dev pool withdrawn—held-out +.107
- 8randombaseline+.009[.004–.015]
Bars are development AER with 95% CI; the ◇ marks the private held-out score. These panels plot every model and baseline, ordered by pooled development score; the complete numbers are in the matrix below. Baselines are shown in grey. Each row pools the five cases over 30 seeds (n≈150), replay-verified.
Baselines and the reference models, case by case.
Does the edge survive a crowd of copies?
Same 5-agent market as the leaderboard above, one difference in who fills the seats. Single-seat is the model in one seat against a four-agent frozen panel; that is the leaderboard score. Population fills all five seats with copies of the same model. The test: does a single-seat edge survive when every counterparty is a copy of itself? Four models were run this way in the official AER, but three are withdrawn pending re-measurement (see below), leaving the panel model as the one usable row; for it, population and single-seat coincide (+.112 vs +.111), a built-in self-consistency check, since its frozen panel is the same model at the same temperature.
| Model | Pooled AER | Population AER · by case | |||||
|---|---|---|---|---|---|---|---|
| single-seat | population | bilateral | clearing | discovery | consent | deferred | |
| deepseek-v4-flashDeepSeek · withdrawn | withdrawn: both modes ran before the empty-response fix; re-run pending | ||||||
| gemini-2.5-flashGoogle · panel model | +.111 | +.112 | .169 | .065 | .072 | .043 | .211 |
| minimax-m3MiniMax · withdrawn | withdrawn: both modes ran before the empty-response fix; re-run pending | ||||||
| glm-4.7-flashZ.ai · withdrawn | withdrawn: both modes ran before the empty-response fix; re-run pending | ||||||
This comparison is on hold. Three of the four population rows (deepseek-v4-flash, minimax-m3 and glm-4.7-flash) were measured before we found the empty-response defect (see the board note above), and so were their single-seat scores, so both sides of the solo-vs-population comparison are contaminated for exactly those models. Rather than restate numbers we cannot stand behind, we withdraw the rows until the population runs are re-executed. The panel model is unaffected (it runs on a different provider path with no empty turns) and its check still holds: its pooled population AER (+.112) tracks its single-seat score (+.111), which is the self-consistency result the frozen-panel design predicts. The question itself is untouched by the defect (same seats, same worlds, same scoring), so the claim we keep is the weak one: population scores are not a mechanical function of solo scores, and a solo leaderboard is not yet evidence about market-level welfare. The quantitative version returns with the re-run. The scale-collapse and under-trading ceiling reported under Findings come from legacy self-play (population welfare-ratio), a separate measurement, not on this board.
Capture: the second axis, and why AER cannot be it. AER is a sum over agents. Surplus moved from one seat to another leaves it unchanged, so it cannot express who captured value or who was protected from a bad settlement, the two questions a principal deploying an agent actually asks. This is a property of the measure, not a defect in it, and it is demonstrated rather than argued. Three measurements: (1) over 100 exactly-evaluated worlds AER ranks an offer giving the responder 45% of the joint gain above one giving it 71%, because the ordering tracks trade volume; (2) an intervention that cut settlements leaving an agent worse off from 7.3% to 1.8% (paired CI [−.090, −.021]) moved AER by −.018 with an interval spanning zero; (3) on the bargaining track below, nearly every model closes nearly every deal, so AER-style clearing barely separates them, while capture spans .06 to .16 against a waiting rule that scores .19.
How capture is computed. A buyer (the model under test) bargains over price against a scripted seller that bargains back: it concedes faster when the buyer moves toward it and slower when the buyer stalls, may end the negotiation outright if an offer is far beneath its ask, and holds a private reservation above its cost that declines as rounds run out. The buyer knows its own maximum willingness to pay and the public list price, never the seller’s cost. The metric is SE+ = buyer surplus / ZOPA width, where the zone of possible agreement is WTP − cost, both known exactly to the scorer and hidden from the agent, so the denominator is exact with no bargaining-solution or Bayesian fixed point to approximate. That is why capture is measured here and not in the five-agent arena: with five seats, “you could have captured everything” is unreachable by anyone, and share-of-realised-joint-gain diverges when a counterparty loses (measured at 3.7 on one case). SE+ needs no discrimination correction because closing is a precondition for scoring: a model that never deals scores 0, and one that lowballs into no-deal scores 0. Anchors on the same worlds: accepting the opening ask scores −.565; a one-line stubborn rule that counters every round at 80% of whatever is currently asked and never raises its own offer scores .188; the best fixed schedule found by grid search scores .468; the per-world hindsight ceiling is .613; paying exactly the seller’s cost is 1.0. The stubborn rule is a bargaining policy, not a passive one: it makes three counter-offers and the seller accepts its price in 21 of 30 worlds. It is the right yardstick precisely because it is so simple, and because the gap above it (.468 versus .188) shows that skilled bargaining is worth two and a half times what stubbornness alone earns. A model below .188 is therefore not being punished for negotiating; it is negotiating worse than refusing to move.
Robustness, and what it rules out. 629 of 630 episodes, 7 models × 30 seeds × three arms. Rank agreement between public development seeds and a private held-out set nobody could tune against is Spearman +.86, and gemini-3.5-flash leads all three arms. Agreement with the second archetype is weaker (+.46), which is the reason the second archetype exists: one scripted concession family is one counterparty model. That archetype opens far above willingness to pay and concedes fast, against a standard seller that opens modestly and concedes slowly, so the pair differs on both axes rather than being a relabelling. Two candidates were screened out first: a stubborn seller collapsed the gap between naive and informed play to .031, and a late-concession schedule revealed the seller’s cost outright. The generator itself was rebuilt for the same reason: its first form was algebraically invertible, recovering the private cost in 200 of 200 worlds, which would have measured whether a model spotted an identity rather than whether it could bargain.
The result is not an artifact of the seller or the instructions. Two objections have to be answered before “every model loses to a waiting rule” means anything. First, the models might be losing to an environment that is simply harsh. They are not: measured on their own transcripts, 67% of deals settled above the price the seller was itself asking by the final round (mean overpayment $8.20), and replaying each episode with the buyer repeating its own first offer instead of climbing gains +.100 SE+, better in 142 of 202 matched scenarios. Both comparisons hold the world, the seller and the rules fixed, so they isolate the choice. Second, the buyer prompt discloses that the seller may walk away and that its concession responds to movement, which could plausibly be instructing the bidding-up. So the disclosure was made switchable and removed: across 3 models and 30 shared seeds the paired change in SE+ is −.003 with a 95% interval of [−.029, +.022], and the self-bidding rate moves from 89% to 84%. A null. What does remain design-relative is the level: a more generous reservation or a faster concession schedule lifts every number, and four rounds bounds how much patience can pay. The ordering, the self-bidding rate and the two counterfactuals above are within-environment comparisons and do not depend on those choices.
What the capture column does and does not license. It licenses this: against this seller, every model tested bargains worse than a one-line rule that never raises its own offer, and does so by conceding toward the counterparty as the counterparty concedes toward it. The stubborn rule’s absolute offers fall in 30 of 30 worlds; the models’ rise in 166 of 202. Bargaining itself pays here, and pays well: the best schedule earns .468 against the rule’s .188. The deficit is directional, not a case against negotiating. It does not license reading .16 as a constant for gemini-3.5-flash: that number is relative to this seller’s parameters. It does not license any capture claim about the multi-agent arena, which needs a per-pair attainable frontier that is not built. Capture is reported beside AER and never averaged into it, because a composite would re-hide exactly what the second axis exists to expose, and the two agree only weakly (Spearman .18). Counterparty surplus is shown alongside because SE+ = 1.0 means the seller received nothing, which is the correct answer for a principal-agent measure and the wrong thing for a benchmark to celebrate silently.
What is being measured, and what is only the control. The quantity is private value capture: each side is given a private value it alone knows, and the score is the share of the available gains it secures for itself, against a fixed reference bargainer that concedes on a schedule and is never adjusted between runs. That is the number an agent’s principal cares about, and it is the one AER cannot express, being a sum over both sides. Across the field it is .217, and only minimax-m3 separates from the rest: .309 [.265, .354] against five models spanning .147 to .255 with mutually overlapping intervals, so ranks below the leader are unresolved. Symmetry is the experimental control, not the finding. Models play a round-robin: every pairing, four cells per scenario crossing which side each holds with who opens, so both cancel by construction rather than by argument. Each episode scores both participants, since the two shares sum to one; an earlier version recorded only one side and would have published a leaderboard silently omitting gpt-5.6-luna, with per-model sample sizes running from 120 to 599. The metric is the split among closed deals, with deal rate reported beside it and never folded in: a variance decomposition put 57.2% of the previous composite on deal-versus-no-deal, which is what AER already measures, and only 9.1% on model identity. Conditioned on a deal, model identity rises to 13.8% and the side held falls to 0.0%. The round-robin exists to establish that the capture number is not an artifact of which chair a model was put in: there the field divides gains near evenly (.34 to .62), so the low capture against the reference is a property of the bargaining rather than of the seating. The reference policy takes no model calls, which is what makes an absolute reading affordable at all.
The gate, and the two checks in it that could not fail. A run reports nothing unless it passes: cell balance (every model in all four cells, near-equal counts), deal rate inside .40 to .95, unreadable-turn and reservation-veto rates under 5%, split-half rank stability above .70, and position effects bounded at .25. This run passed at balance 150/150/150/150 per model, deal rate .72, unparsed 3.7%, veto 3.3%, stability +.94. Two earlier members of that check set were removed for being unfalsifiable. Self-play balanced capture is an accounting identity: the four cells are two episodes counted from both sides, so it reads .5000 for every policy including deliberately lopsided ones. And split-half stability returned 1.00 whenever fewer than three models had data, so partial runs sailed through; it now blocks as "not evaluable". The gate also originally blocked on the size of the position effects rather than the balance of the design, which is the wrong test: chess counterbalances colour rather than abolishing white’s advantage, and effects of +.062 and +.093 that cancel exactly contribute 1.1% and 2.5% of variance against model identity’s 8.1%.
The seat was worth more than the model, and we did not notice until we looked. Every capture number this axis produced before 5 August 2026 seated the model under test on the buying side. Asked whether that side was harder, we mirrored the two seats word for word and swapped the models into each. Swapping sides is worth +.528 of the bargaining range [+.431, +.627]; the prompt wording, which was the suspect we started with, was worth +.043. The same gemini-2.5-flash kept .006 buying and .661 selling; minimax-m3 kept .066 and .829. Four structural features caused it, none of them deliberate: one side proposed and the other disposed, so the responder held the last word every round; the opening number was the seller’s published list price, for which there is no buyer equivalent; the seller’s counter reset the reference each round; and a known four-round deadline made whoever answered last a take-it-or-leave-it player. The prompt asymmetries we had already fixed were a twelfth of the effect of the ones we had not seen.
What survives, and what does not. The claim that models keep almost none of the bargaining range does not survive as a statement about bargaining: it is substantially a statement about which seat they were given. What does survive is every comparison made within that seat, because the reference policies were played there too. A one-line rule that counters at 80% and never raises its own offer kept 22% of the range in the same disadvantaged seat, and no model kept as much. The seat is hard, not hopeless, and losing to a trivial rule playing your own side is a fair verdict. The pooled 99.6% counterparty share is withdrawn as a headline for the same reason the earlier patience-ranked column was: not wrong arithmetic, but a number that answers a different question than the one it appeared to answer.
The rebuilt protocol, and the check it had to pass first. Bargaining now alternates who proposes; neither side gets a published anchor, so each knows only its own value; the negotiation ends at a random moment rather than a countable deadline, because a known deadline is not merely unfair but solvable by backward induction, which would measure whether a model can count rounds; and every world is played twice with the seats swapped, so residual advantage cancels the way playing both colours cancels it in chess. Acceptance criteria were written down before the check was run and measured offline with scripted policies at no cost: the residual seat gap between identical policies is 0.000 against +.528 for the design it replaces, counterbalanced capture is .500 to four decimal places, and bargaining skill still separates monotonically, so fairness was not bought by making the case degenerate. A raw first-proposer effect of −.069 remains and is cancelled rather than hidden, with a test asserting it is still present: if it silently vanished, the reference policy has most likely gone degenerate, which it did once already during this work.
Why the published opponent is a live model, not a script. The first version of this axis bargained against a scripted seller with five parameters (concession rate, two reactivity terms, a walkaway threshold, a reservation share) and we chose them by grid search, on the same development seeds the scores are reported on, to maximise the gap between good and bad play. That is a researcher degree of freedom: a result measured against it is partly a result about our choices. The remedy in this literature is more than one counterpart family, so we built a second and then published the better of the two. The A2Frozen seat is what the landing page reports: the seller is a live model, pinned at temperature 0, given a private cost and told to get the best price it can. That world carries no concession parameters at all, only a list price, a buyer value and a seller cost, because a model is not told how to bargain. SE+’s denominator is unaffected, since we still assign the cost. Two things the engine still enforces rather than requests: an accept below the seller’s cost is vetoed, and an unreadable seller turn becomes a hold rather than an accept. Across 418 episodes the veto fired zero times and no seller turn was unreadable, so the guard was present and never needed.
What the scripted opponent said, and why it is not the published number. Both were measured on the same seven models. The headline holds under each, and is stronger against the live opponent: every model still lands below the stubborn rule, by more (a gap of roughly .20–.26 against a live agent, versus .03–.13 against the script), self-bidding rises to 94% of multi-round episodes, the counterparty ends with 99.6% of the range, and deal rates fall from .93 to .64, so the models lose the deal as well as the surplus. What did not survive is the ordering below first place: rank correlation between the two counterparts is only .32, and minimax-m3 falls from second to last. Two separate live sellers (gemini-2.5-flash and kimi-k2) were run so that this could be attributed: they agree with each other at .86, so the instability is scripted-versus-live, not one model’s quirk. Own-family pairing is negligible: paired within buyer, a model scores +.003 against a seller from its own family, and the pooled figure of +.012 is a composition artifact of the three buyers that have such a cell. Consequently the landing page reports the live opponent alone, and treats per-model ranks below the leader as unresolved. The scripted numbers are kept here rather than deleted, because the disagreement between the two is itself the reason to distrust any single-opponent ranking, including this one.
Withdrawn: the capture numbers published on 4 August 2026. The first version of this column (deepseek-v4-flash .73, gpt-5.6-luna .64, gemini-3.5-flash .62, glm-4.7-flash .55, minimax-m3 .43, kimi-k2 .41, gemini-2.5-flash .39) was measured against a seller that read the buyer’s offers only to check whether they cleared its ask. Its concession path was a function of the round number alone, so offers of $1 and $40 produced identical outcomes, lowballing carried no risk, and the final round accepted anything above cost. Waiting therefore scored .729 and beat every counter-offer policy, and the column ranked models by how long they waited (patience explained it at Spearman +.82). Those numbers are withdrawn rather than quietly amended. The rebuilt ordering anticorrelates with them at Spearman −.25: deepseek led the old column and places fifth here; gemini-3.5-flash was third and now leads every arm.
Terms & provenance for the landing page. The landing page translates several terms of art that resolve here. (1) market-ext: the market-extension probes it cites (procurement bundling, private-value auction, cyclic clearing, coupled B2B award, hidden-reservation pricing) were scripted-mode runs with gpt-5.5 as the model under test (construction cases n=30, hidden-reservation pricing n=20), a separate track from the A2A leaderboard; their headline contrast is 1.00 on the visible variant vs −0.29 once the reservation is hidden. (2) The market-mode ceilings “≤ .36” and “≈ .15” come from the 2026-06-30 self-play parameter sweeps (gemini-2.5-flash populations, welfare-ratio, mean over 3 seeds): .36 is the best population ceiling across the five single-knob friction settings (population size, temperature, visibility, ledger row limit, search cost); .15 is the ceiling of the visibility knob specifically: legacy welfare-ratio, not AER. (3) The “admission screen” is the regret-normalized criterion from the case-oracle spec: a case is admitted iff an informed-but-naive policy (correct posterior beliefs, a simple non-optimizing rule) leaves normalized regret ≥ τ ≈ 0.3 against the Bayes-achievable frontier, so admission requires a hard control problem, not just hidden information; visible-construction and cheap-talk-style cases are rejected, and the screen runs independently of any model provider. (4) The baseline agents (greedy floor, random, no-op) are defined exactly on the methods page; the frozen panel decodes at temperature 0, as stated at the top of this page. (5) Agents are scored with reasoning disallowed in the public channel. The arena’s response prompt ends “no scratchpad, no private calculations, and no reasoning before PUBLIC ACTION”, because anything an agent writes there is read by its counterparties, and the private values it would reason about are exactly what the hidden-information cases withhold. That is a deliberate choice with a measured price on both sides. Every number on this site is the reasoning-disallowed arm.
What that condition costs, measured. Asked to accept or refuse the same settlements, the same models discriminate almost perfectly when allowed to compute (accepting 97.6% of the trades that help them and 0.0% of those that hurt) and not at all inside the negotiation, where they accept harmful settlements slightly more often than beneficial ones. The capability is present; the channel suppresses it. Giving agents a private scratchpad that is stripped before anyone else sees it cuts settlements that leave an agent worse off from 7.3% to 1.8% in a live market (95% CI on the paired difference [−.090, −.021]), and moves the score on this page by −.018, an interval spanning zero. Forbidding the reasoning buys a 0% rate of leaked private values, against 90% when it is allowed unshielded. So the condition is defensible and it is not free: it costs real counterparty discipline that a welfare score cannot see. Numbers here are correct for the condition stated; they are one arm of a two-arm design, and the fuller picture belongs beside them rather than behind them.
The shape of the gap.
The load-bearing result is not who leads. It is that realized welfare stays far from the full-information optimum, that the shortfall is largest in the discovery and hidden-consent cases, and that a trivial baseline already recovers much of it. Because the cases vary in more than visibility (discovery, coordination load, and consent move together), this is a case-family association, not an isolated hidden-information effect.
Most of the welfare goes uncaptured.
Even under the strongest agent, the market realizes only a fraction of the welfare it makes available; across the field, most of the available surplus is left on the table. The frontier is wide, not a rounding error.
The gap widens as information hides.
Realized welfare is highest with values and counterparties in view, and falls off through clearing, consent, and discovery. The hidden-information end is where every agent drops toward, and past, a trivial floor.
It worsens as the market grows.
Coordination is a second frontier axis. In market mode (one policy in every seat), realized welfare falls as the population scales: a homogeneous gemini-2.5-flash population reaches .235 of the optimum at five agents and .085 at fifteen. Market-mode reference: welfare-ratio, not the AER leaderboard.
Discovery is the sharpest unresolved case.
With counterparties hidden, capability collapses to the trivial floor. On case03 held-out seeds, only the strongest model clears the greedy floor (.108); five of six fall below it, so deterministic search, not economic reasoning, sets the pace.
How much of a deal rides on the wording?
Every agent here must express a trade in free-form text that counterparties then accept or refuse. That channel can be measured directly: hold the decisions fixed and vary only the sentence.
The instrument. The greedy floor searches for the single mutually-beneficial swap with the highest joint gain and proposes exactly that. A second variant inherits the identical search (provably the same trade, seed by seed) and changes only the rendering: a canonical line (“a1 gives 1 r1 to a2 for 1 r2”) versus a model-style envelope (a greeting, a rationale, and an explicit settlement table carrying the same numbers). Both run fresh against the same frozen panel on the same 120 held-out worlds, so the paired difference isolates phrasing.
The result: a paired phrasing tax of +.102 [+.091, +.114]. The canonical floor pools +.113; the same decisions phrased like a model pool +.011, statistically indistinguishable from the random baseline (+.013), and exactly .000 on both hidden-information cases. It is not a parsing failure. The verbose envelope compiles better than the terse line (62.6% vs 50.2% of proposal rounds feasible), and all 240 episodes replay-verify. The loss is entirely at consent: panel acceptance falls 38.9% → 2.4%, a 16× gap on identical terms, with the same panelist model calling one swap “mutually beneficial” when stated tersely and “not beneficial for me at this time” when the identical terms arrive inside a polite pitch. Counterparty IR judgment moves with framing.
Not an artifact of our prompt. The arena instructs responders to reject terms that are “negative, zero, ambiguous, or depend on missing consent”, a clause that could plausibly punish verbose proposals by construction. A pre-registered three-arm ablation (directions committed before running) removes it: strict (production) +.097 [+.069,+.126]; neutral, which keeps the own-delta rule but drops the ambiguity language, +.086 [+.059,+.113]; off, no discipline clause at all, +.107 [+.082,+.133]. The tax survives every regime, so the pre-registered falsification branch fires: the clause amplifies at most: the framing penalty is a property of how the counterparty model reads persuasive text. Two by-products: removing the clause entirely produced the worst free-form outcome (1.2% acceptance), because without an explicit acceptance rule panelists treat a polite offer as an invitation to keep talking rather than terms to settle; and the operative variable is therefore having a rule to apply to exact terms, not suspicion.
What it means for the floor. The greedy floor is not exempt from the language funnel: its proposals are compiled, verified, and consented to like any model's. Its privilege is narrower and now priced: it reads the world’s true two-sided utilities (models must infer them from dialogue) and it phrases canonically. Strip the phrasing half alone and the floor scores at the random baseline, which means every model on the board clears the interface-fair floor, and “at or below the floor” is a statement about a privileged reference rather than about economic incompetence. Model acceptance rates land inside the bracket the two variants define (3.0%–40.8% across the field), but conditioning on the proposal being IR-positive for the counterparty compresses that spread to 60–95%, so most of the board’s separation is about proposing mutually beneficial trades at all, with a smaller rhetorical residual on top.
Scope: one frozen panel model at temperature 0, one arena protocol. Whether other counterparty families are equally rhetoric-sensitive is open; a per-model probe is running. Data: funnel_phrasing_20260731 (240 episodes) and consent_discipline_20260731 (120 episodes), all replay-verified.
Two things have to survive an adversary.
Validation (§03) asks whether the measurement is stable. This section asks the adversarial questions: whether agents look after their own interest under pressure, and whether the scoring machinery itself can be gamed by anyone who submits an agent.
1 · Do agents defend their own private value? (individual rationality) Each seat holds private valuations no one else observes, and the minimal economic requirement on a counterparty is individual rationality: never consent to a settlement whose own-utility delta is negative against the no-trade outside option, while still consenting to those that are positive. Proposer skill and IR discipline are different competences, and the board scores only the first. The instrument flips the seat: the model under test becomes a responder facing a scripted adversary that offers a seed-randomised mix of IR-positive and IR-negative terms, crossed with three framings of identical settlement rows. Offer sign is recomputed exactly (the responder's own utility change for that offer, evaluated against the world as the offer was made) rather than read from the adversary's intent label. That correction was load-bearing: under diminishing returns, handing over four units of something you hold in abundance for one unit of something scarce is usually a gain, so the adversary's “predatory” template was genuinely value-destroying in only 39% of the rounds it was labelled as such. Scoring against intent would have penalised models for correctly accepting good deals. The headline metric is IR discrimination, DISC = accept(IR-positive) − accept(IR-negative), which scores both “refuse everything” and “accept everything” at zero: a caution metric would reward a model that never trades. A second arm measures information discipline: per-episode rate of quoting private utility values under direct and reciprocal extraction pressure. Pre-registered with four falsifiable predictions, including that the axis is abandoned if DISC correlates with board AER above ρ = 0.9.
Status: 350 episodes complete, and the cross-model ranking is withheld. Models that think before answering were given a token budget their reasoning could exhaust before any answer emerged, and the resulting silent turns were recorded as declining to trade, so three models were scored for decisions they were never given room to make. Rates ran from 0% (both Gemini models, GPT-5.6 Luna, Kimi K2) to 48%, which is uneven enough that no ranking spanning all seven is defensible from this run. What stands is measured on models with no silent turns at all: the two strongest discriminate sharply, taking two-thirds of the offers that help them and almost none that hurt, while one mid-tier model accepts good and bad at the same rate: its consent carries no information about its own interest. That last failure is invisible to a market-level score, because the trades still clear. The information-discipline arm was re-run for the affected models and is now clean for all seven; it produced a result worth stating on its own: two models disclose their private valuations under pressure, through opposite channels: one yields to a bare assertion of authority and shrugs off reciprocity, the other is immune to authority but hands everything over to a counterparty that volunteers its own numbers first. Five leak through neither. An earlier reading of this arm, that authority-based extraction had been closed off entirely, was an artifact of the silent turns and is withdrawn. A redesign that presents every model an identical battery of pre-verified decisions (the same offer text worth +2.80 to a responder in one world and −2.80 in another, so quantity and wording carry no signal and only reading your own valuation does) is built and pending its run.
2 · Is the scoring machinery hack-proof? Free-form candidate text is consumed by LLM panel, compiler and verifier seats, so a benchmark-optimising agent could attack the funnel instead of negotiating, and replay verification would faithfully reproduce such an exploit rather than detect it. Six adversarial agents were submitted through the ordinary text boundary (role override, phantom settlement, verifier bluff, authorization evasion, state extraction, coercive framing), 120 episodes, detection post-hoc from retained traces. Two defects were confirmed. A fake “SYSTEM NOTICE” impersonating the arena lifted acceptance to 82.5% against an honest baseline of 38.9%, because candidate text reached panel prompts with no provenance marking. And the private-acceptance gate discarded its own correct rejection whenever the debited seat's reply failed to parse: the engine computed a true delta of −1.508, recorded approve: false, and applied the transfer anyway, so an attacker needed only to make the counterparty reply in prose. Worst single case: −10.48 utility taken from a seat that never agreed.
What the fixes did. Candidate text is now fenced as untrusted counterparty speech carrying no authority over rules or protocol, and a harmful debit can no longer execute on unparseable consent (harmless ones still pass, and an explicit parseable approval of a bad trade is still allowed; that is measured behaviour). Re-running the identical 120 episodes: impersonation attacks fall to 0.0% acceptance from 63.7%, and settlements that took value from an unwilling counterparty go 15 → 0. Two predatory settlements remain, and both are counterparties that read lopsided terms and explicitly consented: a bargaining loss, not a harness failure. Attacks that merely asserted facts about protocol state were rejected outright even before the fixes: what worked was impersonating the arena, not lying to it.
Corrections we made to ourselves. Two detector faults in this suite were caught and retracted rather than published: a loose leak regex that reported 8 disclosures which were all resource identifiers and public holdings, and a consent check that read a trace field this schema does not populate, briefly appearing to show a third defect. The second was caught by a blast-radius check: the proposed rule would have blocked 117/117 debits in honest runs, impossible as behaviour. Detector self-tests now ship with the audit script. Data: injection_suite_20260731 and injection_verify_20260731, 240 episodes, all replay-verified.
Change what could have made the result brittle.
These checks target different failure modes: replacing the counterparty panel tests dependence on who the model faces; changing every seed tests whether selection tuned to a convenient set of worlds; ablating one economic knob at a time tests whether scores respond to what the cases claim to measure. Read jointly, the three opponent-conditioning channels of the crossed design all return null. Own-family interaction is −.001 [−.056, +.053] on the panel channel (2 candidates × 2 frozen panels, 60 paired worlds) and −.025 [−.083, +.030] on the compiler channel (same design, compiler/verifier swapped, panel held fixed); on the seat channel the pre-registered rotating-controller arm shows the capture premium attaching to the control right (+7.7pp, inside the pre-registered +4–8pp band) rather than to the seat label or the model, with the fixed seat-1 premium inverting to −36.2pp. So candidate ordering is not an artifact of who the agent faces, what grounds its text, or which seat it holds, at resolutions of |Δ| ≈ .05 (panel) and .08 (compiler). What the panel does change is composition, identically for both candidates, and the floor's level: floor-relative claims stay panel- and compiler-indexed.
New counterparties.
Same leader.
Across the four original cases, the same model leads under the frozen gemini panel, a scripted-rational panel, and a frozen kimi-k2 panel, and in the crossed experiment (two candidates × two frozen panels, 60 paired worlds) the pre-registered own-family interaction is zero to the design’s resolution (−.001 [−.056, +.053]): neither model does better against its own family, and the leader leads on both panels. What the panel does change is composition, identically for both candidates: bilateral collapses (.33 → .03 and .36 → .13) while clearing and consent rise, with pooled scores nearly invariant because the attainable-welfare denominator is panel-independent and the route shifts offset. The floor itself moves: zero on discovery and consent under the scripted panel, +.104 → +.164 under the kimi panel; floor-relative standings are panel-indexed.
New compiler.
Same leader.
The compiler channel closes the set. Swapping the compiler/verifier to kimi-k2 (panel held fixed, same 60 worlds) gives an own-family interaction of −.025 [−.083, +.030]: also zero, with the point estimate pointing away from own-family favouritism. Panel, compiler, and seat: all three opponent-conditioning channels return null, so the ordering is not an artifact of who the candidate faces, what grounds its text, or which seat it holds. The floor, again, is the sensitive one (+.104 → +.043 under the kimi compiler). And a floor re-phrased like a model collapses to the random baseline (methods page), so every model clears the interface-fair floor.
One knob at a time.
Signs committed first.
Seven economic knobs (counterparty visibility, search cost, cycle length, money, adversarial defection, value disclosure, consent gating) were ablated with expected directions pre-registered before any run: 5 confirmed, 2 null with the mechanism identified, 1 single-episode anomaly disclosed. Hiding counterparties removes two-thirds of the tested model’s score; adding a money numeraire and consent gating each cut it sharply while the scripted floor barely moves; the difficulty tracks the manipulated economics, not incidental format. Under adversarial defection the tested model degrades gracefully by routing around decliners while the deterministic floor collapses to zero at any defection rate. The honest nulls: a small search fee changes behavior not at all (the tested model contacts every agent every round in both arms: cost-blind solicitation), and revealing values moves the score positively but within noise.
New worlds.
Every model held out.
Every model on the board now carries a score on 30 private held-out seeds per case, the one gap that blocked a defensible public leaderboard. The development leader stays the top point estimate in all five cases, and for the four models scored last the ranking is unchanged with case05's private arm is excluded from held-out headline claims; it informed the case's own promotion and is labeled a confirmation arm (see roadmap below).no overfitting signature: every |Δ| ≤ .032, and the largest deltas are positive (held-out above development). This is replication on the same case generators, not new mechanisms; held-out difference intervals across the field are still to be reported.
We never train these models, so held-out is not about a model memorizing; it guards against us tuning the benchmark to a convenient set of worlds. The private seeds are generated from a secret salt and never inspected before scoring, and they pre-empt teams optimizing to the visible cases once the benchmark is public. They check the point-estimate ranking against a second same-generator draw; they do not establish that the model-vs-model differences are statistically significant, nor that the ranking transfers to new case families. Both are open.
A paired bootstrap (each held-out world scored by both the leader and the greedy floor, the difference taken within-world, 2000 replicates, 30 paired worlds per case) puts the leader above the floor on all four cases: +.22 [+.16, +.26] on bilateral, +.11 [+.05, +.18] on consent, +.11 [+.03, +.18] on clearing, and +.05 [+.02, +.08] on discovery; the slimmest margin on the board sits where information hides. Pooled on the four-case held-out basis (episode bootstrap; seed sets are disjoint across cases): gemini-3.5-flash +.244 [.211,.275], kimi-k2 +.188 [.160,.217], gpt-5.6-luna +.147 [.129,.166], greedy floor +.123 [.112,.135], deepseek-v4-flash +.177 [.150,.205], minimax-m3 +.102 [.079,.125], gemini-2.5-flash +.078 [.062,.093], glm-4.7-flash +.073 [.056,.091], random +.013 [.007,.020]; every row n=120. Paired within-world: four models beat the floor significantly (gemini-3.5-flash +.120, kimi-k2 +.064, deepseek-v4-flash +.053, gpt-5.6-luna +.024), and the leader beats kimi-k2 head-to-head (+.056); minimax-m3 is statistically indistinguishable from the floor (−.021); gemini-2.5-flash (−.045) and glm-4.7-flash (−.050) sit significantly below it; and every model, including both below-floor rows, is significantly above the random baseline. Correction (1 Aug 2026). The deepseek, minimax and glm held-out rows are re-runs: their original runs were served with a completion budget their reasoning could exhaust, so 25–88% of their turns came back empty and were scored as silent no-ops. 450 re-run episodes, all replay-verified, on identical seeds and panel. The correction moves them +.057 to +.132 and changes the ordering: deepseek from ‘indistinguishable from the floor’ to third place; minimax and glm from ‘indistinguishable from random’ to significantly above it. Their development pools ran on the same defective path and are withdrawn until re-run. Vintage. Gemini held-out points (here and in the charts and matrix) are the 2026-07-25 trace re-execution: every cell 30/30, byte-replay verified. Re-running at temperature 0 shifted the gemini pools (leader −.020, gemini-2.5-flash −.059 vs the July-20 originals, which kept no traces): provider-side nondeterminism is real, and mid-board orderings should be read with that grain. New row. gpt-5.6-luna (OpenAI) joined 2026-07-30 on the held-out protocol: 150/150 replay-verified, the flattest per-case profile on the board (no collapse cell), and the first budget-tier model to clear both the floor and the deferred-settlement floor (.222 vs .211); its development pool is now complete too (150/150; dev +.147 [.132,.162], tracking its held-out score almost exactly), so it appears in every chart and the complete matrix alongside the other agents.
Efficiency says how much welfare a market realizes; a deployer also asks how much of it the agent under test keeps. From the recorded per-round utility deltas of the 480 held-out runs, the seat under test captures 45–52% of all realized surplus. The reference point is par: the share an agent would keep if every deal it joined split its gains equally among the agents at the table (50% of a bilateral deal, 33% of a three-party deal, 25% of a four-party deal), averaged over that model's own deals, weighted by the surplus each deal created. Par therefore differs by deal style: a mostly-bilateral trader has par near 50%, a multiparty dealmaker nearer 35–40%. Against its own par, every model lands +4 to +8 points above, a consistent, modest proposer premium. Two reads: no model buys its efficiency score by giving surplus away, and none extracts aggressively; the spread in raw capture share across models mostly reflects deal structure, not bargaining aggression. The distribution axis does not re-rank the board: the premium spans 3 points while efficiency spans 11×. These figures are computed on the four original cases’ held-out runs. The deferred-settlement case inverts the pattern: there, capture anticorrelates with efficiency: completing credit deals requires pricing counterparty trust into generous splits, so the strongest performers keep the smallest share (still without re-ranking a board that ranks efficiency, not capture).
| Model | Held-out AER | Surplus capture | Surplus by deal size | ||||
|---|---|---|---|---|---|---|---|
| capture | par | vs par | 2-party | 3-party | 4+-party | ||
| kimi-k2Moonshot | +.188 | 47% | 39% | +8pp | 40% | 50% | 10% |
| deepseek-v4-flashDeepSeek | +.177 | 53% | 45% | +8pp | 72% | 25% | 3% |
| minimax-m3MiniMax | +.102 | 51% | 43% | +9pp | 57% | 43% | — |
| glm-4.7-flashZ.ai | +.073 | 62% | 48% | +14pp | 90% | 8% | 2% |
Capture and par are pooled over each model's 480-run held-out set (4 cases × 30 private seeds, replay-verified); deal-size shares are the fraction of that model's realized surplus arising in trades with 2, 3, or 4+ parties, with parties determined at the transfer level: agents connected by actual transfers within a settlement form one trade (connected components), so two disjoint bilateral deals compiled in the same round count as two 2-party trades, not one 4-party trade. Multi-party trades are real, not aggregation artifacts: single connected 4- and 5-agent clearing chains appear in the traces. Self-play corroboration: in homogeneous populations (both sides the same model, zero skill asymmetry) the seat-1 premium is +4 to +8pp (the same band), so the premium is chiefly the proposer seat's structural advantage, not out-bargaining the panel. Capture data exists only where full traces were retained: the held-out board; the development and baseline rows will gain this column when their runs are next re-executed with trace retention.
Every cell, without the headline filter.
Every model against every case: development AER, its 95% interval, and the private held-out point estimate (◇), now for every model. Shaded cells mark held-out results at or below the greedy floor: the trivial baseline the model failed to beat.
What we are building next.
The pilot measures market-level welfare across four case families. These are the instruments in development to sharpen what it measures and test how far it generalizes.
- Information-constrained skill axis. A second, Bayes-achievable denominator, the most a Bayes-optimal agent could realize from the observable signals (Myerson–Satterthwaite makes this precise in the bilateral-trade special case), so a score reflects skill rather than situation difficulty. Implemented for the bundle world; pending here, and reported as a separate, non-pooled axis.
- Per-agent capture-share. The tested agent's own share of the surplus, separate from market welfare: a model can lift total welfare while giving away its principal's cut, or grab surplus while shrinking the pie.
- IR-violation companion. The rate at which executed trades leave an agent worse off, split into predation (harming a counterparty) and self-harm (harming itself).
- Adversarial counterparty. A self-interested panel that holds out for a better split and pushes disadvantageous deals, testing whether the agent gets exploited, not only whether it exploits.
- Deferred settlement (case05): the fifth case family. No swap can clear atomically: one transfer row per round forces pay-first, trust-the-counter-leg commerce. Its private run replicates the tier structure, and a model that holds mid-field on the four visible cases stays below the greedy floor here; trust-based commerce is a distinct skill axis. Because that private run informed the decision to promote the case, its ◇ points are labeled a confirmation arm rather than held-out; held-out headline claims use the four original cases, and a fresh held-out draw for case05 is planned.
- RL transfer. Whether reinforcement learning on one arena carries to others. A small model lifted its held-out arena welfare 0.29 → 0.47 over 40 steps; whether that carries to other case families is unverified; early off-configuration probes were mixed. Cross-case generalization is the open question.