← Back to overview
Agentic exchange economy · the frontier problem

Hidden information is the frontier. No model has crossed it.

The test is how much of a market’s welfare an agent captures when it cannot see everything. With values in view models make real progress; hide them—discovery, private pricing, consent—and capability collapses to where a trivial baseline competes. Characterizing that frontier, not ranking models, is the point.

in view
With values and counterparties visible, models make real progress.
the tractable part
hidden
Discovery, private pricing, and consent stay unsolved—for every model tested.
the frontier
↓ baseline
On the hard cases, capable models are matched by a trivial deterministic search.
the diagnostic
structural
The gap holds across providers and price tiers—a capability frontier, not a model artifact.
not who, but what
01 / The evidence

The gap is structural, not a model story.

Every model faces the same four cases in the same 5-agent market—one seat under test against a four-agent frozen panel (gemini-2.5-flash, temperature 0). The ranking is secondary; what matters is that the hidden-information cases stay far below the visible ceiling for all of them, and a trivial floor still discriminates. The baselines and the Gemini models carry private held-out scores (◇); the cross-provider models are development-only.

AER (Aggregate Efficiency Ratio) is the share of the available gains from trade the agents actually capture — 1.00 reaches the fully efficient outcome, 0 is no improvement over the starting allocation, and a negative value means trades left the group worse off.

Each case isolates one capability—bilateral exchange, clearing, discovery, consent—so the frontier can be located one friction at a time. The integrated arena composes them into a single live market: the cases decompose the economy; the arena is where it must hold together.

01 · b / The field

Every model, the same economy.

Pooled AER over the four cases, 30 development seeds each, with bootstrap 95% intervals. Read it for the shape, not the ranking: a trivial deterministic floor divides the field, and no model closes the hidden-information cases.

  1. 1gemini-3.5-flashGoogle+.209[.176–.243]
  2. 2kimi-k2Moonshot+.156[.131–.183]
  3. 3deepseek-v4-flashDeepSeek+.116[.089–.144]
  4. 4greedy floorbaseline+.104[.092–.116]
  5. 5gemini-2.5-flashGoogle+.083[.065–.101]
  6. 6minimax-m3MiniMax+.028[.014–.045]
  7. 7glm-4.7-flashZ.ai+.014[.005–.025]
  8. 8randombaseline+.008[.003–.015]

Development seeds only; the cross-provider models are dev-only, while the held-out ◇ generalization below covers the baselines and the Gemini models. Baselines are shown in grey. Each row pools the four cases over 30 seeds (n≈120), replay-verified.

01 · c / Per case

Baselines and the Gemini models, case by case.

01 · d / The composed market

Does the edge survive a crowd of copies?

Same 5-agent market as the leaderboard above, one difference in who fills the seats. Single-seat is the model in one seat against a four-agent frozen panel—that is the leaderboard score. Population fills all five seats with copies of the same model. The test: does a single-seat edge survive when every counterparty is a copy of itself? Two models have been run this way in the official AER.

Model Pooled AER Population AER · by case
single-seat population bilateralclearingdiscoveryconsent
deepseek-v4-flashDeepSeek +.116+.129 .108.194.119.094
glm-4.7-flashZ.ai +.014+.014 .016.008.015.023

deepseek holds in a crowd: its pooled population AER (+.129) matches its single-seat pooled score (+.116, from the leaderboard above), and on clearing the population reaches .194—well above deepseek's single-seat clearing of .064 (complete matrix below). glm is weak in both modes. These population runs are official AER: a 5-agent market, n=20 per case, directly comparable to the leaderboard above. The gemini-2.5 scale-collapse and under-trading ceiling reported under Findings come from legacy self-play (population welfare-ratio)—a separate measurement, not on this board.

02 / Findings

The shape of the gap.

The load-bearing result is not who leads. It is that capture stays far from the achievable frontier, that the shortfall widens as information hides, and that a trivial baseline already recovers much of it.

01 / Magnitude

Most of the welfare goes uncaptured.

Even the strongest agent reaches only a fraction of the welfare a market makes available; across the field, most of the achievable surplus is left on the table. The frontier is wide, not a rounding error.

02 / Structure

The gap widens as information hides.

Capture is highest with values and counterparties in view, and falls off through clearing, consent, and discovery. The hidden-information end is where every agent drops toward—and past—a trivial floor.

03 / Scale

It worsens as the market grows.

Coordination is a second frontier axis. In market mode—one policy in every seat—welfare capture falls as the population scales: a homogeneous gemini-2.5-flash population reaches .235 of the optimum at five agents and .085 at fifteen. Market-mode reference: welfare-ratio, not the AER leaderboard.

04 / Failure surface

Discovery is the sharpest unresolved case.

On case03 held-out seeds, gemini-2.5-flash scores .095 versus the greedy floor at .108. The apparent clearing weakness in case02 narrowed away; the supported claim is discovery-specific.

03 / Validation

Change what could have made the result brittle.

Two checks target different failure modes: replacing the counterparty panel tests dependence on who the model faces; changing every seed tests whether selection tuned to a convenient set of worlds.

Four-case panel swap

New counterparties.
Same leader.

Across all four cases, gemini-3.5-flash tops the board under both a frozen-LLM panel and a scripted-rational panel—so the headline ranking is panel-robust. The deterministic floor is not: under the scripted panel it drops to zero on discovery and consent (a bilateral-IR script makes no deals there), and gemini-2.5’s standing relative to the floor flips between panels.

panel sensitivity · four casesleader robust · floor panel-dependent
Private-seed check

New worlds.
Same leader.

Across 30 private seeds per case, gemini-3.5-flash remains first in all four cases. Individual cells can move in either direction—including case02—while the leader-level result generalizes.

unseen worldsheld-out leader · 4 / 4

We never train these models, so held-out is not about a model memorizing—it guards against us tuning the benchmark to a convenient set of worlds. The private seeds are frozen before design and never inspected, they pre-empt teams optimizing to the visible cases once the benchmark is public, and they separate a real ranking from the luck of one 30-world draw.

04 / Complete matrix

Every cell, without the headline filter.

Every model against every case: development AER, its 95% interval, and—for the baselines and the Gemini models—the held-out point estimate (◇). The four cross-provider models are development-only. The shaded cell is the one held-out result that stays below the greedy floor.

05 / Verdict

A deployment verdict, not a score.

AERead scores the welfare an agent captures under realistic frictions—against the frontier that is actually achievable, not a god’s-eye ideal—so the field reads as a readiness verdict, case family by case family.

Ready

Visible bilateral construction.

With values and counterparties in view, the frontier reaches the oracle ceiling. Deterministic search already clears the floor here, and the strongest models add a real margin on top.

full informationsaturated
Not ready

Hidden-information pricing, discovery, and coordination.

When reservation values are private, counterparties are hidden, or the market must self-organize, capability spreads wildly and a trivial deterministic floor stays competitive. This is the unsolved frontier.

hidden informationunsolved
Closeable

A gap, not a ceiling.

The metric charges only the achievable shortfall—the impossible part is carved out. A trivial floor already captures welfare the weaker models miss, and a small model lifted its arena welfare 0.29 → 0.47 with RL. The frontier is reachable, and it moves with training.

reachable · trainablea research target

Scoring under hidden information. AER’s denominator is the achievable welfare frontier—the most a Bayes-optimal agent could capture given only the observable signals—not the god’s-eye optimum. The value no learner can recover (the Myerson–Satterthwaite floor) is carved out, so a low score measures welfare the agent could have captured but did not. Current runs use a conservative fallback that under-credits, never over-credits.

06 / Provenance

A result designed to be replayed.

The page is the readable layer, not the source of truth. Seeded configurations, result summaries, validation outputs, and replay artifacts remain preserved alongside the experiment writeups. The full program spans seven models across six providers, adding Qwen-4B (Alibaba) in the RL and market-mode strands.

Metric

Aggregate efficiency ratio

Pooled AER is the sum of gated realized welfare divided by the sum of available welfare gain. The displayed absolute values use the conservative raw fallback tier.

Execution

Pinned and replay-verified

Four cases × nine agents (three baselines, the two Gemini models, and four cross-provider models) × 30 development seeds, run through the official scoring pipeline. All 1,073 development runs passed replay verification. The constant no-op floor is run for integrity but omitted from the boards above, where it is always zero.

Artifacts

Readable and auditable

Results, seeded configs, and validation outputs are preserved alongside the report. Raw private held-out seeds remain in the source repository.