A2Scripted
One agent, deterministic counterparts, oracle-scored outcomes.
Completing tasks doesn’t tell you whether an LLM can act in an economy. AERead measures what does: the share of a market’s available welfare an agent actually captures, under real frictions.
One agent, deterministic counterparts, oracle-scored outcomes.
The same world with live agents negotiating through a verified ledger.
Discovery, bargaining, consent, and clearing composed in one economy.
Train on capability cells; test transfer in an unseen arena.
A2Scripted isolates a single model against deterministic counterparties and scores the economic outcome with an oracle. It separates basic allocation skill from everything that live interaction adds.
AER (Aggregate Efficiency Ratio) is the share of the available gains from trade the agents actually capture — 1.00 reaches the fully efficient outcome, 0 is no improvement over the starting allocation, and a negative value means trades left the group worse off.
With the market held still, we can ask whether an agent can identify a welfare-improving action before asking it to infer, persuade, or coordinate with anyone else.
Frontier models saturate the visible, full-information construction cells. The open problem moves to strategic inference and pricing when a counterparty's reservation value is private—so fixed scripted construction becomes a floor, not the headline.
A2A replaces deterministic counterparts with live agents that negotiate, reveal information, and settle through a verified ledger. Aggregate efficiency then measures how much of the available welfare an agent captures once the market reacts and information is hidden—which is exactly where capability falls short of the visible ceiling.
A2A swaps A2Scripted's deterministic counterparts for negotiating LLM agents. The AER then measures how much of the available welfare the under-test model captures once counterparts can bargain and hide information.
The arena composes discovery, bargaining, consent, and partial clearing in one economy. It tests whether local agent skill survives the full interaction loop—and whether the result survives new counterparties and unseen seeds. Where the individual cases decompose the economy one capability at a time, the arena composes them back: the same frontier, at market scale.
The compiler and verifier are fixed parts of the environment, not the agent under test, so a score reflects the agent's own behavior. Invalid or unauthorized transfers are rejected before scoring rather than counted as trades.
When counterparties are hidden, the leading margin compresses. On held-out case03, gemini-2.5-flash remains below the deterministic greedy floor (.095 vs .108). The earlier clearing weakness narrows away, so the supported claim is discovery-specific.
A four-case panel swap keeps gemini-3.5-flash first under both a frozen-LLM and a scripted-rational panel, and first across 30 private held-out seeds per case. The leader is panel-robust; the deterministic floor is not—it makes no deals under the scripted panel on discovery and consent.
One model in every seat—a homogeneous population. The one-vs-panel score misses two things: a scale-dependent collapse for some models, and its absence for others.
The gemini and Qwen readings reuse pre-pipeline self-play runs (welfare-ratio, mean over 3 seeds). The deepseek and glm populations are measured in the official AER (n=20 per case), directly comparable to the leaderboard—and they show the collapse is not universal. Market mode stays a distinctive axis, not yet leaderboard-grade.
The RL environment turns capability cells into training tasks, then evaluates on an unseen integrated arena. Reward comes from realized utility and valid exchange, not from matching a preferred answer string.
A scoped Qwen3-4B probe improved clean held-out arena welfare over 40 training steps while keeping completion stable.
A compiler-free re-score preserved and slightly strengthened the trend, supporting a real economic gain rather than better output formatting.
Early evidence · broader arena and curriculum transfer remains openWhat AERead measures, how the market works, and what the pilot does and does not claim.