Notes · 8 October 2026
Activities, not occupations
The fundamental value-creating units of an economy are activities, not occupations. How AERead decides what to measure and distributes its tasks, how that differs from the benchmarks that pick tasks by who is paid, and how AERead helps agents unlock value in the economy.
Two benchmarks now define what “economically valuable” means for AI agents. GDPval takes the sectors that contribute more than 5% of US GDP, the five best-paid digital occupations in each, and thirty tasks per occupation, 1,320 in all, each graded by blinded experts against a professional’s deliverable. Agents’ Last Exam clusters the digital occupations of the federal taxonomy into 13 domains and 55 sub-fields, sources 1,490 task instances from 250 practitioners, and scores each artifact against a verifiable reference. Both pick tasks by where work is paid, and both grade the deliverable. Neither weights a result by what the work is worth.
AERead asks a different question, so it has to pick differently. We benchmark agents’ readiness for value delivery: when an agent acts for a principal against other actors who have their own interests, does the outcome it reaches deliver value to the principal? That question is not organised by occupation. A purchasing agent, a sales manager and a lawyer sit at the same table when a supply contract is signed. It is organised by activity: what is done, with whom, and what depends on it.
A map of the economy by what is done
We built one. Every occupation in the US wage bill (BLS, May 2023: 151.9 million jobs, $9.94 trillion) is assigned to one of ten activities that firms, households and governments perform with other actors, or to “others” when the work transforms or delivers rather than deals. Shares of GDP follow, with the non-labour half of value added allocated in the same proportions; on labour alone the ten activities are 18.9%, which is the floor.
| Activity | Share of GDP | What a world can score |
|---|---|---|
| 1. Procurement and sourcing | 0.5% | all of it |
| 2. Selling and commercial dealing | 9.8% | the negotiated part: quoting, negotiation, account recovery, agents and brokers (5.1%) |
| 3. Financing, credit and settlement | 3.2% | appraisal, credit and structuring terms (2.1%) |
| 4. Hiring and workforce contracting | 1.3% | recruiting, pay and incentive terms (1.0%) |
| 5. Contracting, compliance and disputes | 2.3% | the decisions: terms, consent, settle or escalate |
| 6. Risk transfer and insurance | 0.5% | all of it |
| 7. Operations planning and coordination | 5.9% | multi-party clearing and deferred settlement only |
| 8. Allocating and delegating inside organisations | 10.3% | the mandate every task runs under, not a task |
| 9. Intermediating, verifying and advising | 3.3% | brokerage and advice (0.9%) |
| 10. Administering the public economy | 0.1% | through procurement, auctions and commons, with the state as principal |
| Ten activities | 37.2% | 12.4% passes every test below |
| Everything else: production, care, teaching, hospitality, engineering, creative and administrative work | 62.8% | the deliverable is the question, which GDPval and Agents’ Last Exam already ask |
Hours and decisions
Those shares are shares of what people are paid, and for the activities we care about they understate the stakes by two orders of magnitude, because the value of a deal is not in the hours spent on it. Procurement is 0.5% of GDP as an activity; the purchases it decides come to $23 trillion a year, three-quarters of GDP: $20 trillion of inputs bought by businesses, which GDP leaves out by construction because it counts only value added, and $3.1 trillion of public procurement. Hiring is 1.0% and sets the $15.7 trillion paid in compensation.
An acquisition makes the same point in one deal. Advisory fees run near 1% of deal value ($15.4 billion worldwide in the first half of 2024). The premium negotiated at the table averaged 32 to 33% of the target’s market value for US targets between $1 and $10 billion in 2023 and 2024, and 40 to 44% below $1 billion. And the academic record says where that premium goes: at announcement the combined gain to both companies averages +1.8% of their value, targets gain 16% and acquirers lose 0.7%, so on average the buyer hands the whole anticipated gain to the seller. That is a failure of price discipline, not of the quality of the model or the deck.
There is a general reason the hours lens stays small. When the tasks of a business are complements, a chain that is only as strong as its weakest link, making any one task infinitely productive raises output by no more than that task’s share of it; Chad Jones puts it as “software is about 2% of GDP, if we had infinite software we’d be 2% richer.” Automating transaction work is bounded by its 12.4% share in the same way. A better decision is bounded by something else: the misallocation in the flow it governs.
This is the part of the economy to optimise
These flows are where a better decision has the most to work on. The $23 trillion of purchases, the $15.7 trillion of compensation and the $5.6 trillion of investment are decided by a few hundred billion dollars of judgment, so a dollar of judgment here acts on more value than a dollar of judgment anywhere else in the economy. The value is realised through allocation.
Every one of these markets leaves value on the table because the parties cannot see each other’s information. A buyer who cannot tell a good supplier from a bad one pays for the average, and the good suppliers leave. An acquirer who cannot see the target’s books overpays and later writes the asset down. A worker and a job that fit never meet because neither side could verify the other. A better decision moves the outcome toward what the market would reach if everyone knew everything: the right supplier gets the contract, the right worker the job, the asset the owner who can use it best, the risk the party able to bear it. The distance closed is the value, and it is realised through four channels. Three create it; the fourth, price, moves it.
| Channel | How |
|---|---|
| Allocation | the right trades happen, the wrong ones do not, and assets reach the user who values them most |
| Terms, risk and incentives | risk to the party able to bear it, contracts that elicit effort, contingent terms that close a valuation gap |
| Transaction cost | less spent reaching and enforcing the deal |
| Distribution | the price, the premium, the split of surplus: this one moves value rather than creating it |
We count the value at two levels. The market-level score is the share of the full-information optimum that the market realises; it is immune to transfers by construction and is the headline. The principal-level score is the surplus the agent obtained against the reference, within its mandate; it captures the transfer, because the principal does care who gets the premium. They are reported side by side and never added.
- V
- value realised by better agents, $ per year
- Fc
- flow decided in class c, $ per year: known for every class
- mc
- misallocation in class c, the gap between what is realised and the optimum as a share of the flow: measured world by world
- rc
- share of that gap the agent recovers, in [−1, 1]: the benchmark result
Each case decomposes the gap between what was realised and the optimum and names the stage where it sits, which is how mc is measured. The sign of rc is not automatic, since an agent that is better at extracting or withholding can raise its principal’s take while lowering what the market realises, and one that is too eager or too timid lowers both; that is why the market-level score leads and a share of every task has no deal as the right answer.
What a world can score
We keep an activity when three things hold. Its value is decided rather than produced, so the flow it governs is large against what it costs to perform. There is a counterparty with its own interest and private information, so readiness means bargaining, matching, commitment, protection and refusal rather than output quality. And the result can be scored inside a world at the end of the episode against a computable reference: surplus, cost against the optimum, feasibility, a constraint that was violated.
12.4% of GDP passes all three: procurement, the negotiated part of selling, financing terms, hiring, contracting, risk transfer, brokerage and advice. By wage bill, 82% of that is business judgment and 18% is legal and compliance work, and that is the point rather than a problem. In our worlds judgment is not graded; its consequence is measured. A model is scored by what it obtained for its principal, not by whether an expert would nod at its reasoning.
How the tasks are distributed
Tasks are weighted by the flow a class of transaction decides, not by the hours spent on it, with each transaction counted once and its two seats sharing the weight. Five classes result, and they are also the five headline categories: we do not publish an undecomposed score.
- Business purchases · buyer and seller · $20.0T of inputs bought by businesses39%
- Labour contracts · employer and worker · $15.7T of compensation30%
- Capital and corporate transactions · lender and borrower, acquirer and target · $5.6T of investment11%
- Household purchases, the part a household deals over · household and firm · $3.2T of $20.9T10%
- Public purchases · state and bidder · $3.1T of public procurement10%
Household consumption is counted at the part a household actually deals over. Of the $20.9 trillion spent in 2025, 55% is bought at posted prices, where the decision is the seller’s; 12% is the imputed rent of owner-occupied homes, with no transaction at all; 17% is health care priced between providers and insurers. What remains, $3.2 trillion, is negotiated (leases, vehicles) or chosen on terms and claimed on (financial services and insurance). A firm’s spend is treated differently on purpose: it is decided at the contract and the automated orders execute it, so the whole flow counts.
Inside each class, five rules:
- Every class has tasks at the four stages a deal passes through: finding and qualifying the counterparty, price, structure and terms, and claims or renegotiation after signing, starting at 20, 35, 25 and 20%.
- A quarter to a third of tasks have refusing, walking away or escalating as the best move. In our own runs one model signed 264 of its 1,601 leases at zero rent; another, at its default settings, walked away from deals that were good for its principal. A set with only deal-optimal tasks rewards eagerness, and one with only traps rewards timidity.
- Every task carries the principal’s objective and authority limits, so escalation is a scoreable move rather than a way out.
- Every task records the seat the agent holds and whether the principal is a firm, a household or the state, so results decompose by seat and principal as well as by class and stage.
- Confirmatory runs face scripted, verified counterparties; the harness never judges a move.
Three ways to pick tasks
| GDPval | Agents’ Last Exam | AERead | |
|---|---|---|---|
| Unit | an occupation within a sector | a sub-field of digital occupations | an activity, by transaction class and seat |
| Selection | sectors over 5% of GDP; the top five occupations by wage bill with at least 60% digital tasks | SOC and O*NET occupations with shared software workflows; non-digital sectors excluded | activities whose value is decided, have a counterparty, and can be scored in a world |
| Allocation | 30 tasks per occupation, uniform | by expert availability; one task per sub-field in one tier | by the flow each class decides, floored at 10% |
| Scored by | blinded expert preference against a professional’s deliverable | artifact checks against a reference: gate, then score | the consequence in a world against a computed optimum and the mandate |
| A high score means | the deliverable is as good as a professional’s | the workflow was completed to specification | the principal got the value, within authority |
| Size | 1,320 tasks, 220 public | 1,490 instances, 150 public | v0: five cases; v1 follows the distribution above |
Sources
- Patwardhan et al., GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks, 2025. Sector and occupation selection, 1,320 tasks, expert pairwise grading.
- Sun et al., Agents’ Last Exam, 2026. 13 domains, 55 sub-fields, 1,490 instances, artifact-verified scoring.
- BLS, Occupational Employment and Wage Statistics, May 2023, national table.
- BEA via FRED, 2025 annual: GDP, compensation of employees, personal consumption and its components, gross private domestic investment, intermediate inputs.
- OECD, Government at a Glance 2025, public procurement as a share of GDP: United States 10.1% in 2024.
- Houlihan Lokey, M&A Market Overview, Q3 2024: acquisition premiums for US targets over $100 million, LSEG data.
- LSEG Deals Intelligence, Global M&A Advisory Review, full year 2024: completed M&A advisory fees.
- Jones, AI and Our Economic Future, Stanford Graduate School of Business, 2026: the weak-links model; see also Aghion, Jones and Jones, Artificial Intelligence and Economic Growth, 2017.
- Akerlof, The Market for “Lemons”, Quarterly Journal of Economics 84(3), 1970.
- Hsieh and Klenow, Misallocation and Manufacturing TFP in China and India, Quarterly Journal of Economics 124(4), 2009, Table VI.
- Myerson and Satterthwaite, Efficient Mechanisms for Bilateral Trading, Journal of Economic Theory 29(2), 1983.
- Wallis and North, Measuring the Transaction Sector in the American Economy, 1870–1970, 1986.
- SRS Acquiom, 2025 M&A Deal Terms Study: earnout frequency in private-target deals.
- Andrade, Mitchell and Stafford, New Evidence and Perspectives on Mergers, Journal of Economic Perspectives 15(2), 2001, Table 1.
- SEC Division of Economic and Risk Analysis, Analysis of Merger & Acquisition Activity, June 2025.