A Benchmark Test for LLM Persona Agents in a Macroeconomic Simulation 1.0.0
This model couples a two-asset heterogeneous-agent New Keynesian macroeconomy with a household layer that can be driven in two ways. One is a conventional calibration. The other is a set of large-language-model persona agents built from real survey households in Korea (NaSTaB), the United States (PSID) and Italy (SHIW). Each persona is given its balance sheet, income and prices, and reports consumption, saving and labour supply. A feasibility layer enforces the budget constraint. A conditional-mean surrogate then turns these decisions into a propensity surface. That surface reaches the macroeconomic block through a single channel, the intertemporal marginal propensity to consume matrix. Because of this, a calibrated “twin” can be aligned with the generative layer exactly, and the two can be compared policy by policy.
Both layers are then scored against the held-out consumption of the same households. The model’s role is to serve as a benchmark: it tests whether generative agents add information beyond the balance sheet, and whether that information matches real behaviour.
The deposit contains:
the macroeconomic block;
the persona construction and elicitation pipeline;
the validation gates;
the trained surrogate artifacts, which can be loaded and used without API access (load_surrogate.py);
the article, whose Section 2 is the ODD description, and the TRACE document.
Raw microdata and per-household elicitations are not redistributed.
Release Notes
1.0.0