Computational Model Library

A Benchmark Test for LLM Persona Agents in a Macroeconomic Simulation (1.0.0)

This model couples a two-asset heterogeneous-agent New Keynesian macroeconomy with a household layer that can be driven in two ways. One is a conventional calibration. The other is a set of large-language-model persona agents built from real survey households in Korea (NaSTaB), the United States (PSID) and Italy (SHIW). Each persona is given its balance sheet, income and prices, and reports consumption, saving and labour supply. A feasibility layer enforces the budget constraint. A conditional-mean surrogate then turns these decisions into a propensity surface. That surface reaches the macroeconomic block through a single channel, the intertemporal marginal propensity to consume matrix. Because of this, a calibrated “twin” can be aligned with the generative layer exactly, and the two can be compared policy by policy.

Both layers are then scored against the held-out consumption of the same households. The model’s role is to serve as a benchmark: it tests whether generative agents add information beyond the balance sheet, and whether that information matches real behaviour.

The deposit contains:

the macroeconomic block;
the persona construction and elicitation pipeline;
the validation gates;
the trained surrogate artifacts, which can be loaded and used without API access (load_surrogate.py);
the article, whose Section 2 is the ODD description, and the TRACE document.
Raw microdata and per-household elicitations are not redistributed.

Release Notes

1.0.0

Associated Publications

submitted / under review

A Benchmark Test for LLM Persona Agents in a Macroeconomic Simulation 1.0.0

This model couples a two-asset heterogeneous-agent New Keynesian macroeconomy with a household layer that can be driven in two ways. One is a conventional calibration. The other is a set of large-language-model persona agents built from real survey households in Korea (NaSTaB), the United States (PSID) and Italy (SHIW). Each persona is given its balance sheet, income and prices, and reports consumption, saving and labour supply. A feasibility layer enforces the budget constraint. A conditional-mean surrogate then turns these decisions into a propensity surface. That surface reaches the macroeconomic block through a single channel, the intertemporal marginal propensity to consume matrix. Because of this, a calibrated “twin” can be aligned with the generative layer exactly, and the two can be compared policy by policy.

Both layers are then scored against the held-out consumption of the same households. The model’s role is to serve as a benchmark: it tests whether generative agents add information beyond the balance sheet, and whether that information matches real behaviour.

The deposit contains:

the macroeconomic block;
the persona construction and elicitation pipeline;
the validation gates;
the trained surrogate artifacts, which can be loaded and used without API access (load_surrogate.py);
the article, whose Section 2 is the ODD description, and the TRACE document.
Raw microdata and per-household elicitations are not redistributed.

Release Notes

1.0.0

Version Submitter First published Last modified Status
1.0.0 Taewan You Thu Oct 8 11:09:49 2026 Thu Oct 8 11:09:50 2026 Published

Discussion

This website uses cookies and Google Analytics to help us track user engagement and improve our site. If you'd like to know more information about what data we collect and why, please see our data privacy policy. If you continue to use this site, you consent to our use of cookies.
Accept