one core model · six extensions · uk & us
An open platform for policy analysis and economic research.
The tax-benefit engine behind policyengine.org — what a reform does to each UK or US household, and what it costs — joined to six macro models: the OBR emulator, the Bank of England SVAR, the Fed’s FRB/US, a US HANK, an overlapping-generations model and DEFINE-UK for climate.
Every model here says where it fails.
Each model page ends in a plain verdict, checked against the files that produced it. Each forecast goes on the record before the answer is known, then is scored once it is. The failures are published next to the wins.
the verdicts — three grades, every model gets one
- exact pe-microsim’s household arithmetic, to the penny.
- a benchmark, not a validation The HMRC costing comparison; a household-route multiplier near the OBR’s published 0.3, but a spending multiplier that is 1.0000 by construction.
- no skill boe-svar against a drifting random walk, once the 64 tests are adjusted.
the score — one period on the record so far
UK CPI inflation, 2026Q2
68% band
forecast 2.68%
2.24% 3.16%outturn 2.80%
The forecast and its 68% band come from the round archived on 2026-07-21, before the data arrived. The ONS printed the outturn afterwards.
The outturn landed inside the band. One period inside a band is not a track record.
How the models connect.
The microsimulation is the core: it knows every household, and nothing about the economy around them. Each macro model joins it at one named point, in one of two directions — a reform’s cost goes out to a macro model and feedback comes back (score_reform accepts microsim, obr, og, og+microsim), or a macro path comes in and is pushed down onto households as incidence. Nothing connects macro model to macro model.
Every bridge starts or ends at the microsimulation. How a reform is scored →
-
- what moves
- The cost from the microsimulation goes into the OBR model as a change in household income.
- result
- 1p on the basic rate takes GDP down, building to ~0.35× the revenue raised.
- caveat
- That is close to the OBR’s published 0.3 for income tax — the ~1-by-construction problem is on the spending lever, not this one. But no quarter of the solve converges, so treat the scale as roughly right and single quarters as unreliable.
-
- what moves
- OG-UK’s long-run wage changes feed a second microsimulation run.
- result
- No scored result is published.
- caveat
- The wage change is applied evenly across years, so there is no path from now to then.
-
- what moves
- Four more joins run this way: an economy-wide path goes in, and who gains and loses comes out.
- caveat
- All four are experimental.
united statesNo US macro bridge exists, so a US reform is scored statically.
Six models, six different questions.
Pick by the question you have, not by the model you know. Each line is what that model is for, and what its evidence does and does not support.
obr-macro Tracks the March 2026 EFO to 0.15% GDP MAPE with the anchors held; free-running, 4.48%.
boe-svar Replicates the paper’s GDP decomposition (37.4% vs ~40%), 8pp short on CPI. Against a drifting random walk, no forecast skill survives the 64 tests run.
frb-us Matches the Fed’s own pyfrbus inside its two releases’ disagreement (~1×10−⁸); no predictive claim.
us-hank Solves Auclert et al. (2021) to four decimals — the headline targets are inputs, not results.
psl-og +1pp on the basic rate: GDP −£5.0bn (−0.14%), revenue +0.29pp of GDP by 2030. Targets met by construction; no independent outcome benchmark exists.
define-uk Baseline replicates the manual; scenarios are design-gated and paper-anchored. Deltas only, never levels.
Inspect all seven models, compare them side by side, or read the evidence for each.
Run a hosted model, or use the code directly.
Connect the public MCP server with no PolicyEngine account or API key, use the shared CLI, or call each Python package directly.