one core model · six extensions · uk & us

An open platform for policy analysis and economic research.

The tax-benefit engine behind policyengine.org — what a reform does to each UK or US household, and what it costs — joined to six macro models: the OBR emulator, the Bank of England SVAR, the Fed’s FRB/US, a US HANK, an overlapping-generations model and DEFINE-UK for climate.

boe-svar: UK real GDP growth forecast, median with 68% and 90% credible bands Fan chart. Posterior forecast of UK year-on-year real GDP growth from the boe-svar model, 2026Q2 to 2029Q2, conditioned through 2026Q1. The median path runs from +1.14% to +1.41%. The 90% credible band is widest at 2029Q2, spanning -0.67% to +3.77%; its lower edge is positive for the first 3 quarters and dips below zero from 2027Q1 onward. -1% +0% +1% +2% +3% 2026 2027 2028 2029
UK GDP growth, boe-svar. Raw posterior bands — and against a drifting random walk, the model shows no forecast skill. See the evaluation →
why this one is different

Every model here says where it fails.

Each model page ends in a plain verdict, checked against the files that produced it. Each forecast goes on the record before the answer is known, then is scored once it is. The failures are published next to the wins.

the verdicts — three grades, every model gets one

  1. exact pe-microsim’s household arithmetic, to the penny.
  2. a benchmark, not a validation The HMRC costing comparison; a household-route multiplier near the OBR’s published 0.3, but a spending multiplier that is 1.0000 by construction.
  3. no skill boe-svar against a drifting random walk, once the 64 tests are adjusted.

the score — one period on the record so far

UK CPI inflation, 2026Q2

68% band

forecast 2.68%

2.24% 3.16%

outturn 2.80%

The forecast and its 68% band come from the round archived on 2026-07-21, before the data arrived. The ONS printed the outturn afterwards.

The outturn landed inside the band. One period inside a band is not a track record.

01 — how it fits together

How the models connect.

The microsimulation is the core: it knows every household, and nothing about the economy around them. Each macro model joins it at one named point, in one of two directions — a reform’s cost goes out to a macro model and feedback comes back (score_reform accepts microsim, obr, og, og+microsim), or a macro path comes in and is pushed down onto households as incidence. Nothing connects macro model to macro model.

Every bridge starts or ends at the microsimulation. How a reform is scored →

  1. Reform pe-microsimcost · who pays obr-macromacro feedback Scored result
    what moves
    The cost from the microsimulation goes into the OBR model as a change in household income.
    result
    1p on the basic rate takes GDP down, building to ~0.35× the revenue raised.
    caveat
    That is close to the OBR’s published 0.3 for income tax — the ~1-by-construction problem is on the spending lever, not this one. But no quarter of the solve converges, so treat the scale as roughly right and single quarters as unreliable.

    Score a reform →

  2. Reform OG-UKlong-run wages pe-microsimsecond run No scored result
    what moves
    OG-UK’s long-run wage changes feed a second microsimulation run.
    result
    No scored result is published.
    caveat
    The wage change is applied evenly across years, so there is no path from now to then.

    Open OG-UK →

  3. Economy-wide path pe-microsimwho pays Who gains and loses
    what moves
    Four more joins run this way: an economy-wide path goes in, and who gains and loses comes out.
    caveat
    All four are experimental.

    See every join →

united statesNo US macro bridge exists, so a US reform is scored statically.

02 — what each one is for

Six models, six different questions.

Pick by the question you have, not by the model you know. Each line is what that model is for, and what its evidence does and does not support.

Which model answers which question, by horizon and country A map of the six macro models on a shared time axis running from the recent past through coming quarters and three to five years out to decades. The UK lane holds boe-svar at the recent past for what is driving the economy now, obr-macro at three to five years for scoring a UK reform, and psl-og and define-uk at decades for long-run effects and climate scenarios. The US lane holds frb-us and us-hank in coming quarters, for shock response and for which households absorb a shock by wealth. Models drawn with a solid outline are checked against an external published anchor; psl-og and define-uk are drawn dashed because no independent outcome benchmark exists for them. recent past → now coming quarters 3–5 years decades UK US boe-svar what is driving it now obr-macro score a UK reform psl-og decades ahead define-uk climate scenarios frb-us US shock response us-hank who absorbs it, by wealth checked against a published anchor no independent outcome benchmark
Placed by the horizon each model answers over, and split by country. A solid outline means the model is checked against an external published anchor; a dashed one means no independent outcome benchmark exists for it.

obr-macro Tracks the March 2026 EFO to 0.15% GDP MAPE with the anchors held; free-running, 4.48%.

boe-svar Replicates the paper’s GDP decomposition (37.4% vs ~40%), 8pp short on CPI. Against a drifting random walk, no forecast skill survives the 64 tests run.

frb-us Matches the Fed’s own pyfrbus inside its two releases’ disagreement (~1×10−⁸); no predictive claim.

us-hank Solves Auclert et al. (2021) to four decimals — the headline targets are inputs, not results.

psl-og +1pp on the basic rate: GDP −£5.0bn (−0.14%), revenue +0.29pp of GDP by 2030. Targets met by construction; no independent outcome benchmark exists.

define-uk Baseline replicates the manual; scenarios are design-gated and paper-anchored. Deltas only, never levels.

Inspect all seven models, compare them side by side, or read the evidence for each.

03 — start using it

Run a hosted model, or use the code directly.

Connect the public MCP server with no PolicyEngine account or API key, use the shared CLI, or call each Python package directly.