evidence by model

See the evidence for each model.

Compare results with official benchmarks—and see clearly where no replication ground truth exists.

01 — start here

Validation is a gradient, not a badge.

The five models support different kinds of checks. The verification gradient groups them into three classes; choose a row, then open its evidence below.

The verification gradient: what each model is checked against, and its headline result
ModelVerification classWhat it is checked againstHeadline result
obr-macro Replication with a published anchor OBR Economic and Fiscal Outlook, March 2026; HMRC ready reckoner; ONS outturns Anchored GDP MAPE 0.15%, consumption 0.25%
boe-svar Replication with a published anchor Brignone & Piffer (2025), BoE Macro Technical Paper No. 3; ONS outturns FEVD shares 42.1% / 49.5% vs the paper's ~40% / ~50%
frb-us Replication with a published anchor The Fed's LONGBASE database and pyfrbus solver; Coenen et al. and CBO multiplier ranges Tracking invariant 5.6×10−17 — machine precision
pe-microsim Checked against implemented legislation Implemented UK and US legislation, rule by rule, for a specified household Deterministic household calculation; population estimates add survey and calibration uncertainty
psl-og Calibrated counterfactual — no ground truth ONS and OBR aggregates as calibration targets, not as a validation set Targets met by construction; no published replication exists
Imposed or targeted matches are anchors, not validation. Stronger evidence comes from out-of-sample forecasts, independent reform costings, and solver residuals.

Software fidelity and economic validity are separate.

This audit covers the three macro models. Levels are categorical evidence judgements, not a synthetic score: a model can reproduce its reference code perfectly while still lacking independent forecast or counterfactual validation.

Independent model-quality audit by evidence dimension
Evidence dimensionobr-macroboe-svarfrb-us
Implementation fidelityModerate — incomplete channelsStrong — restrictions and identities gatedStrong — four gated scenarios at the reference-solver noise floor
Predictive validationWeak — raw GDP 5.75%, consumption 9.56% MAPEWeak — the CPI win does not survive a drift benchmark; Bank Rate is the only defensible claimNot assessed — LONGBASE is not a Fed forecast
Identification robustnessNot applicableModerate — proxy data and specification matterNot applicable
Policy-counterfactual validityWeak — one independent tax benchmark; spending multiplier ~1.0 vs the OBR's 0.6Not applicable — reform scoring refusedModerate — published ranges, VAR expectations only
Uncertainty calibrationWeak — sensitivity envelopes pendingModerate — posterior bands, limited coverage historyWeak — bootstrap engine exists, public runs remain deterministic
Vintage reproducibilityModerate — multi-EFO archive pendingModerate — public proxies replace internal seriesModerate — artifacts hashed; multi-vintage tests pending
02 — obr-macro

Against the OBR's own forecast, and HMRC's reckoner.

Anchored to the March 2026 EFO with OBR add-factors, the emulator reproduces GDP to 0.15% MAPE and consumption to 0.25% over 2025Q1–2027Q4 (CI gate: 1%). GDP remains within 0.29% through 2031Q1, but unemployment becomes unreliable after 2027Q4. This fit is by construction; the independent reform costing below is the stronger test. The headline chart and table now use the live March 2026 vintage; the older November 2025 forecast is retained only inside the historical outturn audit.

obr-macro: anchored baseline vs March 2026 EFO, quarterly deviation Line chart. Quarterly percentage deviation of the anchored emulator from the published March 2026 EFO, 2025Q1 to 2027Q4. Real GDP ranges from -0.16% to +0.28% (mean absolute deviation 0.15%); consumption from -0.27% to +0.46% (mean absolute deviation 0.25%). Both series stay well inside the plus or minus 1% band at which continuous integration hard-fails the build, which is off the top and bottom of this frame. real GDP (peak +0.28%) consumption (peak +0.46%) -0.6% -0.3% 0 +0.3% +0.6% 2025Q1 2026Q1 2027Q1 2027Q4
Current March 2026 EFO baseline. CI hard-fails at ±1.00% — off the top and bottom of this frame. MAPE: 0.15% for GDP and 0.25% for consumption over 2025Q1–2027Q4. Computed from papers/obr-macro/figures/fig_anchored_data.csv, regenerated from the March 2026 detailed forecast tables on 21 July 2026.

The independent check is the reform costing. Raising the basic rate of income tax by 1pp from April 2026 scores at £6.46bn in 2026–27 through the PolicyEngine static-costing bridge, against HMRC's ready-reckoner figure of £6.9bn — a deviation of −6.4%; the costing sits inside the £6–8bn range spanned by recent published vintages. The gap widens in later years (−15.6% by 2028–29) where HMRC's figures embed administrative-data fiscal drag that survey microdata capture less fully.

obr-macro: 1p on the basic rate, ours vs HMRC ready reckoner (£bn/yr) Grouped bar chart in billions of pounds per year. PolicyEngine's static costing of a 1 percentage point rise in the UK basic rate of income tax, against HMRC's Direct effects of illustrative tax changes ready reckoner, June 2025 vintage. For the basic rate +1pp, 2026–27 group, ours is 6.46 against HMRC’s 6.90, a deviation of -6.4%. For the basic rate +1pp, 2028–29 group, ours is 6.92 against HMRC’s 8.20, a deviation of -15.6%. The 2028–29 emulator figure is interpolated between the scored endpoints £6.46bn in 2026 and £7.38bn in 2030. 0 2 4 6 8 10 6.46 ours 6.90 HMRC basic rate +1pp, 2026–27 6.92 ours 8.20 HMRC basic rate +1pp, 2028–29
£bn/yr. 2028–29 emulator figure interpolated between scored endpoints (£6.46bn 2026, £7.38bn 2030). PolicyEngine's static costing of a 1pp basic-rate rise against HMRC's Direct effects of illustrative tax changes (June 2025 vintage). Source: obr-macro working paper, comparison table panel B.
obr-macro against the current March 2026 OBR EFO and HMRC's ready reckoner
OursOfficialDeviation
Anchored levels vs EFO March 2026, £bn/qtr
Real GDP, 2025Q1703.8703.4+0.05%
Real GDP, 2027Q4730.6728.6+0.28%
Consumption, 2025Q1429.7429.3+0.09%
Consumption, 2027Q4445.5443.4+0.46%
Basic rate +1pp vs HMRC ready reckoner, £bn/yr
2026–276.466.9−6.4%
2028–29 (interpolated)6.928.2−15.6%
2030 (end of window)7.38≈8.2−10.0%
Tracking error vs the published March 2026 EFO path, MAPE 2025Q1–2027Q4
Anchored (GDP / consumption)0.15% / 0.25%CI gate <1% — passes
Held add-factors (GDP / consumption), 2026Q1–2027Q4 window0.37% / 0.33%6 of 8 computed within band (paper's Nov-2025-vintage scorecard, over its own longer window: 2.2% / 3.6%)
Free-running, raw (GDP / consumption)5.75% / 9.56%over band — report-only, and weak
How much of the OBR emulator scorecard the model actually computes Two stacked bars. Of 21 headline variables in the OBR emulator calibration scorecard, 11 are actually computed by the model and 10 are passthrough, held at the OBR published value and therefore scoring zero error trivially. Of the 11 computed, 3 are fair (real GDP 5.75 per cent, consumption 9.56 per cent, trade balance 0.70 per cent of GDP), 1 is a trivial accounting identity (employment, where both inputs are themselves passthrough), 5 are poor and 2 are off (company profits 79.80 per cent, current account 4.17 per cent of GDP). Four of the 11, or 36 per cent, land within band, and one of those four is the trivial identity, so only three non-trivial computed variables are in band. 21 headline scorecard variables 11 computed 10 passthrough — held at the OBR value of which, the 11 the model computes fair 3 identity 1 poor 5 off 2 Only 4 of the 11 computed variables land within band — 36% — and one of those four is a trivial identity. bands: rates ±1.0pp · net balances ±1.5% of GDP · levels ≤10% MAPE
Raw (un-anchored) calibration against the March 2026 EFO, a single vintage. The two “off” verdicts are bounded by constants the OBR does not publish — the OSHH base constant, the RPI index history, and the overseas rate-of-return normalisers — rather than by live bugs, and are gated against regression rather than tuned to the target. Source: docs/calibration_scorecard.md in the obr-macroeconomic-model repository.

The spending multiplier is ~1.0 by construction, against the OBR's own published 0.6. Under the demand closure a spending shock lands directly in the GDP identity and the behavioural second round is largely inactive, so a £5bn injection returns almost exactly £5bn of GDP. That is roughly a two-thirds overstatement of the impact multiplier relative to the institution being replicated, and it applies to every spending-side figure this page reports. It is the single most important number for anyone reading a policy score off this model.

Three of the channels that would normally damp such a shock are held exogenous in this configuration — imports (so the usual leakage is switched off), Bank Rate (no monetary offset) and CPI. The first two bias the estimated effect upward; the dead profits→dividends channel biases corporation-tax effects downward. These do not cancel, and their net sign is not established. Reform figures here are best read as relative comparisons between structurally identical runs, where the biases largely difference out, rather than as magnitudes.

The honest scorecard, the two dead equations, outturns, and the March 2026 re-anchoring

The honest scorecard. The last row is the one that matters for anyone tempted to read the 0.15% as forecasting skill. The same equations that track the EFO to 0.15% when anchored miss it by 5.75% free-running — precisely the gap the OBR's own add-factor judgement closes in the official process, and the reason reform deltas are always scored against the anchored baseline rather than the raw one. The free-running score is de-seeded, with passthrough variables excluded, and is published report-only. Across the full scorecard only 4 of the 11 computed headline variables land within band (real GDP, consumption, the trade balance, and the trivial employment identity); of the full 21-line scorecard, 10 lines are passthroughs held at the OBR value. The worst line is company profits at 79.80% MAPE on the March baseline (54.57% on the paper's November vintage), which traces to a single unpublished constant in households' operating surplus OSHH — the paper documents and regression-gates it rather than re-tuning it, since tuning it would be fitting to the answer. Other lines that moved on re-anchoring, reported rather than smoothed: the free-running current account widened from 2.76 to 4.17% of GDP and is now over band; RPI improved from 2.03pp to 1.71pp; business investment worsened from 15.48% to 16.12%; the two household-income lines stand at 14.15% and 13.86%.

obr-macro: real GDP level, anchored vs free-running vs the March 2026 EFO (£bn/qtr) Line chart of quarterly real GDP levels in billions of pounds, 2025Q1 to 2027Q4. The published March 2026 EFO path rises from 703.4 to 728.6. The anchored emulator is visually indistinguishable from it, running from 703.8 to 730.6 (mean absolute deviation 0.15 per cent, recomputed here from the plotted series). The free-running emulator, de-seeded and with no add-factors, contracts from 691.0 to 663.1 — a gap that widens to 65 billion pounds, 5.75 per cent mean absolute deviation over the horizon. Free-running and EFO paths from papers/obr-macro/figures/fig_free_running_data.csv; anchored path from papers/obr-macro/figures/fig_anchored_data.csv. Coordinates: value v in billions maps to y = 292 - (v - 660) * 2.95 on a 660 to 740 axis; quarter i of 12 maps to x = 58 + i * 60.545. anchored (0.15% MAD) free-running (5.75% MAD) EFO Mar 2026 660 680 700 720 740 2025Q1 2026Q1 2027Q1 2027Q4
Current March 2026 EFO baseline. Real GDP, £bn/qtr, 2025Q1–2027Q4. The anchored path (0.15% MAPE) sits on top of the EFO; the same equations free-running — de-seeded, no add-factors — contract away from it (5.75% MAPE). Computed from papers/obr-macro/figures/fig_free_running_data.csv and fig_anchored_data.csv, regenerated on 21 July 2026.
Two equations do not fire. The October 2025 model file contains two household-income equations with a bare log(X) left-hand side — log(HHTFA) and log(NDIVHH). The transpiler now parses that form, but both still fail to execute for an independent reason: their exogenous right-hand-side inputs MAJGDP and CORP are absent from the databank entirely, so the right-hand side evaluates non-finite and the solver skips the update. Every OBR figure on this page is computed with those two channels inert. The practical effect is bounded and demonstrable: household dividend income does not respond to a corporate-profits change, so the household-income channel of a corporation-tax reform is muted — a 5pp corporation-tax rise moves FYCPR by −£1,780m while ΔNDIVHH is exactly zero. Reviving the channel needs a sourced CORP series, a calibration decision rather than a code fix.

Forecast versus outturn. Comparing one forecast vintage with another tests agreement, not accuracy. Against ONS outturns published since anchoring, quarter-on-quarter real GDP growth ran 0.1% in 2025Q2, 0.2% in Q3, 0.2% in Q4 and 0.6% in 2026Q1. The emulator's path (0.15, 0.14, 0.25, 0.37) tracks the three 2025 quarters to within 0.06 percentage points, but — like the November EFO it inherits (0.28, 0.20, 0.27, 0.39) — misses the strong 2026Q1 outturn by roughly a quarter of a point. Two caveats govern the reading: this is primarily a test of the OBR's November vintage, the emulator's own contribution being the 0.02–0.13 point gap between the two model rows; and ONS quarterly estimates are themselves revised, so the outturn is a moving target.

obr-macro: quarterly real GDP growth — emulator vs EFO Nov 2025 vs ONS outturn (% q/q) Grouped bar chart, percentage quarter-on-quarter real GDP growth for the four quarters with ONS outturns since anchoring. 2025Q2: emulator 0.15, EFO 0.28, ONS 0.1; 2025Q3: emulator 0.14, EFO 0.20, ONS 0.2; 2025Q4: emulator 0.25, EFO 0.27, ONS 0.2; 2026Q1: emulator 0.37, EFO 0.39, ONS 0.6. The emulator tracks the three 2025 outturns to within 0.06 points; both the emulator and the November EFO it inherits miss the strong 0.6 per cent 2026Q1 outturn by roughly a quarter of a point. Data from papers/obr-macro/figures/fig_outturn_data.csv. Coordinates: value v maps to y = 258 - v * 331.4 on a 0 to 0.7 axis. emulator EFO Nov 2025 ONS outturn 0 0.2 0.4 0.6 0.15 0.28 0.10 2025Q2 0.14 0.20 0.20 2025Q3 0.25 0.27 0.20 2025Q4 0.37 0.39 0.60 2026Q1
November 2025 EFO vintage — the working paper's study; the live baseline is anchored to the March 2026 EFO. % q/q real GDP growth. The emulator tracks the three 2025 outturns to within 0.06pp; both model rows miss the strong 2026Q1 outturn (0.6%) by roughly a quarter point. Computed from papers/obr-macro/figures/fig_outturn_data.csv; the table below carries the exact values.
Quarterly real GDP growth: emulator, EFO vintage, and ONS outturn
Real GDP, % q/qEmulatorEFO Nov 2025ONS outturn
2025Q20.150.280.1
2025Q30.140.200.2
2025Q40.250.270.2
2026Q10.370.390.6

Vintage: re-anchored to the March 2026 EFO. The hosted emulator has been re-anchored from the November 2025 EFO to the OBR's March 2026 forecast, and the headline numbers on this page are computed on that baseline: anchored GDP 0.15% MAPE, consumption 0.25% over 2025Q1–2027Q4, with the anchored horizon extended to 2031Q1 (GDP reproduced to 0.29% at 2031Q1; anchored unemployment is unreliable beyond 2027Q4, drifting to 0.9% against the EFO's 4.1% by 2031Q1). The headline and free-running charts now use March 2026. Only the outturn backtest above retains November 2025, because changing its forecast vintage would erase the historical forecast being tested. The working paper now opens with a dated current-vintage note. Reform effects are differences between structurally identical runs and are insensitive to modest baseline drift, which is why the re-anchoring leaves the £6.46bn/£7.38bn static costing untouched and moves the second-round GDP effect only from −0.057% to −0.058% by 2027Q4.

03 — boe-svar

Against a Bank of England paper, and against outturns.

Against Brignone & Piffer (2025), the one-year share of UK forecast-error variance from global shocks is 42.1% for GDP and 49.5% for CPI. The paper reports roughly 40% and 50%.

boe-svar: global-shock FEVD shares, ours vs Brignone & Piffer (2025) Grouped bar chart. Share of UK forecast-error variance attributed to identified global shocks (world demand, energy and supply) at the one-year horizon. For GDP, our 10,000-draw production run gives 42.1% against the paper's 40.0%; for CPI, 49.5% against 50.0%. GDP sits 2.1 points above the paper's figure, CPI half a point below. The paper's values are approximate. 0% 20% 40% 60% 42.1 ours 40.0 paper UK GDP, 1-yr horizon 49.5 ours 50.0 paper UK CPI, 1-yr horizon
Global shocks = world demand + energy + supply. Ours = 10,000-draw production run; paper values approximate. Global-shock FEVD shares at the one-year horizon: the production run against the paper's published values. Source: papers/boe-svar/figures/comparison_numbers.json and the paper's validation table.
Why the fast CI configuration lands 6–10 points off

The deliberately cheap unweighted CI configuration lands at 49.6% and 43.8% — 6–10 points off, in opposite directions for the two variables. That gap is itself informative rather than embarrassing: the importance weights correct the Arias et al. (2018) zero-restriction sampler towards the uniform-over-rotations posterior, and applying them moves both shares substantially towards the published values, though on the 2026Q1 vintage GDP still lands 2.1 points high. Both configurations sit inside the wide [30, 60]% acceptance band that CI enforces, which is chosen to catch gross regressions without being flaky to Monte-Carlo variation.

boe-svar global-shock FEVD shares against Brignone & Piffer (2025)
Global share, 1-yr FEVDPaperCI config (fast)Production (10k draws)Deviation, CI / production
UK GDP~40%49.6%42.1%+9.6 / +2.1 pp
UK CPI~50%43.8%49.5%−6.2 / −0.5 pp

Out-of-sample. Forecasting seven quarters ahead from a frozen 2024Q2 edge, RMSE is 0.3pp for both GDP growth and CPI inflation. Activity tracks well; the median misses the 2025 inflation peak by up to 0.6pp but captures its reversal.

Fourteen of fourteen outturns inside the 68% band is a calibration failure, not a pass. A correctly calibrated 68% interval should contain roughly 9.5 of 14. Containing all of them means the bands are too wide — the intervals are roughly ±0.95pp around a forecast whose RMSE is 0.3pp, about three times wider than the errors warrant. The model is under-confident here, and this figure was previously presented as a validation win. It is reported now as what it is: evidence of over-dispersion.

Neither reading is strong evidence either way. The fourteen points are seven quarters × two variables from a single forecast origin, so they are heavily correlated across horizon and across variables; the effective number of independent observations is perhaps two or three. Coverage cannot be established from one origin in either direction. The rolling-origin machinery could produce a genuine coverage table across 49 origins and eight horizons; until it does, this page makes no coverage claim.

Rolling-origin audit. A stricter expanding-window test now evaluates 49 historical origins without future-data leakage. Against a no-change benchmark, CPI-level relative RMSE is 0.63 at one quarter and 0.67 at eight quarters (lower is better), while UK GDP sits at 1.06. But a no-change benchmark is too weak for a trending price level: against a random walk with drift the CPI figures become 0.83 and 1.03, and the GDP gap is not statistically significant at any horizon. See section 08 for the full benchmark comparison. The defensible forecasting claim is Bank Rate, not inflation, and the single frozen-edge result must not be read as broad GDP forecast superiority.

boe-svar: out-of-sample forecast fan from the frozen 2024Q2 edge vs ONS outturns Two-panel fan chart. Left panel: year-on-year UK GDP growth; right panel: year-on-year UK CPI inflation. Each shows the posterior median forecast from the frozen 2024Q2 data edge as a line, the 68 per cent credible band as a shaded region over thirteen quarters 2024Q3 to 2027Q3, and ONS outturns for the seven evaluated quarters 2024Q3 to 2026Q1 as dots. All fourteen outturn dots fall inside the 68 per cent band; RMSE 0.32 percentage points for both variables. GDP medians run 1.3, 1.8, 1.3, 0.9, 0.8, 0.9, 0.9 per cent over the evaluated quarters against outturns of 1.0, 1.5, 1.3, 1.4, 1.3, 1.0, 0.9; CPI medians 2.3, 2.7, 2.8, 3.2, 3.2, 3.2, 3.2 against outturns of 2.0, 2.5, 2.8, 3.5, 3.8, 3.4, 3.1. Medians and 68 per cent bands from papers/boe-svar/figures/figure_numbers.json (forecast_table, entries [median, lo68, hi68]); ONS outturns from papers/boe-svar/figures/make_figures.py. Coordinates: GDP panel maps value v to y = 292 - (v + 1) * 59 for the -1 to 3 per cent axis; CPI panel y = 292 - (v - 1) * 59 for the 1 to 5 per cent axis; quarter i of 13 maps to x = 58 + i * 25.667 (GDP) or 416 + i * 25.667 (CPI). median forecast + 68% band ONS outturn frozen 2024Q2 edge at left of each panel -1% 0 +1% +2% +3% 1% 2% 3% 4% 5% GDP growth (YoY, %) CPI inflation (YoY, %) 24Q3 26Q1 27Q3 24Q3 26Q1 27Q3
All fourteen outturns (seven quarters × two variables, 2024Q3–2026Q1) fall inside the 68% credible band, where a calibrated band would contain about 9.5 — the bands are too wide, not the forecast unusually good. RMSE 0.3pp for both. Median forecast and 68% band from the frozen 2024Q2 data edge. Computed from papers/boe-svar/figures/figure_numbers.json (medians and 68% bands) and the ONS outturn series in papers/boe-svar/figures/make_figures.py.
Why there is no official yardstick for those forecast errors

No official counterpart exists for those RMSEs. The authors' companion paper (Staff Working Paper No. 1,165, January 2026) is a methods paper on a different, four-variable SVAR; it publishes no root-mean-squared errors, coverage rates or benchmark comparisons, so the forecast numbers here stand without an official yardstick and are reported as such.

04 — frb-us

Against the Fed's own solver — and against the Fed's own noise floor.

The baseline reproduces the Fed's LONGBASE database across 284 variables and 20 quarters with a maximum error of 5.6×10−17 — machine precision, against a 10−8 CI gate.

Shock responses match pyfrbus 1.0.0 to 6.0×10−9 and 1.1.1 to 1.4×10−8. The Fed's two releases differ by 1.3×10−8, so our differences sit at the reference solver's own numerical noise.

frb-us: residuals against the Fed’s pyfrbus, log scale Horizontal bar chart on a base-10 logarithmic axis of maximum absolute residuals; shorter is closer. ours vs LONGBASE (tracking invariant): 5.6e-17; pyfrbus 1.1.1 vs LONGBASE: 1.1e-08; ours vs pyfrbus 1.0.0 (shock): 6.0e-09; ours vs pyfrbus 1.1.1 (shock): 1.4e-08; pyfrbus 1.1.1 vs 1.0.0 — the Fed’s own two releases: 1.3e-08. The framing comparison is the last row: the Federal Reserve's own two pyfrbus releases disagree with each other by as much as this implementation disagrees with either, so our agreement sits at the scale of the reference implementation's own numerical noise rather than at a chosen tolerance. 1e-18 1e-16 1e-14 1e-12 1e-10 1e-8 ours vs LONGBASE (tracking invariant) 5.6e-17 pyfrbus 1.1.1 vs LONGBASE 1.1e-08 ours vs pyfrbus 1.0.0 (shock) 6.0e-09 ours vs pyfrbus 1.1.1 (shock) 1.4e-08 pyfrbus 1.1.1 vs 1.0.0 — the Fed’s own two releases 1.3e-08
Maximum absolute residuals, log scale — lower is closer. The bottom bar is the framing one: the Fed's own two releases disagree by as much as we disagree with either. On the tracking invariant, pyfrbus 1.1.1 reproduces LONGBASE to 1.1×10−8 against this implementation's 5.6×10−17. Source: papers/frb-us validation tables.

Multipliers. Government-purchases and personal-tax shocks of 1% of GDP land inside every published range they are compared against — the last three rows below.

frb-us against the Fed's pyfrbus and published multiplier ranges
QuantityOursPublishedReading
Tracking invariant, max abs. error5.6×10−17CI gate <10−8; pyfrbus 1.1.1 gives 1.1×10−8machine precision
Shock responses vs pyfrbus 1.0.06.0×10−9at the reference's noise floor
Shock responses vs pyfrbus 1.1.11.4×10−81.1.1 vs 1.0.0: 1.3×10−8the Fed's releases differ as much as we do
Gov. purchases multiplier, yr 1, inertial Taylor rule0.720.7–1.0 (Coenen et al. 2012)inside range
Gov. purchases multiplier, yr 2, fixed funds rate0.99"roughly one" pegged; 1.1–1.2 with accommodationmatches the Board's characterisation
Personal tax cut multiplier, yrs 1–20.22 → 0.320.2–0.4; CBO central ≈0.3inside both ranges

Uncertainty engine. The implementation now supports seeded stochastic simulations that jointly resample the official stochastic-equation residual vector, preserving contemporaneous dependence and reporting failed replications. Model and LONGBASE packages are also verified against separate SHA-256 provenance gates. Published probability results still require a reviewed simulation design; model-consistent expectations remain unavailable.

Two scope caveats: expectations path, and the 2014 FEDS Note comparison

Only the VAR-expectations path was exercised in the version comparison; model-consistent expectations are not implemented here, so forward-guidance-style exercises are out of scope. And published 100bp funds-rate responses are read off the 2014 FEDS Note charts, so those comparisons are approximate by nature — our output-gap trough of −0.50pp against the Note's ≈−0.4pp is within vintage and shock-design differences, not an exact match.

05 — pe-microsim

Against implemented legislation.

This model does not forecast. It deterministically applies encoded UK and US tax-benefit law to a specified household, so its reference is legislation and official guidance. Coverage depends on the upstream rules and tests.

Population estimates also depend on survey data and calibration. The UK run uses enhanced_frs_2023_24; in the 1pp basic-rate check, it is 6.4% below HMRC in year one. Survey versus administrative data, policy assumptions, and HMRC behavioural elasticities explain that separate aggregate gap.

06 — psl-og

No ground truth. Calibration targets instead.

No published replication target exists. The table compares calibration targets with official UK aggregates. Imposed matches are anchors; structural counterfactuals remain conditional on the model's auditable assumptions and code.
psl-og calibration targets against official UK aggregates
QuantityOfficial UK valueSourceModel treatmentDeviation
Government debt-to-GDP95.1% (May 2026); OBR March 2026 forecast 94.4%ONS public sector financesimposed via closure target−0.1pp vs latest ONS (target 95.0%)
Household saving ratio8.9% (2026Q1, down from 9.6%)ONS quarterly sector accountstargeted through β0 by construction
Potential growth≈1.1% / yrOBR EFOimposed, gy = 0.011≈0
Depreciation / capital≈6–7% (CFC / net stock)ONS capital stocks bulletinimposed, δ = 0.065within range
Labour share of income≈0.59–0.60, rising in 2024ONS labour-share series / Blue Book 2025imposed, γ = 0.35 (OG-Core default)≈5pp — the capital share is not yet re-anchored to ONS factor shares
Net capital stock£5.6tn (2024) whole-economy net stock; the model's K maps to the narrower business-capital conceptONS capital stocks bulletinemergent — checked, not imposedconcept-adjusted; no published reconciliation

A zero deviation on an imposed quantity is not independent evidence.

What would count as over-identifying evidence, and the dating and structural caveats

The genuinely over-identifying checks are the ones the steady state must deliver: the capital–output ratio and the interest rate consistent with the calibrated (γ, δ, β) triple. These emerge in an economically sensible range in the deployed configuration, but no systematic published reconciliation of achieved-versus-target moments exists for OG-UK 0.3.2, and the working paper declines to manufacture one. Producing that reconciliation from logged solves is the single most valuable next step for this member.

Two dating caveats: the official column reports figures current at mid-2026 while the deployed calibration was frozen against slightly earlier vintages, and the saving ratio is volatile quarter to quarter, so β is matched to a smoothed level rather than the latest print. Two structural caveats carried over from OG-Core: the ability types and lifetime earnings profiles are still the US tax-microdata estimates, and UK re-estimation is an open task.

07 — side by side

The same questions, answered by other shops.

Published institutional answers sit beside ours; each row states how closely the definitions align.

Published results from other institutions beside ours, with comparability stated per row
QuestionOursTheirsHow comparable
Long-run effect of a 1pp basic-rate rise (psl-og) Not yet produced — the OG-UK side of this comparison is open validation work OBR Working Paper 22, Table 5.1: GDP/person −0.1%, labour supply −0.2%, productivity +0.1% Future benchmark — same scenario and model family; WP 22's figures are the target our run will be judged against, not a match we claim today
Cost of cutting the basic rate 1p, 2026–27 (obr-macro + pe) £6.46bn static HMRC ready reckoner (June 2025): £6.9bn, rising to £8.2bn by 2028–29 Direct, −6.4% in year one — HMRC's figure embeds behavioural response and the OBR's March 2025 forecast
Government-purchases multiplier, year one (frb-us) 0.72 (inertial Taylor rule); 0.99 in year two with the funds rate fixed Ramey (JEP 2019): the literature sits "in a surprisingly narrow range of 0.6 to 1" Direct — both estimates fall inside the survey range
Income-tax multiplier (frb-us, US) 0.22 → 0.32 over years 1–2 OBR's published UK assumption: 0.3 for income tax and NICs Loose — different country and model class; the coincidence is noted, not claimed
Forecast error, 2024Q3 onward (boe-svar) RMSE 0.32pp on quarterly GDP growth and CPI inflation, data frozen at 2024Q2; the 2025 CPI hump missed by up to 0.6pp OBR forecast evaluation (July 2025): one-year GDP growth underestimated by 0.4pp on average since 2010 (external forecasters 0.5pp). Bank of England MPR (Aug 2024): modal CPI for late 2025 around 2.4%, against an outturn peak of 3.8% — a miss of ~1.4pp on the same episode Adjusted — RMSE and mean error are different statistics, and the Bank's modal path is conditioned on market rates; the vintages, at least, are matched
Comparators that could not be verified against a fetchable source — CBO's FRB/US multiplier tables, IFS decile charts, per-band Scottish ready-reckoner figures — are omitted rather than quoted from memory.
08 — forecast accuracy

How do we compare to other forecasters?

Three different statistics answer three different questions, and they are not averaged, ranked, or plotted on one axis: our SVAR's pseudo-out-of-sample accuracy against a naive benchmark (computed here, quarterly), the OBR's and external forecasters' own published real-time errors (cited, annual), and our OBR emulator's tracking of the official forecast (replication accuracy, not forecasting accuracy).

boe-svar against a random walk, computed from our own runs

An expanding-window pseudo-out-of-sample exercise re-fits the BVAR at 49 quarterly origins (2012Q1–2024Q1, data sample 1992Q1–2026Q1) and scores forecasts one to eight quarters ahead against three naive benchmarks. A ratio below 1.0 means the model beats the benchmark.

The benchmark decides the answer. A no-change random walk on a trending log level forfeits the whole trend as forecast error, so beating it on a price index is close to uninformative. Against a random walk with drift — the textbook naive for a trending series — the CPI result largely evaporates: 0.63 becomes 0.83 at one quarter and 0.67 becomes 1.03 at eight, where the model is no longer ahead at all. The pattern in the original numbers, wins only on the two trending price series and ties or losses on the four series where a random walk is a genuinely hard benchmark, was an artefact of the benchmark rather than a finding about the model.

What survives is Bank Rate. It is the one series here that does not trend, so no-change is the right naive for it, and it is the one series that improves under the harder benchmark: 0.79 against drift at one quarter (Diebold–Mariano p = 0.018) and 0.85 at eight. On the evidence assembled here that, not inflation, is the model's defensible forecasting claim.

UK GDP is not beaten by the benchmark either. Its ratio of 1.06–1.09 is not statistically distinguishable from a random walk at any horizon (p = 0.38 to 0.67), and excluding the six origins whose target quarter falls in 2020Q1–2021Q2 it becomes 0.77 at one quarter. Under squared loss the 2020Q2 collapse dominates a 49-origin average. Both the full-sample and the excluding-Covid figures are published; neither is the preferred number.

The AR(1) comparison is reported in the source file but is not usable at long horizons: 94.5% of its eight-step squared error for UK GDP comes from a single origin, because an AR(1) extrapolates the Covid collapse. The apparent 3:1 win against it is arithmetic, not evidence. Every benchmark now carries a worst_origin_mse_share diagnostic for exactly this reason.

boe-svar forecast skill against a random walk with drift, all eight variables Dot matrix of forecast error relative to a random walk with drift, eight variables by eight quarterly horizons, from 49 expanding-window origins spanning 2012Q1 to 2024Q1. A ratio below 1.0 means the model beats the benchmark; filled dots mark differences significant at 5 per cent by a Diebold-Mariano test, hollow dots differences that are not statistically distinguishable. Bank Rate is the only variable that beats the benchmark at every horizon, ranging 0.79 to 0.89, and its one-quarter advantage is significant (p = 0.018). UK CPI starts at 0.83 at one quarter, significant at p = 0.031, but drifts to 1.03 by eight quarters, losing its advantage. World CPI follows the same pattern, 0.94 rising to 1.07. Oil price and CPI energy sit near parity throughout. UK real GDP (1.06 to 1.12) and world GDP (1.07 to 1.10) stay just above parity, but none of those differences is statistically significant. The exchange rate is worst and worsens monotonically from 1.03 to 1.33. Against this harder benchmark only 4 of eight variables beat naive at one quarter and 2 at eight. RMSE ÷ drifting-random-walk RMSE · horizons 1–8 quarters filled = difference significant at 5% (Diebold–Mariano); hollow = not distinguishable 0.8 0.9 1.1 1.2 1.3 1.0 ← model better benchmark better → Bank Rate 0.85 UK CPI 1.03 World CPI 1.07 Oil price 0.98 CPI energy 1.11 UK real GDP 1.12 World GDP 1.10 Exchange rate 1.33
Expanding-window pseudo-out-of-sample error ratios against a random walk with drift, on the level of each series. A driftless walk is too weak a benchmark for a trending series, so this is the fair comparison; the no-change ratios are in the table below. 49 forecast origins, 2012Q1–2024Q1 (the window includes the Covid quarters); estimation sample from 1992Q1, data through 2026Q1; 4 lags. Estimation uses final revised data, so this is pseudo- rather than real-time out-of-sample. No single origin contributes more than 35% of any drift-benchmark squared error. Source: papers/boe-svar/figures/rolling_evaluation.json.
How many of eight variables beat a drifting random walk, by horizon Stacked column chart. Of the model's eight forecast variables, how many have lower root mean squared error than a random walk with drift, at horizons one to eight quarters. The counts are 4, 4, 3, 2, 3, 3, 2, 2 respectively. Each column also separates wins whose difference is statistically significant at 5 per cent by a Diebold-Mariano test from wins that are not: significant wins number 2, 1, 1, 1, 1, 0, 0, 0. From horizon six onward no variable beats the benchmark by a statistically significant margin. At one quarter the two significant wins are Bank Rate and UK CPI; beyond four quarters only Bank Rate remains ahead at all. Of 8 forecast variables, how many beat a random walk with drift 2 4 6 8 0 4 h=1 4 h=2 3 h=3 2 h=4 3 h=5 3 h=6 2 h=7 2 h=8 beats it, and the difference is significant beats it, not significantly does not beat it
Wins are counted at a ratio strictly below 1.0, which is a hard cut: a variable at 0.99 counts as a win and one at 1.01 does not, even though neither is distinguishable from the benchmark. That is why the significance split matters more than the count. Same 49 origins and benchmark as the chart above. Source: papers/boe-svar/figures/rolling_evaluation.json.
boe-svar forecast error relative to two naive benchmarks (ratio; below 1.0 beats the benchmark). Bold marks a win that is statistically significant; * marks any significant difference. All eight variables are shown — the four previously omitted were among the weakest.
variableh=1h=2h=4h=8h=1 ex-Covid
(vs no-change)
vs driftvs no-changevs driftvs no-changevs driftvs no-changevs driftvs no-change
Bank Rate0.79*0.880.85*0.960.89*1.020.851.030.86
UK CPI (level)0.83*0.630.850.620.940.661.030.670.62
World CPI (level)0.940.730.950.681.000.691.070.670.73
Oil price1.011.031.021.051.011.060.981.081.05
CPI energy0.980.980.970.981.031.041.11*1.130.97
UK real GDP (level)1.061.061.091.081.101.091.121.060.77
World GDP (level)1.071.041.101.031.110.931.100.710.51
Exchange rate1.031.041.091.111.191.221.331.411.05

Bank Rate crosses parity beyond h=3 (source file has all eight variables and horizons). These are quarterly, level-basis, pseudo-out-of-sample ratios computed with hindsight-final data — they are not comparable with the real-time annual-growth errors the OBR publishes, which is why they get separate tables.

What the OBR and external forecasters report about themselves

The OBR's Forecast Evaluation Report (July 2025) publishes its own real-time errors on annual growth rates since 2010, beside the median of external forecasters compiled by HM Treasury. Cited, not recomputed; the statistic is the median absolute error on annual rates, so it cannot be compared with the quarterly RMSE ratios above.

Median absolute forecast errors reported in the OBR Forecast Evaluation Report, July 2025 (annual rates, percentage points, forecasts since 2010)
variable · horizonOBRexternal median
Real GDP growth · 1 year ahead0.6pp0.6pp
Real GDP growth · 2 years ahead0.4pp0.4pp
CPI inflation · 1 year ahead0.3pp0.3pp
CPI inflation · 2 years ahead0.9pp0.9pp

Our own frozen-edge experiment is the nearest in-house analogue: the boe-svar RMSE of 0.32pp on quarterly year-on-year GDP growth and CPI inflation over seven quarters from a 2024Q2 data edge (section 03 above). Different statistic (RMSE vs median absolute error), different frequency (quarterly vs annual), one origin against many — so it sits beside the OBR's numbers, not against them. For the Bank of England, the single matched-vintage episode already cited in section 07 — the August 2024 MPR modal CPI path of about 2.4% for late 2025 against an outturn peak of 3.8% — is the only comparison we can source precisely; the Bank does not publish an RMSE-by-horizon table in the MPR.

Our OBR emulator: replication accuracy, not forecast accuracy

The obr-macro numbers in section 02 measure how closely the anchored emulator reproduces the OBR's published March 2026 path (MAPE 0.15% GDP, 0.25% consumption) — that is replication accuracy. It is not an independent forecast: de-seeded and stripped of add-factors, the same model drifts 5.75% below the EFO path over twelve quarters. Any forecast-accuracy credit for obr-macro therefore belongs to the OBR's own forecast, whose errors are the FER numbers above.

Omitted for lack of a defensible source: a NIESR or consensus RMSE table by horizon (NIESR does not publish one in a fetchable form we could verify), a Bank of England RMSE-by-horizon table (not published in the MPR), and an AR(1) benchmark for the SVAR exercise (our stored evaluation includes only the random walk). Each would need a new computation or source before it can appear here.