model 02 — Bank of England structural VAR · boe-svar · UK · hosted

Explain UK growth and inflation.

Decompose UK GDP and inflation into six structural shocks, forecast with credible bands, and explain revisions between quarters.

how far to trust it

Replicates the paper; forecast evidence remains limited.

Two independent checks: replication against the published paper, and — rarer — a 2024Q2-frozen forecast study scored against ONS outturns that arrived since, a genuine out-of-sample test that exists because the data edge is frozen.

This replication against Brignone & Piffer (2025), check by check
checkthis replicationthe paper
Global shocks' share of UK GDP variance (1 year) 42.1% ~40%
Global shocks' share of UK CPI variance (1 year) 49.5% ~50%
IRFs, FEVDs, shocks, historical decompositions Figures 2–6 replicate qualitative match
Out-of-sample forecast error, 2024Q3–2026Q1 0.32pp RMSE for both GDP growth and CPI inflation no published counterpart
68% band coverage over the same seven quarters 14 of 14 outturns inside
Forecast-revision adding-up identity exact per draw (max abs error 5.23e-12) holds by construction
out-of-sample RMSE 0.32pp GDP growth and CPI, 7 quarters from a frozen edge
68% band coverage 14/14 no outturn escaped even the 68% band — though the 49-origin rolling evaluation finds the bands under-cover on average
tests comprehensive estimation, identification, and analysis coverage

The honest reading: the model was right that the 2025 inflation hump would mean-revert and wrong about its peak, by up to 0.6 points — outturns of 3.5 and 3.8% against medians of 3.2%, inside the 68% band but on its upper half. It never saw the Ofgem cap increases, food inflation, or the April 2025 administered price rises. Charts and source data below; full quarterly scorecard in the working paper.

This is an independent replication built from Bank of England Macro Technical Paper No. 3 and public data. It is not produced, maintained, or endorsed by the Bank of England, and its results should not be presented as Bank of England estimates.
evidence — replication

Against a Bank of England paper, and against outturns.

Against Brignone & Piffer (2025), the one-year global-shock shares of UK forecast-error variance land within about a point of the paper's benchmarks.

boe-svar: global-shock FEVD shares, ours vs Brignone & Piffer (2025) Grouped bar chart. Share of UK forecast-error variance attributed to identified global shocks (world demand, energy and supply) at the one-year horizon. For GDP, our 10,000-draw production run gives 42.1% against the paper's 40.0%; for CPI, 49.5% against 50.0%. Both deviations are a percentage point or less. The paper's values are approximate. 0% 20% 40% 60% 42.1 ours 40.0 paper UK GDP, 1-yr horizon 49.5 ours 50.0 paper UK CPI, 1-yr horizon
Global shocks = world demand + energy + supply. Ours = 10,000-draw production run; paper values approximate. Global-shock FEVD shares at the one-year horizon: the production run against the paper's published values. Source: papers/boe-svar/figures/comparison_numbers.json and the paper's validation table.
Why the fast CI configuration lands 6–10 points off

The deliberately cheap unweighted CI configuration lands at 49.6% and 43.8% — 6–10 points off, in opposite directions for the two variables. That gap is itself informative rather than embarrassing: the importance weights correct the Arias et al. (2018) zero-restriction sampler towards the uniform-over-rotations posterior, and applying them moves both shares substantially towards the published values, though on the 2026Q1 vintage GDP still lands 2.1 points high. Both configurations sit inside the wide [30, 60]% acceptance band that CI enforces, which is chosen to catch gross regressions without being flaky to Monte-Carlo variation.

boe-svar global-shock FEVD shares against Brignone & Piffer (2025)
Global share, 1-yr FEVDPaperCI config (fast)Production (10k draws)Deviation, CI / production
UK GDP~40%49.6%42.1%+9.6 / +2.1 pp
UK CPI~50%43.8%49.5%−6.2 / −0.5 pp

Out-of-sample. Forecasting seven quarters ahead from a frozen 2024Q2 edge, RMSE (root-mean-square error) is 0.32pp for both GDP growth and CPI inflation (what the median got right and wrong is in the honest reading above).

Fourteen of fourteen outturns inside the 68% band is a calibration failure, not a pass — a correct 68% interval should contain roughly 9.5 of 14, and the bands are about three times wider than the errors warrant, so this is evidence of over-dispersion, not the validation win it was once presented as. But one frozen origin yields perhaps two or three effectively independent observations, so neither reading carries much weight; the test to weight is the empirical coverage across all 49 rolling origins and eight horizons below.

boe-svar: out-of-sample forecast fan from the frozen 2024Q2 edge vs ONS outturns Two-panel fan chart. Left panel: year-on-year UK GDP growth; right panel: year-on-year UK CPI inflation. Each shows the posterior median forecast from the frozen 2024Q2 data edge as a line, the 68 per cent credible band as a shaded region over thirteen quarters 2024Q3 to 2027Q3, and ONS outturns for the seven evaluated quarters 2024Q3 to 2026Q1 as dots. All fourteen outturn dots fall inside the 68 per cent band; RMSE 0.32 percentage points for both variables. GDP medians run 1.3, 1.8, 1.3, 0.9, 0.8, 0.9, 0.9 per cent over the evaluated quarters against outturns of 1.0, 1.5, 1.3, 1.4, 1.3, 1.0, 0.9; CPI medians 2.3, 2.7, 2.8, 3.2, 3.2, 3.2, 3.2 against outturns of 2.0, 2.5, 2.8, 3.5, 3.8, 3.4, 3.1. Medians and 68 per cent bands from papers/boe-svar/figures/figure_numbers.json (forecast_table, entries [median, lo68, hi68]); ONS outturns from papers/boe-svar/figures/make_figures.py. Coordinates: GDP panel maps value v to y = 292 - (v + 1) * 59 for the -1 to 3 per cent axis; CPI panel y = 292 - (v - 1) * 59 for the 1 to 5 per cent axis; quarter i of 13 maps to x = 58 + i * 25.667 (GDP) or 416 + i * 25.667 (CPI). median forecast + 68% band ONS outturn frozen 2024Q2 edge at left of each panel -1% 0 +1% +2% +3% 1% 2% 3% 4% 5% GDP growth (YoY, %) CPI inflation (YoY, %) 24Q3 26Q1 27Q3 24Q3 26Q1 27Q3
14/14 inside the 68% band — see the over-dispersion caveat above. Frozen 2024Q2 data edge; source: papers/boe-svar/figures/figure_numbers.json.

Rolling-origin audit. A stricter expanding-window test evaluates 49 historical origins without future-data leakage; the full benchmark treatment is in the forecast-accuracy section below. The short version: Bank Rate is the only defensible forecasting claim, and the single frozen-edge result must not be read as broad forecast superiority.

Unemployment satellite. Unemployment is not a VAR variable; its published outlook is an Okun's-law regression of the quarterly change in the ONS unemployment rate on year-on-year GDP growth, fit 1992Q1–2025Q1 with furlough dummies for 2020Q1–2021Q2 (which otherwise attenuate the coefficient about threefold). In a 73-origin rolling test fed by the VAR's own GDP forecasts it beats no-change at horizons 1–4 (relative RMSE 0.82–0.99, excluding furlough targets, not statistically significant) and is clearly worse beyond, so the published path is capped at four quarters; the bands carry GDP-forecast uncertainty only. Full specification and skill table: okun_validation.json in the model repository.

boe-svar: empirical band coverage across 49 origins Line chart of empirical interval coverage by forecast horizon, averaged across the eight model variables, against the nominal 68 and 90 percent levels. 68% band, horizons one to eight: 71%, 68%, 64%, 61%, 60%, 60%, 58%, 57%; 90% band, horizons one to eight: 87%, 81%, 81%, 76%, 76%, 73%, 72%, 72%. Both bands under-cover, and coverage worsens with horizon. The window includes the Covid quarters and the evaluation model carries no Covid dummies, which depresses coverage. 40% 50% 60% 70% 80% 90% 100% 68% band, mean coverage across variables 90% band nominal 68% nominal 90% h1 h2 h3 h4 h5 h6 h7 h8
Empirical interval coverage across 49 expanding-window origins, mean across the eight model variables, by horizon. Predictive bands combine parameter and shock uncertainty (100 posterior draws × 5 paths per origin), percentile intervals on series levels. Both bands under-cover — the 68% band contains outturns 57–71% of the time and the 90% band 72–87%, worsening with horizon. The window includes the Covid quarters and the evaluation model carries no Covid dummies, which depresses coverage; pseudo- rather than real-time out-of-sample, as above. So the frozen-origin over-dispersion reading and this under-coverage reading disagree — a single origin is simply too little data, and the rolling result is the one to weight. Source: papers/boe-svar/figures/coverage_evaluation.json, generated by make_coverage.py.
Why there is no official yardstick for those forecast errors

No official counterpart exists for those RMSEs. The authors' companion paper (Staff Working Paper No. 1,165, January 2026) is a methods paper on a different, four-variable SVAR; it publishes no root-mean-squared errors, coverage rates or benchmark comparisons, so the forecast numbers here stand without an official yardstick and are reported as such.

evidence — forecast accuracy

How does it compare to other forecasters?

Three different statistics answer three different questions, and they are not averaged, ranked, or plotted on one axis: this SVAR's pseudo-out-of-sample accuracy against a naive benchmark (computed here, quarterly), the OBR's and external forecasters' own published real-time errors (cited, annual), and the OBR emulator's tracking of the official forecast (replication accuracy, not forecasting accuracy).

Bank Rate is the only robust forecasting win. CPI loses its advantage against a drift benchmark; GDP does not beat that benchmark over the full sample, and Covid dominates squared error.

Benchmark and sample diagnostics

boe-svar against a random walk, computed from our own runs

An expanding-window pseudo-out-of-sample exercise re-fits the BVAR at 49 quarterly origins (2012Q1–2024Q1, data sample 1992Q1–2026Q1) and scores forecasts one to eight quarters ahead against three naive benchmarks. A ratio below 1.0 means the model beats the benchmark.

The benchmark decides the answer. A no-change random walk on a trending log level forfeits the whole trend as forecast error, so beating it on a price index is close to uninformative. Against a random walk with drift — the textbook naive for a trending series — the CPI result largely evaporates: 0.63 becomes 0.83 at one quarter and 0.67 becomes 1.03 at eight, where the model is no longer ahead at all. The pattern in the original numbers, wins only on the two trending price series and ties or losses on the four series where a random walk is a genuinely hard benchmark, was an artefact of the benchmark rather than a finding about the model.

What survives is Bank Rate. It is the one series here that does not trend, so no-change is the right naive for it, and it is the one series that improves under the harder benchmark: 0.79 against drift at one quarter (Diebold–Mariano p = 0.018) and 0.85 at eight. On the evidence assembled here that, not inflation, is the model's defensible forecasting claim.

UK GDP is not beaten by the benchmark either. Its ratio of 1.06–1.09 is not statistically distinguishable from a random walk at any horizon (p = 0.38 to 0.67), and excluding the six origins whose target quarter falls in 2020Q1–2021Q2 it becomes 0.77 at one quarter. Under squared loss the 2020Q2 collapse dominates a 49-origin average. Both the full-sample and the excluding-Covid figures are published; neither is the preferred number.

The AR(1) comparison is reported in the source file but is not usable at long horizons: 94.5% of its eight-step squared error for UK GDP comes from a single origin, because an AR(1) extrapolates the Covid collapse. The apparent 3:1 win against it is arithmetic, not evidence. Every benchmark now carries a worst_origin_mse_share diagnostic for exactly this reason.

boe-svar forecast skill against a random walk with drift, all eight variables Dot matrix of forecast error relative to a random walk with drift, eight variables by eight quarterly horizons, from 49 expanding-window origins. A ratio below 1.0 means the model beats the benchmark; filled dots mark differences significant at 5 per cent by a Diebold-Mariano test, hollow dots differences that are not statistically distinguishable. Bank Rate runs 0.79 at one quarter to 0.85 at 8; UK CPI runs 0.83 at one quarter to 1.03 at 8; World CPI runs 0.94 at one quarter to 1.07 at 8; Oil price runs 1.01 at one quarter to 0.98 at 8; CPI energy runs 0.98 at one quarter to 1.11 at 8; UK real GDP runs 1.06 at one quarter to 1.12 at 8; World GDP runs 1.07 at one quarter to 1.10 at 8; Exchange rate runs 1.03 at one quarter to 1.33 at 8. Bank Rate is the best at the longest horizon and Exchange rate the worst. Against this harder benchmark only four of eight variables beat naive at one quarter and two at 8. RMSE ÷ drifting-random-walk RMSE · horizons 1–8 quarters filled = difference significant at 5% (Diebold–Mariano); hollow = not distinguishable 0.8 0.9 1.1 1.2 1.3 1.0 ← model better benchmark better → Bank Rate 0.85 UK CPI 1.03 World CPI 1.07 Oil price 0.98 CPI energy 1.11 UK real GDP 1.12 World GDP 1.10 Exchange rate 1.33
Expanding-window pseudo-out-of-sample error ratios against a random walk with drift, on the level of each series. A driftless walk is too weak a benchmark for a trending series, so this is the fair comparison; the no-change ratios are in the table below. 49 forecast origins, 2012Q1–2024Q1 (the window includes the Covid quarters); estimation sample from 1992Q1, data through 2026Q1; 4 lags. Estimation uses final revised data, so this is pseudo- rather than real-time out-of-sample. No single origin contributes more than 35% of any drift-benchmark squared error. Source: papers/boe-svar/figures/rolling_evaluation.json.
How many of eight variables beat a drifting random walk, by horizon Stacked column chart. Of the model's eight forecast variables, how many have lower root mean squared error than a random walk with drift, at horizons one to eight quarters. The counts are 4, 4, 3, 2, 3, 3, 2, 2 respectively. Each column also separates wins whose difference is statistically significant at 5 per cent by a Diebold-Mariano test from wins that are not: significant wins number 2, 1, 1, 1, 1, 0, 0, 0. From horizon 6 onward no variable beats the benchmark by a statistically significant margin. Of 8 forecast variables, how many beat a random walk with drift 2 4 6 8 0 4 h=1 4 h=2 3 h=3 2 h=4 3 h=5 3 h=6 2 h=7 2 h=8 beats it, and the difference is significant beats it, not significantly does not beat it
Wins are counted at a ratio strictly below 1.0, which is a hard cut: a variable at 0.99 counts as a win and one at 1.01 does not, even though neither is distinguishable from the benchmark. That is why the significance split matters more than the count. Same 49 origins and benchmark as the chart above. Source: papers/boe-svar/figures/rolling_evaluation.json.
boe-svar forecast error relative to two naive benchmarks (ratio; below 1.0 beats the benchmark). Bold marks a win that is statistically significant; * marks any significant difference. All eight variables are shown — the four previously omitted were among the weakest.
variableh=1h=2h=4h=8h=1 ex-Covid
(vs no-change)
vs driftvs no-changevs driftvs no-changevs driftvs no-changevs driftvs no-change
Bank Rate0.79*0.880.85*0.960.89*1.020.851.030.86
UK CPI (level)0.83*0.630.850.620.940.661.030.670.62
World CPI (level)0.940.730.950.681.000.691.070.670.73
Oil price1.011.031.021.051.011.060.981.081.05
CPI energy0.980.980.970.981.031.041.11*1.130.97
UK real GDP (level)1.061.061.091.081.101.091.121.060.77
World GDP (level)1.071.041.101.031.110.931.100.710.51
Exchange rate1.031.041.091.111.191.221.331.411.05

Bank Rate crosses parity beyond h=3 (source file has all eight variables and horizons). These are quarterly, level-basis, pseudo-out-of-sample ratios computed with hindsight-final data — they are not comparable with the real-time annual-growth errors the OBR publishes, which is why they get separate tables.

What the OBR and external forecasters report about themselves

The OBR's Forecast Evaluation Report (July 2025) publishes its own real-time errors on annual growth rates since 2010, beside the median of external forecasters compiled by HM Treasury. Cited, not recomputed; the statistic is the median absolute error on annual rates, so it cannot be compared with the quarterly RMSE ratios above.

Median absolute forecast errors reported in the OBR Forecast Evaluation Report, July 2025 (annual rates, percentage points, forecasts since 2010)
variable · horizonOBRexternal median
Real GDP growth · 1 year ahead0.6pp0.6pp
Real GDP growth · 2 years ahead0.4pp0.4pp
CPI inflation · 1 year ahead0.3pp0.3pp
CPI inflation · 2 years ahead0.9pp0.9pp

The nearest in-house analogue is the frozen-edge experiment above: 0.32pp RMSE on quarterly year-on-year GDP growth and CPI inflation over seven quarters. Different statistic, frequency, and origin count, so it sits beside the OBR's numbers, not against them. For the Bank of England, the only precisely sourceable comparison is the matched-vintage episode on the validation index (August 2024 MPR modal CPI of ~2.4% for late 2025 vs an outturn peak of 3.8%); the MPR publishes no RMSE-by-horizon table.

Omitted for lack of a defensible source: NIESR or consensus RMSE tables by horizon, a Bank of England RMSE-by-horizon table, and an AR(1) benchmark for the SVAR exercise. Each needs a new computation or source before it can appear here.
limits

Where this departs from the Bank's model.

Known limits of the structural VAR replication
limitdetail
Frozen data edge Coefficient estimation to 2025Q1; conditioning data to 2026Q1 (refreshed July 2026). Results shift with data revisions. The replication claims are scored on the extended 1992Q1–2025Q1 sample, not the paper's original 1992Q1–2023Q2 window.
Proxied world aggregates The Bank's internal UK-trade-weighted world GDP and CPI are unpublished, so they are rebuilt as chain-weighted US + euro-area + Japan + China aggregates with time-varying UK trade weights.
One ranking does not replicate The paper ranks UK monetary policy as the largest domestic contributor to CPI variance; here UK supply (13.5%) exceeds monetary policy (10.1%). Attributed to the proxy world aggregates and the smaller accepted sample, and documented rather than tuned away.
Assumed lag length and simplified pandemic prior p = 4 is assumed, not selected; the pandemic treatment is simplified to exogenous Covid dummies.
Sign-restriction critiques apply Pointwise medians mix structural models (Fry–Pagan); the Haar prior over rotations is informative about impulse responses (Baumeister–Hamilton). These apply to the Bank's own outputs identically.

every deviation is enumerated in the repository's docs/methodology.md.