model 03 — Bank of England structural VAR · boe-svar · UK · hosted

Explain UK growth and inflation.

Decompose UK GDP and inflation into six structural shocks, forecast with credible bands, and explain revisions between quarters.

how far to trust it

Replicates the paper; forecast evidence remains limited.

Two independent checks: replication against the published paper, and a forecast study frozen at the 2024Q2 data edge and scored against the ONS outturns that arrived since.

Check by check against Brignone & Piffer (2025); the forecast evidence is below
checkthis replicationthe paper
Global shocks' share of UK GDP variance (4 quarters) 37.4% [23.4, 52.2] 68% ~40%
Global shocks' share of UK CPI variance (4 quarters) 42.3% [25.5, 60.3] 68% — about 8pp short ~50%
Identified shocks' total share, UK GDP / UK CPI 76.3% / 78.9% “around 80%”
IRFs, FEVDs, shocks, historical decompositions Figures 2–6 replicate qualitative match
Forecast-revision adding-up identity exact per draw (max abs error 5.23e-12) holds by construction — a regression check, not an agreement
out-of-sample RMSE 0.32pp GDP growth and CPI, 7 quarters from a frozen edge
68% band coverage 14/14 over-dispersion, not a pass — a correct 68% band should hold about 9.5
vs a drifting random walk 0.36min q no variable beats it at any horizon once the 64 tests are adjusted together

The model was right that the 2025 inflation hump would mean-revert and wrong about its peak by up to 0.6pp — outturns of 3.5 and 3.8% against medians of 3.2%, inside the 68% band but on its upper half. It never saw the Ofgem cap increases or the April 2025 administered price rises. Full quarterly scorecard in the working paper.

This is an independent replication built from Bank of England Macro Technical Paper No. 3 and public data. It is not produced, maintained, or endorsed by the Bank of England, and its results should not be presented as Bank of England estimates.
Gated in CI: the replication invariants and the forecast archive's append-only rule. Computed once and refreshed by hand: the 49-origin rolling evaluation and the coverage study. The sign and zero restrictions are imposed by the identification scheme, so that gate protects the software and cannot be evidence about the economy.
evidence — replication

Against a Bank of England paper, and against outturns.

On the paper's own definition — the posterior mean of the per-draw group share of total four-quarter-ahead variance, which leaves about a fifth unexplained — identified global shocks explain 37.4% of UK GDP forecast-error variance against Brignone & Piffer's ~40%, and 42.3% of UK CPI against ~50%.

boe-svar: global-shock FEVD shares, ours vs Brignone & Piffer (2025) Grouped bar chart. Share of UK forecast-error variance attributed to identified global shocks (world demand, energy and supply) four quarters ahead, as a share of total variance — the statistic the paper's Figure 4 plots. For GDP, our production artifact gives 37.4% against the paper's 40.0%, 2.6 points short; for CPI, 42.3% against 50.0%, 7.7 points short. Neither gap is resolvable: on the 600-draw check configuration the 68 per cent posterior band on the same share runs 25 to 56 per cent for GDP and 22 to 47 per cent for CPI, both far wider than the shortfall, and the paper's values are read off its figure and approximate. This comparison and the fan chart elsewhere on this page are different runs: the fan chart is the frozen 2024Q2-edge run of 3,000 draws at seed 20250717. 0% 20% 40% 60% 37.4 ours 40.0 paper UK GDP, 1-yr horizon 42.3 ours 50.0 paper UK CPI, 1-yr horizon
Global-shock (world demand + energy + supply) FEVD shares four quarters ahead: the production run against the paper's approximate published values. No committed artifact records that run's draw count, so none is claimed. Source: papers/boe-svar/figures/comparison_numbers.json.
How noisy that comparison is

The 68% posterior band is about ±14pp wide, so the 8pp CPI shortfall sits inside a single interval: a real gap on the paper's definition, reported as one, but not resolvable at this sample size.

The CI gate uses a cheap ~50-draw configuration whose global GDP share moves about ±6.5pp across sampling seeds, checked against a wide [30, 60]% band. It catches gross regressions; it is not a second opinion on the production run.

Out-of-sample. Seven quarters ahead from the frozen 2024Q2 edge, RMSE is 0.32pp for both GDP growth and CPI inflation.

Fourteen of fourteen outturns inside the 68% band is over-dispersion, not a pass — a correct interval should contain about 9.5 of 14. One frozen origin yields two or three effectively independent observations, so the reading to weight is the 49-origin coverage below, which finds the opposite problem.

boe-svar: out-of-sample forecast fan from the frozen 2024Q2 edge vs ONS outturns Two-panel fan chart. Left panel: year-on-year UK GDP growth; right panel: year-on-year UK CPI inflation. Each shows the posterior median forecast from the frozen 2024Q2 data edge as a line, the 68 per cent credible band as a shaded region over thirteen quarters 2024Q3 to 2027Q3, and ONS outturns for the seven evaluated quarters 2024Q3 to 2026Q1 as dots. All fourteen outturn dots fall inside the 68 per cent band; RMSE 0.32 percentage points on GDP and 0.32 on CPI. GDP medians run 1.3, 1.8, 1.3, 0.9, 0.8, 0.9, 0.9 per cent over the evaluated quarters against outturns of 1.0, 1.5, 1.3, 1.4, 1.3, 1.0, 0.9; CPI medians 2.3, 2.7, 2.8, 3.2, 3.2, 3.2, 3.2 against outturns of 2.0, 2.5, 2.8, 3.5, 3.8, 3.4, 3.1. A dashed rule in each panel marks 2026Q1, the last quarter with an outturn: the six quarters to its right are forecast with nothing yet to check them against, so the fit shown covers seven of the thirteen quarters plotted. This run is 3,000 posterior draws at seed 20250717, 217 of which were accepted by the sign restrictions; it is a different run from the FEVD comparison elsewhere on this page, which comes from the production artifact. Medians and 68 per cent bands from papers/boe-svar/figures/figure_numbers.json (forecast_table, entries [median, lo68, hi68]); ONS outturns from papers/boe-svar/figures/make_figures.py. Coordinates: GDP panel maps value v to y = 292 - (v + 1) * 59 for the -1 to 3 per cent axis; CPI panel y = 292 - (v - 1) * 59 for the 1 to 5 per cent axis; quarter i of 13 maps to x = 58 + i * 25.667 (GDP) or 416 + i * 25.667 (CPI). median forecast + 68% band ONS outturn frozen 2024Q2 edge at left of each panel -1% 0 +1% +2% +3% 1% 2% 3% 4% 5% GDP growth (YoY, %) CPI inflation (YoY, %) 24Q3 26Q1 27Q3 24Q3 26Q1 27Q3 dashed rule = last quarter with an ONS outturn (2026Q1); the 6 quarters to its right are not evaluated
14/14 inside the 68% band — see the over-dispersion caveat above. Frozen 2024Q2 data edge; source: papers/boe-svar/figures/figure_numbers.json.

Rolling-origin audit. Across 49 expanding-window origins, no variable beats a drifting random walk at any horizon once the 64 tests are adjusted together — the benchmark section below. One frozen origin is not evidence of broad forecast superiority.

The svar-unemployment satellite. Unemployment is not a VAR variable: its published outlook comes from a regression of the quarterly change in the ONS rate on year-on-year GDP growth, fitted 1992Q1–2025Q1 with furlough dummies. It beats no-change only at horizons 1–4 (relative RMSE 0.82–0.99, not significant), so the published path is capped there, and its bands carry GDP-forecast uncertainty only. Specification and skill table: unemployment_satellite_validation.json.

boe-svar: empirical band coverage across 49 origins Line chart of empirical interval coverage by forecast horizon, averaged across the eight model variables, against the nominal 68 and 90 percent levels; a vertical bar at each horizon spans the best and worst variable. 68% band, mean coverage at horizons one to eight: 71%, 68%, 64%, 61%, 60%, 60%, 58%, 57%; 90% band, mean coverage at horizons one to eight: 87%, 81%, 81%, 76%, 76%, 73%, 72%, 72%. The 68 per cent band covers 71% at one quarter, just above nominal, and falls below nominal from horizon two onward to 57% at eight; the 90 per cent band under-covers at every horizon, from 87% down to 72%. The mean hides how bad the worst variable is: the lowest cell in the grid is CPI energy at the 68 per cent band, horizon 4, covering 29%. The window includes the Covid quarters and the evaluation model carries no Covid dummies, which depresses coverage. 30% 40% 50% 60% 70% 80% 90% 100% 90% 68% 68% band, mean coverage across variables 90% band vertical bars span the best and worst of the eight variables · lowest cell 29% (CPI energy, 68% band, h4) nominal 68% nominal 90% h1 h2 h3 h4 h5 h6 h7 h8
Empirical interval coverage across 49 expanding-window origins, mean across the eight variables, by horizon. The 68% band over-covers at one quarter (71%) and falls below nominal from two onward; the 90% band under-covers at every horizon. Both worsen with horizon — the opposite of the frozen-origin reading above, and the one to weight. The evaluation model carries no Covid dummies over a window that includes the Covid quarters, which depresses coverage. Source: papers/boe-svar/figures/coverage_evaluation.json.
evidence — forecast accuracy

How does it compare to other forecasters?

Three statistics answer three different questions and are not averaged or ranked: this SVAR's pseudo-out-of-sample accuracy against naive benchmarks (computed here, quarterly), the OBR's and external forecasters' published real-time errors (cited, annual), and the OBR emulator's tracking of the official forecast (replication, not forecasting).

No forecasting win survives the number of tests run. Bank Rate comes closest — 0.79 against a drifting random walk at one quarter, and the smallest of 64 variable×horizon tests at p = 0.018 — but adjusted together (Benjamini–Hochberg) the minimum q is 0.36, so nothing clears a 10% false-discovery rate on any variable at any horizon.

Benchmark and sample diagnostics

boe-svar against a random walk, computed from our own runs

Expanding-window design: 49 quarterly origins (2012Q1–2024Q1), estimation from 1992Q1, data through 2026Q1, horizons 1–8. A ratio below 1.0 beats the benchmark.

The benchmark decides the answer. A no-change walk on a trending log level forfeits the whole trend as error, so beating it on a price index is close to uninformative: UK CPI's 0.63 at one quarter becomes 0.83 against a walk with drift, and 0.67 at eight becomes 1.03. Bank Rate is the one series that does not trend, and the one that improves under the harder benchmark.

Covid dummies matter for GDP, not for inflation. The rolling evaluation omits the six Covid dummies (2020Q1–2021Q2) that every published forecast here carries. On the published specification UK GDP against drift goes 1.06 → 0.99 at one quarter and 1.12 → 0.95 at eight, while CPI and Bank Rate move by less than 0.02. Level with a naive benchmark is still not beating it: UK GDP is statistically inseparable from drift at every horizon (p = 0.33–0.43) in both specifications.

The AR(1) comparison in the source file is not usable at long horizons: 94.5% of its eight-step squared error for UK GDP comes from one origin, because an AR(1) extrapolates the Covid collapse. Every benchmark carries a worst_origin_mse_share diagnostic for that reason.

boe-svar forecast skill against a random walk with drift, all eight variables Dot matrix of forecast error relative to a random walk with drift, eight variables by eight quarterly horizons, from 49 expanding-window origins. Position alone carries better or worse: a dot left of the 1.0 rule means the model beats the benchmark, right of it means the benchmark wins. Filled dots mark differences significant at 5 per cent by a Diebold-Mariano test, hollow dots differences that are not statistically distinguishable; those p-values are pairwise and the grid runs 64 tests, none of which clears a 10 per cent false-discovery rate once adjusted together — the smallest adjusted q-value is 0.36. Bank Rate runs 0.79 at one quarter to 0.85 at 8; UK CPI runs 0.83 at one quarter to 1.03 at 8; World CPI runs 0.94 at one quarter to 1.07 at 8; Oil price runs 1.01 at one quarter to 0.98 at 8; CPI energy runs 0.98 at one quarter to 1.11 at 8; UK real GDP runs 1.06 at one quarter to 1.12 at 8; World GDP runs 1.07 at one quarter to 1.10 at 8; Exchange rate runs 1.03 at one quarter to 1.33 at 8. Bank Rate is the best at the longest horizon and Exchange rate the worst. Against this harder benchmark only four of eight variables beat naive at one quarter and two at 8. The number at the end of each row is that row's ratio at horizon 8. RMSE ÷ drifting-random-walk RMSE · horizons 1–8 quarters filled = significant at 5% pairwise · hollow = not distinguishable large dot and row-end number = h=8 · smallest adjusted q over the 64 cells: 0.36 0.8 0.9 1.1 1.2 1.3 1.0 ← model better benchmark better → Bank Rate 0.85 UK CPI 1.03 World CPI 1.07 Oil price 0.98 CPI energy 1.11 UK real GDP 1.12 World GDP 1.10 Exchange rate 1.33
Error ratios against a random walk with drift, on the level of each series — the fair benchmark for a trending series. Estimated without the Covid dummies the published specification carries; no-change ratios are in the table below. Estimation uses final revised data, so this is pseudo- rather than real-time out-of-sample. The ratios rest on fewer effective observations than 49 origins suggests: in 16 of the 64 cells one origin contributes over 35% of the squared error, peaking at 63% for UK GDP at h=1. Source: papers/boe-svar/figures/rolling_evaluation.json (worst_origin_mse_share.drift).
How many of eight variables beat a drifting random walk, by horizon Column chart, one column per forecast horizon from one to eight quarters. Column height is the number of the model's eight forecast variables whose root mean squared error is below that of a random walk with drift; a dashed reference line across the top marks all eight. The counts are 4, 4, 3, 2, 3, 3, 2, 2, so between 2 and 4 of the 8 beat the benchmark at any horizon and the remaining 4, 4, 5, 6, 5, 5, 6, 6 do not. The shaded lower part of each column is the subset whose difference is significant at 5 per cent on an unadjusted pairwise Diebold-Mariano test: 2, 1, 1, 1, 1, 0, 0, 0. Those p-values are pairwise and the grid runs 64 tests; under a Benjamini-Hochberg adjustment none of them reaches a 10 per cent false-discovery rate and the smallest adjusted q-value is 0.36, so on the adjusted reading the significant-win count is zero at every horizon. At horizons 7 and 8 one variable is significantly worse than the benchmark. Of 8 forecast variables, how many beat a random walk with drift column height = variables with lower RMSE than the benchmark · the rest do not beat it 2 4 6 0 8 all 8 variables 4 h=1 2 4 h=2 1 3 h=3 1 2 h=4 1 3 h=5 1 3 h=6 0 2 h=7 0 2 h=8 0 p<0.05 beats it, p<0.05 unadjusted beats it, not distinguishable Adjusted together, none of the 64 tests clears a 10% false-discovery rate (smallest q = 0.36).
A win is a ratio strictly below 1.0 — a hard cut that counts 0.99 and rejects 1.01 though neither is distinguishable from the benchmark, so the significance split matters more than the count. That significance is pairwise: under a Benjamini–Hochberg adjustment the significant-win count is zero at every horizon. Same origins and benchmark as above. Source: papers/boe-svar/figures/rolling_evaluation.json.
Forecast error relative to two naive benchmarks (below 1.0 beats it). Bold marks a significant win, * any significant difference — unadjusted, so read a star as "worth a second look". Ratios omit the Covid dummies the published specification carries (see the diagnostics above). The ex-Covid column is against no-change — UK GDP's 0.77 there is 1.21 against drift — and neither it nor the full-sample figure is preferred.
variableh=1h=2h=4h=8h=1 ex-Covid
(vs no-change)
vs driftvs no-changevs driftvs no-changevs driftvs no-changevs driftvs no-change
Bank Rate0.79*0.880.85*0.960.89*1.020.851.030.86
UK CPI (level)0.83*0.630.850.620.940.661.030.670.62
World CPI (level)0.940.730.950.681.000.691.070.670.73
Oil price1.011.031.021.051.011.060.981.081.05
CPI energy0.980.980.970.981.031.041.11*1.130.97
UK real GDP (level)1.061.061.091.081.101.091.121.060.77
World GDP (level)1.071.041.101.031.110.931.100.710.51
Exchange rate1.031.041.091.111.191.221.331.411.05

These are quarterly, level-basis ratios computed with hindsight-final data. They are not comparable with the real-time annual-growth errors the OBR publishes, which is why those get a separate table.

What the OBR and external forecasters report about themselves

Cited, not recomputed: the OBR's Forecast Evaluation Report (July 2025) reports its own real-time errors on annual growth rates since 2010, beside the HM Treasury median of external forecasters.

Median absolute forecast errors reported in the OBR Forecast Evaluation Report, July 2025 (annual rates, percentage points, forecasts since 2010)
variable · horizonOBRexternal median
Real GDP growth · 1 year ahead0.6pp0.6pp
Real GDP growth · 2 years ahead0.4pp0.4pp
CPI inflation · 1 year ahead0.3pp0.3pp
CPI inflation · 2 years ahead0.9pp0.9pp

The nearest in-house analogue is the frozen-edge 0.32pp above — a different statistic, frequency and origin count, so it sits beside these numbers, not against them. For the Bank of England the only sourceable comparison is the matched-vintage episode on the validation index; the MPR publishes no RMSE-by-horizon table.

Omitted for lack of a defensible source: NIESR or consensus RMSE tables by horizon, a Bank of England RMSE-by-horizon table, and an AR(1) benchmark for this exercise. The authors' companion paper (Staff Working Paper No. 1,165) is a methods paper on a different SVAR and publishes no errors, so the frozen-edge RMSEs stand without an official yardstick.
limits

Where this departs from the Bank's model.

Known limits of the structural VAR replication
limitdetail
Frozen data edge Estimation to 2025Q1, conditioning to 2026Q1 — the extended sample, not the paper's 1992Q1–2023Q2 window. Results shift with revisions.
Proxied world aggregates The Bank's UK-trade-weighted world GDP and CPI are unpublished, so they are rebuilt as chain-weighted US + euro-area + Japan + China with time-varying UK trade weights.
One ranking does not replicate The paper's prose ranks UK monetary policy as the largest domestic contributor to CPI variance; here UK supply (13.9%) and UK demand (11.4%) both exceed it (8.7%). The paper publishes no FEVD table, so treat this as unresolved rather than a defect; the proxy aggregates are a candidate cause.
Assumed lag length and simplified pandemic prior p = 4 is assumed, not selected; the pandemic is handled by exogenous Covid dummies.
Sign-restriction critiques apply Pointwise medians mix structural models (Fry–Pagan); the Haar prior over rotations is informative about impulse responses (Baumeister–Hamilton). These hit the Bank's own outputs identically.

every deviation is enumerated in the repository's docs/methodology.md.