model 03 — Bank of England structural VAR · boe-svar · UK · hosted
Explain UK growth and inflation.
Decompose UK GDP and inflation into six structural shocks, forecast with credible bands, and explain revisions between quarters.
Replicates the paper; forecast evidence remains limited.
Two independent checks: replication against the published paper, and a forecast study frozen at the 2024Q2 data edge and scored against the ONS outturns that arrived since.
| check | this replication | the paper |
|---|---|---|
| Global shocks' share of UK GDP variance (4 quarters) | 37.4% [23.4, 52.2] 68% | ~40% |
| Global shocks' share of UK CPI variance (4 quarters) | 42.3% [25.5, 60.3] 68% — about 8pp short | ~50% |
| Identified shocks' total share, UK GDP / UK CPI | 76.3% / 78.9% | “around 80%” |
| IRFs, FEVDs, shocks, historical decompositions | Figures 2–6 replicate | qualitative match |
| Forecast-revision adding-up identity | exact per draw (max abs error 5.23e-12) | holds by construction — a regression check, not an agreement |
The model was right that the 2025 inflation hump would mean-revert and wrong about its peak by up to 0.6pp — outturns of 3.5 and 3.8% against medians of 3.2%, inside the 68% band but on its upper half. It never saw the Ofgem cap increases or the April 2025 administered price rises. Full quarterly scorecard in the working paper.
Against a Bank of England paper, and against outturns.
On the paper's own definition — the posterior mean of the per-draw group share of total four-quarter-ahead variance, which leaves about a fifth unexplained — identified global shocks explain 37.4% of UK GDP forecast-error variance against Brignone & Piffer's ~40%, and 42.3% of UK CPI against ~50%.
papers/boe-svar/figures/comparison_numbers.json.How noisy that comparison is
The 68% posterior band is about ±14pp wide, so the 8pp CPI shortfall sits inside a single interval: a real gap on the paper's definition, reported as one, but not resolvable at this sample size.
The CI gate uses a cheap ~50-draw configuration whose global GDP share moves about ±6.5pp across sampling seeds, checked against a wide [30, 60]% band. It catches gross regressions; it is not a second opinion on the production run.
Out-of-sample. Seven quarters ahead from the frozen 2024Q2 edge, RMSE is 0.32pp for both GDP growth and CPI inflation.
Fourteen of fourteen outturns inside the 68% band is over-dispersion, not a pass — a correct interval should contain about 9.5 of 14. One frozen origin yields two or three effectively independent observations, so the reading to weight is the 49-origin coverage below, which finds the opposite problem.
papers/boe-svar/figures/figure_numbers.json.Rolling-origin audit. Across 49 expanding-window origins, no variable beats a drifting random walk at any horizon once the 64 tests are adjusted together — the benchmark section below. One frozen origin is not evidence of broad forecast superiority.
The svar-unemployment satellite. Unemployment is not a VAR variable: its published outlook comes from a regression of the quarterly change in the ONS rate on year-on-year GDP growth, fitted 1992Q1–2025Q1 with furlough dummies. It beats no-change only at horizons 1–4 (relative RMSE 0.82–0.99, not significant), so the published path is capped there, and its bands carry GDP-forecast uncertainty only. Specification and skill table: unemployment_satellite_validation.json.
papers/boe-svar/figures/coverage_evaluation.json.How does it compare to other forecasters?
Three statistics answer three different questions and are not averaged or ranked: this SVAR's pseudo-out-of-sample accuracy against naive benchmarks (computed here, quarterly), the OBR's and external forecasters' published real-time errors (cited, annual), and the OBR emulator's tracking of the official forecast (replication, not forecasting).
No forecasting win survives the number of tests run. Bank Rate comes closest — 0.79 against a drifting random walk at one quarter, and the smallest of 64 variable×horizon tests at p = 0.018 — but adjusted together (Benjamini–Hochberg) the minimum q is 0.36, so nothing clears a 10% false-discovery rate on any variable at any horizon.
Benchmark and sample diagnostics
boe-svar against a random walk, computed from our own runs
Expanding-window design: 49 quarterly origins (2012Q1–2024Q1), estimation from 1992Q1, data through 2026Q1, horizons 1–8. A ratio below 1.0 beats the benchmark.
The benchmark decides the answer. A no-change walk on a trending log level forfeits the whole trend as error, so beating it on a price index is close to uninformative: UK CPI's 0.63 at one quarter becomes 0.83 against a walk with drift, and 0.67 at eight becomes 1.03. Bank Rate is the one series that does not trend, and the one that improves under the harder benchmark.
Covid dummies matter for GDP, not for inflation. The rolling evaluation omits the six Covid dummies (2020Q1–2021Q2) that every published forecast here carries. On the published specification UK GDP against drift goes 1.06 → 0.99 at one quarter and 1.12 → 0.95 at eight, while CPI and Bank Rate move by less than 0.02. Level with a naive benchmark is still not beating it: UK GDP is statistically inseparable from drift at every horizon (p = 0.33–0.43) in both specifications.
The AR(1) comparison in the source file is not usable at long
horizons: 94.5% of its eight-step squared error for UK GDP
comes from one origin, because an AR(1) extrapolates the Covid
collapse. Every benchmark carries a worst_origin_mse_share
diagnostic for that reason.
papers/boe-svar/figures/rolling_evaluation.json (worst_origin_mse_share.drift).papers/boe-svar/figures/rolling_evaluation.json.| variable | h=1 | h=2 | h=4 | h=8 | h=1 ex-Covid (vs no-change) | ||||
|---|---|---|---|---|---|---|---|---|---|
| vs drift | vs no-change | vs drift | vs no-change | vs drift | vs no-change | vs drift | vs no-change | ||
| Bank Rate | 0.79* | 0.88 | 0.85* | 0.96 | 0.89* | 1.02 | 0.85 | 1.03 | 0.86 |
| UK CPI (level) | 0.83* | 0.63 | 0.85 | 0.62 | 0.94 | 0.66 | 1.03 | 0.67 | 0.62 |
| World CPI (level) | 0.94 | 0.73 | 0.95 | 0.68 | 1.00 | 0.69 | 1.07 | 0.67 | 0.73 |
| Oil price | 1.01 | 1.03 | 1.02 | 1.05 | 1.01 | 1.06 | 0.98 | 1.08 | 1.05 |
| CPI energy | 0.98 | 0.98 | 0.97 | 0.98 | 1.03 | 1.04 | 1.11* | 1.13 | 0.97 |
| UK real GDP (level) | 1.06 | 1.06 | 1.09 | 1.08 | 1.10 | 1.09 | 1.12 | 1.06 | 0.77 |
| World GDP (level) | 1.07 | 1.04 | 1.10 | 1.03 | 1.11 | 0.93 | 1.10 | 0.71 | 0.51 |
| Exchange rate | 1.03 | 1.04 | 1.09 | 1.11 | 1.19 | 1.22 | 1.33 | 1.41 | 1.05 |
These are quarterly, level-basis ratios computed with hindsight-final data. They are not comparable with the real-time annual-growth errors the OBR publishes, which is why those get a separate table.
What the OBR and external forecasters report about themselves
Cited, not recomputed: the OBR's Forecast Evaluation Report (July 2025) reports its own real-time errors on annual growth rates since 2010, beside the HM Treasury median of external forecasters.
| variable · horizon | OBR | external median |
|---|---|---|
| Real GDP growth · 1 year ahead | 0.6pp | 0.6pp |
| Real GDP growth · 2 years ahead | 0.4pp | 0.4pp |
| CPI inflation · 1 year ahead | 0.3pp | 0.3pp |
| CPI inflation · 2 years ahead | 0.9pp | 0.9pp |
The nearest in-house analogue is the frozen-edge 0.32pp above — a different statistic, frequency and origin count, so it sits beside these numbers, not against them. For the Bank of England the only sourceable comparison is the matched-vintage episode on the validation index; the MPR publishes no RMSE-by-horizon table.
Where this departs from the Bank's model.
| limit | detail |
|---|---|
| Frozen data edge | Estimation to 2025Q1, conditioning to 2026Q1 — the extended sample, not the paper's 1992Q1–2023Q2 window. Results shift with revisions. |
| Proxied world aggregates | The Bank's UK-trade-weighted world GDP and CPI are unpublished, so they are rebuilt as chain-weighted US + euro-area + Japan + China with time-varying UK trade weights. |
| One ranking does not replicate | The paper's prose ranks UK monetary policy as the largest domestic contributor to CPI variance; here UK supply (13.9%) and UK demand (11.4%) both exceed it (8.7%). The paper publishes no FEVD table, so treat this as unresolved rather than a defect; the proxy aggregates are a candidate cause. |
| Assumed lag length and simplified pandemic prior | p = 4 is assumed, not selected; the pandemic is handled by exogenous Covid dummies. |
| Sign-restriction critiques apply | Pointwise medians mix structural models (Fry–Pagan); the Haar prior over rotations is informative about impulse responses (Baumeister–Hamilton). These hit the Bank's own outputs identically. |
every deviation is enumerated in the repository's
docs/methodology.md.