evidence by model
See the evidence for each model.
Compare results with official benchmarks—and see clearly where no replication ground truth exists.
Validation is a gradient, not a badge.
The five models support different kinds of checks. The verification gradient groups them into three classes; choose a row, then open its evidence below.
| Model | Verification class | What it is checked against | Headline result |
|---|---|---|---|
| obr-macro | Replication with a published anchor | OBR Economic and Fiscal Outlook, March 2026; HMRC ready reckoner; ONS outturns | Anchored GDP MAPE 0.15%, consumption 0.25% |
| boe-svar | Replication with a published anchor | Brignone & Piffer (2025), BoE Macro Technical Paper No. 3; ONS outturns | FEVD shares 42.1% / 49.5% vs the paper's ~40% / ~50% |
| frb-us | Replication with a published anchor | The Fed's LONGBASE database and pyfrbus solver; Coenen et al. and CBO multiplier ranges |
Tracking invariant 5.6×10−17 — machine precision |
| pe-microsim | Checked against implemented legislation | Implemented UK and US legislation, rule by rule, for a specified household | Deterministic household calculation; population estimates add survey and calibration uncertainty |
| psl-og | Calibrated counterfactual — no ground truth | ONS and OBR aggregates as calibration targets, not as a validation set | Targets met by construction; no published replication exists |
Software fidelity and economic validity are separate.
This audit covers the three macro models. Levels are categorical evidence judgements, not a synthetic score: a model can reproduce its reference code perfectly while still lacking independent forecast or counterfactual validation.
| Evidence dimension | obr-macro | boe-svar | frb-us |
|---|---|---|---|
| Implementation fidelity | Moderate — incomplete channels | Strong — restrictions and identities gated | Strong — four gated scenarios at the reference-solver noise floor |
| Predictive validation | Weak — raw GDP 5.75%, consumption 9.56% MAPE | Weak — the CPI win does not survive a drift benchmark; Bank Rate is the only defensible claim | Not assessed — LONGBASE is not a Fed forecast |
| Identification robustness | Not applicable | Moderate — proxy data and specification matter | Not applicable |
| Policy-counterfactual validity | Weak — one independent tax benchmark; spending multiplier ~1.0 vs the OBR's 0.6 | Not applicable — reform scoring refused | Moderate — published ranges, VAR expectations only |
| Uncertainty calibration | Weak — sensitivity envelopes pending | Moderate — posterior bands, limited coverage history | Weak — bootstrap engine exists, public runs remain deterministic |
| Vintage reproducibility | Moderate — multi-EFO archive pending | Moderate — public proxies replace internal series | Moderate — artifacts hashed; multi-vintage tests pending |
Against the OBR's own forecast, and HMRC's reckoner.
Anchored to the March 2026 EFO with OBR add-factors, the emulator reproduces GDP to 0.15% MAPE and consumption to 0.25% over 2025Q1–2027Q4 (CI gate: 1%). GDP remains within 0.29% through 2031Q1, but unemployment becomes unreliable after 2027Q4. This fit is by construction; the independent reform costing below is the stronger test. The headline chart and table now use the live March 2026 vintage; the older November 2025 forecast is retained only inside the historical outturn audit.
papers/obr-macro/figures/fig_anchored_data.csv, regenerated from the March 2026 detailed forecast tables on 21 July 2026.The independent check is the reform costing. Raising the basic rate of income tax by 1pp from April 2026 scores at £6.46bn in 2026–27 through the PolicyEngine static-costing bridge, against HMRC's ready-reckoner figure of £6.9bn — a deviation of −6.4%; the costing sits inside the £6–8bn range spanned by recent published vintages. The gap widens in later years (−15.6% by 2028–29) where HMRC's figures embed administrative-data fiscal drag that survey microdata capture less fully.
| Ours | Official | Deviation | |
|---|---|---|---|
| Anchored levels vs EFO March 2026, £bn/qtr | |||
| Real GDP, 2025Q1 | 703.8 | 703.4 | +0.05% |
| Real GDP, 2027Q4 | 730.6 | 728.6 | +0.28% |
| Consumption, 2025Q1 | 429.7 | 429.3 | +0.09% |
| Consumption, 2027Q4 | 445.5 | 443.4 | +0.46% |
| Basic rate +1pp vs HMRC ready reckoner, £bn/yr | |||
| 2026–27 | 6.46 | 6.9 | −6.4% |
| 2028–29 (interpolated) | 6.92 | 8.2 | −15.6% |
| 2030 (end of window) | 7.38 | ≈8.2 | −10.0% |
| Tracking error vs the published March 2026 EFO path, MAPE 2025Q1–2027Q4 | |||
| Anchored (GDP / consumption) | 0.15% / 0.25% | CI gate <1% — passes | |
| Held add-factors (GDP / consumption), 2026Q1–2027Q4 window | 0.37% / 0.33% | 6 of 8 computed within band (paper's Nov-2025-vintage scorecard, over its own longer window: 2.2% / 3.6%) | |
| Free-running, raw (GDP / consumption) | 5.75% / 9.56% | over band — report-only, and weak | |
docs/calibration_scorecard.md in the obr-macroeconomic-model repository.The spending multiplier is ~1.0 by construction, against the OBR's own published 0.6. Under the demand closure a spending shock lands directly in the GDP identity and the behavioural second round is largely inactive, so a £5bn injection returns almost exactly £5bn of GDP. That is roughly a two-thirds overstatement of the impact multiplier relative to the institution being replicated, and it applies to every spending-side figure this page reports. It is the single most important number for anyone reading a policy score off this model.
Three of the channels that would normally damp such a shock are held exogenous in this configuration — imports (so the usual leakage is switched off), Bank Rate (no monetary offset) and CPI. The first two bias the estimated effect upward; the dead profits→dividends channel biases corporation-tax effects downward. These do not cancel, and their net sign is not established. Reform figures here are best read as relative comparisons between structurally identical runs, where the biases largely difference out, rather than as magnitudes.
The honest scorecard, the two dead equations, outturns, and the March 2026 re-anchoring
The honest scorecard. The last row is the one that
matters for anyone tempted to read the 0.15% as forecasting skill.
The same equations that track the EFO to 0.15% when anchored miss it
by 5.75% free-running — precisely the gap the OBR's
own add-factor judgement closes in the official process, and the
reason reform deltas are always scored against the anchored baseline
rather than the raw one. The free-running score is de-seeded, with
passthrough variables excluded, and is published report-only.
Across the full scorecard only 4 of the 11 computed headline
variables land within band (real GDP, consumption, the trade
balance, and the trivial employment identity); of the full 21-line
scorecard, 10 lines are passthroughs held at the OBR value. The
worst line is company profits at 79.80% MAPE on the
March baseline (54.57% on the paper's November vintage), which
traces to a single unpublished constant in households' operating
surplus OSHH — the paper documents and regression-gates
it rather than re-tuning it, since tuning it would be fitting to the
answer. Other lines that moved on re-anchoring, reported rather than
smoothed: the free-running current account widened from 2.76 to
4.17% of GDP and is now over band; RPI improved from
2.03pp to 1.71pp; business investment worsened from
15.48% to 16.12%; the two household-income
lines stand at 14.15% and 13.86%.
papers/obr-macro/figures/fig_free_running_data.csv and fig_anchored_data.csv, regenerated on 21 July 2026.log(X) left-hand side — log(HHTFA) and
log(NDIVHH). The transpiler now parses that form, but
both still fail to execute for an independent reason: their exogenous
right-hand-side inputs MAJGDP and CORP are
absent from the databank entirely, so the right-hand side evaluates
non-finite and the solver skips the update. Every OBR figure on this
page is computed with those two channels inert. The practical effect
is bounded and demonstrable: household dividend income does not
respond to a corporate-profits change, so the household-income channel
of a corporation-tax reform is muted — a 5pp corporation-tax rise
moves FYCPR by −£1,780m while
ΔNDIVHH is exactly zero. Reviving the channel needs
a sourced CORP series, a calibration decision rather than
a code fix.
Forecast versus outturn. Comparing one forecast vintage with another tests agreement, not accuracy. Against ONS outturns published since anchoring, quarter-on-quarter real GDP growth ran 0.1% in 2025Q2, 0.2% in Q3, 0.2% in Q4 and 0.6% in 2026Q1. The emulator's path (0.15, 0.14, 0.25, 0.37) tracks the three 2025 quarters to within 0.06 percentage points, but — like the November EFO it inherits (0.28, 0.20, 0.27, 0.39) — misses the strong 2026Q1 outturn by roughly a quarter of a point. Two caveats govern the reading: this is primarily a test of the OBR's November vintage, the emulator's own contribution being the 0.02–0.13 point gap between the two model rows; and ONS quarterly estimates are themselves revised, so the outturn is a moving target.
papers/obr-macro/figures/fig_outturn_data.csv; the table below carries the exact values.| Real GDP, % q/q | Emulator | EFO Nov 2025 | ONS outturn |
|---|---|---|---|
| 2025Q2 | 0.15 | 0.28 | 0.1 |
| 2025Q3 | 0.14 | 0.20 | 0.2 |
| 2025Q4 | 0.25 | 0.27 | 0.2 |
| 2026Q1 | 0.37 | 0.39 | 0.6 |
Vintage: re-anchored to the March 2026 EFO. The hosted emulator has been re-anchored from the November 2025 EFO to the OBR's March 2026 forecast, and the headline numbers on this page are computed on that baseline: anchored GDP 0.15% MAPE, consumption 0.25% over 2025Q1–2027Q4, with the anchored horizon extended to 2031Q1 (GDP reproduced to 0.29% at 2031Q1; anchored unemployment is unreliable beyond 2027Q4, drifting to 0.9% against the EFO's 4.1% by 2031Q1). The headline and free-running charts now use March 2026. Only the outturn backtest above retains November 2025, because changing its forecast vintage would erase the historical forecast being tested. The working paper now opens with a dated current-vintage note. Reform effects are differences between structurally identical runs and are insensitive to modest baseline drift, which is why the re-anchoring leaves the £6.46bn/£7.38bn static costing untouched and moves the second-round GDP effect only from −0.057% to −0.058% by 2027Q4.
Against a Bank of England paper, and against outturns.
Against Brignone & Piffer (2025), the one-year share of UK forecast-error variance from global shocks is 42.1% for GDP and 49.5% for CPI. The paper reports roughly 40% and 50%.
papers/boe-svar/figures/comparison_numbers.json and the paper's validation table.Why the fast CI configuration lands 6–10 points off
The deliberately cheap unweighted CI configuration lands at 49.6% and 43.8% — 6–10 points off, in opposite directions for the two variables. That gap is itself informative rather than embarrassing: the importance weights correct the Arias et al. (2018) zero-restriction sampler towards the uniform-over-rotations posterior, and applying them moves both shares substantially towards the published values, though on the 2026Q1 vintage GDP still lands 2.1 points high. Both configurations sit inside the wide [30, 60]% acceptance band that CI enforces, which is chosen to catch gross regressions without being flaky to Monte-Carlo variation.
| Global share, 1-yr FEVD | Paper | CI config (fast) | Production (10k draws) | Deviation, CI / production |
|---|---|---|---|---|
| UK GDP | ~40% | 49.6% | 42.1% | +9.6 / +2.1 pp |
| UK CPI | ~50% | 43.8% | 49.5% | −6.2 / −0.5 pp |
Out-of-sample. Forecasting seven quarters ahead from a frozen 2024Q2 edge, RMSE is 0.3pp for both GDP growth and CPI inflation. Activity tracks well; the median misses the 2025 inflation peak by up to 0.6pp but captures its reversal.
Fourteen of fourteen outturns inside the 68% band is a calibration failure, not a pass. A correctly calibrated 68% interval should contain roughly 9.5 of 14. Containing all of them means the bands are too wide — the intervals are roughly ±0.95pp around a forecast whose RMSE is 0.3pp, about three times wider than the errors warrant. The model is under-confident here, and this figure was previously presented as a validation win. It is reported now as what it is: evidence of over-dispersion.
Neither reading is strong evidence either way. The fourteen points are seven quarters × two variables from a single forecast origin, so they are heavily correlated across horizon and across variables; the effective number of independent observations is perhaps two or three. Coverage cannot be established from one origin in either direction. The rolling-origin machinery could produce a genuine coverage table across 49 origins and eight horizons; until it does, this page makes no coverage claim.
Rolling-origin audit. A stricter expanding-window test now evaluates 49 historical origins without future-data leakage. Against a no-change benchmark, CPI-level relative RMSE is 0.63 at one quarter and 0.67 at eight quarters (lower is better), while UK GDP sits at 1.06. But a no-change benchmark is too weak for a trending price level: against a random walk with drift the CPI figures become 0.83 and 1.03, and the GDP gap is not statistically significant at any horizon. See section 08 for the full benchmark comparison. The defensible forecasting claim is Bank Rate, not inflation, and the single frozen-edge result must not be read as broad GDP forecast superiority.
papers/boe-svar/figures/figure_numbers.json (medians and 68% bands) and the ONS outturn series in papers/boe-svar/figures/make_figures.py.Why there is no official yardstick for those forecast errors
No official counterpart exists for those RMSEs. The authors' companion paper (Staff Working Paper No. 1,165, January 2026) is a methods paper on a different, four-variable SVAR; it publishes no root-mean-squared errors, coverage rates or benchmark comparisons, so the forecast numbers here stand without an official yardstick and are reported as such.
Against the Fed's own solver — and against the Fed's own noise floor.
The baseline reproduces the Fed's LONGBASE database across 284 variables and 20 quarters with a maximum error of 5.6×10−17 — machine precision, against a 10−8 CI gate.
Shock responses match pyfrbus 1.0.0 to
6.0×10−9 and 1.1.1 to
1.4×10−8. The Fed's two
releases differ by 1.3×10−8, so
our differences sit at the reference solver's own numerical noise.
Multipliers. Government-purchases and personal-tax shocks of 1% of GDP land inside every published range they are compared against — the last three rows below.
| Quantity | Ours | Published | Reading |
|---|---|---|---|
| Tracking invariant, max abs. error | 5.6×10−17 | CI gate <10−8; pyfrbus 1.1.1 gives 1.1×10−8 | machine precision |
| Shock responses vs pyfrbus 1.0.0 | 6.0×10−9 | — | at the reference's noise floor |
| Shock responses vs pyfrbus 1.1.1 | 1.4×10−8 | 1.1.1 vs 1.0.0: 1.3×10−8 | the Fed's releases differ as much as we do |
| Gov. purchases multiplier, yr 1, inertial Taylor rule | 0.72 | 0.7–1.0 (Coenen et al. 2012) | inside range |
| Gov. purchases multiplier, yr 2, fixed funds rate | 0.99 | "roughly one" pegged; 1.1–1.2 with accommodation | matches the Board's characterisation |
| Personal tax cut multiplier, yrs 1–2 | 0.22 → 0.32 | 0.2–0.4; CBO central ≈0.3 | inside both ranges |
Uncertainty engine. The implementation now supports seeded stochastic simulations that jointly resample the official stochastic-equation residual vector, preserving contemporaneous dependence and reporting failed replications. Model and LONGBASE packages are also verified against separate SHA-256 provenance gates. Published probability results still require a reviewed simulation design; model-consistent expectations remain unavailable.
Two scope caveats: expectations path, and the 2014 FEDS Note comparison
Only the VAR-expectations path was exercised in the version comparison; model-consistent expectations are not implemented here, so forward-guidance-style exercises are out of scope. And published 100bp funds-rate responses are read off the 2014 FEDS Note charts, so those comparisons are approximate by nature — our output-gap trough of −0.50pp against the Note's ≈−0.4pp is within vintage and shock-design differences, not an exact match.
Against implemented legislation.
This model does not forecast. It deterministically applies encoded UK and US tax-benefit law to a specified household, so its reference is legislation and official guidance. Coverage depends on the upstream rules and tests.
Population estimates also depend on survey data and calibration. The
UK run uses enhanced_frs_2023_24; in the
1pp basic-rate check, it is 6.4% below HMRC in
year one. Survey versus administrative data, policy assumptions, and
HMRC behavioural elasticities explain that separate aggregate gap.
No ground truth. Calibration targets instead.
| Quantity | Official UK value | Source | Model treatment | Deviation |
|---|---|---|---|---|
| Government debt-to-GDP | 95.1% (May 2026); OBR March 2026 forecast 94.4% | ONS public sector finances | imposed via closure target | −0.1pp vs latest ONS (target 95.0%) |
| Household saving ratio | 8.9% (2026Q1, down from 9.6%) | ONS quarterly sector accounts | targeted through β | 0 by construction |
| Potential growth | ≈1.1% / yr | OBR EFO | imposed, gy = 0.011 | ≈0 |
| Depreciation / capital | ≈6–7% (CFC / net stock) | ONS capital stocks bulletin | imposed, δ = 0.065 | within range |
| Labour share of income | ≈0.59–0.60, rising in 2024 | ONS labour-share series / Blue Book 2025 | imposed, γ = 0.35 (OG-Core default) | ≈5pp — the capital share is not yet re-anchored to ONS factor shares |
| Net capital stock | £5.6tn (2024) whole-economy net stock; the model's K maps to the narrower business-capital concept | ONS capital stocks bulletin | emergent — checked, not imposed | concept-adjusted; no published reconciliation |
A zero deviation on an imposed quantity is not independent evidence.
What would count as over-identifying evidence, and the dating and structural caveats
The genuinely over-identifying checks are the ones the steady state must deliver: the capital–output ratio and the interest rate consistent with the calibrated (γ, δ, β) triple. These emerge in an economically sensible range in the deployed configuration, but no systematic published reconciliation of achieved-versus-target moments exists for OG-UK 0.3.2, and the working paper declines to manufacture one. Producing that reconciliation from logged solves is the single most valuable next step for this member.
Two dating caveats: the official column reports figures current at mid-2026 while the deployed calibration was frozen against slightly earlier vintages, and the saving ratio is volatile quarter to quarter, so β is matched to a smoothed level rather than the latest print. Two structural caveats carried over from OG-Core: the ability types and lifetime earnings profiles are still the US tax-microdata estimates, and UK re-estimation is an open task.
The same questions, answered by other shops.
Published institutional answers sit beside ours; each row states how closely the definitions align.
| Question | Ours | Theirs | How comparable |
|---|---|---|---|
| Long-run effect of a 1pp basic-rate rise (psl-og) | Not yet produced — the OG-UK side of this comparison is open validation work | OBR Working Paper 22, Table 5.1: GDP/person −0.1%, labour supply −0.2%, productivity +0.1% | Future benchmark — same scenario and model family; WP 22's figures are the target our run will be judged against, not a match we claim today |
| Cost of cutting the basic rate 1p, 2026–27 (obr-macro + pe) | £6.46bn static | HMRC ready reckoner (June 2025): £6.9bn, rising to £8.2bn by 2028–29 | Direct, −6.4% in year one — HMRC's figure embeds behavioural response and the OBR's March 2025 forecast |
| Government-purchases multiplier, year one (frb-us) | 0.72 (inertial Taylor rule); 0.99 in year two with the funds rate fixed | Ramey (JEP 2019): the literature sits "in a surprisingly narrow range of 0.6 to 1" | Direct — both estimates fall inside the survey range |
| Income-tax multiplier (frb-us, US) | 0.22 → 0.32 over years 1–2 | OBR's published UK assumption: 0.3 for income tax and NICs | Loose — different country and model class; the coincidence is noted, not claimed |
| Forecast error, 2024Q3 onward (boe-svar) | RMSE 0.32pp on quarterly GDP growth and CPI inflation, data frozen at 2024Q2; the 2025 CPI hump missed by up to 0.6pp | OBR forecast evaluation (July 2025): one-year GDP growth underestimated by 0.4pp on average since 2010 (external forecasters 0.5pp). Bank of England MPR (Aug 2024): modal CPI for late 2025 around 2.4%, against an outturn peak of 3.8% — a miss of ~1.4pp on the same episode | Adjusted — RMSE and mean error are different statistics, and the Bank's modal path is conditioned on market rates; the vintages, at least, are matched |
How do we compare to other forecasters?
Three different statistics answer three different questions, and they are not averaged, ranked, or plotted on one axis: our SVAR's pseudo-out-of-sample accuracy against a naive benchmark (computed here, quarterly), the OBR's and external forecasters' own published real-time errors (cited, annual), and our OBR emulator's tracking of the official forecast (replication accuracy, not forecasting accuracy).
boe-svar against a random walk, computed from our own runs
An expanding-window pseudo-out-of-sample exercise re-fits the BVAR at 49 quarterly origins (2012Q1–2024Q1, data sample 1992Q1–2026Q1) and scores forecasts one to eight quarters ahead against three naive benchmarks. A ratio below 1.0 means the model beats the benchmark.
The benchmark decides the answer. A no-change random walk on a trending log level forfeits the whole trend as forecast error, so beating it on a price index is close to uninformative. Against a random walk with drift — the textbook naive for a trending series — the CPI result largely evaporates: 0.63 becomes 0.83 at one quarter and 0.67 becomes 1.03 at eight, where the model is no longer ahead at all. The pattern in the original numbers, wins only on the two trending price series and ties or losses on the four series where a random walk is a genuinely hard benchmark, was an artefact of the benchmark rather than a finding about the model.
What survives is Bank Rate. It is the one series here that does not trend, so no-change is the right naive for it, and it is the one series that improves under the harder benchmark: 0.79 against drift at one quarter (Diebold–Mariano p = 0.018) and 0.85 at eight. On the evidence assembled here that, not inflation, is the model's defensible forecasting claim.
UK GDP is not beaten by the benchmark either. Its ratio of 1.06–1.09 is not statistically distinguishable from a random walk at any horizon (p = 0.38 to 0.67), and excluding the six origins whose target quarter falls in 2020Q1–2021Q2 it becomes 0.77 at one quarter. Under squared loss the 2020Q2 collapse dominates a 49-origin average. Both the full-sample and the excluding-Covid figures are published; neither is the preferred number.
The AR(1) comparison is reported in the source file but is
not usable at long horizons: 94.5% of its eight-step
squared error for UK GDP comes from a single origin, because an AR(1)
extrapolates the Covid collapse. The apparent 3:1 win against it is
arithmetic, not evidence. Every benchmark now carries a
worst_origin_mse_share diagnostic for exactly this reason.
papers/boe-svar/figures/rolling_evaluation.json.papers/boe-svar/figures/rolling_evaluation.json.| variable | h=1 | h=2 | h=4 | h=8 | h=1 ex-Covid (vs no-change) | ||||
|---|---|---|---|---|---|---|---|---|---|
| vs drift | vs no-change | vs drift | vs no-change | vs drift | vs no-change | vs drift | vs no-change | ||
| Bank Rate | 0.79* | 0.88 | 0.85* | 0.96 | 0.89* | 1.02 | 0.85 | 1.03 | 0.86 |
| UK CPI (level) | 0.83* | 0.63 | 0.85 | 0.62 | 0.94 | 0.66 | 1.03 | 0.67 | 0.62 |
| World CPI (level) | 0.94 | 0.73 | 0.95 | 0.68 | 1.00 | 0.69 | 1.07 | 0.67 | 0.73 |
| Oil price | 1.01 | 1.03 | 1.02 | 1.05 | 1.01 | 1.06 | 0.98 | 1.08 | 1.05 |
| CPI energy | 0.98 | 0.98 | 0.97 | 0.98 | 1.03 | 1.04 | 1.11* | 1.13 | 0.97 |
| UK real GDP (level) | 1.06 | 1.06 | 1.09 | 1.08 | 1.10 | 1.09 | 1.12 | 1.06 | 0.77 |
| World GDP (level) | 1.07 | 1.04 | 1.10 | 1.03 | 1.11 | 0.93 | 1.10 | 0.71 | 0.51 |
| Exchange rate | 1.03 | 1.04 | 1.09 | 1.11 | 1.19 | 1.22 | 1.33 | 1.41 | 1.05 |
Bank Rate crosses parity beyond h=3 (source file has all eight variables and horizons). These are quarterly, level-basis, pseudo-out-of-sample ratios computed with hindsight-final data — they are not comparable with the real-time annual-growth errors the OBR publishes, which is why they get separate tables.
What the OBR and external forecasters report about themselves
The OBR's Forecast Evaluation Report (July 2025) publishes its own real-time errors on annual growth rates since 2010, beside the median of external forecasters compiled by HM Treasury. Cited, not recomputed; the statistic is the median absolute error on annual rates, so it cannot be compared with the quarterly RMSE ratios above.
| variable · horizon | OBR | external median |
|---|---|---|
| Real GDP growth · 1 year ahead | 0.6pp | 0.6pp |
| Real GDP growth · 2 years ahead | 0.4pp | 0.4pp |
| CPI inflation · 1 year ahead | 0.3pp | 0.3pp |
| CPI inflation · 2 years ahead | 0.9pp | 0.9pp |
Our own frozen-edge experiment is the nearest in-house analogue: the boe-svar RMSE of 0.32pp on quarterly year-on-year GDP growth and CPI inflation over seven quarters from a 2024Q2 data edge (section 03 above). Different statistic (RMSE vs median absolute error), different frequency (quarterly vs annual), one origin against many — so it sits beside the OBR's numbers, not against them. For the Bank of England, the single matched-vintage episode already cited in section 07 — the August 2024 MPR modal CPI path of about 2.4% for late 2025 against an outturn peak of 3.8% — is the only comparison we can source precisely; the Bank does not publish an RMSE-by-horizon table in the MPR.
Our OBR emulator: replication accuracy, not forecast accuracy
The obr-macro numbers in section 02 measure how closely the anchored emulator reproduces the OBR's published March 2026 path (MAPE 0.15% GDP, 0.25% consumption) — that is replication accuracy. It is not an independent forecast: de-seeded and stripped of add-factors, the same model drifts 5.75% below the EFO path over twelve quarters. Any forecast-accuracy credit for obr-macro therefore belongs to the OBR's own forecast, whose errors are the FER numbers above.