model 07 — ecological stock-flow consistent · define-uk · UK · experimental
Validated where the publications allow it, and no further.
The manual's macro block replicates; the emissions path diverges and is pinned; scenario switch sets are design-gated; the calibration gap to official data is computed. Deltas only, never levels.
The replication gates, check by check.
Every row below is enforced by tests against the cached pinned run
(upstream commit 846081a) and recorded in the adapter's
VALIDATION.md,
so drift fails loudly.
| what is checked | tolerance | status |
|---|---|---|
| Manual Table 4, macro block — growth, unemployment, population 16+, labour force, S1 baseline | population and labour force ±0.01m; growth ±0.31pp; unemployment ±0.2pp | PASS (2026-08-04) — but 2025 growth misses by −0.30pp, clearing by 0.01pp on a gate set after the run |
| Manual Table 4, total emissions, S1 baseline | gated at the observed ratios so further drift fails | DIVERGENCE — pinned; see below |
| Scenario design — each cached scenario toggles exactly its published policy switches | exact flag set per scenario; HOUSING_SUB_RATE=0.4 pinned |
PASS (2026-08-05, 10 tests) |
| FMM-2023 anchors — green public investment GDP peak; baseline 2030 emissions | wide (pre-1.0 vintage): GDP peak 0.7–1.3%; emissions 330–352 MtCO2e | PASS — peak +0.92% vs ≈+1%; 342.8 vs “just under 350”. Regression pins, not agreement tests: a band reaching 352 cannot test “just under 350” |
| Published numeric scenario results for v1.1 | — | CLOSED at the achievable ceiling — none exist; see below |
| Clean-room reimplementation, milestone 1 — §2.2 matrices on the §5 initial values | manual's 4-significant-figure printing precision | PASS (2026-08-04) — 62 tests, 44 of them Table 1/2 row and column identities; 13 defects in the manual pinned as machine-readable MANUAL_GAPS records |
| Calibration vs external observations (ONS/DESNZ/OBR) | computed and regression-tested, not tolerance-gated | COMPUTED — gaps are material; see below |
The baseline against official numbers, computed not prose.
Every row is recomputed from the cached pinned run against pinned external observations and committed as baseline_vs_external.csv; the suite fails on any drift. The annual-mean convention puts 2025 growth at 4.66% where Table 4 prints 4.96% — the same −0.30pp anchoring gap as the Table 4 gate above.
| quantity | model | external | verdict |
|---|---|---|---|
| Real GDP growth, 2025 | 4.66% y/y | 1.31% y/y | Far above outturn — never a forecast |
| Real GDP growth, mean 2025–40 | 2.4% y/y | ≈1.65% y/y | High vs official long-run |
| Unemployment, 2025 | 4.39% | 4.88% | Below outturn by ≈0.5pp |
| Population 16+, 2025 | 55.66m | ≈56m | Consistent |
| Labour force, 2025 | 35.97m | ≈34.5m | Slightly high |
| Total emissions, 2024 | 401.5 MtCO2e/yr | 371 MtCO2e/yr | Above the actuals it should start from |
| Emissions, 2030, vs NDC path | 342.8 MtCO2e/yr | ≈260 MtCO2e/yr | Short of the NDC by design — current-policies baseline |
external sources: ONS national accounts and MGSX (2026-07-26 vintages), ONS population and labour-force estimates, DESNZ 2024 provisional territorial emissions, OBR March 2026 EFO, and the UK NDC — attributed row by row in the committed CSV.
So the levels are not forecasts and are never served as such: the manual says the baseline “should not be seen as a prediction”, and near-term levels are not competitive with the OBR emulator or the BoE SVAR. Only scenario deltas are defensible, carrying the demand-led caveat.
The fiscal multiplier, stated honestly.
The model's cumulative green-public-investment multiplier, on our own
computation — nominal cumulative ΔGDP over ΔSPEND_GVT to 2040 — is
1.78. It is not comparable to the IMF's
1.1–1.5 or the OBR's ≈1.0, and three things push the same
way: those are short-horizon multipliers where ours is cumulative to
2040; the denominator is endogenous (government investment is 12.8% of
SPEND_GVT, and social benefits fall as unemployment does,
shrinking it); and a demand-led closure with Kaldor–Verdoorn
productivity implies a ratio above 1 anyway. The clean impulse-based
version has not been computed, so treat 1.78 as upper-leaning. The
caveats travel with every served delta
(GPI_MULTIPLIER_CAVEATS, gated by test).
Upstream's own Multiplier_Summary.csv is
unusable and never quoted: its multiplier columns are
identical across all eight scenarios and its ΔG matches the cumulative
delta of no variable in the scenario file. The ≈2.4 derived from
it should not be quoted either.
The emissions path diverges from the published table.
The one open divergence: the pinned upstream code runs below the manual's published Table 4 emissions by −3.5% in 2025, −10.2% in 2030 and −23.3% in 2040.
846081a) against the published Table 4. What is gated is the ratio of the two — 0.965 / 0.898 / 0.767, pinned to ±0.02 in the adapter's validation/reference_outputs.json, which is where the percentages above come from; the bar labels round the published levels. Source: validation/figures/data/define_emissions_divergence.csv.
Not an extraction error: the components
(EMIS_NELEC + EMIS_ELEC) sum
exactly to the total and the baseline is identical across scenario
folders, so it is a vintage or calibration gap between the published
table and commit 846081a — raised with the authors as
DEFINE_UK_1.1#1.
No numeric v1.1 scenario results are published. That is the ceiling.
So what is gated is design and self-consistency — the switch sets, and the delta paths against the cached pinned run — plus the two coarse FMM-2023 anchors above. The gate reopens if the authors publish scenario tables or the papers become accessible. Until then: experimental, partial replication, deltas only, and score a reform does not accept this model.