model 07 — ecological stock-flow consistent · define-uk · UK · experimental

Validated where the publications allow it, and no further.

The manual's macro block replicates; the emissions path diverges and is pinned; scenario switch sets are design-gated; the calibration gap to official data is computed. Deltas only, never levels.

how far to trust it

The replication gates, check by check.

Every row below is enforced by tests against the cached pinned run (upstream commit 846081a) and recorded in the adapter's VALIDATION.md, so drift fails loudly.

DEFINE-UK replication gates against the manual and the papers
what is checkedtolerancestatus
Manual Table 4, macro block — growth, unemployment, population 16+, labour force, S1 baseline population and labour force ±0.01m; growth ±0.31pp; unemployment ±0.2pp PASS (2026-08-04) — but 2025 growth misses by −0.30pp, clearing by 0.01pp on a gate set after the run
Manual Table 4, total emissions, S1 baseline gated at the observed ratios so further drift fails DIVERGENCE — pinned; see below
Scenario design — each cached scenario toggles exactly its published policy switches exact flag set per scenario; HOUSING_SUB_RATE=0.4 pinned PASS (2026-08-05, 10 tests)
FMM-2023 anchors — green public investment GDP peak; baseline 2030 emissions wide (pre-1.0 vintage): GDP peak 0.7–1.3%; emissions 330–352 MtCO2e PASS — peak +0.92% vs ≈+1%; 342.8 vs “just under 350”. Regression pins, not agreement tests: a band reaching 352 cannot test “just under 350”
Published numeric scenario results for v1.1 CLOSED at the achievable ceiling — none exist; see below
Clean-room reimplementation, milestone 1 — §2.2 matrices on the §5 initial values manual's 4-significant-figure printing precision PASS (2026-08-04) — 62 tests, 44 of them Table 1/2 row and column identities; 13 defects in the manual pinned as machine-readable MANUAL_GAPS records
Calibration vs external observations (ONS/DESNZ/OBR) computed and regression-tested, not tolerance-gated COMPUTED — gaps are material; see below
This adapter runs the authors' unmodified code at a pinned commit. DEFINE-UK belongs to George, Dafermos, Nikolaidi and co-authors; nothing here is produced or endorsed by them, and no result is theirs.
evidence — calibration

The baseline against official numbers, computed not prose.

Every row is recomputed from the cached pinned run against pinned external observations and committed as baseline_vs_external.csv; the suite fails on any drift. The annual-mean convention puts 2025 growth at 4.66% where Table 4 prints 4.96% — the same −0.30pp anchoring gap as the Table 4 gate above.

DEFINE-UK baseline (cached pinned run) against external observations, from validation/baseline_vs_external.csv
quantitymodelexternalverdict
Real GDP growth, 2025 4.66% y/y 1.31% y/y Far above outturn — never a forecast
Real GDP growth, mean 2025–40 2.4% y/y ≈1.65% y/y High vs official long-run
Unemployment, 2025 4.39% 4.88% Below outturn by ≈0.5pp
Population 16+, 2025 55.66m ≈56m Consistent
Labour force, 2025 35.97m ≈34.5m Slightly high
Total emissions, 2024 401.5 MtCO2e/yr 371 MtCO2e/yr Above the actuals it should start from
Emissions, 2030, vs NDC path 342.8 MtCO2e/yr ≈260 MtCO2e/yr Short of the NDC by design — current-policies baseline

external sources: ONS national accounts and MGSX (2026-07-26 vintages), ONS population and labour-force estimates, DESNZ 2024 provisional territorial emissions, OBR March 2026 EFO, and the UK NDC — attributed row by row in the committed CSV.

So the levels are not forecasts and are never served as such: the manual says the baseline “should not be seen as a prediction”, and near-term levels are not competitive with the OBR emulator or the BoE SVAR. Only scenario deltas are defensible, carrying the demand-led caveat.

evidence — multiplier

The fiscal multiplier, stated honestly.

The model's cumulative green-public-investment multiplier, on our own computation — nominal cumulative ΔGDP over ΔSPEND_GVT to 2040 — is 1.78. It is not comparable to the IMF's 1.1–1.5 or the OBR's ≈1.0, and three things push the same way: those are short-horizon multipliers where ours is cumulative to 2040; the denominator is endogenous (government investment is 12.8% of SPEND_GVT, and social benefits fall as unemployment does, shrinking it); and a demand-led closure with Kaldor–Verdoorn productivity implies a ratio above 1 anyway. The clean impulse-based version has not been computed, so treat 1.78 as upper-leaning. The caveats travel with every served delta (GPI_MULTIPLIER_CAVEATS, gated by test).

Upstream's own Multiplier_Summary.csv is unusable and never quoted: its multiplier columns are identical across all eight scenarios and its ΔG matches the cumulative delta of no variable in the scenario file. The ≈2.4 derived from it should not be quoted either.

evidence — emissions

The emissions path diverges from the published table.

The one open divergence: the pinned upstream code runs below the manual's published Table 4 emissions by −3.5% in 2025, −10.2% in 2030 and −23.3% in 2040.

define-uk: S1 baseline total emissions — pinned run vs manual Table 4 (MtCO2e/yr) Grouped bar chart of S1 baseline total UK emissions in MtCO2e per year, the cached pinned run of upstream commit 846081a against the manual's published Table 4. 2025: pinned run 393, published table 407, a gated ratio of 0.965, a gap of -3.5 per cent; 2030: pinned run 343, published table 382, a gated ratio of 0.898, a gap of -10.2 per cent; 2040: pinned run 249, published table 324, a gated ratio of 0.767, a gap of -23.3 per cent. The pinned code runs below the published table and the gap widens with horizon; what is gated is the ratio of the two, pinned to plus or minus 0.02, so any further drift fails loudly. Bar heights are the levels in validation/figures/data/define_emissions_divergence.csv, rounded to whole MtCO2e; the percentages are the gated ratios themselves, which is why they differ slightly from the ratio of the rounded bars. 0 100 200 300 400 393 pinned run 407 published 2025 (-3.5%) 343 pinned run 382 published 2030 (-10.2%) 249 pinned run 324 published 2040 (-23.3%)
S1 baseline total emissions: the cached pinned run (upstream 846081a) against the published Table 4. What is gated is the ratio of the two — 0.965 / 0.898 / 0.767, pinned to ±0.02 in the adapter's validation/reference_outputs.json, which is where the percentages above come from; the bar labels round the published levels. Source: validation/figures/data/define_emissions_divergence.csv.

Not an extraction error: the components (EMIS_NELEC + EMIS_ELEC) sum exactly to the total and the baseline is identical across scenario folders, so it is a vintage or calibration gap between the published table and commit 846081a — raised with the authors as DEFINE_UK_1.1#1.

limits — the ceiling

No numeric v1.1 scenario results are published. That is the ceiling.

No machine-readable numeric scenario results are published for DEFINE-UK v1.1: the manual stops at the baseline plus scenario design parameters, and the accompanying papers are paywalled with figure-only results. A published-figure replication of the scenarios is not possible today, and every delta this adapter serves carries that status.

So what is gated is design and self-consistency — the switch sets, and the delta paths against the cached pinned run — plus the two coarse FMM-2023 anchors above. The gate reopens if the authors publish scenario tables or the papers become accessible. Until then: experimental, partial replication, deltas only, and score a reform does not accept this model.