Measurement honesty so we optimise SAP-relevant accuracy, not SAP-neutral
misses (ADR-0030 Component Accuracy):
- Add construction_age_band_pm1: an exact-or-adjacent-band hit. Adjacent
RdSAP age bands carry near-identical U-values, so an off-by-one is
~SAP-neutral. Full corpus: exact 78.5% but ±1-band 91.7% (fixture
63.9% -> 83.3%) — most age misses are adjacent.
- Drop window_count from the gate's residual ceilings (cosmetic): the
predicted picture clusters at a mapper-default 4 windows vs actuals 1-21,
but total_window_area (the SAP-relevant signal) stays tight at ~3.4 m2.
Gate: + construction_age_band_pm1 floor 0.8333; window_count no longer gated.
Closes#1222
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Investigated recency-weighting (weight cohort votes by an exponential decay
in cert age). Key finding: it must be SELECTIVE. On the validation corpus it
HURTS permanent categoricals (wall 91.2->89.5, age 78.5->75.7 — discards
still-valid data) but clearly HELPS time-varying ones, where a recent
neighbour reflects the current physical state:
roof_insulation_thickness 56.7 -> 60.7% corpus (+4pp)
29.4 -> 41.2% fixture (+12pp)
So apply a recency-weighted mode only to roof_insulation_thickness (loft
top-ups happen over time); keep the plain mode for permanent categoricals.
tau = 4yr (~2.8yr half-life); falls back to plain mode when no registration
dates are lodged. Gate floor ratcheted 0.2941 -> 0.4118.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
These independent fabric categoricals were template-copied; mode them like
the construction categoricals. Verified mode beats template before applying.
Big fixture win on roof insulation thickness (doubled), floor insulation
neutral-to-positive:
roof_insulation_thickness 14.7% -> 29.4% (gate floor ratcheted up)
floor_insulation 90.6% (unchanged on the fixture)
Glazing type was tried too (+1.6pp on the 40-postcode corpus) but REGRESSED
the 36-target fixture (0.50 -> 0.44) — the gate caught it. Glazing moding is
marginal/noisy, so it's left on the template; revisit with a larger corpus.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The committed CI gate: run the calculator-free leave-one-out scorer over the
frozen anonymised fixture (36 SAP-10.2 targets) and assert each per-component
classification rate / geometry residual is no worse than a committed baseline.
Prediction is deterministic + the fixture frozen, so the numbers reproduce
exactly — a failure is a real regression, never sample noise.
- 19 rate floors + 5 residual ceilings, seeded at the currently-measured
values; they only ever tighten (no-widening ethos on an aggregate).
- Calculator-FREE — component floors are the real gate; the end-to-end
SAP/carbon/PE guards stay out (their floor is the separate API-path
calculator workstream).
- Skips with a message when the fixture is absent.
25 parametrized assertions, all green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>