Commit graph

14 commits

Author SHA1 Message Date
Khalim Conn-Kowlessar
ffaedd8d14 feat(epc-prediction): ±1-band age scoring + window_count cosmetic (#1222)
Measurement honesty so we optimise SAP-relevant accuracy, not SAP-neutral
misses (ADR-0030 Component Accuracy):
- Add construction_age_band_pm1: an exact-or-adjacent-band hit. Adjacent
  RdSAP age bands carry near-identical U-values, so an off-by-one is
  ~SAP-neutral. Full corpus: exact 78.5% but ±1-band 91.7% (fixture
  63.9% -> 83.3%) — most age misses are adjacent.
- Drop window_count from the gate's residual ceilings (cosmetic): the
  predicted picture clusters at a mapper-default 4 windows vs actuals 1-21,
  but total_window_area (the SAP-relevant signal) stays tight at ~3.4 m2.

Gate: + construction_age_band_pm1 floor 0.8333; window_count no longer gated.

Closes #1222

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 10:01:20 +00:00
Khalim Conn-Kowlessar
a5b7310911 feat(epc-prediction): recency-weighted mode for roof insulation (ADR-0029/0030)
Investigated recency-weighting (weight cohort votes by an exponential decay
in cert age). Key finding: it must be SELECTIVE. On the validation corpus it
HURTS permanent categoricals (wall 91.2->89.5, age 78.5->75.7 — discards
still-valid data) but clearly HELPS time-varying ones, where a recent
neighbour reflects the current physical state:
  roof_insulation_thickness  56.7 -> 60.7%  corpus   (+4pp)
                             29.4 -> 41.2%  fixture  (+12pp)

So apply a recency-weighted mode only to roof_insulation_thickness (loft
top-ups happen over time); keep the plain mode for permanent categoricals.
tau = 4yr (~2.8yr half-life); falls back to plain mode when no registration
dates are lodged. Gate floor ratcheted 0.2941 -> 0.4118.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 09:45:22 +00:00
Khalim Conn-Kowlessar
9dd23477ac feat(epc-prediction): cohort-mode roof + floor insulation (ADR-0030)
These independent fabric categoricals were template-copied; mode them like
the construction categoricals. Verified mode beats template before applying.
Big fixture win on roof insulation thickness (doubled), floor insulation
neutral-to-positive:
  roof_insulation_thickness  14.7% -> 29.4%  (gate floor ratcheted up)
  floor_insulation           90.6% (unchanged on the fixture)

Glazing type was tried too (+1.6pp on the 40-postcode corpus) but REGRESSED
the 36-target fixture (0.50 -> 0.44) — the gate caught it. Glazing moding is
marginal/noisy, so it's left on the template; revisit with a larger corpus.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 09:37:45 +00:00
Khalim Conn-Kowlessar
027ee1fba3 refactor(epc-prediction): extract shared leave-one-out scorer + corpus loader (ADR-0030)
"One scorer, two harnesses" (ADR-0030): the committed gate, the local script,
and the future battle-test must run the *same* scoring. Extract it:

- domain/epc_prediction/validation.py — `iter_predictions` (the single
  leave-one-out orchestration: latest-per-address hold-out, SAP-10.2 target
  filter, all-vintage source) + `evaluate_component_accuracy` (calculator-free
  ComponentAccuracy aggregation, the primary signal). Unit-tested.
- harness/epc_prediction_corpus.py — `load_corpus(dir)` IO: corpus dir ->
  Comparable cohorts (maps payloads, carries address + registration_date).

validate_epc_prediction.py now just loads + calls the scorer for the component
section and iterates iter_predictions for the calculator-floored end-to-end.
Identical numbers (181 targets, SAP MAE 6.34) — behaviour-preserving.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 09:12:08 +00:00
Khalim Conn-Kowlessar
275a30a825 feat(epc-prediction): complete component coverage — fabric/glazing/renewables/doors (ADR-0030)
Finish the ADR-0030 Component Accuracy set: roof insulation thickness,
floor insulation, room-in-roof presence, modal glazing type, PV presence,
solar water heating (categoricals) + door count (residual). Presence flags
(room-in-roof, PV, solar) are always-applicable — predicting absence when
present is a real miss.

Template-copied baseline (40-postcode corpus), newly visible:
  floor_insulation         94.0%   solar_water_heating  99.7%
  has_pv                   98.6%   has_room_in_roof     91.9%
  modal_glazing_type       59.0%   <- weak
  roof_insulation_thickness 30.6%  <- weak
  door_count  mean|.| 0.40

compare_prediction now scores 19 categoricals + 5 residuals across every
SAP-load-bearing component group.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 09:00:30 +00:00
Khalim Conn-Kowlessar
cd43c52cf9 feat(epc-prediction): score the heating components (ADR-0030 Component Accuracy)
Heating is the dominant SAP lever (ablating it to actual cut the SAP error
~7 -> ~4.5) yet was entirely unscored. Add the heating group to
compare_prediction's categorical_hits: main fuel / category / control (off
the primary MainHeatingDetail), water-heating fuel / code, has-cylinder,
cylinder insulation, secondary heating (off SapHeating).

Template-copied baseline on the 40-postcode corpus (no predictor change
yet — this just makes the signal visible):
  heating_main_fuel        93.4%
  heating_main_category    92.7%
  water_heating_fuel/code  91.7% / 92.4%
  heating_main_control     62.1%   <- weak
  has_hot_water_cylinder   78.5%
  cylinder_insulation_type 35.8% (n=120)   <- weak
  secondary_heating_type   16.8% (n=125)   <- weak

Fuel/category predict well from the template; controls, cylinder, and
secondary heating are poor and now drive the next predictor slices.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 08:53:15 +00:00
Khalim Conn-Kowlessar
41b5ce5057 refactor(epc-prediction): name-keyed categorical_hits for Component Accuracy (ADR-0030)
ADR-0030 commits Component Accuracy to ~19 categorical components (5 today
+ 8 heating + glazing/renewables). Flat *_correct dataclass fields don't
scale — each needs manual runner wiring. Collapse them into a single
`categorical_hits: dict[str, Optional[bool]]` keyed by component name, which
also matches the runner's name-keyed aggregation (now generic: it tallies
whatever components the comparison reports). No behaviour change; the
classification rates are identical (wall n 578->575 is the 3 certs whose
actual wall is None, now correctly counted as not-applicable via _classify).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 08:50:34 +00:00
Khalim Conn-Kowlessar
fa11df56c2 fix(epc-prediction): dedupe re-lodgements + leak-free leave-one-out (ADR-0029)
The register lists every historical lodgement, so a postcode cohort
contains the same physical address many times (LS61AA: 15 certs / 11
addresses; NG71AA: 15 / 9 — "FLAT 3" appears 3x in each). Two
consequences:

  - Production: a re-lodged neighbour was counting up to 3x towards the
    cohort mode. select_comparables now dedupes candidates to the latest
    cert per address (one comparable per real neighbour) — Comparable
    gains address + registration_date (the register metadata its docstring
    already anticipated, read straight off the cached payload).

  - Validation: leave-one-out leaked — predicting a flat from a near-
    identical re-lodgement of itself. The harness now holds out a whole
    address (excludes every sibling cert) and evaluates on the latest cert
    per address (the best ground truth).

Removing the leak gives the honest numbers (19 distinct addresses):
  wall_construction      93.1% -> 89.5%
  construction_age_band  65.5% -> 52.6%
  roof_construction      79.3% -> 68.4%
  floor_area mean|.|     37.9  -> 52.6 m2
The earlier figures were inflated by self-leakage; these are the real
accuracy to beat.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 00:40:23 +00:00
Khalim Conn-Kowlessar
54a57363f8 feat(epc-prediction): cohort-mode the roof/floor/insulation/age categoricals (ADR-0029)
Only main wall_construction was set to the cohort mode; the other
homogeneous categoricals (wall insulation, construction age band, roof
construction, floor construction) were left as template-copied, so one
median-size template's quirks set them. Apply the same cohort-mode
mechanism to all of them per ADR-0029 decision 4 — the template still
supplies geometry, only the categorical codes move to the mode.

Verified mode beats (or ties) template-copy per categorical before
applying. Smoke corpus (29 leave-one-out) classification rates:
  construction_age_band  55.2% -> 65.5%
  roof_construction      72.4% -> 79.3%
  floor_construction     46.2% -> 84.6%
  wall_insulation_type   93.1% (tie — already template-strong)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 00:31:16 +00:00
Khalim Conn-Kowlessar
ed96df9315 feat(epc-prediction): classify roof/floor/insulation/age categoricals (ADR-0029)
The comparison only scored main wall_construction; everything else the
predictor produces (by template-copy) went unmeasured. Extend
compare_prediction to the rest of the ADR-0029 homogeneous categoricals —
wall insulation type, construction age band, roof construction, floor
construction — and aggregate per-categorical classification rates in the
runner. A categorical hit is "not applicable" (None, excluded from the
denominator) when the actual lodges no value, so absent-roof flats don't
score free wins.

Smoke corpus (29 leave-one-out, all but wall are template-copied today):
  wall_construction      93.1%
  wall_insulation_type   93.1%
  construction_age_band  55.2%   <- loud; candidate for cohort-mode
  roof_construction      72.4%
  floor_construction     46.2%   (n=13)

These numbers drive the next slice (extend cohort-mode coverage).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 00:10:56 +00:00
Khalim Conn-Kowlessar
4fa20ae76b fix(epc-prediction): size-representative template selection (ADR-0029)
Template (the comparable whose structure/geometry is copied wholesale)
was members[0] — an arbitrary draw from the API search order. With floor
area varying widely within a property_type cohort (NG71AA houses span
51-340 m2), this made the copied geometry noisy and systematically large.

Pick the member whose floor area is closest to the cohort median instead,
implementing ADR-0029 decision 4's unimplemented "closest on size"
criterion while keeping the structure coherent (it is still one real
property, so floor dims / windows / parts stay internally consistent for
the calculator).

Smoke corpus (29 leave-one-out predictions):
  floor_area  mean|.| 68.0 -> 37.9 m2  (bias +46.8 -> -3.9)
  window_area mean|.| 11.1 -> 7.3 m2
  parts       mean|.| 1.00 -> 0.38
  SAP |pred-calc - calc(actual)| MAE 7.19 -> 4.86

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 00:05:40 +00:00
Khalim Conn-Kowlessar
f3ad6343a3 feat(epc-prediction): leave-one-out validation harness (ADR-0029)
Pure compare_prediction (TDD): wall-construction classification hit + signed
residuals on floor area, window count, total window area, building-parts count.
Plus validate_epc_prediction.py (IO plumbing): drops each cert from its postcode
cohort, predicts from the rest on guaranteed inputs only, aggregates the metrics,
and reports SAP three ways (pred-calc vs lodged / vs calc-on-actual / vs the
neighbour-mean baseline). Smoke run: wall 90.9%, floor-area mean|·| 42.6 m2 (a
real signal — template-copied floor area is noisy), SAP pred-calc edges baseline.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 23:55:05 +00:00
Khalim Conn-Kowlessar
5e6d2cff16 feat(epc-prediction): EpcPrediction hybrid synthesis (ADR-0029)
predict() copies a representative template comparable's structure (coherent for
the calculator), overrides the homogeneous categorical with the cohort mode
(robust to an atypical template), then applies known Landlord Overrides on top
(a known value wins over the estimate). Proven on wall construction; roof/floor/
insulation/age extend on the same mode+override mechanism, driven next by the
validation harness metrics.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 23:50:07 +00:00
Khalim Conn-Kowlessar
bf6b6fac17 feat(epc-prediction): Comparable Properties selection ladder (ADR-0029)
Pure-domain select_comparables: property type is an always-hard filter; built
form and known Landlord Overrides (e.g. solid brick) are conditioning filters on
the filter-then-relax ladder — applied while >= minimum_cohort survive, relaxed
otherwise (the mixed-street border case degrades gracefully). PredictionTarget
(known inputs) + Comparable (epc + register metadata) + ComparableProperties
(selected cohort). Weighting (recency x similarity) follows in the synthesis slice.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 23:44:57 +00:00