Commit graph

8 commits

Author SHA1 Message Date
Khalim Conn-Kowlessar
41b5ce5057 refactor(epc-prediction): name-keyed categorical_hits for Component Accuracy (ADR-0030)
ADR-0030 commits Component Accuracy to ~19 categorical components (5 today
+ 8 heating + glazing/renewables). Flat *_correct dataclass fields don't
scale — each needs manual runner wiring. Collapse them into a single
`categorical_hits: dict[str, Optional[bool]]` keyed by component name, which
also matches the runner's name-keyed aggregation (now generic: it tallies
whatever components the comparison reports). No behaviour change; the
classification rates are identical (wall n 578->575 is the 3 certs whose
actual wall is None, now correctly counted as not-applicable via _classify).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 08:50:34 +00:00
Khalim Conn-Kowlessar
fa11df56c2 fix(epc-prediction): dedupe re-lodgements + leak-free leave-one-out (ADR-0029)
The register lists every historical lodgement, so a postcode cohort
contains the same physical address many times (LS61AA: 15 certs / 11
addresses; NG71AA: 15 / 9 — "FLAT 3" appears 3x in each). Two
consequences:

  - Production: a re-lodged neighbour was counting up to 3x towards the
    cohort mode. select_comparables now dedupes candidates to the latest
    cert per address (one comparable per real neighbour) — Comparable
    gains address + registration_date (the register metadata its docstring
    already anticipated, read straight off the cached payload).

  - Validation: leave-one-out leaked — predicting a flat from a near-
    identical re-lodgement of itself. The harness now holds out a whole
    address (excludes every sibling cert) and evaluates on the latest cert
    per address (the best ground truth).

Removing the leak gives the honest numbers (19 distinct addresses):
  wall_construction      93.1% -> 89.5%
  construction_age_band  65.5% -> 52.6%
  roof_construction      79.3% -> 68.4%
  floor_area mean|.|     37.9  -> 52.6 m2
The earlier figures were inflated by self-leakage; these are the real
accuracy to beat.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 00:40:23 +00:00
Khalim Conn-Kowlessar
54a57363f8 feat(epc-prediction): cohort-mode the roof/floor/insulation/age categoricals (ADR-0029)
Only main wall_construction was set to the cohort mode; the other
homogeneous categoricals (wall insulation, construction age band, roof
construction, floor construction) were left as template-copied, so one
median-size template's quirks set them. Apply the same cohort-mode
mechanism to all of them per ADR-0029 decision 4 — the template still
supplies geometry, only the categorical codes move to the mode.

Verified mode beats (or ties) template-copy per categorical before
applying. Smoke corpus (29 leave-one-out) classification rates:
  construction_age_band  55.2% -> 65.5%
  roof_construction      72.4% -> 79.3%
  floor_construction     46.2% -> 84.6%
  wall_insulation_type   93.1% (tie — already template-strong)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 00:31:16 +00:00
Khalim Conn-Kowlessar
ed96df9315 feat(epc-prediction): classify roof/floor/insulation/age categoricals (ADR-0029)
The comparison only scored main wall_construction; everything else the
predictor produces (by template-copy) went unmeasured. Extend
compare_prediction to the rest of the ADR-0029 homogeneous categoricals —
wall insulation type, construction age band, roof construction, floor
construction — and aggregate per-categorical classification rates in the
runner. A categorical hit is "not applicable" (None, excluded from the
denominator) when the actual lodges no value, so absent-roof flats don't
score free wins.

Smoke corpus (29 leave-one-out, all but wall are template-copied today):
  wall_construction      93.1%
  wall_insulation_type   93.1%
  construction_age_band  55.2%   <- loud; candidate for cohort-mode
  roof_construction      72.4%
  floor_construction     46.2%   (n=13)

These numbers drive the next slice (extend cohort-mode coverage).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 00:10:56 +00:00
Khalim Conn-Kowlessar
4fa20ae76b fix(epc-prediction): size-representative template selection (ADR-0029)
Template (the comparable whose structure/geometry is copied wholesale)
was members[0] — an arbitrary draw from the API search order. With floor
area varying widely within a property_type cohort (NG71AA houses span
51-340 m2), this made the copied geometry noisy and systematically large.

Pick the member whose floor area is closest to the cohort median instead,
implementing ADR-0029 decision 4's unimplemented "closest on size"
criterion while keeping the structure coherent (it is still one real
property, so floor dims / windows / parts stay internally consistent for
the calculator).

Smoke corpus (29 leave-one-out predictions):
  floor_area  mean|.| 68.0 -> 37.9 m2  (bias +46.8 -> -3.9)
  window_area mean|.| 11.1 -> 7.3 m2
  parts       mean|.| 1.00 -> 0.38
  SAP |pred-calc - calc(actual)| MAE 7.19 -> 4.86

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 00:05:40 +00:00
Khalim Conn-Kowlessar
f3ad6343a3 feat(epc-prediction): leave-one-out validation harness (ADR-0029)
Pure compare_prediction (TDD): wall-construction classification hit + signed
residuals on floor area, window count, total window area, building-parts count.
Plus validate_epc_prediction.py (IO plumbing): drops each cert from its postcode
cohort, predicts from the rest on guaranteed inputs only, aggregates the metrics,
and reports SAP three ways (pred-calc vs lodged / vs calc-on-actual / vs the
neighbour-mean baseline). Smoke run: wall 90.9%, floor-area mean|·| 42.6 m2 (a
real signal — template-copied floor area is noisy), SAP pred-calc edges baseline.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 23:55:05 +00:00
Khalim Conn-Kowlessar
5e6d2cff16 feat(epc-prediction): EpcPrediction hybrid synthesis (ADR-0029)
predict() copies a representative template comparable's structure (coherent for
the calculator), overrides the homogeneous categorical with the cohort mode
(robust to an atypical template), then applies known Landlord Overrides on top
(a known value wins over the estimate). Proven on wall construction; roof/floor/
insulation/age extend on the same mode+override mechanism, driven next by the
validation harness metrics.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 23:50:07 +00:00
Khalim Conn-Kowlessar
bf6b6fac17 feat(epc-prediction): Comparable Properties selection ladder (ADR-0029)
Pure-domain select_comparables: property type is an always-hard filter; built
form and known Landlord Overrides (e.g. solid brick) are conditioning filters on
the filter-then-relax ladder — applied while >= minimum_cohort survive, relaxed
otherwise (the mixed-street border case degrades gracefully). PredictionTarget
(known inputs) + Comparable (epc + register metadata) + ComparableProperties
(selected cohort). Weighting (recency x similarity) follows in the synthesis slice.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 23:44:57 +00:00