`plan-stops-short-of-goal` fires on anything below the goal band — 322/338
on portfolio 814, mostly flats a single band short of B, which drowns the
real hits. Add a sharper, band-keyed check that isolates the genuinely-
stuck tail: a plan that leaves the dwelling still rated E/F/G with the band
UNCHANGED from baseline on an unlimited-budget scenario — the "E->E, no big
movement" worklist. On portfolio 814 / scenario 1271 it flags just 4 (pids
742121, 742210, 742265, 742347) vs 322. No SAP-points threshold; the
strictly-worse case is left to plan-below-baseline-band so the two
partition cleanly.
Also refresh the find-weird-recommendations skill's Phase 4 catalogue:
register the new check in Phase 1, correct the stale 742121 note (it is an
electric maisonette whose HHRSH IS offered but scores -6.3 SAP, so the
Optimiser correctly drops it — not the old mains-gas story), and add the
742265 community-flat findings (HHRSH/ASHP gated out by heat-network
topology; vaulted-ceiling roof with ND insulation).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A focused sibling to audit-ara-portfolio: that skill audits baselines/plans/SAP;
this one audits the *recommendations themselves* — why a measure was or wasn't
offered. Motivated by the portfolio-814 review (Khalim's HHRSH-on-community-
heating, missing-HHRSH, missing-secondary-heating-removal, and a neighbour split).
Adds:
- .claude/skills/find-weird-recommendations/SKILL.md — scan -> neighbour scan ->
live re-model deep-dive -> root-cause -> codify, with a seeded known-bug
catalogue and the query-safety rules inherited from audit-ara-portfolio.
- scripts/audit/anomalies.py: new `plan-stops-short-of-goal` HIGH check — the
default plan ends below the goal band on an unlimited-budget scenario (the
deterministic worklist for "why didn't this get recommended X"). Adds
scenario_budget to the bundle/query so budget-capped scenarios are excluded.
- scripts/audit/neighbour_divergence.py: groups a portfolio by (postcode,
property_type, built_form) and flags effective-SAP outliers vs the cohort
median. Never touches the 26m-row recommendation table, so it is safe
portfolio-wide.
- Tests for both (12 passing).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The `recommendation` table (~26m rows) has no index on `plan_id`, so any query
reaching it via `plan_id` — including the audit's own rollup — seq-scans the
whole table and saturates the shared DB (it blocked the DB during the
portfolio-796 audit). Make the audit safe-by-default:
- statement_timeout (120s) on the audit connection — a hard ceiling so a bad
plan aborts instead of hammering the DB.
- The recommendation rollup (the two solar checks) is now opt-in via
--with-recommendations, and EXPLAIN-gated: it refuses to run (raising
RecommendationScanError) when the plan contains a Seq Scan on recommendation,
which it does on any large portfolio until idx_recommendation_plan_id exists.
- SKILL.md documents the plan_id-no-index trap, the reach-via-property_id /
EXPLAIN-first / confirm-with-user rules, and the index as the real fix.
Verified on 796/1268: default run is bounded and completes (2,952 anomalies over
31,919 properties); --with-recommendations aborts pre-scan portfolio-wide but is
allowed for a single property.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Portfolio-796/scenario-1268 audit found the bulk of MEDIUM/HIGH anomalies
(already-meets-goal-with-works, zero-works-post-differs, plan-score-below-
baseline, and the non-fuel-switch slice of negative-bill-savings) trace to one
root cause: the persisted default plan is stale relative to the live model, so
it resolves on re-model rather than a code change. Confirmed via
run_modelling_e2e on three samples (stored vs live plans differ wholesale).
No new check added — the staleness is already surfaced by the existing checks
and addressed by the override-aware-rebaseline / persistence-fidelity work, so
a new check would only re-flag known divergence. Instead record the expectation
and the "audit the default plan only" rule in the skill Notes so the next
reviewer starts ahead. References kept to durable docs/adr (no PR numbers).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add Phase 6 (self-improve) to audit-ara-portfolio: when a run confirms a
novel systematic problem, codify it as a check — gated on systematic (>=5
props, root-caused), not-already-covered, and /grill-me-pressure-tested.
Each check records provenance (motivating cause + example properties) so the
registry stays sharp and compounds every run.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Script takes an optional --scenario to restrict to one scenario's plans. New
skill drives the full loop: run the deterministic scan, review groups,
deep-dive samples via run_modelling_e2e, characterise sub-classes, and
cross-reference open PRs/ADRs — then proposes new checks to codify.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Closes the mapper-coverage gaps surfaced by the modelling_e2e prediction-cohort
failures (portfolio 796):
- built_form (SAP-16.0): derive from dwelling_type in _normalize_sap_schema_16_x
(Mid-terrace->4, End-terrace->3, Semi-detached->2, Detached->1; flats->modal 4).
ML-only field (SAP calc never reads it) so SAP- and gate-neutral. 5 flat certs
that omitted built_form now map.
- photovoltaic_supply as a measured-array LIST: routed all pre-21 RdSAP mappers
(17.0/17.1/18.0/19.0/20.0.0) through _map_schema_21_pv, whose list branch is now
dict-tolerant (_pv_array_field reads dict OR dataclass). They capture the PV
arrays like 21.0.x instead of raising "'list' object has no attribute
none_or_no_details" and sinking the whole cohort.
- windows-as-dict (16.x): handled in the normalizer (not just windows-as-list).
Genuinely-sparse certs (omit door_count/habitable/glazed_area) remain fail-loud;
the gate-regressing multiple_glazed_proportion default and the recursive
RdSAP-21.0.0 ADR-0028 alignment are left fail-loud + flagged for review (worklist).
+5 regression tests; component-accuracy gate 26/26; 0 new pyright errors.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`.claude/settings.json` was committed to main by mistake — it holds
per-developer permission allow-lists (npx cache paths, /tmp script paths,
and even hardcoded credentials), not shared project config. Mirror the
existing `.claude/settings.local.json` treatment: remove it from the index
and add it to .gitignore so each developer keeps their own local copy.
Claude Code merges settings.json + settings.local.json at runtime, so no
permissions are lost.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Every worklist UPRN now carries schema · engine SAP / lodged · flag. Tally:
64 healthy, 19 MVHR-not-credited (🚩 flag B), 6 heat-pump fuel-39 (🚩 flag A),
4 sparse/NOT MAPPABLE (⛔), 3 Elmhurst-pinned. MVHR is the largest accuracy gap.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- SAP-Schema-16.3: same reduced-field RdSAP shape as 16.2 — generalise the
normaliser to _normalize_sap_schema_16_x and route both 16.2/16.3 through it.
uprn_44012843 maps → SAP 79 (lodged 81).
- SAP-Schema-17.0: structurally identical to the full-SAP 17.1 schema (measured
sap_opening_types), so it parses with the 17.1 dataclass and reuses
from_sap_schema_17_1 with no normalisation. uprn_10023444324 → 80, uprn_
10023444320 → 81.
- Regression tests (16.3 dispatch, 17.0 dispatch) + sap_16_3.json / sap_17_0.json
fixtures; 0 new pyright errors. All 7 e2e UPRNs now map.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
SAP-Schema-16.2 (datatypes/epc/domain/mapper.py):
- 16.2 is structurally an RdSAP-17.1 cert under a different name; add
_normalize_sap_schema_16_2 (field renames + defaults) and dispatch to the
tested from_rdsap_schema_17_1 mapper. uprn_100020933699 maps → SAP 71.
- Honour a "Single glazed" windows description when multiple_glazing_type="ND"
(was defaulting to double) → RdSAP-21 code 5; eng 72→71 (lodged 70).
- 4 regression tests + sap_16_2.json fixture; 0 new pyright errors.
Flat party-wall fix (domain/sap10_calculator/worksheet/heat_transmission.py):
- Full-SAP flats carry flatness in dwelling_type, not property_type, so the
party-wall default fell through to the 0.25 house value instead of the RdSAP
Table-15 flat 0.0. Add _is_flat_or_maisonette_dwelling fallback + regression
test. uprn_10093116529 80→81 (matches the cert's lodged party u_value 0).
Accuracy corpus pins (tests/domain/sap10_calculator/test_real_cert_sap_accuracy.py):
- uprn_10093116543 (SAP-17.1 gas-combi semi): engine 81 (Elmhurst 77; documented
full-SAP→RdSAP residual — measured wall/floor U + PCDB boiler vs RdSAP defaults).
- uprn_10093116529 (SAP-17.1 g/f flat): engine 81 (Elmhurst 78).
devcontainer: add poppler-utils (pdfinfo) for the documents-parser PDF fixtures.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Reduced-field window U: heat_transmission derived the synthesised-window
raw U from u_window(all None) -> the 2.5 placeholder regardless of glazing.
Now routes the (uniform) glazing_type code through u_window (RdSAP Table 24)
so e.g. double pre-2002 reads 2.8, not 2.5. Only the pre-SAP10 reduced-field
path is affected (21.0.1 certs carry per-window U upstream) — the RdSAP-21.0.1
corpus gauge is unchanged at 66.9% within-0.5.
test_real_cert_sap_accuracy: pin uprn_10002468137 (RdSAP-17.1, all-electric
storage heaters) at SAP 61, validated against Elmhurst on identical inputs
(dual off-peak immersion, 110 L cylinder, 2 baths). Our engine reproduces
Elmhurst's fuel cost to the penny; lodged 55 is the old SAP-2012 schema.
Tooling to grow the accuracy corpus:
- scripts/fetch_real_life_epc_sample.py — capture a cert by UPRN into the corpus.
- scripts/compare_epc_paths.py — diff gov-API vs Elmhurst-summary EpcPropertyData
and run both through the engine, localising mapper vs calculator differences.
- skill validate-cert-sap-accuracy — the end-to-end loop (capture -> Elmhurst
inputs -> human builds -> compare -> reconcile -> pin in the test).
- skill epc-to-elmhurst-rdsap-inputs reference: corrected immersion (code 1=dual),
cylinder size (code 2 = Normal/110 L), and bath-count (WWHRS sub-tab) mappings.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>