Tier-2 (full national bulk streaming) is deferred. The near-term scale
validation is a Tier-1.5: a few-thousand-cert anonymised corpus stored in
S3 (too large to commit, far more stable than the 36-target gate fixture),
pulled to a temp dir and run through the same load_corpus +
evaluate_component_accuracy. Reuses the committed-fixture machinery wholesale
— only the data source differs. One scorer, three data sources (committed
fixture / S3 corpus / bulk stream).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Records the grilling-session decisions amending ADR-0029's validation:
- Source cohort keeps all cert vintages (components are agnostic of the SAP
methodology that rated them); only the held-out validation TARGET is
restricted to SAP 10.2. Amends ADR-0029 decision 5 ("pre-SAP10 dropped").
- Component Accuracy (predicted vs API actual components) is the primary,
calculator-independent signal. calc(predicted) vs calc(actual) rejected
(circular ground truth, hides calculator error); neighbour-mean-lodged-SAP
baseline rejected (mixes SAP versions). calc(predicted) vs API-lodged
SAP/carbon/PE kept as a secondary, calculator-floored guard.
- Two tiers: committed anonymized fixture (ratcheting CI gate) + bulk-export
national battle-test on harness/epc_bulk.py + harness/cohort.py, emitting
accuracy + a failure taxonomy, re-baselining the gate floors.
CONTEXT.md: Comparable Properties corrected to all-vintage source; new
Component Accuracy term. ADR-0029 Validation section marked superseded.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>