assessment-model/docs/runbooks/recommendation-denormalization-backfill.md
Khalim Conn-Kowlessar 9bf36ad4b6 perf(backfill): faster, resumable recommendation denormalization + ops tooling
Overhaul the recommendation plan_id/material_* backfill for the 21M-row prod
table (see docs/runbooks/recommendation-denormalization-backfill.md):

- Single combined pass: set plan_id + all four material_* columns in one write
  per row (was two full passes, rewriting every row twice). Guard + COALESCE keep
  it idempotent and resumable — done rows are never rewritten.
- Session tuning: synchronous_commit=off, tunable work_mem / maintenance_work_mem.
- Range + phase controls (BACKFILL_START_ID/END_ID/SKIP_CARDINALITY/ONLY_FINALIZE)
  to run disjoint id-range workers in parallel when the disk isn't the bottleneck.
- Percent-complete + ETA in the progress log.
- Finalize builds the two new indexes plain by default (faster in a quiet
  window) with a CONCURRENTLY toggle.

The real speed-up in prod was dropping the 3 secondary indexes for the backfill
then rebuilding — the table is packed (fillfactor 100) so updates can't go HOT
and otherwise maintain every index per row. Index drop/rebuild + vacuum are kept
as standalone SQL so ops controls exactly when the app loses/regains them:

- sql/drop-recommendation-backfill-indexes.sql
- sql/rebuild-recommendation-backfill-indexes.sql (1GB maintenance_work_mem — safe
  on the 4GB instance)
- sql/vacuum-recommendation.sql

run-sql.ts runs .sql files statement-by-statement, autocommitted (so VACUUM /
CREATE INDEX CONCURRENTLY work), for environments without the psql CLI. Wired up
as npm run db:run-sql and backfill:recommendation-restore.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 10:00:46 +00:00

5.9 KiB

Runbook: recommendation denormalisation backfill

Populates recommendation.plan_id and recommendation.material_* from the plan_recommendations / recommendation_materials join tables, as the first step toward making those columns the source of truth and dropping the join tables. See ADR 0001 for why this runs outside drizzle.

The table is ~21m rows. This runbook is the fast procedure.

Why it's slow by default, and the lever

recommendation is packed (fillfactor 100), so an UPDATE can't do a HOT update — Postgres writes a new row version and inserts into every index on the table for all ~21m rows. That index write-amplification (and its WAL) is what drains IO.

The lever: drop the secondary indexes, backfill, rebuild them. With them gone, each updated row only maintains the primary key (~4 index inserts → 1). Combined with the single-pass rewrite (plan_id + material in one UPDATE, not two) and synchronous_commit = off, this turns a ~1-day run into a few hours.

DROP INDEX is near-instant — the cost people mean is the rebuild, which is why we do it once, at the end, in a low-usage window.

What the app needs to keep working

Only the 3 secondary indexes. The half-filled plan_id/material_id columns are NULL-tolerant (the read paths already handle NULLs), and the new plan_id/material_id indexes aren't used by anything yet. So at any stopping point, rebuild the 3 secondary indexes and the app is back to today's behaviour — a 50%-done backfill is invisible to it.

Prerequisites

  • npm run migration:migrate has applied 0222/0224 (the column + NOT VALID FK adds).
  • .env.local points at the target DB (DB_HOST/DB_PORT/DB_USERNAME/DB_PASSWORD/DB_NAME).
  • A low-usage window.

Files

  • Script: src/app/db/backfill-recommendation-denormalization.ts (npm run backfill:recommendation-denormalization)
  • Drop: src/app/db/sql/drop-recommendation-backfill-indexes.sql
  • Rebuild:src/app/db/sql/rebuild-recommendation-backfill-indexes.sql

Procedure A — one window (a few hours, start to finish)

# 1. Drop the secondary indexes (app reads degrade to seq scans until step 4)
psql "$DATABASE_URL" -f src/app/db/sql/drop-recommendation-backfill-indexes.sql

# 2. Run the full backfill + finalize (builds the 2 new indexes, validates FKs)
npm run backfill:recommendation-denormalization

# 3. Rebuild the secondary indexes — app back to normal
psql "$DATABASE_URL" -f src/app/db/sql/rebuild-recommendation-backfill-indexes.sql

# 4. Reclaim the ~2x bloat in-place updates leave behind, refresh stats
psql "$DATABASE_URL" -c "VACUUM (ANALYZE) recommendation;"

DATABASE_URL is illustrative — use whatever psql connection matches .env.local.


Procedure B — split across windows (50% now, finish later)

Each window is self-contained; between windows the app is 100% normal.

Every window except the last:

psql "$DATABASE_URL" -f src/app/db/sql/drop-recommendation-backfill-indexes.sql
BACKFILL_FINALIZE=false npm run backfill:recommendation-denormalization   # Ctrl-C any time; it's resumable
psql "$DATABASE_URL" -f src/app/db/sql/rebuild-recommendation-backfill-indexes.sql

BACKFILL_FINALIZE=false skips the new-index build + FK validate — there's no point building the plan_id/material_id indexes until the data is all in.

The final window (when the per-batch logs / the NULL summary show plan_id is down to ~0 and only genuinely material-less rows remain):

psql "$DATABASE_URL" -f src/app/db/sql/drop-recommendation-backfill-indexes.sql
npm run backfill:recommendation-denormalization       # FINALIZE defaults on: builds new indexes, validates FKs
psql "$DATABASE_URL" -f src/app/db/sql/rebuild-recommendation-backfill-indexes.sql
psql "$DATABASE_URL" -c "VACUUM (ANALYZE) recommendation;"

Cost of splitting: you rebuild the secondary indexes once per window. Two of the three are partial (tiny), so each restore is effectively one full-size index build (recommendation_property_id_idx, a few minutes).

Interrupting mid-run

Ctrl-C is safe. The script commits per batch, so completed batches persist. Re-run to continue — the IS NULL guards skip every row already done. If you're worried about app performance while stopped, rebuild the secondary indexes first.


Alternative — no index juggling (simpler, slower)

If you'd rather not touch indexes at all, just run the tuned single pass with the indexes left in place and stop/resume freely:

BACKFILL_FINALIZE=false npm run backfill:recommendation-denormalization

Slower per row (index write-amplification stays), but the app is always fully indexed and there's zero drop/rebuild ceremony. Finalize on the last run.


Tunables (env)

Var Default Notes
BACKFILL_BATCH_SIZE 100000 rows scanned per committed batch
BACKFILL_SLEEP_MS 0 inter-batch pause; raise to throttle IO if not in a quiet window
BACKFILL_SYNC_COMMIT off on to keep synchronous_commit (safe to leave off — run is resumable)
BACKFILL_MAINTENANCE_WORK_MEM 1GB memory for the finalize index builds
BACKFILL_WORK_MEM 128MB memory for the batch joins
BACKFILL_INDEX_CONCURRENTLY (off) true builds the new indexes CONCURRENTLY (non-blocking, slower)
BACKFILL_FINALIZE (on) false skips new-index build + FK validate (partial windows)

Verifying completion

The script prints remaining NULLs at the end:

  • plan_id NULL → should trend to ~0 (investigate any residue: recommendations with no plan_recommendations row).
  • material_id NULL → expected for recommendations that have no material; not an error.

After the backfill is fully done

A later migration can SET NOT NULL on plan_id and drop the join tables, gated on this backfill having completed with zero unexpected NULLs. Out of scope here.