Model/docs/adr/0060-bulk-document-download-package-model.md
Khalim Conn-Kowlessar 3c9b75cbce Hand-pick selection by landlord_property_id, not property.id
uploaded_files has no property_id — matching is on landlord_property_id — so the
property_ids path was pointless indirection (look property.id up in property just
to translate back to landlord_property_id). The FE holds landlord_property_id per
row anyway.

task.inputs hand-pick key is now landlord_property_ids: str[] (was property_ids:
int[]). resolve_selection unions two landlord_property_id sources — project_codes
expanded via hubspot_deal_data, and the hand-picked ids taken as given — and no
longer queries the property table at all. Trigger-only change; the domain plan,
orchestrator and matching already work in landlord_property_id.

ADR-0060, CONTEXT.md and the request schema updated to match.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-08 12:04:35 +00:00

7.2 KiB

status
accepted (builds on ADR-0055, ADR-0059)

Bulk Document Download builds one capped, best-effort Download Package per request

Users need to pull the documents held in uploaded_files for many properties at once — per property, the latest file of each Document Type — as a single archive, without clicking through the UI file by file. The request is initiated in the front end, can span a hand-picked set of properties or one or more whole HubSpot projects (by project code), and finishes with a link emailed to the requester (ADR-0059).

uploaded_files links to a property by uprn or landlord_property_id (both nullable) and has no property_id column; file_type (the Document Type) is itself nullable; and files live across arbitrary buckets (s3_file_bucket per row).

Decision

A request produces exactly one Download Package: a ZIP of one folder per property, each holding the latest file of each Document Type — built best-effort, size-capped at the trigger, on the app-owned-task + attach-mode lane (ADR-0055).

  • Trigger & lifecycle (ADR-0055). The front end creates the app-owned tasks row and writes the selection config{project_codes?: str[], landlord_property_ids?: str[], portfolio_id?: int} — into tasks.inputs (TEXT/JSON, FE-owned), then calls the FastAPI route with only the task_id. A large selection therefore never travels in an HTTP body. The route reads the selection from tasks.inputs and resolves it to the distinct landlord_property_id set — the union of every property in the named HubSpot project_codes (read from hubspot_deal_data, where the project↔property grain lives) and the hand-picked landlord_property_ids. landlord_property_id is the selection key throughout because uploaded_files is matched on it (it has no property_id). At least one of the two must be provided; portfolio_id is optional and used only to name the package. The route then caps the set, resolves the recipient email from the authenticated user (ADR-0059), pins the resolved recipe (landlord_property_ids, recipient_email, package_name) onto one pre-created sub_task's inputs, and drops one SQS message (task_id, sub_task_id). The new applications/bulk_document_download Lambda runs in attach mode, reading that recipe from the sub_task; TaskOrchestrator owns status + roll-up. (task.inputs = the FE's selection; sub_task.inputs = the backend's resolved recipe.)
  • One package = one sub_task, streamed. No fan-out. A Download Package can be several GB, so it is never held whole in memory: each document is read from its own s3_file_bucket, written into an on-disk ZIP in /tmp, and released; the finished archive is multipart-uploaded from disk (S3Client.upload_file). The Lambda's ephemeral storage (/tmp) is raised accordingly (up to 10 GB) — see Consequences. One artifact, one URL, one email.
  • Selection cap at the route. Reject > N properties synchronously with a legible "narrow your selection" error (property count, not a document/byte query — cheap, no S3 on the hot path). A /tmp size budget inside the Lambda is a backstop that fails the sub_task with a clear reason if a pathological selection still overflows. N starts at a conservative, tunable value.
  • Matching on landlord_property_id. Files are gathered by landlord_property_id; the missing property_id on uploaded_files is a known future gap, out of scope here. Property display info (the folder name) is enriched from the hubspot deals data (repositories/hubspot_deals/).
  • Layout. One folder per property named by human-readable address (hubspot-enriched), with landlord_property_id appended for uniqueness; inside, one file per Document Type = the newest by s3_upload_timestamp. Rows with a null Document Type are skipped and listed in the run output.
  • Best-effort failure contract. Build from whatever resolves. Skipped properties (no documents) and skipped null-type files are recorded in sub_task.outputs. The run fails only on an infrastructure error (S3/DB/zip/upload/email) or when the whole selection yields zero documents. The email carries an "N properties, M documents, X skipped" summary.
  • Delivery — both channels. The presigned URL (60-minute expiry) and the skip-summary are written to sub_task.outputs (the FE already polls task status and can show the link) and emailed (ADR-0059).
  • A dedicated exports bucket. Packages are written to a separate DOCUMENT_EXPORTS_BUCKET, not DATA_BUCKET — the generated archives are a distinct class of artifact (transient, user-facing, wide-read source but single-writer) and keeping them out of the shared data bucket keeps IAM and any future retention policy cleanly scoped to exports. No lifecycle rule is imposed on DATA_BUCKET. Retention on the exports bucket is an open decision (the packages are transient, so an expiry is likely warranted, but it is left to a follow-up rather than baked in here).

Considered options

  • Fan out into per-batch sub_tasks / multiple ZIPs (like the modelling run). Rejected: the user would receive N links / N emails and "the download" would stop being one file; the capped-single-package model keeps the artifact and the notification singular.
  • Cap on resolved document count / bytes. Rejected for v1: it adds synchronous DB + S3 HEAD work to the trigger and couples the route to the packaging logic; a property-count cap is cheap and predictable.
  • Strict "all-or-nothing" packaging. Rejected: one unreadable row or a property with no documents would deny the entire package; best-effort + reporting matches "include documents and properties where they exist".
  • Match on uprn. Deferred: landlord_property_id is the agreed key and hubspot enrichment is keyed to the deal/property; revisit if/when the property_id gap on uploaded_files is closed.

Consequences

  • New DDD pieces: applications/bulk_document_download/ (thin handler + trigger body), orchestration/bulk_document_download_orchestrator.py, packaging rules in domain/, a landlord_property_id-keyed "latest-per-Document-Type" query on the uploaded-file repository, a generate_presigned_url + multipart upload on S3Client, and the email port/adapter from ADR-0059. The only backend/ touch is the trigger route.
  • The stored, pinned property_id set makes a run reproducible even if the portfolio changes between trigger and execution.
  • N, the Lambda's ephemeral-storage (/tmp) size, and its memory are tunable knobs, not load-bearing invariants. /tmp must be raised beyond the 512 MB default (up to 10 GB) to hold a multi-GB archive — a new terraform knob on the shared Lambda modules, backward-compatible (every other Lambda keeps 512 MB).
  • A new S3 bucket (DOCUMENT_EXPORTS_BUCKET) is provisioned for the packages, with its own IAM (single-writer, presign-read). Its retention policy is a deliberate follow-up, not set here.
  • Closing the uploaded_files.property_id gap later would let matching move off landlord_property_id without changing the package model.