uploaded_files has no property_id — matching is on landlord_property_id — so the property_ids path was pointless indirection (look property.id up in property just to translate back to landlord_property_id). The FE holds landlord_property_id per row anyway. task.inputs hand-pick key is now landlord_property_ids: str[] (was property_ids: int[]). resolve_selection unions two landlord_property_id sources — project_codes expanded via hubspot_deal_data, and the hand-picked ids taken as given — and no longer queries the property table at all. Trigger-only change; the domain plan, orchestrator and matching already work in landlord_property_id. ADR-0060, CONTEXT.md and the request schema updated to match. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
7.2 KiB
| status |
|---|
| accepted (builds on ADR-0055, ADR-0059) |
Bulk Document Download builds one capped, best-effort Download Package per request
Users need to pull the documents held in uploaded_files for many properties
at once — per property, the latest file of each Document Type — as a single
archive, without clicking through the UI file by file. The request is initiated
in the front end, can span a hand-picked set of properties or one or more whole
HubSpot projects (by project code), and finishes with a link emailed to the
requester (ADR-0059).
uploaded_files links to a property by uprn or landlord_property_id
(both nullable) and has no property_id column; file_type (the Document
Type) is itself nullable; and files live across arbitrary buckets
(s3_file_bucket per row).
Decision
A request produces exactly one Download Package: a ZIP of one folder per property, each holding the latest file of each Document Type — built best-effort, size-capped at the trigger, on the app-owned-task + attach-mode lane (ADR-0055).
- Trigger & lifecycle (ADR-0055). The front end creates the app-owned
tasksrow and writes the selection config —{project_codes?: str[], landlord_property_ids?: str[], portfolio_id?: int}— intotasks.inputs(TEXT/JSON, FE-owned), then calls the FastAPI route with only thetask_id. A large selection therefore never travels in an HTTP body. The route reads the selection fromtasks.inputsand resolves it to the distinctlandlord_property_idset — the union of every property in the named HubSpotproject_codes(read fromhubspot_deal_data, where the project↔property grain lives) and the hand-pickedlandlord_property_ids.landlord_property_idis the selection key throughout becauseuploaded_filesis matched on it (it has noproperty_id). At least one of the two must be provided;portfolio_idis optional and used only to name the package. The route then caps the set, resolves the recipient email from the authenticated user (ADR-0059), pins the resolved recipe (landlord_property_ids,recipient_email,package_name) onto one pre-createdsub_task'sinputs, and drops one SQS message (task_id,sub_task_id). The newapplications/bulk_document_downloadLambda runs in attach mode, reading that recipe from the sub_task;TaskOrchestratorowns status + roll-up. (task.inputs= the FE's selection;sub_task.inputs= the backend's resolved recipe.) - One package = one sub_task, streamed. No fan-out. A Download Package can
be several GB, so it is never held whole in memory: each document is read
from its own
s3_file_bucket, written into an on-disk ZIP in/tmp, and released; the finished archive is multipart-uploaded from disk (S3Client.upload_file). The Lambda's ephemeral storage (/tmp) is raised accordingly (up to 10 GB) — see Consequences. One artifact, one URL, one email. - Selection cap at the route. Reject
> Nproperties synchronously with a legible "narrow your selection" error (property count, not a document/byte query — cheap, no S3 on the hot path). A/tmpsize budget inside the Lambda is a backstop that fails the sub_task with a clear reason if a pathological selection still overflows.Nstarts at a conservative, tunable value. - Matching on
landlord_property_id. Files are gathered bylandlord_property_id; the missingproperty_idonuploaded_filesis a known future gap, out of scope here. Property display info (the folder name) is enriched from the hubspot deals data (repositories/hubspot_deals/). - Layout. One folder per property named by human-readable address
(hubspot-enriched), with
landlord_property_idappended for uniqueness; inside, one file per Document Type = the newest bys3_upload_timestamp. Rows with a null Document Type are skipped and listed in the run output. - Best-effort failure contract. Build from whatever resolves. Skipped
properties (no documents) and skipped null-type files are recorded in
sub_task.outputs. The run fails only on an infrastructure error (S3/DB/zip/upload/email) or when the whole selection yields zero documents. The email carries an "N properties, M documents, X skipped" summary. - Delivery — both channels. The presigned URL (60-minute expiry) and the
skip-summary are written to
sub_task.outputs(the FE already polls task status and can show the link) and emailed (ADR-0059). - A dedicated exports bucket. Packages are written to a separate
DOCUMENT_EXPORTS_BUCKET, notDATA_BUCKET— the generated archives are a distinct class of artifact (transient, user-facing, wide-read source but single-writer) and keeping them out of the shared data bucket keeps IAM and any future retention policy cleanly scoped to exports. No lifecycle rule is imposed onDATA_BUCKET. Retention on the exports bucket is an open decision (the packages are transient, so an expiry is likely warranted, but it is left to a follow-up rather than baked in here).
Considered options
- Fan out into per-batch sub_tasks / multiple ZIPs (like the modelling run). Rejected: the user would receive N links / N emails and "the download" would stop being one file; the capped-single-package model keeps the artifact and the notification singular.
- Cap on resolved document count / bytes. Rejected for v1: it adds synchronous DB + S3 HEAD work to the trigger and couples the route to the packaging logic; a property-count cap is cheap and predictable.
- Strict "all-or-nothing" packaging. Rejected: one unreadable row or a property with no documents would deny the entire package; best-effort + reporting matches "include documents and properties where they exist".
- Match on
uprn. Deferred:landlord_property_idis the agreed key and hubspot enrichment is keyed to the deal/property; revisit if/when theproperty_idgap onuploaded_filesis closed.
Consequences
- New DDD pieces:
applications/bulk_document_download/(thin handler + trigger body),orchestration/bulk_document_download_orchestrator.py, packaging rules indomain/, alandlord_property_id-keyed "latest-per-Document-Type" query on the uploaded-file repository, agenerate_presigned_url+ multipart upload onS3Client, and the email port/adapter from ADR-0059. The onlybackend/touch is the trigger route. - The stored, pinned
property_idset makes a run reproducible even if the portfolio changes between trigger and execution. N, the Lambda's ephemeral-storage (/tmp) size, and its memory are tunable knobs, not load-bearing invariants./tmpmust be raised beyond the 512 MB default (up to 10 GB) to hold a multi-GB archive — a new terraform knob on the shared Lambda modules, backward-compatible (every other Lambda keeps 512 MB).- A new S3 bucket (
DOCUMENT_EXPORTS_BUCKET) is provisioned for the packages, with its own IAM (single-writer, presign-read). Its retention policy is a deliberate follow-up, not set here. - Closing the
uploaded_files.property_idgap later would let matching move offlandlord_property_idwithout changing the package model.