Case study
Genomics platform infrastructure
I built Python, Airflow, and PostgreSQL systems for laboratory workflows, genomic processing, and customer data delivery.
Context
The work
The Genomics Platform supports the software connecting laboratory work, genomic analysis, quality checks, and data delivery. Although many of our clinical and research programs needed the same basic capabilities, each also came with its own workflows, data requirements, and delivery agreements.
I worked on the shared platform layer, mainly in Python, Airflow, PostgreSQL, and cloud infrastructure, building common workflow pieces where the underlying behavior was stable while keeping necessary program differences explicit. For context, Broad has delivered more than 3 million samples for over 1,400 groups across more than 50 countries (these are institutional figures; my own scope is described below).
What I owned
My role
- Built reusable sample-data delivery, onboarding, validation, and orchestration capabilities
- Led hands-on implementation of core Python and Airflow platform features
- Defined PostgreSQL models connecting sample, laboratory, pipeline, and delivery activity
- Authored technical scopes and evaluated architecture options with laboratory, clinical, and engineering partners
- Automated cloud-data retention policies to replace repetitive manual cleanup
My role remained hands-on as my responsibilities grew: I wrote technical scopes, worked through implementation options with my manager and partner teams, split larger initiatives into work that teammates and I could take on, and documented system behavior in our GitHub Wiki so that important operating knowledge was not limited to the people who had written the code.
Project example
Customer data delivery
Originally, customer data moved through Terra Data Repository, a platform developed by Broad and Google. That path fit teams already working in the Google Cloud ecosystem, but it could also require customers to establish additional cloud access and workflows before receiving their data.
I owned much of the hands-on modernization of this delivery path. I built Airflow workflows that modeled complex, customer-specific packages combining sample data and metadata, then delivered those packages either through Terra or directly to customer-designated cloud buckets. By keeping the packaging and orchestration shared while making the destination configurable, we avoided maintaining a separate code path for every customer and gave customers a more direct way to receive data in infrastructure they already used.
Project example
Sample traceability
Tracing an individual sample often meant reconstructing its path manually from logs after an error surfaced. That became especially difficult when the same sample carried different identifiers in upstream and downstream laboratory or analysis systems.
I extended our data models and PostgreSQL schema to store those identifiers and workflow statuses, making a sample's state queryable instead of leaving it scattered across log files. I then worked with the data-warehouse team to shape the new data for ingestion and visualization, which later supported customer-level views of pipeline efficiency.
- Sample identifiersMap names across systems
- Laboratory statePersist upstream activity
- Pipeline stateMake processing status queryable
- Delivery stateConnect downstream handoffs
Evidence
Scale
- Pipeline scale
- 94K+ samples
- Program scope
- 5+ programs
- Delivery reach
- 7+ partners
The pipelines supported more than 94,000 samples across 5+ clinical and research programs and delivered patient-sample data for 7+ government, industry, and academic partners. These figures describe the pipelines and delivery network supported by the core features I led, including work forNIH eMERGE and the Broad Institute's COVID-19 response. For COVID-19, I helped operate and enhance a production system within a broader testing operation that delivered more than 36 million tests over 3.3 years and approximately 5% of U.S. tests at peak.
Reflection
Lessons from five years on the platform
A large part of the work was deciding where reuse should stop. Mapping dependencies and business rules gave us a way to compare programs, identify behavior we could consolidate, preserve differences that were genuinely necessary, and remove conditional paths that had accumulated without adding much value. I documented those decisions so teammates, leadership, and my future self could understand how the system was expected to behave.
It was easy for every new customer request to become another branch in the code, so I came to prefer configuration over customer-specific logic and straightforward implementations over clever ones. Feedback from partner teams was still essential, but we had to distinguish genuine requirements from additions that would make the platform harder to maintain without materially improving it.
The platform also could not depend on one person knowing how everything worked. Automated tests, audit trails, documentation, pair programming, and mentoring helped spread that knowledge across the team, while sustainable workloads made it possible for us to keep delivering and operating the system over time.