Case study

Clinical genomics data pipelines

I built data pipelines and validation tools for an NIH-funded clinical genomics study.

Organization
Broad Institute
Genomics Platform · NIH eMERGE
Role
Software and data-infrastructure contributor
Timeframe
During 2021–2026 Broad tenure

Context

The study

eMERGE studied how genomic information could be implemented across a national clinical research network. The scientific program determined which risk scores were meaningful, but we still needed an operational system to process results, apply quality controls, and support reporting across multiple sites; my contribution was in that software and data infrastructure rather than in the design of the study or development of its polygenic risk scores.

155K+genetic risk results
25K+participants
10clinical sites

What I owned

My role

  • Implemented cloud functions that ingested, processed, and delivered sample data
  • Defined Python data models and database schemas representing sample data and workflow state
  • Provisioned and maintained supporting cloud resources through Terraform
  • Built validation and error-reporting behavior into the processing workflow
  • Expanded unit and integration testing, CI/CD automation, and production debugging practices

Project example

Validating incoming data

Incoming sample data arrived as JSON with identifiers, metadata, status, and project-specific fields. Legitimate variation between projects made a single rigid format impractical, but accepting every variation would allow malformed data to travel farther through the workflow before it failed.

I defined JSON schemas for the major sample, customer, and project combinations, allowing explicitly identified nondeterministic fields while validating the rest of each payload through conditional logic in Airflow. Invalid inputs failed early and produced a clear message in Slack, giving our team the exact sample and validation details needed to work with the upstream team on a correction before the error propagated downstream.

Data access

Moving data into Terra

I automated ingestion of polygenic risk score data into Terra Data Repository through cloud-function workflows, giving analysis teams a repeatable path for accessing, modifying, and exporting data within Terra for further processing and delivery. The broader program included more than 155,000 genetic risk results, 25,000 participants, and 10 clinical sites; my contribution was the supporting pipeline infrastructure and co-authorship of the resulting Nature Medicine publication.

Clinical review

Giving reviewers a place to inspect results

I implemented a web interface that brought manually maintained Excel tracking data together with pipeline results so that clinical reviewers, including geneticists, could inspect them within a secure workflow. In the process, I introduced UI/UX and frontend practices to a team whose experience was mainly in backend development, and came to understand that a scientifically correct result is only one part of a reliable clinical system; the surrounding processing, quality control, review, and reporting steps also have to work together.