Senior Director – Real World Data (RWD) Architect - Engineer
Pharma
Job description
At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us. Purpose: The Director RWD Data Engineering is a hand-on technical role that leads the end-to-end engineering vision for Lilly’s real-world data (RWD) infrastructure. This individual leads design and executes the scalable, cloud-native pipelines and data products that allow HEOR, SDIA, Statisticians, Medical, and Clinical teams to generate evidence faster, more reproducibly, and at greater scientific depth than is possible through traditional vendor engagements. This job involves a depth of understanding of the multi-modal RWD ecosystem across Lilly’s therapeutic areas to contextualize and drive RWD to the necessary end-user data products by leading the creation of sophisticated data engineering products, creating processes for automation of data ingestion and product creation, and leading special projects for Global Medical Affairs – Health Economics and Outcomes functions and the broader enterprise. Further, this position will be responsible for identifying and advocating standard processes across the data asset lifecycle, working closely with other data domain and analytics leaders. Collaborating closely with multi-functional teams, you will lead the technical implementation of data products, ensuring scalability, reliability, and performance. The ideal candidate possesses deep expertise in data engineering, strong problem-solving skills, and a passion for leveraging data to drive business outcomes. This position reports to HEOR Central and is embedded within the BIA organization and works in close partnership with HEOR, SDIA, Statisticians, Medical, and Clinical teams. Responsibilities : This job description is intended to provide a general overview of the job requirements at the time it was prepared. The job requirements of any role/position can change over time and can include additional responsibilities not specifically described in the job description. Consult with your supervisor regarding your actual job responsibilities and any related duties that might be required for the role/position. Lead the design, development, and implementation of cloud-native data products and high-throughput data pipelines that transform raw real-world data into scalable, reliable, analysis-ready assets supporting analytics, reporting, and evidence generation. Lead the Analytic Data Products Strategy to deliver key data assets that enable streamlined, compliant execution and analytics. Own the end-to-end lifecycle of RWD data products, from requirements gathering and prototyping through production deployment and optimization, ensuring scalability, reliability, performance, and reproducibility across cloud environments (e.g., Databricks, AWS S3, Azure Data Lake). Build, optimize, and maintain ETL/ELT ingestion and transformation pipelines for large-scale, multi-modal RWD — including claims, complex EHR data, and other linked healthcare datasets — handling data volumes ranging from tens of millions to billions of records. Implement and manage lakehouse-style data architectures (e.g., medallion bronze/silver/gold patterns) using Databricks and cloud object storage (AWS S3, ADLS) to produce versioned, partitioned, and audit-ready data assets. Write and maintain reusable, version-controlled transformation logic incorporating healthcare coding and terminology standards (e.g., ICD-10/ICD-9, NDC, RxNorm, SNOMED, CPT/HCPCS, LOINC) to produce domain-level datasets such as demographics, diagnoses, treatments, procedures, encounters, and labs. Optimize SQL and distributed processing workloads (e.g., Spark-based jobs) for performance across very large datasets, applying partitioning, indexing, predicate pushdown, denormalization, and other optimization strategies appropriate to analytical workloads. Translate analytic, business, and research requirements into reproducible data extraction and transformation logic, supporting cohort construction, temporal logic, and consistent reuse of RWD across teams. Apply deep understanding of healthcare data structures and standards when engineering data products, ensuring datasets are fit for purpose for downstream analytics and compliant with scientific, regulatory, and audit expectations. Establish and implement standard engineering practices and methodology across the data asset lifecycle, including automated data ingestion, data quality checks, integrity testing, validation, monitoring, alerting, and documentation from source table to analysis-ready output. Lead CI/CD pipeline setup, code review, and testing standards, ensuring all transformation code is version-controlled, tested, and deployable in a reproducible manner. Collaborate closely with multi-functional partners — data scientists, statisticians, analytics leaders, and other technical teams — to understand business and technical requirements and develop documentation of RWD engineering standards, transformation templates, code list repositories, and pipeline performance guidelines. Provide technical consultation to collaborators on appropriate use of data products and underlying RWD assets, including structural limitations of specific data sources, join strategies, and performance considerations; develop source-specific training materials for HEOR scientists, SDIA, and statisticians. Develop and implement KPIs to measure system performance, efficiency and pull through to program impact. Create an inclusive culture where producing and maintaining high-quality data is a core discipline. Technical Skills: Core Engineering Skills High proficiency in SQL optimization across cloud platforms — complex joins, window functions, query tuning, workload management — on AWS Redshift, Databricks SQL, Snowflake, or BigQuery. Python fluency: pandas, PySpark, Polars for large-scale data manipulation; workflow orchestration with Apache Airflow, Prefect, or Dagster for production pipeline scheduling and monitoring — including Databricks Workflows for orchestrating multi-task jobs within the Lakehouse. Distributed computing: Apache Spark (PySpark), Dask, or Ray — ability to write, tune, and debug distributed jobs processing multi-terabyte datasets across partitioned cloud storage, including Databricks clusters with auto-scaling and spot instance optimization. Cloud data engineering: hands-on pipeline development on AWS (S3, Glue, Redshift, EMR), Azure (ADLS, Synapse, ADF), or GCP (BigQuery, Dataflow) — as well as Databricks on any major cloud (AWS, Azure, or GCP) using Unity Catalog for cross-workspace governance — not just configuration. Delta Lake / Apache Iceberg: time travel, schema evolution, upsert/merge operations, partition optimization — building versioned, ACID-compliant data assets at scale; Delta Lake experience ideally hands-on wi
Stand out for this role
NoxPharm tailors your CV to this exact job description — matching the keywords recruiters and ATS systems screen for. Built for pharma & life sciences.
Tailor my CV now — free to trySimilar Pharma jobs
DO NOT APPLY- TEST 1
Thermo Fisher Scientific — Carlsbad, California, USA
Mechanical Assembler
Thermo Fisher Scientific — Eindhoven, Netherlands
CRA (Level II)
Thermo Fisher Scientific — 2 Locations
Application Scientist
Thermo Fisher Scientific — Shanghai, China
DO NOT APPLY- TEST 4
Thermo Fisher Scientific — Remote, United States
Biostatistician II
Thermo Fisher Scientific — Beijing, China