Principal Data Engineer, Biologics Discovery

Johnson & Johnson 3 Locations Updated 19 September 2026
PharmaBiotechCROMedTechQuality Assurancepythoncdmemacroinformazure

Job description

At Johnson & Johnson, we believe health is everything. Our strength in healthcare innovation empowers us to build a world where complex diseases are prevented, treated, and cured, where treatments are smarter and less invasive, and solutions are personal. Through our expertise in Innovative Medicine and MedTech, we are uniquely positioned to innovate across the full spectrum of healthcare solutions today to deliver the breakthroughs of tomorrow, and profoundly impact health for humanity. Learn more at jnj.com . As guided by Our Credo, Johnson & Johnson is responsible to our employees who work with us throughout the world. We provide an inclusive work environment where each person is considered as an individual. At Johnson & Johnson, we respect the diversity and dignity of our employees and recognize their merit. Job Function: Data Analytics & Computational Sciences Job Sub Function: Data Engineering Job Category: Scientific/Technology All Job Posting Locations: Raritan, New Jersey, United States of America, Spring House, Pennsylvania, United States of America, Titusville, New Jersey, United States of America Job Description: Our expertise in Innovative Medicine is informed and inspired by patients, whose insights fuel our science-based advancements. Visionaries like you work on teams that save lives by developing the medicines of tomorrow. Join us in developing treatments, finding cures, and pioneering the path from lab to life while championing patients every step of the way. Learn more at https://www.jnj.com/innovative-medicine About the opportunity Johnson & Johnson Innovative Medicine is seeking a Principal Data Engineer dedicated to our Biologics Discovery organization. This is a high-leverage role responsible for shaping how discovery data is structured, connected, and made AI-ready across Biologics Discovery. The role serves as the bridge between scientific workflows, data consumers, and technology partners, ensuring that discovery data products support scientific research, analytics, machine learning, and agentic workflows. This position will be based at one of our office locations in either Spring House, PA (strongly preferred), Titusville, NJ, or Raritan, NJ. (No remote option.) Why this role matters: High-quality, well-governed scientific data is foundational to our vision for AI-enabled biologics discovery. This role provides senior technical leadership within Biologics Discovery, translating scientific needs into data products, scientific data models, and requirements, and working with enterprise data and technology partners to ensure discovery data is trusted, connected, and AI-ready as the broader ecosystem evolves. Position Summary As a Principal Data Engineer, you will lead the design of discovery data products, scientific data models (schemas, entities, and relationships), and integration requirements that enable discovery data to be acquired, connected, harmonized, and delivered across the Biologics Discovery ecosystem. You will work closely with scientists and AI/ML teams to translate their needs into durable, reusable, and AI-ready data assets. Working in close partnership with enterprise Data Strategy & Products and Technology teams, you will ensure discovery data needs are represented in enterprise standards and that those standards are effectively applied within Biologics Discovery. You will help shape the future-state discovery data ecosystem while delivering near-term value through trusted data products, harmonized data, and metadata practices that support long-term interoperability and reuse. Key Responsibilities Discovery Data Products & Integration Design and deliver AI-ready discovery data products that support ML, AI, and insight generation across Biologics Discovery, applying agile delivery practices to respond to evolving scientific needs. Define and lead the delivery of scalable integration requirements, transformation patterns, and data schemas that support discovery data acquisition, harmonization, and downstream analytics, working with scientific, data, and technology stakeholders to enable reliable data exchange across systems. Translate scientific and analytical requirements from discovery teams into data product specifications, data contracts, acceptance criteria, and delivery requirements, in partnership with scientists, AI/ML teams, and technology partners. Define access and data consumption patterns that enable analytics, modeling, and agentic AI workflows, aligned with industry data standards and frameworks. Catalog discovery instruments, data types, and data sources to inform data product prioritization and sustainable integration approaches. Data Stewardship & Organizational Impact Establish, champion, and drive adoption of standards and best practices for discovery data products, including data quality, provenance, lineage, reproducibility, metadata, and documentation. Partner with ontology, data architecture, platform, and AI teams to ensure discovery data products are connected, discoverable, and suitable for advanced analytics, ML, and agentic AI applications. Apply FAIR data principles, so data products are reusable, scalable, and interoperable. Serve as a thought leader in scientific data architecture, harmonization, and AI-ready data practices across the organization. Qualifications: Required Degree in Computer Science, Data Science, Engineering, or a related computational field. 8+ years (Bachelor's), 5+ years (Master's), or 3+ years (Ph.D.) of experience designing and delivering data products, data models, and analytics-ready datasets within pharmaceutical, biotechnology, or life sciences organizations. Deep proficiency in Python and SQL, with hands-on experience designing and implementing reusable and scalable data assets on cloud data platforms (e.g., Snowflake, AWS, Azure, BigQuery) to support analytics, machine learning, and AI-driven workflows. Experience applying FAIR data principles, metadata management, controlled vocabularies, data lineage, and provenance practices in scientific data environments. Demonstrated technical leadership through architecture reviews, mentorship, code reviews, or leadership of complex technical initiatives. Proven ability to lead complex cross-functional technical initiatives and establish data standards across multiple stakeholder groups. Strong communication and the seniority to represent the team credibly in architecture, data, and strategy forums. Preferred Prior technical mentorship or leadership responsibilities. Experience with ontologies, semantic technologies, knowledge graphs, or ontology-driven data architectures. Experience in biologics discovery, high-throughput experimentation, or external partner (CRO/CDMO) data integration.

Stand out for this role

NoxPharm tailors your CV to this job description by aligning your experience with the role requirements and terminology. Built for pharma & life sciences.

Tailor my CV now — free to try