Advisor - Data Architect, Data Foundry
Pharma
Job description
At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us. Location: San Francisco, CA Reports to: Lead, Data Architecture (R9), Architecture4Insight Overview Lilly Small Molecule Discovery is purpose-built to create molecules that make life better for people. Discovery Technology and Platforms (DTP) accelerates molecule discovery by building optimized foundational platforms, streamlining lab operations through advanced technologies and data connectivity, and investing in novel capabilities. Data Foundry is a multidisciplinary team within DTP that enables AI-native drug discovery through four integrated pillars: Architecture4Insight (data infrastructure and scientific software), Methods4Insight (analytical and computational methods), Automation & Scale4Insight (lab automation and agentic workflows), and Preparedness4Insight (data governance and readiness). These pillars empower every Lilly scientist to make optimal decisions by providing seamless access to data, insights, and AI-driven capabilities—serving both human scientists and autonomous AI agents. Position Summary We are seeking Data Architects at multiple levels to design and build the data infrastructure that makes AI-native drug discovery possible. You will create the schemas, ontologies, data models, knowledge graphs, and platform architectures that transform raw scientific data into machine-actionable, FAIR-compliant, insight-ready assets—serving both discovery scientists and autonomous AI agents. This role is the foundation of Architecture4Insight . Everything the software engineering team builds—pipelines, APIs, prototypes—depends on the data models and platform architecture this team designs. You will work with deep knowledge of scientific data (chemical, biological, HTE, automation-generated) to create custom-fit solutions, then partner with Tech@Lilly to scale and maintain them. The role spans three focus areas depending on expertise: data modeling & ontologies , data platform & lakehouse architecture , and knowledge graph & specialized data systems . You will independently design schemas, select technologies, and make build-vs-buy recommendations for their domain. Responsibilities Data Modeling & Ontologies Design and implement data models, schemas, and ontologies for chemical, biological, and automation-generated data that serve discovery workflows across the portfolio. Define and maintain controlled vocabularies, metadata standards, and FAIR-compliant data frameworks in partnership with Preparedness4Insight. Implement semantic data standards (RDF, OWL, SPARQL) and ontology engineering practices to create interoperable, machine-readable scientific data. Data Platform & Lakehouse Architecture Design and implement data lakehouse architecture using modern platforms (Databricks, Snowflake, or equivalent), including data storage patterns, partitioning strategies, and query optimization. Build and optimize ETL/ELT pipelines using Spark, dbt, or similar tools to transform raw scientific data into analytical and ML-ready formats. Implement real-time and streaming data integration (Kafka, Kinesis, event-driven patterns) connecting LIMS, instruments, and lab automation systems to the data infrastructure. Knowledge Graph & Specialized Data Systems Design and implement knowledge graphs (Neo4j, Amazon Neptune, TigerGraph) that capture molecular, target, pathway, and experimental relationships across the discovery landscape. Architect specialized data solutions: array databases (TileDB) for genomics/imaging, document stores (MongoDB) for experimental records, and vector databases for embedding-based retrieval supporting ML and RAG workflows. Build query and traversal patterns that enable scientists and AI agents to ask relational questions across the entire data landscape. Cross-Functional Partnership Partner with scientific software engineers to ensure data architectures are implementable, performant, and well-documented. Collaborate with Methods4Insight to design data structures that support analytical model training, deployment, and evaluation. Work with Tech@Lilly to define scaling strategies, ensure enterprise compliance, and transition data architectures to production-grade management. Contribute to build-versus-buy-versus-adopt decisions by evaluating commercial and open-source data platforms against Data Foundry requirements. Basic Requirements M.S. or PhD in Computer Science, Data Science, Bioinformatics, Computational Biology, Information Science, or related STEM field MS (with 6+ years ) and PhD (with 2+ years) of data architecture, data engineering, or scientific informatics experience. Deep expertise in at least one of the focus areas: relational databases, data modeling and ontology engineering, data platform and lakehouse architecture (Databricks, Snowflake, Spark), or knowledge graph and specialized database systems (Neo4j, Neptune, MongoDB, TileDB) Preferred Qualifications Working familiarity with multiple database paradigms — relational, graph, document, columnar, key-value — and strong SQL skills. Understanding of scientific data types and experimental workflows in life sciences or pharma (chemical, biological, HTE data). Strong communication skills with ability to translate data architecture concepts for both technical and scientific audiences. Familiarity with cloud platforms (AWS, Azure, or GCP) and modern data integration patterns. Pharmaceutical or biotech research industry experience, particularly in discovery data management or research informatics. Experience with semantic web technologies: RDF, OWL, SPARQL, Protégé, or equivalent ontology engineering tools. Hands-on experience with graph databases (Neo4j, Neptune, TigerGraph) and knowledge graph design patterns for scientific data. Data lakehouse architecture experience: Databricks (Delta Lake, Unity Catalog), Snowflake, or equivalent; ETL/ELT with Spark, dbt. Experience with streaming/real-time data platforms (Kafka, Kinesis, Flink) and event-driven architectures. Familiarity with LIMS, ELN systems (e.g., Benchling), and laboratory instrument data integration. Experience with vector databases (Pinecone, Weaviate, pgvector) and embedding-based retrieval for ML/RAG applications. Array databas
Stand out for this role
NoxPharm tailors your CV to this exact job description — matching the keywords recruiters and ATS systems screen for. Built for pharma & life sciences.
Tailor my CV now — free to trySimilar Pharma jobs
DO NOT APPLY- TEST 1
Thermo Fisher Scientific — Carlsbad, California, USA
Mechanical Assembler
Thermo Fisher Scientific — Eindhoven, Netherlands
CRA (Level II)
Thermo Fisher Scientific — 2 Locations
Application Scientist
Thermo Fisher Scientific — Shanghai, China
DO NOT APPLY- TEST 4
Thermo Fisher Scientific — Remote, United States
Biostatistician II
Thermo Fisher Scientific — Beijing, China