Director, Molecular AI & Federated Learning
Pharma
Job description
At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us. Organization Overview Lilly Catalyze360 is a comprehensive approach to enabling the early-stage biotech ecosystem by democratizing access to infrastructure, expertise , and resources. Through its interconnected pillars—Lilly Ventures, Lilly Gateway Labs, Lilly ExploR&D , and Lilly TuneLab —Catalyze360 strategically removes barriers that traditionally block bold science from becoming life- changing medicines, providing biotechs with flexible combinations of capital, physical lab space, R&D capabilities, AI/ML tools, and decades of enterprise learning . Lilly TuneLab is an artificial intelligence and machine learning (AI/ML) platform that provides biotech companies access to drug discovery models trained on years of Lilly's research data. Lilly estimates that this first release of AI models includes proprietary data obtained at a cost of over $1 billion , representing one of the industry's most valuable datasets used to train an AI system available to biotechnology companies . By integrating advanced in silico modelling and federated learning, we connect pioneering machine learning algorithms, substantial computational power, exclusive datasets, and Lilly’s domain-specific knowledge to drive innovation in drug discovery and facilitate access to optimal therapies for patients. Job Summary The Director, Molecular AI & Federated Learning is a senior technical leadership role within the TuneLab platform, setting the technical vision that unites privacy-preserving federated learning with generative small-molecule design. This position pairs deep expertise in medicinal chemistry, ADMET prediction, and molecular optimization with advanced capabilities in federated foundation models and multi-task learning, and is responsible for the predictive and generative models that accelerate small-molecule lead optimization and candidate selection across the TuneLab federated network. As a technical director, the role leads through vision, methodological rigor, and mentorship—guiding scientists and shaping research strategy across internal teams and external biotech partners—rather than through formal people management. Key Responsibilities Technical Vision & Research Strategy: Set the technical direction for federated learning and molecular AI across TuneLab—defining a research agenda that unifies privacy-preserving foundation models, multi-task learning, and generative small-molecule design, and aligning it with platform and portfolio priorities. Technical Leadership & Mentorship: Serve as a principal technical authority and mentor for data scientists and engineers—guiding experimental design, reviewing methods and code, and raising the scientific bar across the team, while influencing technical decisions across disciplines internally and with external partners. Federated Foundation Models: Architect novel deep learning architectures (e.g., Transformer and graph neural network–based) for large-scale federated pre-training on unlabeled or partially labeled data distributed across multiple partner sources. Semi-Supervised & Self-Supervised Learning: Advance state-of-the-art semi-supervised and self-supervised methods (e.g., contrastive learning, masked auto-encoding) tailored to the constraints of federated learning, such as communication bottlenecks and data heterogeneity. Federated Optimization & Aggregation: Develop robust, communication-efficient aggregation strategies (e.g., FedAvg, FedProx, SCAFFOLD) that remain stable for large, complex models and handle non-IID data across clients. Scalability, Simulation & Performance: Profile and optimize the computational performance—memory, latency, and communication cost—of federated training and inference for scale, and build high-fidelity simulation environments to test, debug, and benchmark federated strategies before real-world deployment. Federated Multi-Task Learning: Architect multi-task learning models that leverage shared representations across related endpoints to improve predictive performance and data efficiency in a federated ecosystem, where each client may hold data for only a subset of tasks. Data & Task Heterogeneity: Design algorithms that address extreme task and feature heterogeneity across clients—personalized models, meta-learning, and gradient-aggregation methods robust to non-IID data—and apply regularization that prevents negative transfer while encouraging positive knowledge sharing. Downstream Adaptation & Validation: Create efficient protocols for fine-tuning and adapting pre-trained federated models to specific downstream tasks, and establish rigorous validation frameworks with appropriate per-task metrics and fairness assessment across clients and tasks. Small Molecule Property Prediction: Build multi-task models for small-molecule properties—including ADMET endpoints, solubility, permeability, metabolic stability, and off-target liabilities—across diverse chemical representations (SMILES, graphs, 3D conformations). Generative Chemistry Models: Design and deploy state-of-the-art generative models (VAEs, diffusion models, flow matching, autoregressive models) for de novo design, lead optimization, and scaffold hopping that respect synthetic accessibility and drug-likeness constraints. ADMET-Driven, Multi-Objective Design: Develop integrated prediction–generation pipelines that optimize molecules simultaneously across multiple ADMET properties while maintaining target potency, using multi-objective optimization and Pareto-front exploration. Chemical Space & Synthetic Feasibility: Implement efficient exploration of synthetically accessible chemical space—reaction-aware generation, retrosynthetic-planning integration, and fragment-based design—collaborating with synthetic chemists to ensure generated molecules are practically synthesizable. Structure-Activity & Representation Learning: Learn and exploit structure–activity relationships from sparse, noisy federated bioactivity data—including matched molecular pair analysis and activity-cliff prediction—and develop self- and semi-supervised molecular representations that generalize to novel chemical series. Interpretability & Scientific Insight: Apply explainability (XAI) techniques to complex multi-task and molecular models to understand predictions and uncover relationships between endpoints, generating novel scientific insight while respecting IP and competitive boundaries across federated partners. Be
Stand out for this role
NoxPharm tailors your CV to this exact job description — matching the keywords recruiters and ATS systems screen for. Built for pharma & life sciences.
Tailor my CV now — free to trySimilar Pharma jobs
Mechanical Assembler
Thermo Fisher Scientific — Eindhoven, Netherlands
CRA (Level II)
Thermo Fisher Scientific — 2 Locations
Application Scientist
Thermo Fisher Scientific — Shanghai, China
Biostatistician II
Thermo Fisher Scientific — Beijing, China
Sr Project Mgr
Thermo Fisher Scientific — Beijing, China
Programmer Analyst
Thermo Fisher Scientific — Guangdong, China