Principal SRE Engineer
Pharma
Job description
At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us. About the technology organization Technology at Lilly builds and operates mission-critical digital products and platforms that support the discovery, development, and delivery of medicines that make life better for people around the world. Our teams operate in highly regulated, high-availability environments, where operational excellence, reliability, and quality are non-negotiable. Within Technology at Lilly, the Digital Core organization applies a product, platform, and reliability-first mindset, ensuring that operational capabilities scale sustainably across the enterprise. About the Team Technology at Lilly builds and maintains capabilities using pioneering technologies like the most prominent tech companies. What differentiates Lilly IT is that we redefine what's possible through tech to advance our purpose, creating medicines that make life better for people around the world, including data-driven drug discovery, connected clinical trials, resilient enterprise platforms, and intelligent digital operations. We hire the best technology professionals from a variety of backgrounds, so they can bring an assortment of knowledge, skills, and diverse thinking to deliver creative solutions in every area of our business. The Digital Core team leads Lilly's transformation into the Digital and AI era. They inspire digitally empowered teams to new ways of working and accelerate innovation and agility. This team powers and advances the entire company by building and maintaining world-class technology capabilities and platforms. The Reliability Engineering team is the engineering-first function that owns the stability, observability, and operational quality of a multi-application production estate. It operates in close partnership with the engineering team that builds the agentic automation platform, and is in active transition from human-executed operations to engineering-led, agent-assisted reliability. Role summary You build the self-healing automation, author the runbooks, and turn root-cause analysis into durable engineering fixes that let the production estate heal itself instead of paging a human. You work within the standards the Senior Principal SRE Engineer sets — SLOs, error budgets, observability — and you're the one who encodes them into working automation and documented procedure. This is a hands-on individual-contributor role focused on execution and codification rather than cross-estate reliability strategy. You decide, in partnership with the Senior Principal SRE Engineer, which recurring patterns warrant a self-healing investment versus a documented manual runbook, and you build whichever is right. You are an individual contributor. You do not manage people. You partner daily with the Senior Principal SRE Engineer, the agentic automation engineering team, and Operations on validating outcomes. Success is measured by self-healing coverage, runbooks authored and adopted, reduced recurrence of known failure modes, and the safety record of every automation you sign off. What you'll be doing 1) Self-healing automation & resilience patterns Design and build self-healing automation — circuit breakers, graceful degradation, automated remediation — for the failure modes that recur most across the estate. Run resilience or chaos testing to validate that self-healing patterns behave correctly before they're trusted in production. Continuously expand self-healing coverage as new failure modes are identified and proven safe to automate. Partner with the Senior Principal SRE Engineer on which failure modes justify self-healing investment versus a documented manual runbook. 2) Runbook authorship & validation Author and validate the remediation runbooks for the production estate: safe execution order, rollback steps, and exception handling for every documented fix. Keep the runbook library current as systems, dependencies, and failure modes evolve, retiring runbooks that no longer apply. Define and apply the graduation criteria that let a runbook move from human-executed to agent-assisted to autonomous. 3) RCA to durable fix Lead or contribute to root-cause analysis for significant incidents, and drive the blameless postmortem process to a durable engineering fix — not just a narrative. Convert recurring incident patterns into codified runbooks and, where appropriate, self-healing automation. Track fix effectiveness against recurrence, and escalate to the Senior Principal SRE Engineer when a fix needs broader engineering investment. Participate in high-severity incident response, including acting as incident commander for escalations within your area. 4) Partnership with agentic automation & operations Partner with the Agentic Automation Engineering team on which fixes are safe to hand off as agent-assisted remediations, and on the confidence thresholds and human-in-the-loop boundaries that keep them safe. Sign off on agent graduation criteria (accuracy over volume, zero P1/P2 caused) before an automation moves to a higher autonomy tier. Partner with Operations on outcome validation, feeding what's learned back into the runbook library and self-healing patterns. 5) Incident response & regulated-environment practice Ensure runbooks and self-healing automation meet Lilly's change-control, audit, and validated-environment standards. Document procedures so that audit evidence falls out of normal operation, not a special exercise. Mentor other reliability and automation engineers on runbook quality and self-healing design. Contribute proven patterns back to the broader reliability practice, in partnership with the Senior Principal SRE Engineer and Senior Architect. How you will succeed : At the principal engineering level for reliability, success is defined by the durability and safety of what you build: Be recognized as the engineer who turns incidents into durable fixes, not repeat pages. Demonstrate measurable growth in self-healing coverage and runbook adoption, with falling recurrence of known failure modes. Maintain a clean safety record: automations you sign off don't cause P1/P2 incidents. Build runbooks and automation that make good practice the default, not a personal habit. Your Basic Qualifications: Bachelor's degree in Computer Science, Information Technology, or a related technical engineering discipline, including Software Engineering, Computer Engineering, Information Systems, Cybersecurity, Infor
Stand out for this role
NoxPharm tailors your CV to this job description by aligning your experience with the role requirements and terminology. Built for pharma & life sciences.
Tailor my CV now — free to trySimilar Pharma jobs
Specialty Representative, Metabolic - Syracuse, NY
AbbVie — Syracuse, us
Senior Manager, Robotics and Vision Systems
AbbVie — North Chicago, us
Associate DevSecOps Engineer, Platform and Tooling
AbbVie — Mettawa, us
Principal Research Scientist I Data
AbbVie — North Chicago, us
Associate Specialist, Scientific Compliance
AbbVie — North Chicago, us
Specialty Representative, Migraine - Fairfax, VA
AbbVie — Fairfax, us