Senior Principal SRE Engineering
PharmaClinical ResearchRegulatory AffairsQuality Assurancegcpemacroinformazureaws
Job description
At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us. About the technology organization Technology at Lilly builds and operates mission-critical digital products and platforms that support the discovery, development, and delivery of medicines that make life better for people around the world. Our teams operate in highly regulated, high-availability environments, where operational excellence, reliability, and quality are non-negotiable. Within Technology at Lilly, the Digital Core organization applies a product, platform, and reliability-first mindset, ensuring that operational capabilities scale sustainably across the enterprise. About the Team Technology at Lilly builds and maintains capabilities using pioneering technologies like the most prominent tech companies. What differentiates Lilly IT is that we redefine what's possible through tech to advance our purpose, creating medicines that make life better for people around the world, including data-driven drug discovery, connected clinical trials, resilient enterprise platforms, and intelligent digital operations. We hire the best technology professionals from a variety of backgrounds, so they can bring an assortment of knowledge, skills, and diverse thinking to deliver creative solutions in every area of our business. The Digital Core team leads Lilly's transformation into the Digital and AI era. They inspire digitally empowered teams to new ways of working and accelerate innovation and agility. This team powers and advances the entire company by building and maintaining world-class technology capabilities and platforms. The Reliability Engineering team is the engineering-first function that owns the stability, observability, and operational quality of a multi-application production estate. It operates in close partnership with the engineering team that builds the agentic automation platform, and is in active transition from human-executed operations to engineering-led, agent-assisted reliability. Role Summary: As the Senior Principal SRE Engineering Lead, you are the senior-most engineering authority for the reliability of the supported production estate. You own the bar for what reliable means in this organization: service-level objectives, error-budget governance, observability standards, and the engineering practices that protect both production and the team's engineering time. This role is the judgment layer between an agent's recommendation and a production change. You combine engineering rigor with operational pragmatism, and you decide which patterns surfaced from production warrant durable engineering investment. You hold the SRE engineering bar within the function — the cross-pillar reference architecture is owned by the Senior Architect, but the practice, standards, and engineering judgment of reliability at this site are yours. You are a senior individual contributor. You do not manage people. You partner with the Reliability Leader, the Senior Principal Tech Shift Lead in the same pillar, the Senior Architect, and the senior engineering principals in the agentic automation team. Success is measured by organizational impact, sustained reliability outcomes, and the ability to scale reliability through systems and people, not heroics. What you'll be doing: 1) Reliability strategy and SLO governance Define service-level objectives and indicators across the supported production estate, tiered by application risk and business impact. Govern error-budget burn: when to slow change, when to invest in durable fixes, when to accept the budget. Drive the adoption of reliability reporting and the disciplined operating cadence that makes SLOs real, not decorative. Hold the SLO and error-budget conversation with product and application owners — including the conversation about what their service needs to change to meet the bar. 2) Engineering standards, observability, and self-healing Establish observability and instrumentation standards as the contract every supported application must meet — and hold the line on them. Set the bar for infrastructure-as-code, continuous-delivery hardening, and deployment safety across the supported estate. Define the engineering work that flows from incidents and root-cause analyses into durable production change, and the architectural patterns (self-healing runbooks, graceful degradation, circuit breakers) that reduce repeat failure. 3) Incident learning, durable fixes, and partnership with agentic automation Drive blameless postmortem culture; ensure root-cause analyses produce engineering work, not just narrative. Govern which patterns from production warrant durable engineering investment, against the error-budget regime. Partner with the production operations team in the same pillar on which patterns from the field warrant engineering attention; partner with the agentic automation engineering team on which fixes become safe agent-assisted remediations, and on the confidence thresholds, guardrails, and human-in-the-loop boundaries that make those remediations safe in production. Lead high-severity incident response as incident commander when escalation reaches this seat, and coach the team to handle the rest. 4) Compliance, security, and regulated-environment readiness Ensure reliability practices comply with Lilly standards and applicable regulatory requirements, and engineer them so that audit evidence falls out of normal operation. Promote secure operational practices, auditability, and validated-environment-friendly engineering as part of the standards the team holds. Act as a trusted technical leader in regulated and validated environments. 5) Technical leadership and talent development Set the engineering bar through standards, expectations, and role modeling. Mentor senior reliability engineers in the pillar; build the bench for sustained team growth and develop the next layer of principal engineers. Influence engineering, product, and platform leaders through credibility and outcomes rather than authority. Contribute to the evolution of enterprise-wide reliability practices in partnership with the Senior Architect and peer technical leaders. How you will succeed At the senior-most engineering individual-contributor level for reliability, success is defined by breadth of impact and sustained outcomes: Be recognized as the senior reliability authority for your area. Demonstrate measurable, sustained improvements such as: reduced major incidents, fewer recurring failures, improved time-to
Stand out for this role
NoxPharm tailors your CV to this job description by aligning your experience with the role requirements and terminology. Built for pharma & life sciences.
Tailor my CV now — free to trySimilar Pharma jobs
NA & GCSO Performance Center Senior Manager
Johnson & Johnson — 4 Locations
Associate Director, MAA & RWE
Johnson & Johnson — Raritan, New Jersey, United States of America
Associate Director, Global Market Access Analytics & RWE
Johnson & Johnson — 2 Locations
Senior Manager Technical Product Management
Johnson & Johnson — Beerse, Antwerp, Belgium
医学经理,免疫疾病领域临床开发
Johnson & Johnson — Shanghai, China
Cluster Finance Senior Analyst
Johnson & Johnson — Prague, Czechia