Associate Director — Reliability Engineering & Operations

Eli Lilly IN: Hyderabad Updated 31 August 2026
PharmaClinical ResearchRegulatory AffairsQuality Assurancegcppythonemacroinformazure

Job description

At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us. Associate Director — Reliability Engineering & Operations Job description About Lilly At Lilly, everything we do starts with patients. We unite caring with discovery to make life better for people around the world. Headquartered in Indianapolis, Indiana, our global team of over 50,000 employees work with urgency and purpose to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. We bring our best to this work because people depend on it. If you’re driven by purpose and determined to make a meaningful difference for patients, we invite you to bring your skill and your commitment to Lilly. About Tech@Lilly At Lilly, technology is not a support function. It is how a global medicine company operates, innovates, and delivers. Lilly in Hyderabad builds the capabilities that make this possible, cloud platforms, AI systems, and automation at enterprise scale, all in service of a purpose that makes this technology work genuinely distinctive, from advancing drug discovery to enabling connected clinical trials to keeping a global medicine company running at the standard patients deserve. Experience 10+ years Location Hyderabad (Onsite) Employment Type Full-time Job Family M1 — Engineering Management / Site Reliability / Production Engineering About the technology organization Technology at Lilly builds and operates mission-critical digital products and platforms that support the discovery, development, and delivery of medicines that make life better for people around the world. Our teams operate in highly regulated, high-availability environments, where operational excellence, reliability, and quality are non-negotiable. Within Tech@Lilly, the Digital Core organization applies a product, platform, and reliability-first mindset, ensuring that operational capabilities scale sustainably across the enterprise. About the Team Tech@Lilly builds and maintains capabilities using pioneering technologies like the most prominent tech companies. What differentiates Lilly IT is that we redefine what's possible through tech to advance our purpose, creating medicines that make life better for people around the world, including data-driven drug discovery, connected clinical trials, resilient enterprise platforms, and intelligent digital operations. We hire the best technology professionals from a variety of backgrounds, so they can bring an assortment of knowledge, skills, and diverse thinking to deliver creative solutions in every area of our business. The Digital Core team leads Lilly's transformation into the Digital and AI era. They inspire digitally empowered teams to new ways of working and accelerate innovation and agility. This team powers and advances the entire company by building and maintaining world-class technology capabilities and platforms. Role summary As Associate Director of Reliability Engineering, you lead the engineering team accountable for the stability, observability, and operational quality of a multi-application production estate. You manage a team of senior reliability and production engineers — site reliability leads, production operations leads, principal production engineers, and quality engineers — and you own the operating rhythm that turns production signal into durable engineering work. This is a first-line engineering management role with real production weight. You are the named owner for the reliability of your estate, the escalation point above shift leads during major incidents, and the manager who decides what your team stops doing manually and starts doing through code. You spend your time on people, priorities, and the engineering decisions that make reliability stick — not on tickets, status reports, or escalation theatre. What you'll be doing 1) Lead the reliability engineering team Manage, coach, and develop a team of senior reliability and production engineers across site reliability, production operations, and quality engineering. Own hiring, performance, career development, and succession for the team; build a strong bench and credible engineering pathways for both engineers and team leads. Run the operating cadence — standups, on-call reviews, incident reviews, weekly reliability reviews — that keeps the team focused on outcomes, not noise. Protect engineering time: ensure operational load is bounded, on-call is sustainable, and durable engineering work is funded, scheduled, and shipped. 2) Own reliability outcomes for the estate Hold the team accountable to service-level objectives, error-budget policy, and the engineering standards that make them real, not decorative. Govern on-call quality: page volume, time-to-acknowledge, time-to-recover, repeat-offender rate, and the ratio of toil to engineering work — and act on the numbers. Be the named-owner escalation point above shift leads during major incidents. Run command discipline, drive communications to senior stakeholders, and ensure root-cause analyses produce engineering work — not narrative. Own the operational risk posture of the supported estate: what's fragile, what's improving, what needs investment, what should be decommissioned or returned to vendor — and carry that view into the portfolio conversation. 3) Drive the transition from manual operations to engineering Partner with the platform and automation engineering team on which production patterns become agent-assisted or fully automated remediations, and which require platform-side investment. Identify, prioritize, and sponsor the engineering work that retires toil at the source — instrumentation, self-healing runbooks, infrastructure-as-code coverage, deployment hardening. Hold a clear, defended point of view on what your team should stop doing manually in the next 6–12 months, what that requires in platform investment, and what it unlocks in reduced cost-to-serve. Track and report the steady-state-ops curve: are we measurably moving from human-executed to engineering-executed work, quarter over quarter? 4) Compliance, security, and audit readiness Set and enforce the team's engineering standards for observability, alerting, change management, and incident response, and hold the line on them under delivery pressure. Ensure operational practices comply with applicable regulatory requirements (GxP, SOX, or equivalent), and produce credible audit evidence as a by-product of normal engineering work — not as a separate scramble. Partner with security, validation, and change governance teams on controls that protect production without slowing engineering down. 5) Stakeholder partnership and global delivery Partner with application owners, product managers, and platform leaders on the reliability of their services — what the SLOs say, what the error budget allow

Stand out for this role

NoxPharm tailors your CV to this exact job description — matching the keywords recruiters and ATS systems screen for. Built for pharma & life sciences.

Tailor my CV now — free to try