Lead Automation Engineer
PharmaClinical ResearchQuality Assurancepythonemacroinform
Job description
At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us. About the technology organization Technology at Lilly builds and operates mission-critical digital products and platforms that support the discovery, development, and delivery of medicines that make life better for people around the world. Our teams operate in highly regulated, high-availability environments, where operational excellence, reliability, and quality are non-negotiable. Within Tech@Lilly, the Digital Core organization applies a product, platform, and reliability-first mindset, ensuring that operational capabilities scale sustainably across the enterprise. About the Team Tech@Lilly builds and maintains capabilities using pioneering technologies like the most prominent tech companies. What differentiates Lilly IT is that we redefine what's possible through tech to advance our purpose — creating medicines that make life better for people around the world — including data-driven drug discovery, connected clinical trials, resilient enterprise platforms, and intelligent digital operations. The Digital Core team leads Lilly's transformation into the Digital and AI era. Lilly Capability Centre India (LCCI), Hyderabad, is Lilly's premier Global Technology Hub, harnessing data, AI, analytics, and digital solutions to revolutionize healthcare and improve patient outcomes worldwide. Role summary You build the automation and the observability around it that turns a validated fix into a production-safe, autonomous action. Sitting inside Agentic Automation Engineering, you take remediation procedures authored and validated by SRE — and automation opportunities surfaced through discovery, telemetry, and alert correlation — and codify them into automation that runs safely, is fully instrumented, respects rollback and exception handling, and earns broader autonomy over time. This is a hands-on individual-contributor role focused on build quality, codification rigor, and instrumentation depth. You work daily with SRE's runbook library and graduation criteria, with the observability stack that proves an automation is behaving as designed, and with Operations' outcome validation, so automation only ever runs what's been proven safe. Success is measured by automations shipped and codified, runbooks translated into working automation, the quality of the signals you emit alongside them, zero P1/P2 incidents caused by automation you've built, and the pace at which your automation earns broader autonomy. What you'll be doing 1) Automation build — from validated fix to production automation Build production-grade automation that executes validated remediation steps end to end — detection through action — with safe rollback built in. Take automation opportunities and known-fix candidates surfaced through discovery, telemetry, and alert correlation, and turn them into shipped, tested automations. Maintain and extend the existing automation portfolio (scripted orchestration, platform tooling, and, where appropriate, RPA) as new opportunities are approved. Working proficiency with Ansible, Terraform, or equivalent IaC/configuration management tooling for building safe, repeatable remediation and provisioning automation. 2) Runbook codification Translate SRE-authored remediation procedures into codified, executable automation steps: safe execution order, rollback steps, and exception handling. Partner with SRE to keep the runbook library and its automated counterparts in sync as remediation procedures evolve. Document each codified runbook clearly enough that another engineer, or an agent, can execute or extend it without relying on tribal knowledge. 3) Observability & instrumentation — first-class, not an afterthought Design and implement the metrics, logs, traces, and events that make every automation's behavior explainable in production — inputs, decisions, actions taken, rollbacks, outcomes. Define and maintain SLIs/SLOs for the automations you own, and wire them into dashboards and alerts that Operations and SRE actually use. Partner with the platform and alerting teams on signal correlation and alert tuning so automations trigger on clean, high-confidence signal rather than noise. Build the feedback loops that push automation outcome data into the knowledge base and confidence model driving what gets automated next. 4) SRE pattern implementation Apply SRE patterns in what you build: error-budget awareness, blameless-postmortem-driven fixes, and graduation-criteria-aware rollout. Implement confidence thresholds and staged autonomy — an automation earns broader scope only as it demonstrates accuracy over volume, per agreed graduation criteria. Build in the checks and evidence trail that let an automation prove it has caused zero P1/P2 incidents before it's trusted with more scope. 5) Cross-team partnership & autonomy graduation Partner with SRE to validate that an automation meets graduation criteria (accuracy over volume, zero P1/P2 caused) before it moves to a higher autonomy tier. Work with Operations to review outcome validation and incident feedback, closing the loop between what ran and what should run differently next time. Support ticket volume and shift flexibility as operational demand requires, in line with the team's protected build-time model. How you will succeed Be recognized as a dependable builder whose automations run safely, roll back cleanly, are observable end-to-end, and rarely need rework. Demonstrate measurable throughput: runbooks codified, automations shipped, instrumentation delivered, and time-to-production for each. Ship automation that graduates to higher autonomy tiers on schedule, with zero P1/P2 incidents caused along the way. Build codified runbooks and telemetry clear enough that others can extend your work — and diagnose its behavior — without you in the room. What you should bring Required 10+ years of hands-on automation engineering experience, building and maintaining production automation or scripted remediation in an enterprise IT operations environment. Strong, hands-on observability skills : designing SLIs/SLOs, instrumenting code and workflows with metrics, structured logs and distributed traces, and building dashboards and alerts that drive action. Practical experience with at least one major observability stack (Prometheus/Grafana/OpenTelemetry, Splunk, Datadog, Dynatrace, or equivalent). Alerting maturity : designing signal-based alerts, correlation rules, and noise-reduction strategies; comfort tuning alerts based on real incident data rather than intuition. Demonstrated experience turning documented fixes or remediation procedures into reliable, repeatable automation — not just one-off scripts. Practical scripting/programming ability (Python, PowerShell, Bash, or similar) sufficient to build, test, and maintain production-grade automation and i
Stand out for this role
NoxPharm tailors your CV to this job description by aligning your experience with the role requirements and terminology. Built for pharma & life sciences.
Tailor my CV now — free to trySimilar Pharma jobs
Record to Report PEC Analyst (Korean Speaker)
Roche — Petaling Jaya
Officer - Administration
Roche — Shanghai
(高级)治疗领域专员 - 乳腺肿瘤治疗领域 - 长沙(伊赫莱专岗)
Roche — Changsha
Statistical Scientist
Roche — 2 Locations
Technical Specialist RCSC (m/w/d) - Next Generation Sequencing (NGS)
Roche — Mannheim
P&C Business Partner (Advisory) - Spanish speaking
Roche — Budapest