Manager, Site Reliability Engineer - Data Platforms

Amgen India - Hyderabad Updated 24 September 2026
PharmaBiotechQuality Assurancesaspythonemacroinformazure

Job description

Career Category Engineering Job Description Join Amgen’s Mission of Serving Patients At Amgen, if you feel like you’re part of something bigger, it’s because you are. Our shared mission-to serve patients living with serious illnesses-drives all that we do. Since 1980, we’ve helped pioneer the world of biotech in our fight against the world’s toughest diseases. With our focus on four therapeutic areas -Oncology, Inflammation, General Medicine, and Rare Disease- we reach millions of patients each year. As a member of the Amgen team, you’ll help make a lasting impact on the lives of patients as we research, manufacture, and deliver innovative medicines to help people live longer, fuller happier lives. Our award-winning culture is collaborative, innovative, and science based. If you have a passion for challenges and the opportunities that lay within them, you’ll thrive as part of the Amgen team. Join us and transform the lives of patients while transforming your career. Manager, Site Reliability Engineer - Data Platforms About Amgen Amgen harnesses the best of biology and technology to fight the world’s toughest diseases, and make people’s lives easier, fuller and longer. We discover, develop, manufacture and deliver innovative medicines to help millions of patients. Amgen helped establish the biotechnology industry more than 40 years ago and remains on the cutting-edge of innovation, using technology and human genetic data to push beyond what’s known today. About the Role AWS Site Reliability Engineer with strong cloud architecture expertise to design, build, automate, and operate secure, resilient, scalable, and cost-efficient AWS platforms. You will combine AWS architecture leadership with practical Site Reliability Engineering-writing Infrastructure as Code, developing automation, improving observability, solving complex production problems, and strengthening operational excellence. This role also includes responsibility for leading an SRE pod and managing assigned engineers. However, it is primarily a hands-on technical role. The successful candidate will remain a key technical contributor who leads through architecture, implementation, production ownership, incident response, coaching, and example. What you will do Roles & Responsibilities: AWS Architecture and Platform Engineering Design secure, highly available, scalable, and cost-efficient AWS architectures for EDSE applications and data platforms. Define reusable platform services, Infrastructure as Code modules, reference architectures, and engineering guardrails. Make and document architectural decisions across networking, identity, compute, containers, storage, security, observability, and resilience. Lead architecture and production-readiness reviews, addressing reliability, security, performance, operability, and cost risks. Reliability and Production Operations Establish and improve SRE practices, including SLIs, SLOs, error budgets, availability targets, capacity planning, and operational-readiness criteria. Use operational data to identify recurring failures, performance bottlenecks, capacity risks, and opportunities to reduce manual toil. Participate in the on-call rotation and lead the technical response to complex or high-severity production incidents. Conduct blameless post-incident reviews and ensure corrective actions produce durable engineering improvements. Automation, Delivery, and Observability Build reusable, tested infrastructure and operational automation using Terraform and programming languages such as Python. Automate provisioning, configuration, validation, deployment, recovery, compliance checks, and routine operational activities. Strengthen CI/CD and GitOps practices through automated testing, deployment controls, progressive delivery, and reliable rollback mechanisms. Standardize metrics, logs, traces, dashboards, and actionable alerting to improve issue detection, diagnosis, and recovery. Resilience, Security, and Cost Efficiency Design and validate backup, high-availability, and disaster-recovery strategies aligned with defined business and recovery objectives. Conduct restore tests, failover exercises, resilience reviews, and controlled game-day scenarios. Embed least-privilege access, encryption, network segmentation, secrets management, vulnerability remediation, and policy-based controls into platform designs. Improve AWS cost efficiency through right-sizing, resource lifecycle management, tagging, usage analysis, and architecture optimization without compromising reliability or security. Pod and People Leadership Lead the SRE pod by setting technical direction, prioritizing work, managing operational commitments, and driving delivery to completion. Serve as a senior AWS and SRE advisor, partnering with application engineering, data engineering, cybersecurity, architecture, product, and FinOps teams. Manage and mentor assigned engineers through regular feedback, one-on-one discussions, technical coaching, design and code reviews, and career-development support. The expected allocation of responsibilities is: Approximately 75-80% hands-on technical work , including AWS architecture, coding, Infrastructure as Code, automation, design reviews, production troubleshooting, observability, performance improvement, and incident response. Approximately 20-25% pod and people leadership , including prioritization, work allocation, delivery coordination, mentoring, one-on-one discussions, performance feedback, career development, and removing team blockers. The individual will be expected to move comfortably between architecture decisions, hands-on implementation, production operations, incident leadership, stakeholder communication, and people development. The balance may vary temporarily during major incidents, critical releases, or important delivery milestones. What we expect of you Basic Qualifications and Experience: Master’s or Bachelor’s degree in computer science or engineering field and 9 to 12 years of relevant experience, including substantial hands-on experience in AWS cloud engineering, platform engineering, infrastructure engineering, DevOps, or Site Reliability Engineering. Prior experience leading a technical pod, engineering squad, or small team while continuing to contribute hands-on. Must-Have Skills: Significant hands-on experience designing, implementing, and operating business-critical production workloads on AWS, including experience with EKS, Sagemaker, Bedrock, VPC, PrivateLink, S3, EC2, KMS, CloudWatch, CloudTrail, Secrets Manager, Lambda, and RDS. Deep knowledge of AWS architecture across networking, identity and access management, security, compute, containers, storage, monitoring, and resilience. Strong experience with Infrastructure as Code using Terraform, CloudFormation, AWS CDK, or comparable technologies. Proficiency in at least one programming or scripting language, such as Python, Go, Java, TypeScript, Bash, or PowerShell. Ability to write maintainable, production-quality infrastructure code, automation, and operational tooling. Practical experience applying SRE principles, including SLIs and SLOs, actionable alerting, incident response, root-cause analysis, post-incident improvement, and toil reduction. Experience building or operating containerized platforms using Docker and Kubernetes,

Stand out for this role

NoxPharm tailors your CV to this job description by aligning your experience with the role requirements and terminology. Built for pharma & life sciences.

Tailor my CV now — free to try