Site Reliability Engineer – Data Platforms
PharmaBiotechQuality Assurancesasemacroinformaws
Job description
Career Category Engineering Job Description Join Amgen’s Mission of Serving Patients At Amgen, if you feel like you’re part of something bigger, it’s because you are. Our shared mission-to serve patients living with serious illnesses-drives all that we do. Since 1980, we’ve helped pioneer the world of biotech in our fight against the world’s toughest diseases. With our focus on four therapeutic areas -Oncology, Inflammation, General Medicine, and Rare Disease- we reach millions of patients each year. As a member of the Amgen team, you’ll help make a lasting impact on the lives of patients as we research, manufacture, and deliver innovative medicines to help people live longer, fuller happier lives. Our award-winning culture is collaborative, innovative, and science based. If you have a passion for challenges and the opportunities that lay within them, you’ll thrive as part of the Amgen team. Join us and transform the lives of patients while transforming your career. About Amgen Amgen harnesses the best of biology and technology to fight the world’s toughest diseases, and make people’s lives easier, fuller and longer. We discover, develop, manufacture and deliver innovative medicines to help millions of patients. Amgen helped establish the biotechnology industry more than 40 years ago and remains on the cutting-edge of innovation, using technology and human genetic data to push beyond what’s known today. About the Role EDSE Data Platforms is seeking a Site Reliability Engineer who is equally comfortable improving production systems in AWS and establishing how services should be operated. Our SRE team manages AWS platform operations and the infrastructure supporting applications across EDSE. As our service portfolio grows, we need an experienced individual contributor who can create clear, practical operating arrangements across SRE, application teams, Security, business stakeholders, and other service partners. This is a deliberately balanced role, with responsibilities split approximately: 50% service governance, operational readiness, resilience, and process leadership 50% hands-on AWS site reliability engineering The balance will be measured over a typical quarter and may vary during major incidents, application handovers, continuity exercises, audits, and delivery periods. This is not a policy-only, PMO, or ITSM process-administration role. You will remain hands-on with AWS infrastructure, automation, observability, incident response, and reliability improvement while creating lightweight, enforceable operational controls. What you will do Roles & Responsibilities: Service governance and operating model - approximately 50% Establish and operate a consistent application handover process, with readiness criteria covering ownership, architecture, dependencies, monitoring, access, security findings, runbooks, recovery, known risks, knowledge transfer, hypercare, and formal acceptance. Identify incomplete or unsafe handovers and recommend deferring acceptance. Ensure exceptions are documented, time-bound, and approved by the appropriate business, service, security, or risk owner. Maintain a service catalogue and operating profile for every supported application, including criticality, owners, recurring tasks, dependencies, support hours, access requirements, backup coverage, recovery objectives, and escalation paths. Define clear responsibilities, decision rights, support boundaries, and segregation of duties across SRE, application engineering, Product, Security, business owners, vendors, and other stakeholders. Agree service indicators and internal objectives, proposed SLAs, support hours, severity definitions, response and restoration expectations, maintenance windows, recovery expectations, and escalation paths. Lead business-continuity and disaster-recovery planning, including business-impact analysis, dependency mapping, recovery procedures, technical recovery tests, and remediation tracking. Establish repeatable security-review and audit-readiness processes covering technical evidence, control self-assessments, access reviews, vulnerability remediation, and audit findings while preserving the independence of Security and Audit. Define standards and ownership for essential operational knowledge, including service profiles, architecture and dependency records, runbooks, incident playbooks, recovery procedures, access models, and known-risk registers. Lead regular service and reliability reviews, turning operational performance, incidents, recovery readiness, security actions, and risks into prioritised improvements with accountable owners. Hands-on AWS site reliability engineering - approximately 50% Operate and improve secure, reliable, scalable, and cost-effective application infrastructure on AWS. Build and maintain infrastructure as code, delivery pipelines, and automation that removes repetitive operational effort. Improve observability through meaningful metrics, logs, traces, dashboards, alerts, and service-health reporting. Participate in on-call and incident response, including diagnosis, mitigation, recovery, communication, escalation, and blameless post-incident reviews. Improve availability, performance, capacity, resilience, backup and recovery, patching, infrastructure lifecycle management, and security posture. Partner with application teams on architecture, production readiness, deployment safety, application resilience, and operational quality. Apply SRE practices such as SLIs, SLOs, error budgets, automation, and learning from failure to reduce repeat incidents, alert fatigue, manual toil, and recovery time. What we expect of you Basic Qualifications and Experience: Master's or Bachelor’s degree in Computer Science, Engineering, Information Systems, or related field 5-8 years of relevant experience Must-Have Skills: Strong understanding of AWS infrastructure, networking, identity and access management, security, monitoring, backup, and resilience patterns with depth in the services most relevant to data platforms: IAM, VPC, PrivateLink, S3, EC2, KMS, CloudWatch, CloudTrail, Secrets Manager, and STS. Working knowledge of EKS, Lambda, Glue, EMR, RDS will be good to have. Experience in operating business-critical production services on AWS. Strong hands-on experience in Site Reliability Engineering, Platform Engineering, Cloud Operations, or Infrastructure Engineering. Practical experience with infrastructure as code, CI/CD, operational automation, monitoring, observability, and incident response. Demonstrated experience establishing or improving service handovers, support models, operational ownership, service levels, business continuity, disaster recovery, or security-control processes. Experience defining and testing recovery objectives, recovery plans, and technical recovery procedures. Understanding of security controls, audit evidence, access governance, vulnerability management, risk acceptance, and segregation of duties. Ability to convert unclear cross-team responsibilities into practical, agreed, and measurable operating arrangements. Strong facilitation, negotiation, documentation, and stakeholder-management skills. Ability to influence technical, business, security, and leadership stakeholders without relying on direct reporting authority. A pragmatic approach to governance that improves accountability and reliability without introducing unne
Stand out for this role
NoxPharm tailors your CV to this job description by aligning your experience with the role requirements and terminology. Built for pharma & life sciences.
Tailor my CV now — free to trySimilar Pharma jobs
Associate, Project Management (Pharmaceutical Labeling)
Open Scientific — Hauppauge, us
Maintenance Coordinator
Open Scientific — Bohemia, us
Manufacturing Operators - Pharmaceutical
Open Scientific — Melville, us
Pharmaceutical Mechanics - Packaging, Production, Maintanence
Open Scientific — Hauppauge, us
Warehouse - Pharmaceutical
Open Scientific — Hauppauge, us
Pharmaceutical Machine Operators
Open Scientific — Hauppauge, us