Staff HPC Infrastructure Engineer
Job description
Company Description Guardant Health is a leading precision oncology company focused on guarding wellness and giving every person more time free from cancer. Founded in 2012, Guardant® is transforming patient care and accelerating new cancer therapies by providing critical insights into what drives disease through its advanced blood and tissue tests, real-world data and AI analytics. Guardant tests help improve outcomes across all stages of care, including screening to find cancer early, monitoring for recurrence in early-stage cancer, and treatment selection for patients with advanced cancer. For more information, visit guardanthealth.com and follow the company on LinkedIn , X (Twitter) and Facebook . Guardant's HPC team builds and operates the computational technology backbone of the company. This includes scalable data storage that holds petabytes of genomics data, high-performance compute clusters running a custom bioinformatics pipeline in production and R&D environments, and the software infrastructure that hosts an ecosystem of services for internal data processing and external data integration. The HPC engineering team is looking for a Staff-level engineer with broad, all-round HPC competency and specialized depth in one or more of the following: Red Hat-family OS management, networking, storage, Kubernetes, or Slurm. Experience applying these skills both on-premise and in the cloud is a plus, as we continue to evolve our HPC footprint. This role carries technical depth and cross-team representation for HPC compute, networking, and storage-integration initiatives, partnering closely with our engineering team and our managed service provider (MSP) as we scale operations. To support Guardant Health's fast growth over the next few years, we need a strong technical engineer who can help maintain and grow the HPC infrastructure through this expansion while working closely with corporate IT, SQA, and DevOps/SRE teams. Strong teamwork and the ability to manage multiple in-flight, cross-functional projects at once are essential for success in this role. About the Role You enjoy an agile, very fast paced and highly technical environment. You are a self-driven, accomplished technologist who strives to continually improve your skills as the HPC landscape evolves and the computational infrastructure scales. You are dedicated to engineering excellence yet pragmatic and flexible. You have the ability to maintain the day-to-day support SLA while running various key projects that move the business forward. You are comfortable being the deepest technical voice in the room on your declared specialism, even among more senior colleagues, while operating as a peer to the other engineers on the team. Essential Duties and Responsibilities Broad HPC Skills (all-round) · Manage multiple HPC clusters and cluster file systems · Integrate cloud bursting as part of the HPC abstraction work · Research, develop, and implement the next generation HPC solutions · Troubleshoot the production system stack down to source code level, e.g shell scripts, Python, and others · Maintain, monitor, and support the infrastructure environment and/or facilities · Use and maintain enhanced production monitoring and addition capability · Support improvements for increased system reliability and performance · Support multiple systems or applications of medium to high complexity complexity defined by size, technology used, and system feeds and interfaces) with multiple concurrent users, ensuring control, integrity, and accessibility · Support systems at remote locations, including internationally · Mentor junior engineers on HPC best practices · Work with offsite consultants to maintain the infrastructure · Work with vendors to troubleshoot, upgrade, and repair systems as needed · Represent HPC infrastructure networking and storage-integration topics in cross-functional planning with networking, SQA, DevOps/SRE, and the MSP · Set up and ownership supporting XDMoD instances for HPC metric and monitoring · Participate in a 24/7 on-call rotation Networking · Act as the technical peer for HPC networking and interconnect initiatives with the dedicated networking engineer · Collaborate on design, performance tuning, and troubleshooting of HPC Ethernet · Work with enterprise networking on integration of HPC systems with the bandwidth-on-demand system that connects our sites and cloud infrastructure · Work with the networking infrastructure team to manage and optimize connectivity to and from HPC systems and global locations Storage · Act as the technical peer for the architecture and integration strategy for. HPC storage in partnership with the dedicated storage engineer and MSP · Serve as a technical point of contact for the MSP storage relationship and help define and evolve SLAs, validate delivery, and escalate technical issues · Support the transition of day-to-day storage operations to the MSP without loss of performance or reliability Required Qualifications Bachelor’s degree in Computer Science or a related field with 8–12 years of relevant experience; Master’s degree with 6–8 years of relevant experience; or PhD with 3–5 years of relevant experience Strong experience in systems and/or infrastructure engineering, including Linux/Unix administration and TCP/IP networking. Hands-on experience with automation tools, such as Ansible or equivalent technologies. Experience with high-performance networking technologies, such as InfiniBand, RoCE, RDMA, or equivalent, including troubleshooting in production environments. Experience supporting large-scale data storage and high-performance computing (HPC)/compute environments. Experience working with both on-premise and cloud-based infrastructure, such as AWS, Google Cloud Platform (GCP), Azure, or similar environments. Experience developing and supporting software release, operations, and infrastructure automation processes and toolsets. Strong experience creating and maintaining system administration and technical documentation. Preferred Qualifications Cisco Certified Network Professional (CCNP) certification Experience with Arista and compatible networking, up to and including 400 Gb/s links</li
Stand out for this role
NoxPharm tailors your CV to this exact job description — matching the keywords recruiters and ATS systems screen for. Built for pharma & life sciences.
Tailor my CV now — free to trySimilar pharma jobs
Mechanical Assembler
Thermo Fisher Scientific — Eindhoven, Netherlands
CRA (Level II)
Thermo Fisher Scientific — 2 Locations
Application Scientist
Thermo Fisher Scientific — Shanghai, China
Biostatistician II
Thermo Fisher Scientific — Beijing, China
Sr Project Mgr
Thermo Fisher Scientific — Beijing, China
Programmer Analyst
Thermo Fisher Scientific — Guangdong, China