Senior Cloud Platform Engineer (AWS), AI Infrastructure - Evinova
PharmaClinical ResearchQuality Assurancepythonemacroaws
Job description
WHY JOIN US? Evinova is a health-tech business focused on accelerating better health outcomes by advancing digital transformation across the life sciences sector. By combining science-based expertise, evidence-led rigor, and deep human insight, we design digital solutions that enable healthcare to work better for everyone. Operating at the intersection of healthcare, technology, data, and analytics, we are helping unlock the full potential of digital health, transforming how clinical research is conducted, how care is delivered, and how patients experience healthcare. Our solutions are built to scale, driving efficiency, improving decision-making, and ultimately delivering better outcomes for patients worldwide. At Evinova, we are driven by a shared purpose to transform health through data and digital innovation. Our teams collaborate across disciplines to solve complex challenges, continuously learning and evolving in a fast-paced, high-impact environment. We also recognize the importance of flexibility and balance. Our ways of working support both individual needs and team collaboration. To foster connection and collaboration, employees are expected to work from the office three days per week , creating opportunities for in-person teamwork, innovation, and meaningful connection. This role is located in the Greater Toronto Area and follows a hybrid work model. Candidates must reside within commuting distance of the GTA or be willing to relocate for this opportunity. Introduction to Role The Machine Learning and Artificial Intelligence Operations team (ML/AI Ops) is the cloud platform engineering team responsible for building and operating the infrastructure that enables our AI Engineers and Data Scientists to deploy Generative AI applications reliably, securely, efficiently and at scale. As a Senior Cloud Platform Engineer on the ML/AI Ops team, you will design, build and operate the AWS platform that runs our production Generative AI, agentic AI and conversational AI workloads. Most of the code you write will be AWS CDK (TypeScript and/or Python) that creates the infrastructure for other teams to build on. This is a cloud and platform engineering role rather than an AI application development role. You will partner closely with the engineers and Data Scientists who build agents and models, and provide the infrastructure, deployment patterns, model access, observability and operational capabilities they need to move solutions from experimentation into reliable production environments. You will work across AWS infrastructure, infrastructure as code (IaC), Amazon Bedrock AgentCore, Amazon ECS, CI/CD, SageMaker Unified Studio, AI gateways, observability, scalability, reliability, security, governance and cost optimization. Your work will establish reusable platform capabilities that allow teams across Evinova to deploy and operate solutions faster and more reliably while meeting the requirements of a highly regulated pharmaceutical environment. Accountabilities Cloud Platform Engineering Design, build and operate scalable AWS cloud platform capabilities for production ML/AI and Generative AI workloads. Create reusable infrastructure, tooling and deployment patterns that enable AI Engineers and Data Scientists to independently deploy and operate their applications. Write and maintain AWS IaC primarily AWS CDK in TypeScript and/or Python, including reusable CDK constructs that other teams consume. Build and operate containerized workloads using Amazon ECS and AgentCore. Develop reusable platform capabilities across compute, networking, IAM, secrets management, storage, model access and workload isolation. Build and maintain CI/CD and GitOps workflows that enable safe, automated and repeatable deployments across environments. Partner with engineers and Data Scientists to transition prototypes and research workloads into resilient, production-grade services. Build self-service capabilities and automation that improve developer experience and reduce operational toil. Reliability, Scalability & Operational Excellence Engineer platform capabilities that improve the availability, scalability, resiliency and performance of production GenAI workloads. Design and implement autoscaling, load balancing, retries, timeouts, fallback, rate limiting and failure-recovery strategies. Establish monitoring, alerting, SLOs, runbooks and production-readiness standards for ML/AI workloads. Troubleshoot complex production issues across AWS infrastructure, container, application and model-provider layers. Automate operational processes and proactively identify opportunities to improve platform reliability and performance. Drive cloud and model cost optimization through capacity management, workload optimization and data-driven analysis. AI Gateway & Model Access Build and operate a centralized AI gateway that routes requests across model providers. Provide secure, reliable and governed access to foundation models through Amazon Bedrock, OpenAI, Anthropic, Microsoft Foundry, Gemini Enterprise Agent Platform (formerly Vertex AI) and other model platforms. Implement model/provider routing, fallback, authentication, rate limiting, quotas and cost controls. Enable AI teams to evaluate and change model providers without tightly coupling their applications to individual model endpoints. Provide the compute, networking, storage and runtime infrastructure to reliably operate agentic AI, RAG and conversational AI workloads in production. Observability, Governance & Cost Optimization Build platform-level observability for GenAI workloads, including token consumption, latency, throughput, errors, model/provider performance and cost attribution by team and application. Implement standardized tracing, logging, metrics and alerting capabilities that can be adopted across AI applications. Integrate observability technologies such as Amazon CloudWatch, OpenTelemetry, Datadog and Splunk. Build the telemetry and data pipelines that AI teams use to evaluate and monitor production LLM behavior. Build appropriate security, auditability and governance controls for AI workloads operating within a regulated environment. Support compliance with applicable industry standards and practices, including Good Clinical Practice and Good Machine Learning Practice and other GxP related standards and practices. Representative Projects Build a library of AWS CDK constructs that gives an AI team a production-ready Amazon ECS service, with IAM, secrets, networking and observability, from a single import. Deploy and operate a centralized AI gateway on Amazon ECS, with multi-provider routing, fallback, quotas and rate limiting. Build multi-account CI/CD with CDK Pipelines or GitOps, with safe, staged deployments across environments. Implement token cost attribution by team and application with SLO and burn-rate alerts. Migrate container workloads from Amazon EKS to Amazon ECS without loss of reliability. Essential Skills/Experience Minimum of 4 years of hands-on experience in a platform engineering, infrastructure, site reliability engineering (SRE) or DevOps role where infrastructure code was your main output. Deep hands-on AWS cloud engineering experience, including designing, deploying and operating production cloud-native infrastructure. Strong experience with AWS services such as Bedrock, ECS, IAM, VPC, load balancing
Stand out for this role
NoxPharm tailors your CV to this job description by aligning your experience with the role requirements and terminology. Built for pharma & life sciences.
Tailor my CV now — free to trySimilar Pharma jobs
Principal Scientist, Small Molecule Analytical Chemistry
Merck & Co — 2 Locations
Customer Excellence Associate Manager - Remote
Stryker — 5 Locations
MSL - Imunologia (Dermatologia) - Belo Horizonte/MG
AbbVie — Belo Horizonte, br
IRDP- International Recruitment and Development Intern Program 2026--MedTech
Johnson & Johnson — Shanghai, China
Territory Manager (d/m/w) im Außendienst Electrophysiology - Region Bayern
Johnson & Johnson — Norderstedt, Schleswig-Holstein, Germany
Technical Product Owner - Regulatory Excellence Platform Lead
Johnson & Johnson — Hyderabad, Andhra Pradesh, India