Lead AI Operations Engineer

Johnson & Johnson 3 Locations Updated 29 September 2026
PharmaMedTechRegulatory AffairsQuality Assuranceheorpythonemacroinformazure

Job description

At Johnson & Johnson, we believe health is everything. Our strength in healthcare innovation empowers us to build a world where complex diseases are prevented, treated, and cured, where treatments are smarter and less invasive, and solutions are personal. Through our expertise in Innovative Medicine and MedTech, we are uniquely positioned to innovate across the full spectrum of healthcare solutions today to deliver the breakthroughs of tomorrow, and profoundly impact health for humanity. Learn more at jnj.com . As guided by Our Credo, Johnson & Johnson is responsible to our employees who work with us throughout the world. We provide an inclusive work environment where each person is considered as an individual. At Johnson & Johnson, we respect the diversity and dignity of our employees and recognize their merit. Job Function: Technology Product & Platform Management Job Sub Function: Technical Product Management Job Category: Scientific/Technology All Job Posting Locations: Lisbon, Portugal, Madrid, Spain, Milano, Italy Job Description: We are recruiting for a Lead AI Operations Engineer based in Milan - Italy ; Madrid ; Spain or Lisbon ; Portugal The AI Operations Engineer is responsible for shaping, designing, implementing, and continuously improving the enterprise capabilities required to operate AI applications and AI agents safely, reliably, transparently, and cost-effectively at scale. The role combines hands-on AI platform engineering with Site Reliability Engineering, DevSecOps, LLMOps, AgentOps, FinOps, security, and compliance practices. The AI Operations Engineer builds reusable operational capabilities across observability, runtime controls, cost management, auditability, incident response, and production support. The role works closely with the Agent Factory, AI Engineering, Data Platforms, Cloud Infrastructure, Cybersecurity, Privacy, Risk, Quality, and Responsible AI stakeholders. The role does not own the end-to-end lifecycle management of agents or AI products. Agent design, development, functional evaluation, release content, product evolution, and retirement decisions remain with the Agent Factory and the relevant AI product teams. The AI Operations Engineer provides the shared operational platform, telemetry, controls, and guardrails that enable those teams to run AI solutions in production. Key Responsibilities AI Observability & Production Reliability Design and implement end-to-end observability for AI applications and agents, including prompts, responses, model calls, tool calls, retrieval steps, decision paths, latency, failures, token consumption, and session context. Establish common telemetry and distributed tracing across agent workflows, APIs, data services, vector stores, model endpoints, and external tools. Build operational dashboards and alerts covering availability, latency, errors, reliability, quality signals, policy violations, consumption, and service health. Define service-level indicators, service-level objectives, error budgets, alert thresholds, and operational readiness criteria for production AI services. Enable trace-based debugging, incident reconstruction, and controlled session replay while protecting confidential or sensitive information in logs. Monitor retrieval quality, data freshness, model and prompt regressions, anomalous agent loops, degraded tool performance, and unexpected runtime behavior. Lead technical root-cause analysis for AI platform and runtime incidents and convert findings into preventive controls, automation, and engineering improvements. LLMOps & AgentOps Platform Enablement Engineer reusable pipelines, templates, and controls for configuration, prompt, model, and agent-component versioning across environments. Implement automated technical gates for deployment readiness, including integration tests, regression checks, operational validation, security checks, and observability coverage. Enable controlled rollout patterns such as canary releases, feature flags, model or provider routing, fallback strategies, and technical rollback mechanisms. Provide common operational tooling that supports multiple models, frameworks, clouds, and agent patterns without creating a separate operating process for each solution. Integrate functional evaluation signals supplied by the Agent Factory or AI product teams into deployment gates and runtime monitoring, while functional quality ownership remains with those teams. Maintain reusable runbooks, reference implementations, engineering standards, and paved-road patterns for production operation and support. AI FinOps & Consumption Efficiency Create transparent metering, allocation, and showback capabilities by product, agent, workflow, model, environment, and business unit where the required identifiers are available. Monitor token consumption, model utilization, repeated or runaway loops, retrieval overhead, infrastructure usage, and cost per successful transaction or workflow. Implement budgets, thresholds, anomaly alerts, and runtime guardrails to detect and contain unexpected consumption. Partner with Agent Factory and product teams to optimize model selection, routing, context size, caching, batching, retries, and tool usage while preserving agreed quality and compliance requirements. Define operational unit economics and provide evidence for capacity planning, optimization priorities, and platform investment decisions. Security, Governance & Compliance Engineering Embed security, privacy, Responsible AI, and compliance controls into the shared AI runtime and operational toolchain in line with enterprise policies and approved risk frameworks. Implement identity, role-based access control, least privilege, managed identities, secrets management, and segregation of duties for agents, tools, services, and operators. Engineer runtime controls for prompt injection, jailbreak attempts, unauthorized tool use, excessive permissions, data leakage, unsafe execution paths, and anomalous access patterns. Design privacy-aware logging, retention, redaction, and access patterns for prompts, responses, memory, traces, and audit evidence. Provide auditable records linking versions, configurations, identities, actions, approvals, policy decisions, and operational outcomes. Automate policy checks and evidence collection where feasible, partnering with Cybersecurity, Privacy, Quality, Legal, Risk, and Responsible AI stakeholders for control definition and approval. Support threat modelling, security testing, incident response, remediation, and continuous control improvement for the AI platform. Operational Service Management & Enablement Define operating processes for monitoring, support, incident management, problem management, change management, escalation, and service recovery. Create production-readiness checklists, servic

Stand out for this role

NoxPharm tailors your CV to this job description by aligning your experience with the role requirements and terminology. Built for pharma & life sciences.

Tailor my CV now — free to try