Director, Data Integration Engineer
PharmaRegulatory AffairsQuality Assuranceheormedicalveevagdpemacroraveinform
Job description
ROLE SUMMARY This role owns data integration engineering for the Medical Affairs AI Acceleration portfolio. It ensures all solutions are integrated successfully to source data. Working as an individual-contributor technical expert within the Engineering organization, this role is directly accountable for the data pipelines and integration patterns that connect Medical Affairs AI solutions to existing enterprise data sources and analytic platforms. It ensures each new AI capability has reliable, well-governed access to the data it needs without duplicating or re-architecting the underlying data platforms. This is a hands-on engineering role that develops integration software to be leveraged in any Medical Affairs AI solution. The role partners closely with Solution Architecture to implement the data integration standards and RAG/vector-database patterns defined at the architecture level, and serves as the go-to data engineering expert for Product Management and UX/Experience Design partners seeking to understand what data is available, where it lives, and how reliably it can be surfaced. The Director, Data Integration Engineer owns the design, build, and operation of the data pipelines and integration layer that connect Medical Affairs' AI-enabled products and platforms to existing data sources. Reporting to the Senior Director, Engineering, this role translates solution architecture and product requirements into reliable, governed, production-grade data pipelines that are leveraged in solutions the team is delivering for Medical Affairs. This is a technical, hands-on role responsible for designing the integration and data-access patterns for each AI solution and building, testing and operating them. This role partners closely with Build Engineers, Solution Architecture, and Product Management to ensure data readiness keeps pace with AI delivery. ROLE RESPONSIBILITIES Data Integration Architecture & Pipeline Engineering Design and build data pipelines and integration patterns connecting core enterprise systems to Medical Affairs' AI and analytics platforms and data sources. Build and maintain targeted data pipelines that extract, transform, and serve the specific data each AI solution needs, prioritizing reuse across solutions. Establish and follow data pipeline coding standards for solution-level integration work, aligning with the data models and cataloging practices maintained by the centralized data organization. Monitor and manage data pipeline latency, ensuring each AI solution receives data within the timeliness thresholds its use case requires. AI & Data Enablement Partner with Solution Architecture to implement data integration patterns supporting RAG pipelines, vector databases, GraphRAG (graph-structured retrieval context for LLMs), and other AI/ML data access patterns. Ensure data feeding Agentic AI and LLM-based systems is well-governed, accurately labeled, and monitored for quality and drift. Apply automation techniques (AI/ML, low-code/no-code tooling where appropriate) to accelerate data delivery. Support context-aware and context-driven AI models by ensuring underlying data structures capture the necessary business context. Data Governance, Quality & Compliance Collaborate and partner with the centralized Commercial AI Data Strategy team to align on data profiling, sourcing and investigation. Apply data governance and data cataloging best practices, maintaining data dictionaries, lineage documentation, and playbooks. Ensure data integration practice complies with Pfizer data privacy and regulatory standards (GDPR, HIPAA, GxP as applicable). Own identification and classification of personal information (PI/PII) flowing through integration pipelines, and apply masking, tokenization, or de-identification before sensitive data is stored in a vector database or made accessible to AI/ML systems. Cross-Functional Partnership & Delivery Partner with the Senior Director, Engineering and Build Engineers to ensure data pipelines are delivered in step with product and platform build cycles. Partner with Solution Architecture to ensure data integration design aligns with enterprise architectural standards and reuse patterns. Continuous Improvement Drive best practices and world-class data engineering capability, staying current with modern data platform technology (e.g., Snowflake, graph databases) and AI-enabled data tooling. Establish a culture of high performance, transparency, and continuous improvement within the data integration discipline. BASIC QUALIFICATIONS Candidate demonstrates a breadth of diverse leadership experiences and capabilities including: the ability to influence and collaborate with peers, develop and coach others, oversee and guide the work of other colleagues to achieve meaningful outcomes and create business impact. Bachelor's degree in Computer Science, Data Analytics or related field. 8+ years of hands-on data engineering or application integration experience, including building data pipelines and APIs that connect applications to existing data platforms. Advanced proficiency in SQL and hands-on experience building ETL/ELT pipelines. Experience consuming and integrating with modern data platform technologies (e.g., Snowflake, or equivalent), including querying, API-based access, and working within an existing data governance framework. Experience with API design and integration patterns (REST/GraphQL, event-driven or streaming integration) for connecting applications to backend data sources. Experience with data analytics automation and business process automation using AI, ML, low-code, or no-code tools. Experience supporting RAG pipelines, vector databases, or other AI/ML data access patterns. Solid understanding of Agile delivery and CI/CD practice. Experience identifying and protecting personal information (PI/PII) in data pipelines, including applying data masking, tokenization, or de-identification techniques before sensitive data reaches AI-accessible stores such as vector databases. Good knowledge of data governance and data cataloging best practices. Experience partnering with engineering, architecture, and business stakeholders in a matrixed delivery model. Excellent communication and stakeholder management skills. Familiarity with data privacy standards and pharma industry compliance practices (GDPR, HIPAA, GxP). Direct experience integrating with Veeva CRM, Salesforce Life Sciences/Marketing Cloud, or other Medical Affairs systems via API. Familiarity with GraphRAG patterns — using graph-structured data to organize retrieval context for LLMs — as a complement to standard vector-based RAG. PREFERRED QUALIFICATIONS Experience in pharmaceutical, life sciences, or another regulated industry; Medical Affairs or commercial life sciences data experience a significant plus. Experience with API gateway/middleware technologies and event-streaming platforms for real-time application integration. Advanced degree in a technical field. Non-Standard Work Schedule, Travel, or Environment Requirements Must be able to travel to Pfizer offices, vendor offices, and other t
Stand out for this role
NoxPharm tailors your CV to this job description by aligning your experience with the role requirements and terminology. Built for pharma & life sciences.
Tailor my CV now — free to trySimilar Pharma jobs
Mechanical Assembler
Thermo Fisher Scientific — Eindhoven, Netherlands
CRA (Level II)
Thermo Fisher Scientific — 2 Locations
Application Scientist
Thermo Fisher Scientific — Shanghai, China
Biostatistician II
Thermo Fisher Scientific — Beijing, China
Sr Project Mgr
Thermo Fisher Scientific — Beijing, China
Programmer Analyst
Thermo Fisher Scientific — Guangdong, China