Job Description
Must Have
Role: Senior Associate - Data Engineer (Life Sciences)
Job Description:
Key Responsibilities
Data Pipeline Architecture: Design, build, and maintain automated ETL/ELT pipelines to migrate and synchronize data across Salesforce LSC, Veeva CRM, and Vault CRM.
Integration & Synchronization: Implement robust API-based integrations (REST/SOAP) to ensure data consistency between Veeva Vault and enterprise master data management (MDM) systems.
Data Quality & Compliance: Develop automated validation scripts in Python to ensure all clinical and commercial data meets GxP and 21 CFR Part 11 standards.
Schema & Model Management: Lead the design of scalable data schemas and models that align with life-sciences-specific standards like FHIR.
Workflow Orchestration: Utilize tools like Apache Airflow to schedule and monitor complex data processing tasks, ensuring high availability of analytics-ready data.
Technical Skills & Qualifications
Education: Bachelor’s or Master’s degree in Computer Science, Information Systems, or a related technical field.
Programming: Expert-level Python and SQL (PostgreSQL, Redshift, or Snowflake).
Salesforce/Veeva Stack:
Salesforce Bulk and REST APIs.
Veeva Vault Loader and Vault API for high-volume data migration.
Salesforce Data Cloud for unifying patient and HCP profiles.
Infrastructure: Experience with cloud-based big data technologies such as Apache Spark, AWS Glue, or Databricks.
CI/CD: Proficient in version control (Git) and deployment pipelines for data-as-code.
Experience Required:
Total Professional Experience: 4–5 years in a Data Engineering or Data Systems role.
Domain Expertise: Minimum 2 years of direct experience working within the Life Sciences (Pharmaceutical, Biotech, or MedTech) industry.
Platform-Specific Experience:
Salesforce: 3+ years of experience with Salesforce platforms or Health Cloud, including deep knowledge of its data models and objects.
Veeva Systems: 2+ years of hands-on experience with Veeva CRM, Veeva Vault (Clinical, Quality, or MedComms), and Vault CRM.
Core Technical Experience:
3+ years of advanced Python development (e.g., Pandas, PySpark) for ETL/ELT pipeline construction.
Proven experience integrating CRM platforms with downstream Data Warehouses (e.g., Snowflake, Redshift) using API-led patterns.
Good to have
Preferred Certifications (Optional)
Salesforce Ecosystem:
Salesforce Certified Data Cloud Consultant
Salesforce Certified Platform Data Architect
Salesforce Certified Administrator
Veeva Systems:
Certified Veeva CRM Administrator
Vault CRM Administrator
Vault Platform Administrator
Veeva CDB Clinical Data Programmer
General Data Engineering:
Cloud Platform Professional (e.g., AWS Certified Data Engineer – Associate or Google Professional Data Engineer)
Databricks/Snowflake Certification
EQUAL OPPORTUNITY
Indegene is proud to be an Equal Employment Employer and is committed to the culture of Inclusion and Diversity. We do not discriminate on the basis of race, religion, sex, colour, age, national origin, pregnancy, sexual orientation, physical ability, or any other characteristics. All employment decisions, from hiring to separation, will be based on business requirements, the candidate’s merit and qualification. We are an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, colour, religion, sex, national origin, gender identity, sexual orientation, disability status, protected veteran status, or any other characteristics.