Job Description
We are a technology-led healthcare solutions provider. We are driven by our purpose to enable healthcare organizations to be future-ready. We offer accelerated, global growth opportunities for talent that’s bold, industrious, and nimble. With Indegene, you gain a unique career experience that celebrates entrepreneurship and is guided by passion, innovation, collaboration, and empathy. To explore exciting opportunities at the convergence of healthcare and technology, check out www.careers.indegene.com Looking to jump-start your career? We understand how important the first few years of your career are, which create the foundation of your entire professional journey. At Indegene, we promise you a differentiated career experience. You will not only work at the exciting intersection of healthcare and technology but also will be mentored by some of the most brilliant minds in the industry. We are offering a global fast-track career where you can grow along with Indegene’s high-speed growth. We are purpose-driven. We enable healthcare organizations to be future ready and our customer obsession is our driving force. We ensure that our customers achieve what they truly want. We are bold in our actions, nimble in our decision-making, and industrious in the way we work.
Role Overview: The Data Engineer will be responsible for designing, developing, and optimizing large scale data pipelines and analytical solutions supporting Life Sciences Commercial Operations. This role requires deep expertise in Databricks (including Genie), AWS cloud services, and PySpark, with strong domain knowledge of pharmaceutical/biotech commercial datasets such as IQVIA, Symphony, DDD, NPA, Xponent, Claims, Specialty Pharmacy, and CRM/Sales data.
The ideal candidate will partner closely with Commercial Analytics, Data Science, and Business Stakeholders to deliver high quality, scalable data products that enable sales insights, targeting, forecasting, incentive compensation, and field performance reporting.
Must Have
Must Have
Key Responsibilities
• Databricks Development
o Build, optimize, and maintain data pipelines using Databricks, Delta Lake, and Genie powered workflows.
o Implement LLM based automation and query generation using Databricks Genie for data exploration and business user self service.
o Develop notebooks, jobs, workflows, and ML ready datasets within the Databricks environment.
• AWS Cloud Engineering
o Design and manage data ingestion, storage, and processing using AWS services such as S3, Glue, Lambda, EMR, Redshift, and Athena.
o Implement secure, scalable architectures following AWS best practices (IAM, VPC, encryption, monitoring).
• PySpark & Big Data Processing
o Build distributed data transformation pipelines using PySpark for high volume commercial datasets.
o Optimize PySpark jobs for performance, cost efficiency, and reliability.
o Implement unit testing, CI/CD, and code versioning using Git based workflows.
• Life Sciences Commercial Data Expertise
o Ingest, harmonize, and model datasets including:
IQVIA (Xponent, DDD, NPA, LAAD, NSP)
Claims (medical, pharmacy)
Specialty Pharmacy data feeds
Sales & CRM (Veeva, Salesforce)
Roster, Territory, Alignment datasets
o Build commercial data marts supporting:
Sales reporting and dashboards
Targeting and segmentation
Incentive compensation
Forecasting and demand analytics
Field performance insights
• Cross Functional Collaboration
o Work with Commercial Analytics, Data Science, IT, and Business teams to translate requirements into scalable data solutions.
o Ensure data quality, lineage, governance, and compliance with industry standards (HIPAA, GxP, SOC2).
Required Qualifications
• 5–10+ years of experience in Data Engineering or Big Data Analytics.
• Strong hands on experience with Databricks, Genie, Delta Lake, and MLflow.
• Advanced proficiency in PySpark, Python, SQL, and distributed computing.
• Deep experience with AWS (S3, Glue, Lambda, EMR, Redshift, IAM).
• Proven experience working with Life Sciences Commercial datasets.
• Strong understanding of data modeling, ETL/ELT, and cloud architecture.
• Experience with CI/CD, Git, DevOps, and automated workflow orchestration.
Good to have
Preferred Qualifications
• Experience with Databricks Unity Catalog and Lakehouse architecture.
• Familiarity with LLM based automation or AI assisted analytics.
• Experience supporting Commercial Operations, Sales Leadership, and Field Teams.
• Knowledge of incentive compensation methodologies and targeting algorithms.
• Experience with BI tools (Tableau, Power BI, Qlik).
Role Competencies
• Strong analytical and problem solving skills.
• Excellent communication and documentation abilities.
• Ability to work in fast paced, cross functional environments.
• High attention to detail and commitment to data accuracy.