Experience
4 - 7 yrs
Job Location
Pune, India
Vacancy
3
Designation
Data Scientist
Job Type
Not specified
Job Description
Role & responsibilities
Model Development & Deployment: Build, train, and deploy advanced machine learning, deep learning, and AI models optimized for large-scale healthcare data aggregation, multi-modal biomedical datasets, and pre-clinical/clinical data structures.
- NLP & Biomedical Ontologies: Design and implement proprietary Natural Language Processing (NLP) pipelines and integrate biomedical ontologies to mine, parse, and extract structural insights from complex scientific publications, clinical trial registries, patents, and unstructured medical text.
- Cross-Functional Collaboration: Partner closely with engineering, product management, and business teams to translate complex data science research prototypes into scalable, high-impact commercial software products and production platform features.
- Data Quality & Compliance: Maintain stringent data quality, integrity, and privacy standards across all computational pipelines, ensuring systems are strictly compliant with global healthcare and pharmaceutical regulations (e.g., HIPAA, GDPR).
- Innovation & Tooling Evaluation: Evaluate and introduce novel analytical methodologies, scalable cloud-based ecosystems (AWS/GCP), and best-in-class frameworks to constantly improve life sciences data management and model lifecycle performance (MLOps).
- Technical Mentorship: Provide guidance and mentorship to junior and mid-level data scientists, cultivating a culture of technical excellence, continuous learning, and ethical data utilization.
- Stakeholder Insights: Translate multi-layered analytics into clear data visualizations and regular progress updates for key stakeholders to support commercial, clinical, and regulatory decision-making.
Preferred candidate profile
Education: Masters degree or Ph.D. in a quantitative discipline such as Computer Science, Data Science, Statistics, Bioinformatics, Computational Biology, or a related field. Extensive domain exposure within life sciences, digital health, or pharma is required.
- Years of Experience: 4+ years of hands-on professional experience in data science or AI R&D roles
- Technical Toolkit: Profound technical mastery in Python, R, machine learning frameworks (e.g., PyTorch, TensorFlow, Scikit-Learn), NLP architectures (Transformers, LLMs, BERT variants), and database systems (SQL, NoSQL, Graph Databases).
- Cloud & Production MLOps: Proven experience developing and shipping models within scalable cloud architectures (AWS, GCP, or Azure) and using containerization (Docker, Kubernetes).
- Regulated Data Handling: Hands-on experience working with highly sensitive, regulated datasets and a clear understanding of data governance structures under HIPAA or GDPR.
- Communication: Strong communication skills with an ability to clearly articulate complex computational models and statistical anomalies to both technical peers and business-oriented teams.
Preferred Attributes
- Domain Context: Strong knowledge of drug discovery pipelines, clinical workflows, or real-world data curation (EHR, claims data, omics datasets).
- Contributions: A record of peer-reviewed publications, patents, or open-source contributions in biomedical AI, NLP, or health-tech platforms.
- Agility: Prior experience performing under rapid growth pressures within a fast-paced health-tech startup environment.
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.