Machine Learning Lead

Process9
Posted on
Process9 logo

Experience
4 - 8 yrs
Job Location
Gurugram, India
Vacancy
2
Designation
Lead Machine Learning Engineer
Job Type
Not specified

Job Description

Key Responsibilities

  • Model Training & Fine-Tuning: Build, fine-tune, and optimize state-of-the-art NLP, LLM, Speech, and Vision models for scheduled Indian languages, utilizing parameter-efficient methods (LoRA, QLoRA, PEFT).
  • Indic Tokenization & Linguistics: Architect custom tokenizers and text-normalization pipelines to address the "fertility problem" in Devanagari, Dravidian, and other regional scripts, ensuring low-latency and cost-effective model inference.
  • Multimodal System Design: Develop robust OCR engines capable of parsing complex script geometries (conjoint consonants, Shirorekha, vowel modifiers) and integrate them into document intelligence pipelines.
  • Speech Engineering: Deploy and scale robust STT (Speech-to-Text) and TTS (Text-to-Speech) pipelines capable of handling heavy code-mixing (e.g., Hinglish, Tanglish), regional accents, and localized dialects.
  • Vernacular Guardrails & Evaluation: Establish culturally contextual benchmark datasets and implement safety guardrails.
  • Production Deployment (MLOps): Package and serve models using high-throughput frameworks (vLLM, Triton, ONNX) optimized for GPU environments, minimizing computational overhead for massive cross-lingual workloads.
  • Vernacular Fraud & Anomaly Detection: Architect risk-scoring systems and anomaly detection models capable of identifying fraud patterns in native scripts and code-mixed formats.

Essential Qualifications & Technical Skills

  • Education: Bachelors or Master's degree in Computer Science, Mathematics, Statistics, or a closely related quantitative field.
  • Experience: 4+ years of professional experience building and deploying machine learning models in production environments, with a proven track record in Indian Language NLP, Speech, or Anomaly Detection.
  • Programming: Expert-level proficiency in Python and standard ML frameworks (PyTorch, TensorFlow).
  • Indic AI Stack: Direct, hands-on experience with specialized Indic frameworks and datasets (e.g., AI4Bharat's IndicTrans2/IndicWhisper, Bhashini API, Kathbath, Sarvam-105B, or Aksharantar).
  • Fraud Stack: Proficiency in tabular/graph-based ML toolkits (XGBoost, LightGBM, PyTorch Geometric) and handling highly imbalanced target variables (SMOTE, class weights).
  • NLP & LLMs: Deep understanding of Transformer architectures, sequence-to-sequence modeling, cross-lingual embeddings, vector databases (Milvus, Pinecone, Qdrant), and quantization tools (bitsandbytes, GPTQ).
  • Speech & Vision Processing: Experience processing raw audio signals (grapheme-to-phoneme conversion, spectrogram analysis) or document structures using OCR networks (CRAFT, DBNet, LayoutLM).
  • Handling Code-Mixing: Proven ability to build models that gracefully parse text or speech containing heavy code-switching (mixed Latin/regional scripts, multi-language grammar).

No Referrers Available

There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.