Experience
10 - 15 yrs
Salary (CTC)
₹25L - ₹30L
Job Location
Chennai, India
Vacancy
1
Designation
Senior Data Scientist
Job Type
Not specified
Job Description
Role & responsibilities
Senior Data Scientist
We are looking for a Senior Data Scientist with deep expertise in machine learning, intelligent document processing, and cloud-native AI pipelines. You will lead the design and delivery of production-grade AI/ML components within a governed, human-in-the-loop automation platform, working across the full spectrum from raw data ingestion to model deployment and continuous improvement.
Key Responsibilities
- Architect and implement multi-tier document extraction pipelines combining deterministic parsing, OCR, and large language model (LLM) inference, applying the right technique for each document category.
- Design and train document classification models (e.g., using Amazon SageMaker) for routing documents to appropriate extraction strategies.
- Develop and optimize prompt engineering patterns for LLM-based extraction (Amazon Bedrock), including confidence scoring, source-region anchoring, and hallucination detection strategies.
- Build and maintain data validation frameworks: schema validation, cross-field consistency checks, and configurable confidence thresholds per field and document type.
- Design canonical data models and normalization pipelines that consolidate heterogeneous source data into a unified schema with full field-level lineage.
- Lead the design of agentic AI components (extraction orchestration, layout detection, exception analysis) with clearly defined capability envelopes, escalation logic, and human approval gates.
- Define and track extraction accuracy metrics; drive continuous improvement loops using structured human feedback.
- Collaborate with solution architects to ensure ML components integrate cleanly with serverless AWS infrastructure (Lambda, Step Functions, S3, Glue, DynamoDB, EventBridge).
- Mentor junior team members and establish best practices for model governance, change management, and reproducibility.
Machine Learning & AI
- Document classification using supervised ML (SageMaker or equivalent)
- LLM prompt engineering for structured data extraction (Bedrock, OpenAI, or similar)
- OCR post-processing: table reconstruction, merged cell handling, multi-line value normalization (Amazon Textract or equivalent)
- Confidence scoring design and threshold calibration
- Cosine similarity and document fingerprinting for template matching
- Agentic AI orchestration patterns with governed fallback and escalation design
Data Engineering & Python
- Python: pandas, openpyxl, boto3 production-grade, not notebook-only
- ETL pipeline design for structured (Excel/CSV), semi-structured (PDF tables), and unstructured document types
- Schema design and canonical data modeling
- Data quality frameworks: completeness detection, duplicate identification, cross-field validation
- SQL and relational data modeling (PostgreSQL preferred)
Cloud - AWS
- AWS Lambda, Step Functions, S3, SageMaker, Bedrock, Textract, Glue, DynamoDB, EventBridge, SQS, KMS
- Serverless-first architecture patterns; cost-efficient compute design for seasonal/batch workloads
Governance & MLOps
- Model registry, versioning, and change board processes
- Audit trail design: immutable lineage from source document to output
- CI/CD integration for ML pipeline components
- RBAC and data security in multi-tenant cloud environments
Experience
- 12+ years in data science or ML engineering roles
- At least 2 production deployments involving intelligent document processing or NLP pipelines
- Demonstrated experience designing human-in-the-loop systems, not just fully automated models
Preferred candidate profile
- Power BI or Amazon QuickSight for operational dashboards
- React or familiarity with audit workbench UI requirements for human-in-the-loop review queues
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
