Experience
4 - 6 yrs
Salary (CTC)
₹14.1L - ₹15.9L
Job Location
Pune, India
Vacancy
1
Designation
Senior Data Engineer
Job Type
Not specified
Job Description
Job Purpose
To effectively design, develop, and manage data solutions using ETL technologies such as Azure Databricks (ADB) , Azure Data Factory (ADF) and SQL
Duties and Responsibilities
Data Engineering Platform Development
- Build scalable pipelines using Azure Databricks (PySpark, SQL)
- Develop and orchestrate ETL workflows using Azure Data Factory
- Work with Delta Lake architecture (BronzeSilverGold layers)
- Enable real-time and batch data processing pipelines
- Enable data exposure via APIs for BI and downstream systems
CI/CD DevOps
- Implement CI/CD pipelines for data and AI workflows
- Automate deployments across environments (Dev, QA, Prod)
- Ensure version control and reproducibility
Database Proficiency: Strong knowledge of SQL and experience with relational databases like SQL Server, MySQL, etc.
KEY RESPONSIBILITIES
- Translate business requirements into technical solutions in collaboration with the PMO team.
- Own end-to-end delivery of data projects, ensuring on-time execution and adherence to quality standards.
- Design technical architecture and guide development efforts for enhancements and new projects.
- Develop and maintain robust ETL pipelines and data integration modules across systems.
- Ensure high data quality, data anomaly resolution of critical process issues.
- Monitor and resolve performance bottlenecks in data workflows and programs.
- Establish best practices, standard operating procedures, and drive their implementation across teams.
- Act as a liaison with business users and product managers to support daily data needs and strategic initiatives.
- Coordinate with internal and external development teams to troubleshoot and resolve issues efficiently.
- Manage workload through effective planning, prioritization, and progress tracking.
Key Decisions / Dimensions
- Define semantic layer design and metric definitions
- Prioritize data vs AI optimization trade-offs
- Handle production issues with RCA and long-term fixes
- Drive architectural decisions for lakehouse + Data integration
Major Challenges
- Ensuring Data Delivery within TAT
- Driving adoption of GenAI-based BI over traditional dashboards
- Balancing performance, cost, and scalability
- Managing dependencies across data engineering, AI, and business teams
Required Qualifications and Experience
REQUIRED SKILLS EXPERIENCE
Must Have
- Azure Databricks PySpark, SQL, Delta Lake
- Semantic Modeling Metrics Layer design
- Hands-on with Databricks workflows
- Pyspark (Pandas, PySpark, FastAPI)
- Azure Data Factory (ADF) for ETL pipelines
- Strong SQL and data modeling skills
Good to Have
- Cosmos DB / MongoDB (NoSQL concepts)
- Azure Data Explorer (KQL)
DATA STACK (MANDATORY FOR SCREENING)
- Databricks Lakehouse - PySpark, SQL, Delta Lake
- AI for BI - Databricks Genie, Genie Rooms, Instructions, Agents
- ETL Orchestration - Azure Data Factory
- Programming - Pyspark
- Cloud Platform - Azure (Preferred)
- DevOps - CI/CD Pipelines, Git
