Experience
5 - 10 yrs
Job Location
Gurugram, India
Vacancy
1
Designation
Data Engineer
Job Type
Not specified
Job Description
Job Summary
We are seeking an experienced Data Engineer (5-8 years) to design, build, and operate scalable, production-grade data platforms on AWS. In this role you will own the full lifecycle of data pipelines - from ingestion and orchestration through transformation, storage, and governance - using a modern lakehouse stack built on Apache Iceberg, AWS Glue, and Snowflake. You will partner with data scientists, analysts, and platform teams to deliver reliable, well-governed, and cost-efficient data products that power analytics and business decisions across Insurance Systems.
Responsibilities
- Design, build, and maintain scalable, fault-tolerant data pipelines and ETL/ELT processes for structured, semi-structured, and unstructured data.
- Own end-to-end pipeline orchestration - scheduling, dependency management, retries, SLAs, and observability - using a data pipeline orchestrator.
- Build and manage lakehouse data assets on Apache Iceberg, including partitioning, schema evolution, time-travel, and table maintenance (compaction, snapshot expiry).
- Develop and operate large-scale data processing jobs using Apache Spark, including AWS Glue Spark jobs.
- Model, load, and optimize data in Snowflake, and govern data assets using data catalogs (AWS Glue Data Catalog and Snowflake Horizon).
- Architect and implement data solutions on AWS using native services (S3, Glue, EMR, Lambda, Athena, Kinesis, Redshift, IAM).
- Ensure data quality, integrity, lineage, and security across all systems and environments.
- Integrate data from diverse sources including APIs, relational databases, streaming feeds, and third-party tools.
- Monitor, troubleshoot, and tune pipelines for performance, reliability, and cost efficiency.
- Contribute to platform architecture, design reviews, code reviews, and engineering best practices.
- Mentor junior engineers and help drive standards across the data engineering function.
Qualifications
- 5-8 years of hands-on data engineering experience building production data pipelines.
- Strong proficiency in Python for data engineering and automation.
- Advanced SQL and strong experience with relational databases (e.g., PostgreSQL, MySQL).
- A data pipeline orchestrator is mandatory - proven experience operating a workflow orchestration tool (e.g., Apache Airflow, Dagster, or equivalent) in production.
- Hands-on experience with Apache Spark for large-scale distributed data processing.
- Production experience with Apache Iceberg (or an equivalent open table format) for lakehouse storage.
- Hands-on experience with AWS Glue (ETL jobs and the Glue Data Catalog).
- Experience with Snowflake as a cloud data warehouse.
- Experience with data cataloging and governance using AWS Glue Data Catalog and Snowflake Horizon.
- Strong experience with AWS as the primary cloud platform, including S3, EMR, Lambda, Athena, Kinesis, Redshift, and IAM.
- Solid understanding of data modeling, data warehousing, and lakehouse architecture patterns.
- Working knowledge of REST APIs and data integration techniques.
- Strong problem-solving, analytical, and debugging skills.
Preferred / Good-to-Have
- Experience with Dagster as a data pipeline orchestrator (strongly preferred).
- Experience with containerization and orchestration (Docker, Kubernetes / Amazon EKS).
- Exposure to CI/CD pipelines and infrastructure-as-code (e.g., Terraform, AWS CloudFormation/CDK).
- Experience with streaming / real-time data (Kafka, Amazon Kinesis, Spark Structured Streaming).
- Familiarity with data observability and quality frameworks (e.g., Great Expectations, dbt tests).
- Knowledge of data governance, security, and compliance best practices.
- Experience within financial services or insurance data domains.
Technology Stack at a Glance
- Cloud Platform - AWS (S3, Glue, EMR, Lambda, Athena, Kinesis, Redshift, IAM)
- Orchestration - Data pipeline orchestrator required (e.g., Airflow); Dagster good-to-have
- Processing - Apache Spark, AWS Glue
- Lakehouse / Storage - Apache Iceberg on Amazon S3
- Data Warehouse - Snowflake
- Data Catalog / Governance - AWS Glue Data Catalog, Snowflake Horizon
- Languages - Python, SQL
Soft Skills
- Strong communication and cross-functional collaboration skills.
- Ability to work in a fast-paced, agile environment.
- Self-driven with a proactive, ownership-oriented mindset.
- Ability to mentor peers and communicate technical concepts to non-technical stakeholders.
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.