Job Description
We are seeking a skilled and motivated Data Engineer with expertise in PySpark to join our data engineering team. You will be responsible for building scalable data pipelines, transforming large volumes of data, and supporting data analytics and machine learning initiatives across the organization.
Key Responsibilities:-
Design, develop, and maintain large-scale data pipelines using PySpark and Apache Spark.
-
Work with structured and unstructured data from various sources including databases, APIs, and streaming platforms.
-
Optimize ETL processes for performance, scalability, and reliability.
-
Collaborate with data scientists, analysts, and other engineers to understand data requirements and deliver solutions.
-
Implement data quality checks and validation procedures.
-
Support batch and real-time data processing needs.
-
Monitor and troubleshoot production data pipelines and jobs.
-
Document technical processes, architecture, and pipeline workflows.
