Experience
5 - 8 yrs
Job Location
Bengaluru, India
Vacancy
1
Designation
Data Engineer
Job Type
Not specified
Job Description
PFB the updated JD for Data Engineer role.
Kindly share the Quality profiles for Pune location.
ECMS : 518472
Location : Pune
Relevant Experience : 5+ Years
Total years of experience: 7+ Years.
Billing Rate : 9000 to 10000 INR per day
Were looking for a Big Data Lead Engineer to:
engineer reliable data pipelines for sourcing, processing, distributing, and storing data in different ways, using cloud (Azure) data platform infrastructure effectively.
transform data into valuable insights that inform business decisions, making use of our internal data platforms and applying appropriate analytical techniques.
develop, train, and apply data engineering techniques to automate manual processes, and solve challenging business problems.
ensure the quality, security, reliability, and compliance of our solutions by applying our digital principles and implementing both functional and non-functional requirements.
build observability into our solutions, monitor production health, help to resolve incidents, and remediate the root cause of risks and issues.
understand, represent, and advocate for client needs.
Codify best practices, methodology and share knowledge with other engineers in Banking
have a continuous improvement mindset, who is always on the look out for ways to automate and reduce time to market for deliveries.
Your expertise
Extensive experience in building Data Processing pipelines using Apache Spark/Databricks with Python
and PySpark
Good Knowledge of inner working on Apache Spark. Structured streaming is a plus.
Deep understanding of Python and its ecosystem, principles and tooling that helps to write production
grade applications e.g. PEP8, MyPy, PyLint, Pytest.
Good knowledge of data design patterns and methodologies to build a data lake, based Azure cloud
stack e.g. ADLSv2.
Experience in creating data structures optimized for storage and various query patterns for DeltaLake,
Parquet, Avro
Deep understanding of the SDLC using Gitlab, Github and knowledge of CI/CD is a plus.
Good to have working experience in cloud (Azure is preferrable). Knowledge of (Kafka or Event Hub) is
plus
Datalakehouses using medallion architecture. Knowledge of DataMesh principles is a plus.
Ability to debug using tools Spark UI, Ganglia UI, expertise in Optimizing Spark Jobs
The ability to work across structured, semi-structured, and unstructured data, extracting information and
identifying linkages across disparate datasets.
Experience of building applications using Polars, Pandas, Numpy is plus.
Experience of building microservices on Kubernetes is plus
Experience in traditional data warehousing concepts (Kimball Methodology, Star Schema, SCD2).
Experience in orchestration tools like Azure Databricks Workflow, Apache Airflow is plus.
Ability to clearly communicate complex solutions.
Strong problem solving and analytical skills.
Working experience in Agile methodologies (SCRUM)
A proven team player with strong leadership skills, who can work in a collaborative way across business units, teams and regions.
ECMS ID#
518472
Number of Openings*
3
Duration of contract*
6 months
Total Yrs. of Experience*
5 to 8 yrs
Domain*
Financial Services
Detailed JD
Were looking for a Big Data Lead Engineer to:
engineer reliable data pipelines for sourcing, processing, distributing, and storing data in different ways, using cloud (Azure) data platform infrastructure effectively.
transform data into valuable insights that inform business decisions, making use of our internal data platforms and applying appropriate analytical techniques.
develop, train, and apply data engineering techniques to automate manual processes, and solve challenging business problems.
ensure the quality, security, reliability, and compliance of our solutions by applying our digital principles and implementing both functional and non-functional requirements.
build observability into our solutions, monitor production health, help to resolve incidents, and remediate the root cause of risks and issues.
understand, represent, and advocate for client needs.
Codify best practices, methodology and share knowledge with other engineers in Banking
have a continuous improvement mindset, who is always on the look out for ways to automate and reduce time to market for deliveries.
Your expertise
Extensive experience in building Data Processing pipelines using Apache Spark/Databricks with Python
and PySpark
Good Knowledge of inner working on Apache Spark. Structured streaming is a plus.
Deep understanding of Python and its ecosystem, principles and tooling that helps to write production
grade applications e.g. PEP8, MyPy, PyLint, Pytest.
Good knowledge of data design patterns and methodologies to build a data lake, based Azure cloud
stack e.g. ADLSv2.
Experience in creating data structures optimized for storage and various query patterns for DeltaLake,
Parquet, Avro
Deep understanding of the SDLC using Gitlab, Github and knowledge of CI/CD is a plus.
Good to have working experience in cloud (Azure is preferrable). Knowledge of (Kafka or Event Hub) is
plus
Datalakehouses using medallion architecture. Knowledge of DataMesh principles is a plus.
Ability to debug using tools Spark UI, Ganglia UI, expertise in Optimizing Spark Jobs
The ability to work across structured, semi-structured, and unstructured data, extracting information and
identifying linkages across disparate datasets.
Experience of building applications using Polars, Pandas, Numpy is plus.
Experience of building microservices on Kubernetes is plus
Experience in traditional data warehousing concepts (Kimball Methodology, Star Schema, SCD2).
Experience in orchestration tools like Azure Databricks Workflow, Apache Airflow is plus.
Ability to clearly communicate complex solutions.
Strong problem solving and analytical skills.
Working experience in Agile methodologies (SCRUM)
A proven team player with strong leadership skills, who can work in a collaborative way across business
units, teams and regions
engineer reliable data pipelines for sourcing, processing, distributing, and storing data in different ways, using cloud (Azure) data platform infrastructure effectively.
transform data into valuable insights that inform business decisions, making use of our internal data platforms and applying appropriate analytical techniques.
develop, train, and apply data engineering techniques to automate manual processes, and solve challenging business problems.
ensure the quality, security, reliability, and compliance of our solutions by applying our digital principles and implementing both functional and non-functional requirements.
build observability into our solutions, monitor production health, help to resolve incidents, and remediate the root cause of risks and issues.
understand, represent, and advocate for client needs.
Codify best practices, methodology and share knowledge with other engineers in Banking
have a continuous improvement mindset, who is always on the look out for ways to automate and reduce time to market for deliveries.
Your expertise
Extensive experience in building Data Processing pipelines using Apache Spark/Databricks with Python
and PySpark
Good Knowledge of inner working on Apache Spark. Structured streaming is a plus.
Deep understanding of Python and its ecosystem, principles and tooling that helps to write production
grade applications e.g. PEP8, MyPy, PyLint, Pytest.
Good knowledge of data design patterns and methodologies to build a data lake, based Azure cloud
stack e.g. ADLSv2.
Experience in creating data structures optimized for storage and various query patterns for DeltaLake,
Parquet, Avro
Deep understanding of the SDLC using Gitlab, Github and knowledge of CI/CD is a plus.
Good to have working experience in cloud (Azure is preferrable). Knowledge of (Kafka or Event Hub) is
plus
Datalakehouses using medallion architecture. Knowledge of DataMesh principles is a plus.
Ability to debug using tools Spark UI, Ganglia UI, expertise in Optimizing Spark Jobs
The ability to work across structured, semi-structured, and unstructured data, extracting information and
identifying linkages across disparate datasets.
Experience of building applications using Polars, Pandas, Numpy is plus.
Experience of building microservices on Kubernetes is plus
Experience in traditional data warehousing concepts (Kimball Methodology, Star Schema, SCD2).
Experience in orchestration tools like Azure Databricks Workflow, Apache Airflow is plus.
Ability to clearly communicate complex solutions.
Strong problem solving and analytical skills.
Working experience in Agile methodologies (SCRUM)
A proven team player with strong leadership skills, who can work in a collaborative way across business
units, teams and regions
Mandatory skills
Azure databricks, Azure data factory, Python, Pyspark
BGCheck (Pre onboarding Or Post onboarding)
Before Onboarding
Any client prerequisite BGV Agency*
FADV
Is there any working in shifts from standard Daylight (to avoid confusions post onboarding) *
Regular timings
