Experience
5 - 8 yrs
Salary (CTC)
₹7L - ₹10.4L
Job Location
Hyderabad, India
Vacancy
1
Designation
AWS Data Engineer
Job Type
ONSITE
Job Description
Position Description:
Founded in 1976, CGI is among the largest independent IT and business consulting services firms in the world. With 94,000 consultants and professionals across the globe, CGI delivers an end-to-end portfolio of capabilities, from strategic IT and business consulting to systems integration, managed IT and business process services and intellectual property solutions. CGI works with clients through a local relationship model complemented by a global delivery network that helps clients digitally transform their organizations and accelerate results. CGI Fiscal 2024 reported revenue is CA$14.68 billion and CGI shares are listed on the TSX (GIB.A) and the NYSE (GIB). Learn more at cgi.com.
Job Title: AWS Data Engineer / PySpark Developer
Position: Senior Software Engineer
Experience: 5- 8 Years
Category: Software Development/ Engineering
Shift: General
Main location: Hyderabad, Bangalore, Chennai
Position ID: J0726-2568
Employment Type: Full Time
Education Qualification: Bachelor's degree in Computer Science or related field or higher with minimum 4 years of relevant experience.
We are looking for an experienced AWS Data Engineer / PySpark Developer with around 5 years of hands-on experience in AWS, Python, PySpark, SQL, and NoSQL databases. The candidate should have strong experience in developing scalable data solutions, cloud-based ETL pipelines, production support, debugging, CI/CD, and containerized environments.
The ideal candidate will possess strong problem-solving and analytical skills, along with the ability to collaborate effectively with technical teams, business stakeholders, and cross-functional teams in an Agile environment.
Job Title: AWS Data Engineer / PySpark Developer
Position: Senior Software Engineer
Experience: 5- 8 Years
Category: Software Development/ Engineering
Shift: General
Main location: Hyderabad, Bangalore, Chennai
Position ID: J0726-2568
Employment Type: Full Time
Education Qualification: Bachelor's degree in Computer Science or related field or higher with minimum 4 years of relevant experience.
We are looking for an experienced AWS Data Engineer / PySpark Developer with around 5 years of hands-on experience in AWS, Python, PySpark, SQL, and NoSQL databases. The candidate should have strong experience in developing scalable data solutions, cloud-based ETL pipelines, production support, debugging, CI/CD, and containerized environments.
The ideal candidate will possess strong problem-solving and analytical skills, along with the ability to collaborate effectively with technical teams, business stakeholders, and cross-functional teams in an Agile environment.
Your future duties and responsibilities:
Design, develop, and maintain scalable data engineering and ETL solutions using Python and PySpark.
Develop and optimize PySpark applications for large-scale data processing.
Troubleshoot performance issues, data quality problems, job failures, and other challenges in PySpark-based applications.
Develop data pipelines using AWS Glue and implement workflow orchestration using AWS Step Functions and/or Apache Airflow.
Work with AWS services including IAM, Lambda, Glue, Redshift, and CloudWatch.
Develop and optimize complex SQL queries for data extraction, transformation, and analysis.
Work with NoSQL databases such as MongoDB and MongoDB Atlas.
Implement data solutions that meet scalability, reliability, security, and performance requirements.
Monitor production data pipelines and troubleshoot failures using AWS CloudWatch and application logs.
Perform root-cause analysis for production incidents and document issues, resolutions, and preventive actions.
Develop and maintain unit and automation tests using frameworks/tools such as JUnit, Cucumber, and Selenium, as applicable.
Implement and maintain CI/CD pipelines using Jenkins and follow established deployment processes.
Work with Docker and Kubernetes, including AWS container orchestration platforms such as EKS/ECS.
Participate in deployment activities across development, testing, and production environments.
Follow DevOps best practices for source control, automated builds, testing, deployment, monitoring, and release management.
Develop dashboards and reports using Power BI, preferably, to provide meaningful insights into business and operational data.
Participate in Agile/Scrum ceremonies, sprint planning, daily stand-ups, retrospectives, and backlog discussions.
Collaborate with developers, architects, QA teams, DevOps engineers, business analysts, and other stakeholders.
Communicate technical issues, solutions, project status, and risks effectively to technical and non-technical stakeholders.
Develop and optimize PySpark applications for large-scale data processing.
Troubleshoot performance issues, data quality problems, job failures, and other challenges in PySpark-based applications.
Develop data pipelines using AWS Glue and implement workflow orchestration using AWS Step Functions and/or Apache Airflow.
Work with AWS services including IAM, Lambda, Glue, Redshift, and CloudWatch.
Develop and optimize complex SQL queries for data extraction, transformation, and analysis.
Work with NoSQL databases such as MongoDB and MongoDB Atlas.
Implement data solutions that meet scalability, reliability, security, and performance requirements.
Monitor production data pipelines and troubleshoot failures using AWS CloudWatch and application logs.
Perform root-cause analysis for production incidents and document issues, resolutions, and preventive actions.
Develop and maintain unit and automation tests using frameworks/tools such as JUnit, Cucumber, and Selenium, as applicable.
Implement and maintain CI/CD pipelines using Jenkins and follow established deployment processes.
Work with Docker and Kubernetes, including AWS container orchestration platforms such as EKS/ECS.
Participate in deployment activities across development, testing, and production environments.
Follow DevOps best practices for source control, automated builds, testing, deployment, monitoring, and release management.
Develop dashboards and reports using Power BI, preferably, to provide meaningful insights into business and operational data.
Participate in Agile/Scrum ceremonies, sprint planning, daily stand-ups, retrospectives, and backlog discussions.
Collaborate with developers, architects, QA teams, DevOps engineers, business analysts, and other stakeholders.
Communicate technical issues, solutions, project status, and risks effectively to technical and non-technical stakeholders.
Required qualifications to be successful in this role:
Must Have Skills:
Programming: Python
Big Data: PySpark, Apache Spark
Database: SQL, MongoDB, MongoDB Atlas
AWS: IAM, AWS Glue, Lambda, Redshift, CloudWatch
Orchestration: AWS Step Functions, Apache Airflow
Containers: Docker
Container Orchestration: Kubernetes, Amazon EKS/ECS
CI/CD: Jenkins and DevOps practices
Testing: JUnit, Cucumber, Selenium
Reporting/Visualization: Power BI preferred
Development Practices: Agile/Scrum, SDLC, code review, testing, deployment, and production support
Preferred Experience
Hands-on experience building and supporting AWS-based data pipelines.
Strong understanding of PySpark optimization, debugging, and performance tuning.
Experience troubleshooting production failures and performing root-cause analysis.
Experience with enterprise CI/CD and deployment processes.
Understanding of cloud security concepts, particularly AWS IAM and access management.
Experience working with distributed data processing and large datasets.
Knowledge of monitoring, logging, alerting, and operational support.
Experience working in an Agile delivery environment.
Key Competencies
Strong problem-solving and debugging skills.
Excellent analytical and troubleshooting abilities.
Good understanding of software development and deployment processes.
Strong communication and collaboration skills.
Ability to work effectively with cross-functional and geographically distributed teams.
Ability to understand business requirements and translate them into technical solutions.
Strong ownership and accountability for deliverables and production issues.
Role Expectations
The candidate should be comfortable working across the complete data engineering lifecycle, including:
Requirement Analysis Development Unit Testing Code Review CI/CD Deployment Monitoring Production Support Root Cause Analysis Documentation
The role requires a combination of hands-on technical development, cloud data engineering, production troubleshooting, DevOps practices, and stakeholder communication.
Programming: Python
Big Data: PySpark, Apache Spark
Database: SQL, MongoDB, MongoDB Atlas
AWS: IAM, AWS Glue, Lambda, Redshift, CloudWatch
Orchestration: AWS Step Functions, Apache Airflow
Containers: Docker
Container Orchestration: Kubernetes, Amazon EKS/ECS
CI/CD: Jenkins and DevOps practices
Testing: JUnit, Cucumber, Selenium
Reporting/Visualization: Power BI preferred
Development Practices: Agile/Scrum, SDLC, code review, testing, deployment, and production support
Preferred Experience
Hands-on experience building and supporting AWS-based data pipelines.
Strong understanding of PySpark optimization, debugging, and performance tuning.
Experience troubleshooting production failures and performing root-cause analysis.
Experience with enterprise CI/CD and deployment processes.
Understanding of cloud security concepts, particularly AWS IAM and access management.
Experience working with distributed data processing and large datasets.
Knowledge of monitoring, logging, alerting, and operational support.
Experience working in an Agile delivery environment.
Key Competencies
Strong problem-solving and debugging skills.
Excellent analytical and troubleshooting abilities.
Good understanding of software development and deployment processes.
Strong communication and collaboration skills.
Ability to work effectively with cross-functional and geographically distributed teams.
Ability to understand business requirements and translate them into technical solutions.
Strong ownership and accountability for deliverables and production issues.
Role Expectations
The candidate should be comfortable working across the complete data engineering lifecycle, including:
Requirement Analysis Development Unit Testing Code Review CI/CD Deployment Monitoring Production Support Root Cause Analysis Documentation
The role requires a combination of hands-on technical development, cloud data engineering, production troubleshooting, DevOps practices, and stakeholder communication.
Skills:
- AWS Glue
- AWS Lambda
- AWS S3
- PowerBuilder
- Python
- Docker
- Kubernetes
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
