Experience
3 - 8 yrs
Job Location
Bengaluru, India
Vacancy
1
Designation
Site Reliability Engineer
Job Type
ONSITE
Job Description
Job Summary:
This role contributes to Credit Production Engineerings mission of providing comprehensive 24/7 live production support for PayPals Credit Engineering organization. The engineer serves as part of the dedicated production support team embedded within Credit Engineering, supporting the complete product Credit portfolio of products including Revolving Credit, Closed-Ended, and Merchant Lending products. They collaborate within CPEs global coverage model across APAC and North America time zones, participating in incident response and monitoring activities while working under guidance to improve operational excellence. Job Description:
Essential Responsibilities:
- Actively monitor and analyze system metrics to ensure the availability, performance, and reliability of digital platforms and applications.
- Diagnose and resolve complex system issues, perform root cause analysis, and implement long-term fixes to prevent recurrence.
- Create and maintain automation scripts, tools, and processes to streamline operations, reduce manual effort, and enhance reliability.
- Configure and improve monitoring and alerting tools to provide actionable insights into system health and performance.
- Analyze system usage trends and forecast future resource requirements to ensure scalability and prevent capacity-related issues.
- Work with development teams to design and implement reliable, fault-tolerant systems, incorporating best practices for high availability.
- Ensure smooth and reliable software releases by managing continuous integration and continuous delivery (CI/CD) pipelines.
- Perform failure simulations, such as chaos engineering, to identify weaknesses and improve system robustness.
- Develop and maintain detailed documentation for system configurations, operational procedures, and incident handling workflows.
- Partner with development, operations, and product teams to enhance system architecture for improved performance, reliability, and scalability.
Minimum Qualifications:
- 1+ years relevant experience and a Bachelor s degree OR Any equivalent combination of education and experience.
Preferred Qualifications
- 3+ years of experience in Production Engineering, Site Reliability Engineering, or similar roles
- Strong problem-solving skills with the ability to debug complex, distributed, multi-tier applications
- Proven experience leading and driving resolution of high-severity production incidents, including coordinating across teams
- Demonstrated ability to identify recurring issues and drive long-term fixes and operational improvements
- Strong understanding of microservices architecture and the software development lifecycle (SDLC); proficiency in at least one programming language (e.g., Java, Python) to debug issues and collaborate with development teams
- Hands-on experience with monitoring and alerting tools (e.g., Splunk, Datadog, Nagios, Kibana)
- Solid understanding of databases, including SQL and stored procedures (BigQuery, Oracle, PostgreSQL)
- Experience working in cloud environments (AWS preferred)
- Proficiency in Unix/Linux systems and shell scripting
- Experience with batch job schedulers or workflow orchestration tools (e.g., Control-M, Airflow, UC4)
- Familiarity with incident management and collaboration tools (JIRA, Confluence, ServiceNow)
- Strong verbal and written communication skills, with the ability to influence and collaborate across cross-functional teams
Standout Qualifications (Nice to Have)
- Experience building automation to reduce operational toil (e.g., scripting, tooling, runbook automation)
- Experience mentoring junior engineers or leading incident reviews/postmortems
Travel Percent:
0
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
