Experience
4 - 9 yrs
Salary (CTC)
₹23.7L - ₹32.2L
Job Location
Bengaluru, India
Vacancy
1
Designation
Senior Site Reliability Engineer
Job Type
Not specified
Job Description
Job Summary
Wells Fargo is seeking a Senior Systems Operations Engineer.
Responsibilities- Lead or participate in managing all installed systems and infrastructure within the Systems Operations functional area
- Contribute in increasing system efficiencies and lowering the human intervention time on related tasks
- Review and analyze moderately complex operational support systems, application software, and system management tools to ensure the highest levels of systems and infrastructure availability
- Work with vendors and other technical personnel for problem resolution
- Lead team to meet technical deliverables while leveraging solid understanding of technical process controls or standards
- Collaborate with vendors and other technical personnel to resolve technical issues and achieve highest levels of systems and infrastructure availability
- 4+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
- 4 years of application development, implementation and Support experience (JAVA / .Net )
- 3+ years of strong experience on production monitoring tools such as Splunk/Google Cloud Logging(GCL), AppDynamics etc.
- 2+ years of experience working on any Cloud environments IBM OpenShift, Azure, GCP, Pivotal Cloud Foundry(PCF)
- Strong hands-on experience with OpenShift and container orchestration platform like Kubernetes.
- Experience with on-premise to cloud migration, including hybrid environments.
- 2+ years of strong experience working on Java, React, Mongo DB, Kafka, Spring Microservices, API development.
- Proficiency in scripting and automation using languages like Python, Bash etc.
- Expertise using observability tools such as Prometheus, Thousand Eyes and Grafana.
- 2+ years of experience in one or a combination of the following: Agile, Kanban, or Lean methodology
- Strong in Oracle and SQL Database queries, procedures, functions, stored procedures, performance tuning, PL/SQL.
- Excellent verbal, written, and interpersonal communication skills
- Outstanding problem solving and decision making skills
- Ability to help triage complex technical support issues In-depth technical solution knowledge, including installation, configuration, performance tuning
- Strong analytical skills with high attention to detail and accuracy
- Strong leadership skills, ability to lead and drive platform transformation initiatives in the team by taking up more ownership and accountability.
- Ability to develop partnerships and collaborate with other business and functional area
- Core responsibilities for this Site Reliability Engineer (SRE) position are:
- Proficient in software development leveraging OOPS Programming concepts (JAVA / .Net), React, Mongo DB, Kafka, Spring Microservices, API development.
- Responsible for Driving SRE Initiatives (Code Analysis and Solutioning, Automation, Instrumentation, Resiliency, Reliability and Stability Improvement).
- Automate key SRE metrics and processes including customer impact, % availability of critical business flows, SLO/SLI adherence, error budget, alerting/notification systems, and to reduce time to recovery.
- Enhance end to end application or system observability by creating new alerts and developing dashboards using the monitoring/log analysis/analytic tools such as Splunk, AppDynamics, Elastic Search, PowerBI, Tableau etc
- Use Splunk, AppDynamics and other monitoring tools to identify, analyze and resolve complex production issues in real-time.
- Collaborate with teams to migrate on-premise environments to cloud platforms like IBM OpenShift, Azure or GCP.
- Onboard critical customer journeys to assess the availability of critical business flows, identify service level objectives and indicators, instrument applications for observability, onboard to CI/CD pipeline taking advantage of continuous testing, introduce continuous inspection, continuous improvement, and conduct destructive testing to reach 99.995% availability for the firms critical products and services leading to higher customer satisfaction and customer experience.
- UnderstandsApplication logging standards and able to engineer health check and monitoring solutions
- Strong knowledge on RDBMS and ability to design Entity Relationship in conjunction to development work
- Proficient with Database Programming, SQL, Stored Procedures
- Incident Management: Triage incidents, engage partner teams, provide status updates and facilitate business user communication.
- Problem Management: Ticket management for daily tasks and efforts that are brought to support attention. Root cause analysis.
- Monitoring: Implementation of Alerts and Configuration - Customize alerting tools based on application specific thresholds. Enable business transaction monitoring.
- Automation of manual production support activities Daily Health Check, BCP Fail over, Daily/Weekly Reporting, Application Pool recycling and Start/Stop scripts etc.
- Identify bottlenecks in Batch processing and provide solutions to improve performance
- Support communication efforts with application teams.
