TymblHub

© 2026 TymblHub

Senior Site Reliability Engineer

NCR Voyix
Posted on
NCR Voyix logo

Experience
7 - 11 yrs
Salary (CTC)
₹20L - ₹25L
Job Location
Chennai, India
Vacancy
1
Designation
Senior Site Reliability Engineer
Job Type
ONSITE

Job Description

Senior Site Reliability Engineer (DevOps/SRE Unified Observability)

We are seeking a Senior Site Reliability Engineer (Unified Observability) to support the F1 Next Generation Customer Unified Observability initiative. This role will design, build, and operate a unified observability platform delivering end-to-end visibility across NCR Voyix Restaurants, Retail, and Payments environments.

The engineer will help establish a single operational view spanning infrastructure, applications, Kubernetes platforms, cloud services, customer experience, and business transactions—enabling proactive operations, faster incident response, and improved service reliability.

Key Responsibilities

Design and implement enterprise observability solutions across Azure, GCP, Kubernetes (AKS/GKE), and hybrid environments.
Define and implement observability standards for monitoring, logging, distributed tracing, and telemetry, aligned to SRE and DevOps best practices.
Build and maintain enterprise dashboards for real-time visibility into platform, application, and customer health.
Implement and operationalize SLIs, SLOs, error budgets, and reliability KPIs across critical services.
Integrate observability with ServiceNow.
Drive automation-first operations for SRE workflows using Ansible, Chef, Rundeck, Terraform, Bash, and PowerShell.
Partner with Product Engineering, Infrastructure, Security, and Operations teams to improve reliability and operational readiness.
Enable DevOps delivery and reliability through CI/CD pipelines using GitHub Actions and GitOps practices.

Required Skills & Experience

Strong hands-on experience in Site Reliability Engineering (SRE) and DevOps.
Deep expertise in Kubernetes platforms, especially AKS and/or GKE.
Strong cloud operations experience in Microsoft Azure and Google Cloud Platform (GCP).
Strong infrastructure automation and IaC expertise using Terraform.
Proficiency in scripting/programming for automation (Python, Go, Bash, PowerShell).
Strong experience in modern CI/CD implementation with GitHub Actions.
Hands-on GitOps knowledge with FluxCD and ArgoCD.
Strong observability stack expertise, including:
Appdynamics
Dynatrace
Splunk Observability
OpenTelemetry
Monitoring, logging, distributed tracing, and telemetry engineering
Experience improving alert quality, noise reduction, and incident response effectiveness.
Practical understanding of SRE operations: availability, performance, resiliency, capacity, and incident management.

Preferred / Nice-to-Have
Exposure to enterprise integration patterns with ITSM tools (e.g., ServiceNow).
Basic understanding of LLM models and practical use of LLM-assisted tooling to accelerate SDLC delivery (e.g., code generation, troubleshooting, runbook enhancement, and test acceleration)


No Referrers Available

There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.