Experience
5 - 10 yrs
Job Location
Pune, India
Vacancy
1
Designation
Site Reliability Engineer
Job Type
ONSITE
Job Description
You will ensure the stability, performance, and reliability of our cloud-native Hospitality Revenue Management System (RMS) a multi-tenant SaaS platform serving hotels globally.
This role involves deep collaboration with engineering to fix reliability issues in Java microservices , build resilient cloud infrastructure, and optimize end-to-end performance of forecasting, rate publishing, and PMS/CRS integrations.
What you ll be doing...
- Reliability & Platform Resilience
- Define and manage SLIs/SLOs for RMS services (latency, availability, job success, data freshness).
- Improve reliability of Java-based microservices by:
- Fixing resource leaks, threading issues, timeouts, connection pool exhaustion
- Improving resilience patterns (retries, backoff, circuit breakers
- Profiling memory/CPU bottlenecks
- Resolve performance issues impacting:
- Forecast generation cycles
- Rate push APIs
- PMS/CRS connector API calls.
- AWS Cloud Engineering
- Operate and optimize AWS services:
- EC2, EKS (Kubernetes), S3, CloudFront
- RDS for SQL Server
- Lambda, SNS/SQS, API Gateway
- VPC, Security Groups, NAT, Transit Gateway
- Implement cost optimization (compute efficiency, right-sizing, egress reduction).
- Infrastructure as Code (Terraform)
- Build, version, and maintain infrastructure using Terraform:
- Reusable modules
- Remote state + state locking
- Drift management
- Multi-environment automation
- Enforce GitOps principles and automated provisioning.
- Observability & Incident Response (Datadog Mandatory)
- Build end-to-end observability using Datadog , including:
- Metrics, logs, traces
- JVM profiling (GC, heap, threads)
- Datadog APM for Java services
- Datadog Synthetics for API uptime (PMS, CRS, OTA)
- Datadog RUM (if applicable)
- Configure SLO-based alerting to reduce noise and improve signal quality.
What you ll bring to us
- 5+ years in SRE /Production Engineering roles.
- Strong hands-on expertise in AWS cloud infrastructure.
- Terraform experience building production-grade infrastructure.
- Strong understanding of Java internals:
- Experience with Datadog for observability (APM, metrics, logs, dashboards).
- Strong SQL Server skills (tuning, indexing, query optimization).
- Proficient in scripting (Python/Bash/Go).
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
