TymblHub

© 2026 TymblHub

Site Reliability Engineer

SAS
Posted on
SAS logo

Experience
5 - 10 yrs
Job Location
Pune, India
Vacancy
1
Designation
Site Reliability Engineer
Job Type
ONSITE

Job Description

You will ensure the stability, performance, and reliability of our cloud-native Hospitality Revenue Management System (RMS) a multi-tenant SaaS platform serving hotels globally.
This role involves deep collaboration with engineering to fix reliability issues in Java microservices , build resilient cloud infrastructure, and optimize end-to-end performance of forecasting, rate publishing, and PMS/CRS integrations.
What you ll be doing...
  • Reliability & Platform Resilience
    • Define and manage SLIs/SLOs for RMS services (latency, availability, job success, data freshness).
    • Improve reliability of Java-based microservices by:
      • Fixing resource leaks, threading issues, timeouts, connection pool exhaustion
      • Improving resilience patterns (retries, backoff, circuit breakers
      • Profiling memory/CPU bottlenecks
      • Resolve performance issues impacting:
      • Forecast generation cycles
      • Rate push APIs
      • PMS/CRS connector API calls.
      • AWS Cloud Engineering
        • Operate and optimize AWS services:
          • EC2, EKS (Kubernetes), S3, CloudFront
          • RDS for SQL Server
          • Lambda, SNS/SQS, API Gateway
          • VPC, Security Groups, NAT, Transit Gateway
        • Implement cost optimization (compute efficiency, right-sizing, egress reduction).
    • Infrastructure as Code (Terraform)
      • Build, version, and maintain infrastructure using Terraform:
        • Reusable modules
        • Remote state + state locking
        • Drift management
        • Multi-environment automation
      • Enforce GitOps principles and automated provisioning.
    • Observability & Incident Response (Datadog Mandatory)
      • Build end-to-end observability using Datadog , including:
        • Metrics, logs, traces
        • JVM profiling (GC, heap, threads)
        • Datadog APM for Java services
        • Datadog Synthetics for API uptime (PMS, CRS, OTA)
        • Datadog RUM (if applicable)
      • Configure SLO-based alerting to reduce noise and improve signal quality.
What you ll bring to us
  • 5+ years in SRE /Production Engineering roles.
  • Strong hands-on expertise in AWS cloud infrastructure.
  • Terraform experience building production-grade infrastructure.
  • Strong understanding of Java internals:
  • Experience with Datadog for observability (APM, metrics, logs, dashboards).
  • Strong SQL Server skills (tuning, indexing, query optimization).
  • Proficient in scripting (Python/Bash/Go).

No Referrers Available

There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.