TymblHub

© 2026 TymblHub

Site Reliability Engineer

Grid Dynamics
Posted on
Grid Dynamics logo

Experience
3 - 8 yrs
Job Location
Central Remote, Remote
Vacancy
1
Designation
Site Reliability Engineer
Job Type
ONSITE

Job Description

  • Cross-team teamwork, build and maintain relationships with the customer teams, the user community, architects, and engineering teams, jointly work on key deliverables ensuring production scalability and stability
    Effective Root cause analysis of major production incidents and developing learning documentation.
  • Plan and perform capacity expansion and upgrades in timely manner avoiding any scaling issues and bugs.
  • Automation of repetitive tasks to reduce manual effort and avoid Human errors.
  • Tune alerting and setup observability to proactively identify the issues and performance problems.
  • Work closely with L-3 teams in reviewing new use cases, cluster hardening techniques for building a robust and reliable platforms.
  • leverage Devops tools, disciplines( Incident, problem and change management) and standards in day to operations.
Essential functions
  • 3+ years of experience with modern middleware technologies . These might include (Tomcat, Apache, Springboot, SQS, JBoss, IBM MQ, IBM DataPower, Hazelcast, Flink, Connect Direct, SSL)
  • Understanding of Linux/Unix systems, networking, cloud platforms (AWS, Azure, GCP), containerization (Kubernetes, Docker), and infrastructure-as-code tools ( Terraform , Ansible ).
  • Proficiency with monitoring tools ( Prometheus , Grafana, Datadog, etc.), logging systems (ELK stack, Splunk ), and tracing tools (Jaeger, Zipkin).
  • Proven track record of automating complex tasks and processes to improve efficiency and reliability using Python, Go, Java, or similar.
Qualifications
  • 2 or more years of work experience with a Bachelor s Degree or more than 2 years of work experience with an Advanced Degree (e.g. Masters, MBA, JD, MD) Hands on experience working as a Payment System SRE in managing cross platforms.
  • Excellent Python programming skills for automation requirement for repetitive dev-ops tasks Person will be responsible to perform SRE and Engineering activities Payment platforms Understanding of Linux, networking, CPU, memory and storage.
  • Knowledge on Java and Python is good to have.
  • Excellent interpersonal, verbal, and written communication skills.
Would be a plus
  • Cloud & System Architecture: Design scalable, resilient systems across hybrid cloud platforms (AWS, GCP, Azure).
  • AI/ML Operations: Support and optimize ML model deployment pipelines and monitoring systems.
  • Observability & Performance: Master advanced monitoring, tracing, and performance optimization techniques.
  • Automation & Intelligence: Build smart alerting systems and automated remediation workflows.
  • Distributed Systems: Design and maintain globally distributed payment processing systems.