Job Description
Job Description Site Reliability Engineer (SRE) for Lavu Production Systems US time
Location: US Time/ Hyderabad, India
Department: DevOps
Shift: PST time zone
About Lavu
Lavu is a global leader in mobile point-of-sale (mPOS) systems for restaurants and bars. We empower thousands of restaurant owners in over 60 countries to run their businesses efficiently. Our platform handles millions of transactions, synchronizes critical operational data in real-time, and integrates with various third-party partners.
The Role
We are looking for a Site Reliability Engineer who treats "uptime" as a religion. In this role, you won't just be fixing servers; you will be building the automated immune system that keeps our platform running during the Friday night dinner rush. You will bridge the gap between Development and Operations, ensuring that our infrastructure scales as fast as our customer base.
If you enjoy debugging complex distributed systems, automating manual toil, and ensuring that a restaurant in Canda can process a payment at the exact same moment a bar in New York closes a tab, we want to talk to you.
Key Responsibilities
- Keep the Lights On: Maintain 99.99% availability for Lavus core transaction services. You will own the health of our production environment.
- Infrastructure as Code: Manage and provision our cloud infrastructure (AWS) using Terraform and Ansible.
- Observability: Build and maintain monitoring dashboards (Datadog/New Relic/Prometheus) that alert us to problems before the customer calls support.
- Incident Response: Participate in the on-call rotation. When things break, you lead the triage, fix the issue, and write the blameless Post-Incident Review (PIR) to ensure it never happens again.
- Database Reliability: Optimize and maintain our data persistence layers (MySQL /Redis). You will troubleshoot slow queries and manage replication health.
- Eliminate Toil: Identify repetitive manual tasks (e.g., deployments, certificate renewals) and write Python/Go/Bash scripts to automate them out of existence.
- Security: Work with the security team to implement best practices for network security, IAM roles, and data encryption.
What Were Looking For
- Experience: 3+ years in SRE, DevOps, or Systems Engineering roles.
- Cloud Native: Deep expertise with AWS (EC2, RDS, Elastic Cache, Lambda, VPC).
- Coding Skills: You can write clean, maintainable code to glue systems together.
- Containerization: Strong experience with Docker and orchestration via Kubernetes (EKS).
- Database Chops: You understand how relational databases work at scale. You know what a slow query log is and how to fix it.
- The Mindset: You remain calm under pressure. You approach failure as a learning opportunity, not a blame game.
- B.S. Computer Science, or related fields
Bonus Points
- Experience with PHP/LAMP stack environments
- Background in the FinTech or Payments industry (understanding PCI compliance).
- Experience implementing Chaos Engineering practices.
Why Join Lavu?
- Impact: Your code runs in thousands of restaurants. You directly protect the livelihoods of business owners.
- Modernization: We are actively evolving our stack. You will get to help architect the next generation of our infrastructure.
- Culture: A collaborative, and remote-friendly team
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.