Site Reliability Engineer - Product Support

BreachLock
Posted on
BreachLock logo

Experience
2 - 3 yrs
Job Location
Pune, India
Vacancy
1
Designation
Site Reliability Engineer
Job Type
Not specified

Job Description

BreachLock is standing up a new Reliability Engineering team and we are looking for a motivated SRE to join at the ground floor. Working alongside a Senior SRE and the Reliability Lead, you will help build the observability, monitoring, and operational foundations that underpin BreachLocks PTaaS and vulnerability management products.

This role is grounded in core SRE practice alert triage, incident response, runbook authoring, and operational automation. As you grow into the role, you will also gain exposure to our internal AI-powered support automation system built on a RAG architecture, with mentorship from the Senior SRE.

If you are curious, technically sharp, and want to grow into a senior SRE or platform engineering role within a cybersecurity product company, this is the right environment.

Platform Monitoring & Incident Response

Monitor and triage alerts from Grafana Cloud, Checkly, and load balancer logs; escalate per defined runbooks

Participate in the on-call rotation as a supporting responder; document incident observations for post-incident reviews

Assist in maintaining availability and performance dashboards covering uptime, HTTP error rates, and API latency (p50/p95/p99)

Monitor scan job success rates, PTaaS report delivery timelines, and auth success rates; proactively flag anomalies

Contribute to writing and maintaining runbooks and operational playbooks under Senior SRE guidance

Automation & Engineering Support

Automate recurring operational tasks using Python scripts and Bash to reduce manual toil for routine checks and alert responses

Assist with OpenTelemetry instrumentation tasks on Python (Django/Flask) and Node.js services

Support deployment validation and rollback procedures in CI/CD pipelines (GitHub Actions / Cloud Build)

Maintain documentation for infrastructure configuration, alert thresholds, and incident log history

Actively participate in weekly team syncs, reliability reviews, and learning sessions with the Senior SRE

Cloud & Infrastructure

2 3 years of hands-on experience with GCP or a major cloud provider (AWS, Azure); GCP strongly preferred

Comfort with Linux command line, basic networking (DNS, TCP/IP, HTTP), and cloud compute fundamentals

Working knowledge of containerization with Docker; basic Kubernetes concepts; GKE exposure is a plus

Monitoring & Observability

Hands-on experience with at least one monitoring platform (Grafana, Datadog, CloudWatch, or equivalent)

Ability to read and interpret dashboards, understand alert configurations, and identify metric anomalies

Basic understanding of log aggregation and structured logging; exposure to Loki or ELK is a bonus

Development & Scripting

Working proficiency in Python scripting for automation and operational tooling

Familiarity with Git version control and pull request workflows

Basic understanding of REST APIs and HTTP; exposure to GraphQL is a plus

Soft Skills

Strong written communication clear incident updates, runbook entries, and status reports

Eagerness to learn SRE principles SLOs, error budgets, toil reduction and absorb new technologies quickly with mentorship

Calm under pressure during incident response; follows process and asks good questions

NICE TO HAVE

Familiarity with AI / LLM concepts understanding of how large language models are queried and integrated into applications

Exposure to RAG (Retrieval-Augmented Generation) systems how embedding pipelines, vector databases (ChromaDB, Pinecone, or equivalent), and LLM response generation fit together

Any hands-on experience with vector databases or LLM APIs (OpenAI, Anthropic, or open-source models) through coursework, personal projects, or internship

Google Cloud Associate Cloud Engineer certification or currently in-progress

Understanding of cybersecurity concepts or experience in a security-adjacent product environment

Exposure to Terraform at a basic level (reading/modifying existing configurations)

Experience with Checkly, Prometheus, or Grafana Cloud stack in any capacity

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

No Referrers Available

There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.