Job Description
Job Description:
Role: Site Reliability Engineer
Location: Mumbai, Bangalore, Pune,Hyderabad
Experience: 4-9 yrs
Primary Skill : DevOps with Kubernetes and SRE background.
Secondary Skill : OpenTelemetry, Splunk/Grafana
Role Summary:
We are seeking a dynamic and skilled expert in SRE principles, OpenTelemetry, Telegraf, Splunk/Prometheus/ Grafana to lead and support our transition from Splunk to a modern observability stack. The ideal candidate will have strong expertise in architecting and implementing observability solutions using these tools and will play a critical role in both the migration process and ongoing management of the observability environment.
Key Responsibilities:
Design and implement monitoring solutions using OpenTelemetry and Telegraf.
Set up, configure, and manage Prometheus and Grafana environments.
Integrate OpenTelemetry with Prometheus and Grafana to ensure seamless observability across applications and infrastructure.
Lead the migration from Splunk to the new observability stack, ensuring a smooth transition.
Develop and maintain Grafana dashboards and alerts to provide real-time insights into system performance.
Design, implement, and maintain OpenTelemetry instrumentation in our applications and services.
Develop and manage observability pipelines that collect, process, and visualize telemetry data (traces, metrics, and logs).
Collaborate with developers and DevOps teams to ensure best practices in observability, monitoring, and incident response.
Analyse telemetry data to identify performance bottlenecks and improve system reliability.
Contribute to the development of documentation, tutorials, and best practices around OpenTelemetry and observability tools.
Stay up to date with the latest OpenTelemetry developments and trends in the observability landscape.
Must-Have Skills:
Strong experience with OpenTelemetry and Telegraf for monitoring and data collection.
Deep understanding of observability and monitoring best practices.
Experience with traces, spans, and distributed tracing using OpenTelemetry.
Configuration of Loki (for logs) and Tempo (for traces) in Grafana.
Scripting and automation skills (Python, Shell, etc.) for task automation and environment management.
Proven experience with OpenTelemetry or similar observability frameworks (e.g., Prometheus, Jaeger, Grafana).
Strong understanding of cloud-native architecture, microservices, and container orchestration (e.g., Kubernetes).
Proficiency in one or more programming languages (e.g., Go, Java, Python, JavaScript).
Familiarity with distributed tracing concepts and techniques.
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
