Experience
7 - 11 yrs
Job Location
Pune, India
Vacancy
1
Designation
Architect
Job Type
Not specified
Job Description
Operational Empathy & Developer Enablement Partner closely with feature and product engineers to embed observability into the development lifecycle
Translate complex logs and telemetry into clear, actionable Grafana dashboards that help teams understand system behavior and blast radius
Full-Stack Observability & Forensics Lead distributed tracing initiatives by correlating frontend exceptions (Sentry) with backend logs and traces (VictoriaLogs/OpenSearch), enabling a seamless end-to-end (north-to-south) view of system health
Telemetry Gap Identification & Instrumentation Proactively identify blind spots in logging, metrics, and traces
Implement custom instrumentation across service layers to capture high-cardinality data while maintaining system performance
Incident Response Automation Design and build Python-based automation tools to reduce Time to Truth during incidents by automating log aggregation, telemetry correlation, and diagnostic reporting
Service Level Ownership Advocate for meaningful reliability metrics by defining and refining SLIs and SLOs that truly reflect user experience and satisfaction, balancing rapid feature delivery with long-term system stability
Job Requirements
Technical Skills & Experience Expert-level experience with OpenSearch and VictoriaLogs, including indexing strategies and optimization for high-volume log ingestion and querying
Strong hands-on expertise with Grafana, building intuitive dashboards that clearly communicate system behavior and incident patterns
Python: Advanced proficiency for building automation scripts, diagnostic tooling, and observability glue code
TypeScript & Lua: Working familiarity required
You should be comfortable reading and understanding service codebases to trace request flows end-to-end (deep expertise not required on day one)
Experience with Sentry, including performance monitoring and profiling capabilities to proactively identify regressions and bottlenecks
Additional Expectations
Strong analytical and problem-solving mindset with a passion for system reliability and visibility
Ability to collaborate effectively with application engineers, platform teams, and incident responders
Clear communication skills to translate complex system data into actionable insights for diverse stakeholders
Translate complex logs and telemetry into clear, actionable Grafana dashboards that help teams understand system behavior and blast radius
Full-Stack Observability & Forensics Lead distributed tracing initiatives by correlating frontend exceptions (Sentry) with backend logs and traces (VictoriaLogs/OpenSearch), enabling a seamless end-to-end (north-to-south) view of system health
Telemetry Gap Identification & Instrumentation Proactively identify blind spots in logging, metrics, and traces
Implement custom instrumentation across service layers to capture high-cardinality data while maintaining system performance
Incident Response Automation Design and build Python-based automation tools to reduce Time to Truth during incidents by automating log aggregation, telemetry correlation, and diagnostic reporting
Service Level Ownership Advocate for meaningful reliability metrics by defining and refining SLIs and SLOs that truly reflect user experience and satisfaction, balancing rapid feature delivery with long-term system stability
Job Requirements
Technical Skills & Experience Expert-level experience with OpenSearch and VictoriaLogs, including indexing strategies and optimization for high-volume log ingestion and querying
Strong hands-on expertise with Grafana, building intuitive dashboards that clearly communicate system behavior and incident patterns
Python: Advanced proficiency for building automation scripts, diagnostic tooling, and observability glue code
TypeScript & Lua: Working familiarity required
You should be comfortable reading and understanding service codebases to trace request flows end-to-end (deep expertise not required on day one)
Experience with Sentry, including performance monitoring and profiling capabilities to proactively identify regressions and bottlenecks
Additional Expectations
Strong analytical and problem-solving mindset with a passion for system reliability and visibility
Ability to collaborate effectively with application engineers, platform teams, and incident responders
Clear communication skills to translate complex system data into actionable insights for diverse stakeholders
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.