TymblHub

© 2026 TymblHub

Site Reliability Architect

LotusFlare
Posted on

Experience
7 - 11 yrs
Job Location
Pune, India
Vacancy
1
Designation
Architect
Job Type
Not specified

Job Description

Operational Empathy & Developer Enablement Partner closely with feature and product engineers to embed observability into the development lifecycle

Translate complex logs and telemetry into clear, actionable Grafana dashboards that help teams understand system behavior and blast radius

Full-Stack Observability & Forensics Lead distributed tracing initiatives by correlating frontend exceptions (Sentry) with backend logs and traces (VictoriaLogs/OpenSearch), enabling a seamless end-to-end (north-to-south) view of system health

Telemetry Gap Identification & Instrumentation Proactively identify blind spots in logging, metrics, and traces

Implement custom instrumentation across service layers to capture high-cardinality data while maintaining system performance

Incident Response Automation Design and build Python-based automation tools to reduce Time to Truth during incidents by automating log aggregation, telemetry correlation, and diagnostic reporting

Service Level Ownership Advocate for meaningful reliability metrics by defining and refining SLIs and SLOs that truly reflect user experience and satisfaction, balancing rapid feature delivery with long-term system stability

Job Requirements

Technical Skills & Experience Expert-level experience with OpenSearch and VictoriaLogs, including indexing strategies and optimization for high-volume log ingestion and querying

Strong hands-on expertise with Grafana, building intuitive dashboards that clearly communicate system behavior and incident patterns

Python: Advanced proficiency for building automation scripts, diagnostic tooling, and observability glue code

TypeScript & Lua: Working familiarity required

You should be comfortable reading and understanding service codebases to trace request flows end-to-end (deep expertise not required on day one)

Experience with Sentry, including performance monitoring and profiling capabilities to proactively identify regressions and bottlenecks

Additional Expectations

Strong analytical and problem-solving mindset with a passion for system reliability and visibility

Ability to collaborate effectively with application engineers, platform teams, and incident responders

Clear communication skills to translate complex system data into actionable insights for diverse stakeholders

No Referrers Available

There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.