Job Description
Job Title: Engineer - AWS, Azure, GCP, AKS
Experience: 6+ Years
Education: Any Graduate
Location: Mumbai
Job Description
Role Overview
This role acts as the critical operational bridge between:
-
The in-house monitoring product development and QE team(s),
-
Internal ITSM/ticketing teams using ServiceNow,
-
Multiple enterprise DBA organizations across:
-
Oracle Database
-
Open-source database platforms (MySQL, PostgreSQL, MongoDB, Cassandra, etc.)
-
Microsoft SQL Server
Key Responsibilities
Production Operations Monitoring Support
-
Provide operational ownership and production support for the in-house enterprise monitoring platform.
-
Monitor health, performance, alerting quality, and operational stability of monitoring services.
-
Analyze monitoring gaps, false positives, missed alerts, and operational inefficiencies.
-
Ensure monitoring coverage across Oracle, Open-source, and SQL Server database environments.
Incident Escalation Management
-
Act as the operational point-of-contact during production incidents involving monitoring failures, alerting gaps, or infrastructure issues.
-
Coordinate incident triage across:
-
DBA teams,
-
Monitoring development teams,
-
Infrastructure teams,
-
Service management teams.
-
Drive bridge calls and ensure effective stakeholder communication during critical outages.
-
Perform root cause analysis (RCA) and post-incident operational reviews.
ServiceNow Ticket Workflow Coordination
-
Work with ServiceNow for:
-
Operational escalations,
-
Service requests.
-
Review ticket quality and ensure operational accuracy of issue classification and routing.
-
Improve ticket workflows between DBA teams and monitoring platform support teams.
-
Collaborate with internal support organizations to streamline escalation processes.
Cross-Functional DBA Collaboration
-
Collaborate closely with enterprise DBA teams supporting:
-
Oracle Database
-
MySQL
-
PostgreSQL
-
MongoDB
-
Apache Cassandra
-
Microsoft SQL Server
-
Cloud services (AWS, AZURE, GCP)
-
Understand operational monitoring requirements specific to each database technology.
-
Work with DBAs to validate alert thresholds, event correlation, and monitoring accuracy.
-
Serve as the operational liaison between DBAs and monitoring team developers and QE.
Operational Excellence Reliability Engineering
-
Identify recurring operational pain points and recommend automation opportunities.
-
Improve alert quality, event correlation, and monitoring reliability.
-
Participate in operational readiness reviews for new monitoring features.
-
Help define operational standards, playbooks, and escalation procedures.
Monitoring Observability Engineering
-
Support enterprise observability initiatives involving:
-
Metrics,
-
Events,
-
Alerting,
-
Dashboards,
-
Health monitoring,
-
Incident correlation.
-
Work with both commercial and in-house monitoring systems.
-
Analyse operational telemetry to identify systemic reliability concerns.
DevOps CI/CD Enablement
-
Collaborate with engineering teams to improve CI/CD pipelines.
-
Implement deployment strategies (blue-green, canary, rolling updates).
-
Advocate for reliability-focused design patterns.
Security Compliance
-
Ensure infrastructure adheres to security standards and compliance requirements.
-
Participate in vulnerability assessments and remediation.
Required Technical Skills
-
Strong production support and operations experience in enterprise environments.
-
Strong experience with cloud platforms (AWS, Azure, or GCP).
-
Expertise in monitoring observability tools (e.g., Prometheus, Grafana, Datadog, or in-house tools).
-
Working knowledge of:
-
ServiceNow
-
Incident workflows,
-
Escalation management,
-
Operational support models.
-
Exposure to database technologies including:
-
Oracle Database
-
Microsoft SQL Server
-
MySQL
-
PostgreSQL
-
NoSQL ecosystems preferred.
-
Strong understanding of:
-
Linux systems,
-
Infrastructure monitoring,
-
Alerting concepts,
-
Production operations.
-
Experience supporting 24x7 enterprise production environments.
Preferred Qualifications
-
Experience working with in-house monitoring or observability product teams.
-
Familiarity with SRE/DevOps operational practices.
-
Exposure to enterprise event management systems.
-
Knowledge of automation/scripting (Python, Shell, PowerShell).
-
Experience handling high-severity production incidents.
Critical Non-Technical Skills
An ideal candidate must demonstrate:
Operational Intuition
-
Ability to detect operational anomalies early.
-
Strong troubleshooting instinct and pattern recognition.
-
Fearless Communication
-
Ability to speak confidently during incidents and escalations.
-
Comfortable engaging senior stakeholders and multiple technical teams.
-
Cross-Team Collaboration
-
Ability to coordinate effectively across DBA teams, support organizations, and development groups.
-
Calmness Under Pressure
-
Structured decision-making during high-severity incidents.
-
Ownership Mindset
-
Drives issues to closure rather than relying solely on assigned ownership boundaries.
-
Investigative Curiosity
-
Continuously analyses why operational failures occur and how they can be prevented.
Preferred Qualifications
-
Certifications in cloud platforms (AWS/Azure/GCP).
-
Familiarity with SRE/DevOps operational practices.
-
Exposure to enterprise event management systems.
-
Knowledge of automation/scripting (Python, Shell, PowerShell).
