TymblHub

© 2026 TymblHub

Software Engineer II

7 Eleven
Posted on
7 Eleven logo

Experience
7 - 12 yrs
Salary (CTC)
₹21.8L - ₹25.2L
Job Location
Bengaluru, India
Vacancy
1
Designation
Software Engineer II
Job Type
ONSITE

Job Description

Job Summary

Software Engineer II

6 - 8 Years

Bangalore

As a Senior AI Platform RunOps Engineer, you will own day-2 operations for the AI platform. This role is focused on availability, incident response, platform support, runbooks, operational governance, release readiness, service health, auditability, and continuous improvement across models, agents, gateways, MCP services, and supporting platform infrastructure.

Responsibilities

  • Own the operational health of AI platform services, including availability, support readiness, monitoring, and operational escalations.
  • Build and maintain production runbooks, support procedures, incident response workflows, and service recovery patterns for AI platform components.
  • Monitor and improve platform performance across latency, throughput, error rates, token usage, and operational cost signals.
  • Drive release readiness, environment promotion, rollback safety, and operational acceptance criteria for AI platform changes.
  • Support the production operations of gateways, model-serving layers, agent runtimes, MCP services, and platform integrations.
  • Implement and maintain operational controls around secrets, audit logging, identity, service access, and compliance evidence.
  • Partner with platform and security teams on DR, failover, backup, resilience, and infrastructure recovery planning.
  • Operate enterprise platform capabilities including AKS, managed identities, Key Vault integrations, SSO mappings, and certificate management where required.
  • Provide production support inputs for AI platform governance, change management, onboarding, and service lifecycle decisions.
  • Support Apigee X operational integration for AI-facing APIs, including traffic controls, monitoring, onboarding, and secure exposure patterns.

Required qualifications

  • 7+ years of experience in production operations, SRE, platform support, DevOps, or cloud infrastructure engineering.
  • Strong hands-on experience supporting distributed systems in production, including Linux systems, containers, CI/CD, and cloud-native services.
  • Strong experience in monitoring, logging, ing, incident response, and service health management.
  • Experience writing and maintaining operational runbooks, troubleshooting guides, and recovery procedures.
  • Experience with identity, access control, audit logging, secrets handling, and security-aware platform operations.
  • Experience supporting AI, data, or analytics platforms in production environments is strongly preferred.
  • Strong scripting or development capability in Python and other automation-friendly languages.
  • Ability to collaborate effectively across engineering, security, support, and business teams during operational events.

Preferred qualifications

  • Experience with enterprise AI platform operations, including model-serving services, agent runtimes, prompt governance, or evaluation systems.
  • Experience with Databricks, Dataiku, MLflow, Airflow, or similar platforms requiring production-grade operational governance.
  • Experience with AKS or Kubernetes-backed platform operations, including autoscaling, certificate management, identity delegation, and cluster troubleshooting.
  • Experience with Apigee X or similar API management platforms for production operations and secure service exposure.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.