Experience
8 - 13 yrs
Job Location
Hyderabad, India
Vacancy
1
Designation
Platform Engineer
Job Type
Not specified
Job Description
Job Title
AI Orchestration / Platform Engineer
Job Details - Number of positions: 1
- Level: Senior / Lead Platform Individual Contributor
- Primary locations: Hyderabad or Noida preferred; exceptional onshore candidates may be considered
- Target / alternate titles: AI Platform Engineer; Agent Platform Engineer; LLMOps Engineer; GenAI Platform Engineer; AI DevOps Engineer; AI Site Reliability Engineer; ML Platform Engineer
- Core keywords: agent orchestration, workflow engine, multi-agent systems, model gateway, model routing, LLMOps, platform engineering, Kubernetes, containers, CI/CD, infrastructure as code, policy as code, observability, OpenTelemetry, queues, state, retries, resilience, cost telemetry, SRE
- Recruiter red flags: Traditional DevOps profile with no AI runtime understanding; framework-only orchestration without production platform depth; manual deployments; no observability or incident ownership; cannot explain model routing, state, retries, evaluation, or AI-specific controls.
Role purpose: Build and operate the orchestration and runtime capabilities that allow agentic applications to move reliably from development into production. The role will integrate agent workflows with existing cloud, security, policy, CI/CD, and observability capabilities, supporting multi-step execution, model routing, state, resiliency, evaluation, and production operations without rebuilding the client's established platform foundations.
Client and delivery context: The client already has delivery pipelines, policy-based architecture, security controls, and production review processes. The engineer must integrate with and enhance those capabilities. The platform should support multiple models, clouds, frameworks, and developer-productivity ecosystems without forcing avoidable lock-in. The engineer will support shared patterns used across different business-process agents and may work across two parallel delivery tracks. Platform work must be pragmatic and outcome-driven, with emphasis on enabling engineers to release and operate reliable agents quickly.
Primary ownership: Agent workflow and runtime orchestration, including state, routing, queues, retries, timeouts, scheduling, persistence, and exception management. Model gateway and routing patterns, provider abstraction, policy-based selection, fallback, rate limits, quotas, and cost controls. CI/CD, environment promotion, configuration, secrets, infrastructure integration, release validation, rollback, and operational readiness. Application and platform observability, reliability engineering, incident response, capacity, performance, and production support.
Responsibilities - Design and implement orchestration patterns for multi-step agents, multi-agent collaboration, deterministic workflows, long-running tasks, approvals, and event-driven execution.
- Implement durable state, checkpoints, queues, retries, backoff, idempotency, timeouts, compensation, dead-letter handling, and recovery for agent workflows.
- Integrate model gateways and routing logic that select models based on capability, sensitivity, latency, cost, availability, or policy requirements.
- Build deployment and release patterns for agent services, orchestration components, prompts, tool definitions, configurations, and evaluation assets.
- Integrate with existing CI/CD, infrastructure-as-code, policy-as-code, secrets, identity, container, and cloud-runtime capabilities.
- Create observability for agent traces, model calls, tool calls, workflow state, token usage, cost, latency, errors, quality signals, and dependency health.
- Implement automated quality and security gates, including unit and integration tests, evaluation suites, policy checks, vulnerability scans, and rollback criteria.
- Optimize runtime performance, concurrency, throughput, caching, context use, model selection, and infrastructure cost.
- Support production incidents, root-cause analysis, capacity planning, resiliency testing, disaster recovery, and operational runbooks.
- Develop reusable platform templates, SDKs, reference pipelines, dashboards, and onboarding guidance for agent-development teams.
- 8+ years of platform, DevOps, SRE, cloud, distributed-systems, or software-engineering experience.
- Production experience supporting AI/ML, LLM, agentic, workflow, or high-scale distributed application platforms.
- Strong experience with containers, Kubernetes or equivalent runtimes, CI/CD, infrastructure as code, configuration, secrets, and automated deployment.
- Understanding of agent orchestration concepts including state, checkpoints, retries, timeouts, queues, long-running tasks, human approvals, and failure recovery.
- Experience with logging, metrics, distributed tracing, OpenTelemetry or equivalent observability, alerting, dashboards, and incident response.
- Strong scripting or development skills in Python, Go, Java, TypeScript, or comparable languages.
- Ability to integrate platform capabilities with security, identity, policy, data, network, and enterprise approval requirements.
- Experience balancing reliability, delivery speed, latency, throughput, portability, and operating cost.
- Experience with model gateways, multi-model routing, provider abstraction, fallback, quotas, or AI cost controls.
- Experience with LangGraph, Temporal, Airflow, Argo Workflows, Durable Functions, Step Functions, Kubernetes operators, or equivalent orchestration technologies.
- Experience with AI tracing and evaluation platforms, prompt/model registries, feature flags, canary releases, and regression gates.
- Experience with Terraform, Pulumi, Helm, GitOps, policy engines, service mesh, event streaming, and API gateways.
- Experience operating platforms across AWS, Azure, GCP, hybrid, or private environments.
- Experience in regulated enterprises or systems with sensitive proprietary data and formal production controls.
- Kubernetes, Docker, serverless or managed container platforms; Terraform/Pulumi/Helm/GitOps; GitHub Actions, GitLab, Jenkins, Azure DevOps, or equivalent; LangGraph, Temporal, Airflow, Argo, Step Functions, Durable Functions, or equivalent orchestration; model gateways and provider APIs; queues and events; OpenTelemetry, Prometheus, Grafana, cloud monitoring, AI tracing and evaluation tools. Exact products are flexible.
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
