Job Description
Job Summary
Kaufman Rossin is seeking an experienced AI Developer based in Bengaluru to play a central role in building KRs AI application capability on Azure. This is a full-stack AI engineering role - spanning LLM integration, RAG architecture, API development, Azure-native deployment, local/self-hosted model operations, and responsible AI implementation - with a focus on delivering real, production-grade solutions that create measurable value for KRs advisory practice and its clients.
The ideal candidate is deeply technical, intellectually curious, pragmatic about shipping working software, and acutely aware of the cost and security implications of every architectural decision. You will work directly from well-defined user stories, collaborate with the Azure Architect on model and infrastructure decisions, and leverage Claude Code and AI-assisted development tools as a standard part of your engineering workflow.
Key Responsibilities
AI Application Development
- Design, build, and deploy scalable AI-powered applications on Azure infrastructure - including internal productivity tools for KR staff and client-facing solutions for RAS and advisory engagements.
- Develop and implement Retrieval-Augmented Generation (RAG) pipelines using Azure OpenAI Service, Azure AI Search, and Azure AI Foundry - handling document ingestion, chunking strategies, embedding generation, vector indexing, and query-time retrieval.
- Build and maintain LLM-powered features: document intelligence, contract analysis, automated summarization, regulatory content extraction, conversational agents, and structured data extraction from unstructured sources.
- Integrate a broad range of AI models into application workflows - including Azure OpenAI (GPT-4o, o1, o3), Anthropic Claude (via Azure or API), Meta Llama, Mistral, and other open-source models - selecting the right model for each use case based on capability, cost, latency, and data sensitivity requirements.
- Stand up, configure, and maintain local/internal AI model deployments using Ollama, vLLM, or Azure AI Foundry local compute - for use cases requiring on-premises processing, data residency compliance, or cost-controlled inference at scale.
- Leverage Claude Code and other AI-assisted development tools throughout the engineering workflow - using agentic coding capabilities to accelerate development, improve code quality, generate tests, and explore solution options faster.
- Architect multi-turn conversational interfaces and agentic workflows using Azure AI Foundry Agent Service, Semantic Kernel, or LangChain - for use cases requiring reasoning, tool use, or multi-step orchestration.
Security Responsible AI
- Apply KRs AI governance framework in every build - content filtering via Azure AI Content Safety, DLP controls, usage audit logging, data sensitivity classification, and Conditional Access enforcement for AI-powered tools.
- Implement security best practices across all AI application components - Entra ID authentication, RBAC, Managed Identity for service-to-service access, secret management via Azure Key Vault, and Private Endpoint connectivity - treating security as a first-class design requirement, not an afterthought.
- Implement responsible AI practices in every deliverable - input/output validation, PII detection, content filtering, audit logging, and explainability documentation for regulated advisory contexts.
- Document model behavior, known limitations, bias considerations, and risk mitigations for each AI feature delivered - producing responsible AI summaries that satisfy KRs governance requirements for client-facing deployments.
- Implement secure, multi-tenant data isolation patterns in AI applications - ensuring client data, internal data, and model context are correctly scoped and never cross tenant boundaries.
Cost-Conscious Architecture FinOps
- Design AI application architectures with cost efficiency as a first-class concern - selecting models, deployment tiers, and inference patterns that deliver the required capability at the lowest sustainable cost.
- Configure and manage Azure AI Foundry deployments - model selection, deployment slots, token quotas, rate limiting, and cost monitoring - ensuring efficient and cost-controlled model consumption.
- Evaluate build vs. buy vs. local model decisions for each AI use case - weighing commercial API costs against self-hosted model operational costs, latency trade-offs, and data sensitivity requirements.
- Integrate Azure Monitor and Application Insights into AI applications - instrumenting latency, token usage, error rates, and custom AI quality metrics to support ongoing cost and performance monitoring.
- Produce regular AI cost reports - token consumption by application, model cost per query, and cost-per-outcome metrics - giving IT leadership visibility into AI spend and ROI.
Azure Infrastructure Deployment
- Deploy AI applications and services on Azure using containerized architectures - Azure Container Apps, AKS, or Azure App Services - depending on workload requirements and scale.
- Build and maintain CI/CD pipelines in Azure DevOps for AI application workloads - automating build, test, security scan, and deployment stages through to production.
- Implement API layers for AI capabilities using Azure API Management, Azure Functions, or FastAPI - exposing AI features to front-end applications and internal integrations in a secure, versioned, and observable way.
- Apply infrastructure-as-code practices (Bicep or Terraform) for AI service provisioning - ensuring environments are reproducible, auditable, and aligned with KRs Azure landing zone governance.
Data Integration
- Build data ingestion and preprocessing pipelines that prepare structured and unstructured data for AI consumption - document parsing (PDF, Word, Excel), OCR via Azure AI Document Intelligence, and schema normalization.
- Integrate AI features with KRs enterprise data sources and application ecosystem - Azure Data Lake Storage, APIs, and internal platforms - through well-designed, maintainable integration layers.
- Partner with KRs Architecture team on data flow design, storage tier selection, and retrieval performance optimization for knowledge bases and document stores.
Collaboration Delivery
- Work directly from user stories, acceptance criteria, and BA specifications - asking clarifying questions proactively and flagging technical constraints or cost implications early rather than late.
- Participate actively in sprint ceremonies - planning, standups, reviews, and retrospectives - contributing estimates, surfacing blockers, and demoing completed AI features to business stakeholders with clear, non-technical explanations.
- Collaborate with KRs Architecture team on architectural decisions, design reviews, model selection, and cost/security trade-offs.
- Produce and maintain clear technical documentation: API references, architecture decision records (ADRs), deployment runbooks, and model integration guides - so every feature is maintainable beyond the original developer.
- Mentor junior developers on AI development patterns, Azure service usage, cost optimization, and responsible AI practices.
Required
- 3+ years of hands-on software development, with at least 2 years focused on AI/ML application development in a production environment.
- Demonstrated experience building production applications with LLMs across multiple model families - OpenAI, Anthropic Claude, Meta Llama, Mistral, or similar.
- Hands-on experience building RAG pipelines - document ingestion, embeddings, vector search, chunking strategies, and retrieval optimisation.
- Experience standing up and operating local or self-hosted AI model deployments (Ollama, vLLM).