TymblHub

© 2026 TymblHub

AI Software Lead PyTorch & CUDA Runtime (Next-Gen Accelerator)

SanDisk India Device Design Centre
Posted on
SanDisk India Device Design Centre logo

Experience
8 - 13 yrs
Job Location
Bengaluru, India
Vacancy
1
Designation
Lead Software Engineer
Job Type
ONSITE

Job Description

Job Description

Role Overview

We are looking for a Software Lead (8+ years experience) to own the runtime and neural network (NN) layer of a next-generation AI accelerator platform. This role focuses on designing, optimizing, and implementing NN operators and developing new ops using CUDA/custom runtime APIs to deliver high-performance execution on custom AI hardware.

Key Responsibilities

  • Design and optimize NN operators for performance-critical workloads
  • Develop new NN ops using CUDA/custom runtime APIs
  • Drive runtime-level optimizations across compute, memory, and scheduling
  • Own runtime NN layer interfaces and execution model
  • Implement and optimize operator fusion (e.g., matmul + bias + LayerNorm) for efficient hardware utilization
  • Identify and resolve performance bottlenecks across the stack
  • Collaborate with compiler, PyTorch framework, and low-level SW teams

Impact

  • Own how efficiently AI workloads execute on the platform
  • Drive performance, scalability, and hardware utilization through optimized runtime and NN ops design

Qualifications

Required Qualifications

  • 8+ years in systems software / runtime / performance engineering
  • Strong experience with:
    • PyTorch / TensorFlow / JAX or similar frameworks
    • NN operator/kernel development and optimization
    • operator fusion and graph-level optimizations
  • Hands-on expertise in:
    • C/C++ and CUDA (or similar low-level programming)
    • runtime systems and execution engines
  • Strong understanding of:
    • memory hierarchy, data movement, and parallel execution

Preferred Qualifications

  • Experience with GPU/NPU or custom AI accelerators
  • Familiarity with XLA / MLIR / compiler-runtime interaction
  • Experience optimizing LLM or large-scale DL workloads

Disclaimer : This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

No Referrers Available

There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.