TymblHub

© 2026 TymblHub

Sr. ML Platform Engineer

CrowdStrike
Posted on
CrowdStrike logo

Experience
12 - 17 yrs
Job Location
Bengaluru, India
Vacancy
1
Designation
Senior Machine Learning Engineer
Job Type
ONSITE

Job Description

Platform Reliability & Debugging: Diagnose and resolve issues across Ray, Spark, Airflow, MLflow, JupyterHub, Kubeflow, and SLURM Perform root cause analysis on production incidents affecting training and inference pipelines Debug performance bottlenecks, resource contention, memory leaks, and scheduling conflicts Develop debugging tools and diagnostic frameworks
System Optimization & Performance: Profile and optimize Ray clusters and Spark jobs on K8s and Cloud (EMR/Dataproc) Troubleshoot JupyterHub spawner issues, kernel crashes, and resource allocation Optimize SLURM job scheduling, GPU allocation, and HPC cluster utilization
Infrastructure & Monitoring: Build observability solutions and automated health checks Develop runbooks, alerting workflows, and incident response procedures Maintain platform stability metrics (SLAs, error rates, latency)
Collaboration: Partner with ML and ML Platform engineers to resolve workflow issues Conduct post-mortems and mentor on debugging techniques
What Youll Need:
  • 12+ years in distributed systems engineering
  • 5+ years debugging ML platforms in production
  • Deep expertise in 3+ one of: Ray, Spark, JupyterHub, SLURM, K8 Performance profiling, optimization, and capacity planning
Technical Skills (Expertise in at least one):
  • Distributed ML: Ray, Spark, SLURM, Jupyter Ecosystem (debugging failures, performance tuning)
  • ML Platforms: Airflow, MLflow, JupyterHub (troubleshooting core components) Infrastructure: Kubernetes, Docker, AWS/GCP/Azure/OCI
  • Observability: Profiling tools, distributed tracing, Prometheus, Grafana, log aggregation
  • Programming: Expert Python debugging, multi-language proficiency, Linux/Unix
What Sets You Apart: Open-source ML infrastructure contributions Experience with high-throughput inference systems and reducing MTTR Published debugging guides or tools Chaos engineering and GPU/CUDA debugging experience On-call and incident management experience
Benefits of Working at CrowdStrike:
  • Market leader in compensation and equity awards
  • Comprehensive physical and mental wellness programs
  • Competitive vacation and holidays for recharge
  • Paid parental and adoption leaves
  • Professional development opportunities for all employees regardless of level or role
  • Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections
  • Vibrant office culture with world class amenities
  • Great Place to Work Certified across the globe

No Referrers Available

There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.