Job Description
MTS SYSTEMS DESIGN ENGINEER
THE ROLE:
We are looking for a dynamic, energetic Lead / Staff Software Engineer to join our growing team in AI (Artificial Intelligence) group. In this role, the individual will be responsible for developing AI/ML specific C/C++ kernels and dataflow schedules for AMD Ryzen Processors built on XDNA Neural Processor Units (NPU) to map LLMs, Stable Diffusion networks on NPU. As a C++ Kernel Developer, you will play a crucial role in designing, optimizing, and implementing machine learning kernels specifically tailored for vector processors. Your work will directly impact the efficiency, speed, and accuracy of our machine learning models.
KEY RESPONSIBILITIES:
- Kernel Development:
- Design and implement highly optimized C++ kernel library for NPU/GPU.
- Collaborate with the research and software teams to integrate these kernels into the existing software stack.
- Vector Processor Optimization:
- Work closely with hardware engineers to understand the architecture of VLIW vector core units such as MAC, GeMM, and non-linear functions.
- Develop vectorized code that leverages SIMD (Single Instruction, Multiple Data) and ILP (instruction level parallelism) for maximum performance.
- Profile and analyze the performance of existing kernels.
- Identify bottlenecks and optimize critical sections for better throughput.
- Develop CPU models for the ML operators in C++/ Python to validate accuracy.
- Write unit tests and integration tests to ensure correctness and reliability.
- Validate kernel performance across different hardware platforms.
- Document design specs for new kernels and the performance improvements.
- Follow coding guidelines, use tools like git to maintain code and create pull-requests, and documentation.
- Collaborate with cross-functional teams, including machine learning researchers and software engineers.
PREFERRED EXPERIENCE:
- Excellent C/C++ and Python coding skills
- Good understanding of SIMD/Tensor/VLIW processor architecture to exploit parallelism.
- Experience with vectorized programming (SIMD) and parallel computing.
- Familiarity with machine learning frameworks (e.g., TensorFlow, PyTorch) is a plus.
- Experience with silicon bring-up and pre-silicon validation on Emulation platforms is a plus.
- Knowledge of low-level hardware details (cache hierarchy, memory access patterns) is desirable.
- Excellent problem-solving skills and a passion for performance optimization.
ACADEMIC CREDENTIALS:
BS/Masters/PhD degree in Computer Science, Electrical Engineering, or a related field with around 10/8/5 year experience respectively.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
