Role Overview
We are looking for a highly skilled GPU compute, MLIR compiler, and kernel optimization engineer with deep expertise in GPU compute, MLIR-based code generation, and end-to-end performance optimization for AI workloads. In this role, you will design, optimize, and deploy high-performance GPU compute kernels, build and extend MLIR compiler backends, and collaborate closely with ML, runtime, and hardware teams to push the limits of performance on modern GPU architectures.
Responsibilities
- Develop and optimize GPU compute kernels targeting OpenCL and Vulkan compute backends for high-throughput AI/ML workloads.
- Design, build, and extend MLIR dialects across multiple abstraction levels including frontend, graph-level, tensor IR, and runtime/low-level dialects.
- Implement and maintain MLIR-based compiler passes and transformations such as tiling, fusion, bufferization, and vectorization.
- Conduct profiling and bottleneck analysis of compiled kernels using GPU counters and vendor-specific profilers.
- Build and maintain GPU runtime infrastructure for both OpenCL and Vulkan, including memory management and command buffer orchestration.
- Develop and extend code generation pipelines for automatic lowering from tensor IR to efficient kernels.
- Implement performance-critical schedules including tiling, loop fusion, and parallelism within MLIR-based backends.
- Collaborate with framework teams to optimize model lowering for computer vision and LLM workloads.
- Design and implement robust compiler and runtime components using modern C/C++.
Requirements
- Strong hands-on experience with MLIR framework, including custom dialects and lowering pipelines.
- Deep expertise across MLIR abstraction levels (TOSA, StableHLO, Linalg, Tensor, Vector, MemRef, SCF, GPU, and LLVM).
- Strong hands-on experience in OpenCL programming and runtime management.
- Solid understanding of Vulkan compute programming and runtime internals.
- Strong understanding of GPU architecture, memory hierarchies, and async compute.
- Proficiency in C/C++ for system-level development.
- Experience with kernel profiling and bottleneck analysis on GPU platforms.
- Strong background in machine learning fundamentals (CV and LLM workloads).
- Bachelor's degree with 4+ years, Master's with 3+ years, or PhD with 2+ years of relevant Systems Engineering experience.
Nice to Have
- Hands-on experience with IREE, TVM, XLA, or LLVM.
- Familiarity with IREE's compiler and runtime architecture.
- Experience contributing to open-source MLIR or IREE projects.
- Knowledge of quantization and mixed-precision inference.
- Exposure to multi-target compilation (CPU, GPU, NPU).
- Familiarity with cross-vendor GPU profiling tools like ARM Streamline, Qualcomm Snapdragon Profiler, or Intel VTune.
Skills
- MLIR
- C/C++
- OpenCL
- Vulkan
- GPU Architecture