Role Overview
We are seeking a AI/ML Performance Engineer to drive performance analysis, optimization, and validation of AI workloads across Neural Processing Units (NPU), and GPUs in complex SoC platforms with a robust background in the field of deep learning (DL) to validate the performance of state-of-the-art low-level perception (LLP) as well as end-to-end AD models. The role focuses on AI/ML inference performance, latency, throughput, power efficiency, and system-level integration, working closely with architecture, compiler, runtime, and framework teams.
Responsibilities
- Analyze and optimize AI/ML workload performance on NSP, NPU, and GPU across single and multi-accelerator systems.
- Identify bottlenecks across compute, memory bandwidth/latency, NoC, DMA, SMMU, and scheduling.
- Drive optimizations for latency-critical inference, high-throughput pipelines, and concurrent multi-VM execution.
- Perform roofline, utilization, and memory access pattern analysis using hardware counters.
- Benchmark and optimize CNNs, Transformers, VLMs, and diffusion models for real-world and customer use cases.
- Optimize model execution using operator fusion, graph partitioning, quantization, mixed precision, tiling, and batching.
- Collaborate with system software teams on drivers, runtimes, memory allocation, power, thermal, and QoS constraints.
- Analyze performance under concurrent, virtualized, and safety-critical or real-time environments.
- Build and use benchmarks, micro-benchmarks, profilers, and regression tools to derive actionable insights.
- Partner with architecture, compiler, SDK, and AI framework teams to influence HW/SW design and resolve performance escalations.
Requirements
- Bachelor's degree in Computer Science Engineering, Information Systems, or related field and 1+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
- Good at software development with exposure to AI coding tools and excellent analytical, development, and problem-solving skills.
- Experience in embedded software (Linux, Android, RTOS).
- Experience in system performance domain with focus on neural accelerators and GPU.
- Knowledge on CPU, NPU, GPU, NOC, DDR architecture.
- Strong understanding of Machine Learning fundamentals.
- Knowledge in neural network quantization, compression, pruning algorithms, deep learning kernel/compiler optimization.
- Strong communication skills.
Skills
- Deep Learning
- Embedded Software
- NPU
- GPU
- Python