Role Overview
Aion is an enterprise AI platform providing a full-stack solution for building, fine-tuning, and deploying AI at scale. We are looking for a hardware engineer passionate about building the infrastructure that powers large-scale AI systems, including servers, GPUs, networking, and high-performance computing environments.
Responsibilities
- Design and build hardware platforms optimized for AI training and inference workloads.
- Evaluate, integrate, and validate servers, GPUs, networking equipment, and storage systems.
- Optimize compute, memory, storage, networking, and GPU performance for AI workloads.
- Build reliable and fault-tolerant hardware infrastructure and develop validation and diagnostics processes.
- Support hardware deployment and develop automation for provisioning and monitoring.
- Collaborate across hardware, software, and AI engineering teams to improve system performance.
Requirements
- 4+ years of experience in hardware engineering, systems engineering, or data center infrastructure.
- Strong understanding of server architecture, CPUs, GPUs, memory, and storage.
- Experience with AI infrastructure including NVIDIA GPUs or AMD accelerators.
- Knowledge of PCIe, NVLink, InfiniBand, and Ethernet.
- Experience with Linux systems, hardware diagnostics, and firmware updates.
- Familiarity with rack-scale deployments and data center operations.
- Knowledge of automation and scripting using Python or Bash.
Skills
- NVIDIA GPUs
- Linux
- Python
- InfiniBand
- PCIe