Role Overview
At PhoenixAI, you’ll work at the intersection of cutting-edge ML and Robotics innovation, helping bring frontier models to production. From deploying optimized foundation models to building low-latency inference infrastructure, your work will define how engineers interact with intelligence. This role is ideal for those excited to make research real and shape the future of design through Physical AI.
Responsibilities
- Fine-tune, evaluate, and deploy LLMs and foundation models for specialized chip engineering tasks.
- Build scalable inference systems, agentic workflows, and RAG pipelines to support intelligent, interactive design tools.
- Optimize model performance through batching, caching, quantization, and streaming to achieve sub-second latency.
- Collaborate closely with AI scientists to operationalize experimental techniques and deliver robust, production-grade systems.
- Contribute to architecture decisions around scalable, reliable ML infrastructure in a cloud-native environment.
Requirements
- Hands-on experience working with LLMs, including prompt tuning, fine-tuning, and evaluation.
- Deep expertise in Python, PyTorch, and the surrounding ML tooling ecosystem.
- Familiarity with orchestration frameworks and agentic workflows (e.g., LangChain, LlamaIndex).
- Strong understanding of vector search, semantic embeddings, and RAG systems.
- Solid grasp of scalable ML architecture and the principles behind high-availability inference systems.
Skills
- Python
- PyTorch
- LLMs
- LangChain
- RAG
Bonus Points
- Experience working directly with ML researchers or in applied research settings.
- Startup experience or contributions to high-velocity, cross-functional teams.
- Open-source contributions to ML tooling or infrastructure projects.
- Experience deploying models in cloud environments (e.g. AWS SageMaker, Bedrock, vLLM, or SGLang).