Role Overview
As an AI Engineer, you will own the AI Infrastructure Architecture, designing and implementing asynchronous multi-agent orchestration and managing end-to-end latency from user messages to AI responses. You will build resilient inference pipelines, implement intelligent request routing, and migrate critical AI conversation flows from monoliths to dedicated services.
Responsibilities
- Design and implement asynchronous multi-agent orchestration
- Own end-to-end latency from user message to AI response
- Build resilient inference pipelines that gracefully degrade under load
- Implement intelligent request routing and load balancing for AI workloads
- Migrate critical AI conversation flow from monolith to dedicated services
- Implement WebSocket/streaming infrastructure for real-time chat
- Design circuit breakers and fallback strategies for AI model failures
- Build comprehensive observability for AI system performance
- Optimize credit data retrieval and caching strategies
Requirements
- 3 to 5 years building production systems handling >10k concurrent users
- Proven experience with async/event-driven architectures
- Hands-on experience scaling ML/AI inference in production
- Deep understanding of caching strategies (Redis, in-memory, CDN)
- Experience with message queues and real-time communication protocols
- Experience with AI model serving frameworks (TensorFlow Serving, Triton, etc.)
- Understanding of AI inference optimization (batching, caching, model quantization)
- Knowledge of conversation state management and context handling
Skills
- Python
- Redis
- TensorFlow Serving
- WebSocket
- System Design
Growth Path
- Direct impact on customer subscription retention through performance
- Exposure to cutting-edge AI infrastructure challenges
- Ownership of technical decisions affecting revenue-generating conversations
- Path to leading AI platform team as you scale