Role Overview
We are looking for an AI Engineer with a strong focus on Large Language Models (LLMs) and Generative AI to design, build, and deploy intelligent, LLM-powered systems. You'll work on prompt engineering, RAG pipelines, agentic workflows, and fine-tuning, while building robust backend services in Java and Python — with React front-end knowledge as a plus to help ship end-to-end AI features.
Responsibilities
- Design, build, and deploy LLM-powered applications — chatbots, copilots, agents, RAG systems, and summarization/extraction pipelines
- Develop and optimize prompt engineering strategies (few-shot, chain-of-thought, structured output, function/tool calling) for production use cases
- Build and maintain Retrieval-Augmented Generation (RAG) pipelines — embeddings, chunking strategies, vector databases (Pinecone, Weaviate, FAISS, pgvector)
- Fine-tune and evaluate LLMs (open-source and closed-source) using techniques like LoRA/QLoRA, instruction tuning, and RLHF where applicable
- Design and implement agentic workflows using frameworks like LangChain, LangGraph, LlamaIndex, or custom orchestration
- Build backend services and APIs in Java and/or Python to serve LLM applications at scale
- Implement evaluation frameworks for LLM outputs — accuracy, hallucination rate, latency, cost, and safety metrics
- Manage prompt versioning, model routing, and caching for cost/performance optimization across multiple LLM providers (OpenAI, Anthropic, open-source models)
- Collaborate with front-end developers (or contribute directly in React) to surface LLM features in user-facing products
- Monitor deployed LLM systems for drift, degradation, and safety/guardrail violations; iterate on mitigation strategies
- Implement guardrails, content filtering, and responsible AI practices (bias mitigation, prompt injection defense, data privacy)
- Stay current with the fast-moving LLM/GenAI landscape (new models, techniques, tooling) and assess applicability to the business
Requirements
- Bachelor's or Master's degree in Computer Science, Machine Learning, or related field (or equivalent practical experience)
- Strong programming skills in Java and Python
- Hands-on experience with LLM APIs (OpenAI, Anthropic Claude, Gemini, etc.) and open-source LLMs (Llama, Mistral, etc.)
- Practical experience with prompt engineering and RAG architecture
- Experience with vector databases and embedding models
- Familiarity with LLM orchestration frameworks: LangChain, LangGraph, LlamaIndex, Semantic Kernel, or similar
- Solid understanding of transformer architecture and how LLMs work under the hood
- Experience building and consuming REST/gRPC APIs for AI-serving backends
- Experience with cloud platforms (AWS, GCP, Azure) and containerization (Docker, Kubernetes)
Skills
- Python
- Java
- LangChain
- LLMs
- React