Role Overview
We are looking for a passionate Generative AI Engineer with hands-on experience in building production-ready LLM applications, RAG pipelines, AI backend services, and intelligent document or vision systems. The ideal candidate should be comfortable working across model fine-tuning, prompt engineering, vector databases, and scalable API development.
Responsibilities
- Design, develop, and deploy LLM/SLM-based applications using LangChain, LangGraph, OpenAI, and Hugging Face.
- Build and optimize RAG pipelines with vector databases such as Pinecone or FAISS.
- Fine-tune transformer models using LoRA, Unsloth, and Hugging Face Transformers.
- Develop scalable backend services using Python, FastAPI, REST APIs, and GraphQL.
- Implement async processing and performance optimization for high-traffic AI services.
- Work on OCR, document AI, speech-to-text, and computer vision pipelines where required.
- Deploy AI workloads using Docker, GCP Vertex AI, AWS S3, and Kubernetes.
- Collaborate with product, frontend, and engineering teams to deliver production-grade AI systems.
- Maintain technical documentation, version control, and deployment workflows.
Requirements
- 1+ year of experience in Generative AI / AI Engineering.
- Strong proficiency in Python.
- Experience with LangChain, OpenAI APIs, Hugging Face, RAG, and Prompt Engineering.
- Knowledge of Vector Embeddings, Pinecone, or FAISS.
- Experience with FastAPI, REST APIs, and async programming.
- Familiarity with Docker and cloud platforms (GCP or AWS).
- Understanding of model quantization, ONNX, and deployment optimization is a plus.
Nice to Have
- Computer Vision (Face Recognition, Liveness Detection)
- Whisper / Speech-to-Text systems
- Stable Diffusion or image generation models
- GraphQL experience
- Kubernetes exposure
Skills
- Python
- LangChain
- FastAPI
- Pinecone
- Docker