Role Overview
We are seeking an experienced AI Data Engineer to design, develop, and deploy cutting-edge LLM-based applications. This role requires strong engineering expertise combined with hands-on experience in Generative AI, RAG pipelines, and agentic systems.
Responsibilities
- Design and develop LLM-based applications using single-agent or simple multi-agent patterns for business use cases.
- Build and maintain RAG pipelines: data ingestion, chunking, embeddings, retrieval, and response generation.
- Implement prompt engineering techniques (prompt templates, chaining, basic tool/function calling).
- Develop backend services/APIs for AI applications using Python frameworks (FastAPI / Flask / Streamlit).
- Integrate AI solutions with enterprise systems, databases, and APIs.
- Apply basic guardrails and validation checks to improve response quality and reduce hallucination.
- Work with Data Engineering teams to ensure data quality, pipeline efficiency, and proper documentation.
- Collaborate with MLOps teams for deployment, monitoring, and iterative improvements.
- Document solutions, reusable components, and best practices.
Requirements
- 4–6 years total experience, with 1+ year hands-on experience in GenAI / LLM-based applications.
- Strong hands-on experience with LLMs (Claude, OpenAI, etc.), RAG pipelines, and GPT + Agentic AI implementation.
- Experience with frameworks like LangChain, LangGraph, or similar.
- Deep understanding of LLM limitations, evaluation, and optimisation strategies.
- Strong Python/Pyspark engineering expertise (production-grade development) with proven API integration experience.
- Deep data analysis experience and handling large volumes of data.
- Experience with Fabric/Azure Databricks/Snowflake data engineering integration.
- Good exposure to Cloud platforms (Azure/AWS/GCP) and SQL.
- Familiarity with Containers, CI/CD, and monitoring.
- Prior experience in Data Engineering (ETL/ELT, pipelines, orchestration), Data Science/ML lifecycle (especially NLP), or Analytics engineering.
- Good-to-have: Exposure to model fine-tuning (LoRA/PEFT), LLM output evaluation, and familiarity with the Azure AI / Azure OpenAI / AI Search ecosystems.