Role Overview
We are looking for a talented Data Scientist with expertise in Generative AI, Large Language Models (LLMs), NLP, and Machine Learning to join our growing AI team. The ideal candidate will have hands-on experience building LLM-powered applications, developing RAG solutions, fine-tuning foundation models, and deploying AI solutions at scale.
Responsibilities
- Design, develop, and deploy end-to-end AI/ML solutions.
- Build LLM-powered applications using LangChain, LlamaIndex, OpenAI, Azure OpenAI, or similar frameworks.
- Develop and optimize prompt engineering strategies to improve model accuracy and reliability.
- Implement Retrieval-Augmented Generation (RAG) pipelines using vector databases such as FAISS, Pinecone, Chroma, or Weaviate.
- Fine-tune and evaluate LLMs including GPT, Llama, Mistral, Claude, and Gemini.
- Implement structured output validation and hallucination mitigation techniques using Pydantic and output parsers.
- Build NLP solutions for:
- Named Entity Recognition (NER)
- Text Classification
- Sentiment Analysis
- Information Extraction
- Document Summarization
- Question Answering
- Semantic Search
- Process and analyze large-scale unstructured data from documents, PDFs, emails, and other enterprise sources.
- Design, train, evaluate, and deploy supervised and unsupervised machine learning models.
- Perform feature engineering, model selection, hyperparameter tuning, and model validation.
- Build and maintain scalable ML pipelines from data ingestion to production deployment.
- Monitor model performance and address data drift through continuous optimization.
- Develop REST APIs using FastAPI or Flask.
- Work closely with engineering, product, and business teams to deliver AI-driven solutions.
- Perform exploratory data analysis (EDA) and communicate insights through data visualization and reporting.
Requirements
- Bachelor's or Master's degree in Computer Science, Data Science, Statistics, Mathematics, or a related field.
- 4–7 years of hands-on experience in Data Science, Machine Learning, or AI roles.
- Strong proficiency in Python (Pandas, NumPy, Scikit-learn).
- Experience with LLM frameworks such as LangChain, LlamaIndex, OpenAI, or Azure OpenAI.
- Hands-on experience with Prompt Engineering, RAG, and LLM application development.
- Strong knowledge of NLP libraries including Hugging Face Transformers, spaCy, and NLTK.
- Experience with Deep Learning frameworks such as PyTorch or TensorFlow.
- Experience with Pydantic for structured outputs and data validation.
- Proficiency in SQL and Vector Databases such as FAISS, Pinecone, Chroma, or Weaviate.
- Experience working with cloud platforms (AWS, Azure, or GCP).
- Familiarity with Git, Docker, CI/CD, FastAPI, and Flask.