Role Overview
At JoshTalks AI, we believe voice will become the primary interface between humans and machines. We build benchmarks, datasets, and AI systems that power some of the world's most widely used speech models. This is a full-time, paid, on-site internship in Gurgaon for students graduating in 2027.
Responsibilities
- Deliver processed speech and audio datasets via cloud-based container pipelines using Docker and Kubernetes.
- Design and manage end-to-end data delivery workflows and build ETL pipelines.
- Design and run evaluations for ASR and speech-to-speech systems.
- Fine-tune speech recognition models such as Whisper and wav2vec2.
- Experiment with multilingual, code-switched, and noisy speech data.
Requirements
- Students in B.Tech/B.E. graduating in 2027 (CSE, EE, AI/ML, or related fields).
- Strong interest in speech, audio, NLP, or multimodal AI.
- Hands-on experience with PyTorch, TensorFlow, or Hugging Face Transformers.
- Experience with cloud platforms like AWS, GCP, or Azure and container tools.
Nice to Have
- Open-source contributions, GitHub projects, or Kaggle experience.
- Experience with multilingual or low-resource speech data.
- Familiarity with ETL workflows, data versioning (DVC, MLflow), or orchestration (Airflow, Prefect).
Skills
- Python
- PyTorch
- Docker
- Kubernetes
- AWS