Role Overview
At JoshTalks AI, we believe voice will become the primary interface between humans and machines. We build benchmarks, datasets, and AI systems that power some of the world's most widely used speech models. This is a full-time, paid internship (6-12 months) where you will directly contribute to work that influences the global speech AI ecosystem.
Responsibilities
- Deliver processed speech and audio datasets via cloud-based container pipelines using Docker and Kubernetes.
- Design and manage end-to-end data delivery workflows from raw ingestion to structured output.
- Build and maintain ETL pipelines to prepare data for model training and evaluation at scale.
- Design and run evaluations for ASR and speech-to-speech systems.
- Benchmark leading speech models to identify real-world strengths, weaknesses, and failure modes.
- Fine-tune speech recognition models such as Whisper or wav2vec2.
- Experiment with multilingual, code-switched, accented, and noisy speech data.
Requirements
- Students in B.Tech/B.E. graduating in 2027 (CSE, EE, AI/ML, or related fields).
- Strong interest in speech, audio, NLP, or multimodal AI.
- Hands-on experience in fine-tuning speech or language models (Whisper, wav2vec2, HuBERT, etc.).
- Experience building speech-driven pipelines, classifiers, or assistants.
- Proficiency with PyTorch, TensorFlow, or Hugging Face Transformers.
- Experience with cloud platforms (AWS, GCP, Azure) and container tools (Docker, Kubernetes).
Skills
- PyTorch
- TensorFlow
- Docker
- Kubernetes
- AWS