Role Overview
Dataclap builds AI data pipelines — annotation, RLHF, human-in-the-loop, and MLOps — for clients across North America and Europe. We're looking for a Computer Vision Developer who can go beyond running a single off-the-shelf model: someone who has stitched multiple models into working pipelines, and who has trained or fine-tuned models rather than only consuming APIs. You'll design and build vision systems that power model-assisted labeling, automated QA of annotations, and delivery pipelines for our clients' datasets.
Responsibilities
- Design and build multi-model computer vision pipelines (e.g. detection → tracking → segmentation → classification / OCR) that run reliably at scale.
- Fine-tune and train custom models from open-source checkpoints to hit client-specific accuracy and edge-case requirements.
- Evaluate and benchmark candidate models, select the right architecture for each task, and document trade-offs.
- Build model-assisted labeling and auto-QA tooling to accelerate our annotation and HITL workflows.
- Handle the full lifecycle: data preparation, augmentation, training, validation, error analysis, and iteration.
- Optimize models for inference — quantization, ONNX/TensorRT export, batching — and package them for deployment.
- Set up experiment tracking, versioning, and reproducible training runs.
- Collaborate with annotation, delivery, and DevOps teams.
Requirements
- 3+ years of total software/ML engineering experience, with at least 1–2 years working specifically in computer vision.
- Hands-on experience with multiple computer vision models (e.g., YOLO, Faster R-CNN, DETR, SAM, Mask R-CNN, U-Net, ResNet, EfficientNet, ViT).
- Demonstrated experience building pipelines that chain multiple models together.
- Proven experience training custom models or fine-tuning from open-source models.
- Strong Python and solid experience with PyTorch and/or TensorFlow and OpenCV.
- Comfort with the data side: dataset curation, augmentation, and evaluating with metrics like mAP, IoU, and F1.
- Ability to read a recent CV paper or model repo and get it running.
Skills
- Python
- PyTorch
- TensorFlow
- OpenCV
- ONNX
Nice to Have
- Experience with vision-language / multimodal models (VLMs).
- Inference optimization and edge deployment experience (ONNX, TensorRT, quantization, distillation).
- MLOps exposure: Docker, experiment tracking (Weights & Biases, MLflow), model versioning, CI/CD.
- Familiarity with annotation platforms (CVAT, Label Studio).
- Cloud experience (AWS / GCP / Azure).
Benefits
- Health insurance
- Leave encashment
- Provident Fund