🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Job Description Data Scientist Computer Vision & Vision Language Models (VLM) Location: Pune, Maharashtra, India Experience: 5+ Years Total Experience | Minimum 2+ Years Relevant Experience in Computer Vision / Multimodal AI Education Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Data Science, Computer Engineering, or a related field. Technical Skills Machine Learning & Deep Learning Strong understanding of Machine Learning and Deep Learning fundamentals. Strong understanding of Transformer architectures and attention mechanisms. Practical knowledge of Large Language Models (LLMs) and Vision Language Models (VLMs). Hands-on experience fine-tuning deep learning models. Computer Vision Strong understanding of image processing and computer vision fundamentals. Experience working with image datasets and visual data pipelines. Familiarity with OCR, document understanding, object detection, image classification, or related computer vision tasks. Multimodal AI Experience working with Vision Language Models and multimodal systems. Understanding of model evaluation, benchmarking, and performance optimization techniques. Experience with parameter-efficient fine-tuning methods such as LoRA, QLoRA, or PEFT. Programming & Frameworks Solid Python programming skills. Hands-on experience with: PyTorch Hugging Face Transformers OpenCV CUDA/GPU-based training environments Preferred Qualifications Experience working with engineering drawings, technical drawings, CAD documents, or manufacturing data. Exposure to vector graphics, PDFs, SVGs, CAD formats, or technical document processing. Experience with synthetic data generation techniques. Hands-on experience with: Docker MLflow Databricks Model deployment workflows Experience working with cloud-based AI/ML platforms. Knowledge of MLOps best practices and experiment management. Desired Candidate Profile The ideal candidate: Has a robust foundation in Computer Vision and Deep Learning. Understands how transformer-based multimodal models work. Is comfortable experimenting with new architectures and research papers. Can independently analyze datasets and identify prospects for improvement. Enjoys solving challenging AI problems using both research and engineering approaches. Has a passion for building practical AI solutions from cutting-edge research. .