About the Role

We are looking for a Robotics Developer with strong experience in camera-based perception, deep learning, and modern vision-language models.

You will work on real-world autonomous robots operating in dynamic environments. The role involves deploying and optimizing camera systems, building perception pipelines, integrating deep-learning models, and exploring Vision-Language Models (VLMs) and Vision-Language-Action models (VLAs) for robotic decision-making and interaction.

This is a hands-on engineering role requiring experience moving systems from prototypes and research environments to reliable deployment on physical robots.

Key Responsibilities

  • Develop, deploy, and maintain camera-based perception systems on autonomous robots.
  • Integrate monocular, stereo, depth, USB, MIPI, GMSL, or similar camera systems.
  • Build and optimize camera pipelines using tools such as GStreamer, OpenCV, ROS, and NVIDIA multimedia frameworks.
  • Diagnose camera, driver, synchronization, latency, frame-drop, and hardware-interface issues.
  • Develop deep-learning models for object detection, segmentation, tracking, depth estimation, scene understanding, and anomaly detection.
  • Train, fine-tune, evaluate, and deploy neural-network models on edge-compute platforms.
  • Work with Vision-Language Models for scene interpretation, semantic reasoning, operator assistance, and robot-task understanding.
  • Explore and integrate Vision-Language-Action models and multimodal policies for robotic applications.
  • Build datasets, annotation pipelines, evaluation frameworks, and model-monitoring tools.
  • Optimize models using TensorRT, ONNX, CUDA, quantization, pruning, or other acceleration techniques.
  • Integrate perception and AI components with ROS-based navigation and control systems.
  • Profile and optimize CPU, GPU, memory, latency, and power consumption on embedded hardware.
  • Conduct testing on physical robots and support field deployments.
  • Collaborate with robotics, controls, navigation, hardware, cloud, and operations teams.
  • Document system architecture, deployment procedures, experiments, and performance results.

Required Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Robotics, Electronics, Electrical Engineering, Artificial Intelligence, or a related field.
  • 2 or more years of professional experience in robotics, computer vision, autonomous systems, or applied AI.
  • Strong programming skills in Python and C++.
  • Experience deploying and debugging camera systems on real hardware.
  • Strong understanding of image processing, camera calibration, coordinate transformations, and computer-vision fundamentals.
  • Hands-on experience with deep-learning frameworks such as PyTorch or TensorFlow.
  • Experience with computer-vision models, including object detection, segmentation, classification, or tracking.
  • Experience with ROS or ROS 2.
  • Familiarity with Linux development and debugging.
  • Experience deploying models on NVIDIA GPUs or edge-compute platforms.
  • Understanding of model training, validation, benchmarking, and dataset management.
  • Strong debugging and problem-solving skills.

Preferred Qualifications

  • Experience with Vision-Language Models such as CLIP, LLaVA, Qwen-VL, InternVL, Florence, or similar multimodal models.
  • Experience with Vision-Language-Action models, imitation learning, behaviour cloning, reinforcement learning, or robot foundation models.
  • Experience fine-tuning multimodal models using LoRA, QLoRA, instruction tuning, or related techniques.
  • Experience with NVIDIA Jetson platforms, CUDA, TensorRT, DeepStream, or JetPack.
  • Experience with GStreamer and hardware-accelerated video pipelines.
  • Experience with MIPI CSI, GMSL, FPD-Link, RealSense, stereo, or depth cameras.
  • Experience with multimodal sensor fusion involving cameras, LiDAR, radar, IMU, or odometry.
  • Familiarity with transformer architectures and large-language-model inference.
  • Experience with Docker, Git, CI/CD, and cloud-based model-training workflows.
  • Experience deploying autonomous robots in outdoor, industrial, warehouse, airport, campus, or public environments.
  • Exposure to SLAM, localization, navigation, motion planning, or robot-control systems.

Technical Skills

  • Languages: Python, C++
  • Robotics: ROS, ROS 2, TF, sensor integration
  • Computer Vision: OpenCV, camera calibration, detection, segmentation, tracking
  • Deep Learning: PyTorch, TensorFlow, transformers
  • Multimodal AI: VLMs, VLAs, vision transformers, language-conditioned policies
  • Deployment: ONNX, TensorRT, CUDA, quantization
  • Video Systems: GStreamer, DeepStream, V4L2
  • Platforms: NVIDIA Jetson, embedded Linux, GPU-based edge systems
  • Tools: Docker, Git, Linux debugging and profiling tools

What We Value

  • A strong hands-on engineering mindset.
  • Ability to debug across software, operating systems, drivers, cameras, and hardware.
  • Interest in deploying AI on real robots rather than working only with offline datasets.
  • Ability to balance model accuracy with latency, compute, reliability, and operational constraints.
  • Ownership of projects from experimentation through field deployment.
  • Comfort working in a fast-moving robotics environment where systems are tested under real-world conditions.

What You Will Work On

  • Camera-based perception for autonomous mobile robots.
  • Small-object and obstacle detection in complex environments.
  • Multimodal scene understanding using VLMs.
  • Language-conditioned robot behaviours and VLA-based research.
  • Edge deployment and optimization of deep-learning models.
  • Reliable perception pipelines for continuously operating robots.
  • Tools for data collection, model evaluation, diagnostics, and field debugging.
Job Category: Robotics Developer – Perception and AI
Job Type: Full Time
Job Location: India Noida

Apply for this position

Allowed Type(s): .pdf, .doc, .docx