About the Role
We are looking for a Robotics Developer with strong experience in camera-based perception, deep learning, and modern vision-language models.
You will work on real-world autonomous robots operating in dynamic environments. The role involves deploying and optimizing camera systems, building perception pipelines, integrating deep-learning models, and exploring Vision-Language Models (VLMs) and Vision-Language-Action models (VLAs) for robotic decision-making and interaction.
This is a hands-on engineering role requiring experience moving systems from prototypes and research environments to reliable deployment on physical robots.
Key Responsibilities
- Develop, deploy, and maintain camera-based perception systems on autonomous robots.
- Integrate monocular, stereo, depth, USB, MIPI, GMSL, or similar camera systems.
- Build and optimize camera pipelines using tools such as GStreamer, OpenCV, ROS, and NVIDIA multimedia frameworks.
- Diagnose camera, driver, synchronization, latency, frame-drop, and hardware-interface issues.
- Develop deep-learning models for object detection, segmentation, tracking, depth estimation, scene understanding, and anomaly detection.
- Train, fine-tune, evaluate, and deploy neural-network models on edge-compute platforms.
- Work with Vision-Language Models for scene interpretation, semantic reasoning, operator assistance, and robot-task understanding.
- Explore and integrate Vision-Language-Action models and multimodal policies for robotic applications.
- Build datasets, annotation pipelines, evaluation frameworks, and model-monitoring tools.
- Optimize models using TensorRT, ONNX, CUDA, quantization, pruning, or other acceleration techniques.
- Integrate perception and AI components with ROS-based navigation and control systems.
- Profile and optimize CPU, GPU, memory, latency, and power consumption on embedded hardware.
- Conduct testing on physical robots and support field deployments.
- Collaborate with robotics, controls, navigation, hardware, cloud, and operations teams.
- Document system architecture, deployment procedures, experiments, and performance results.
Required Qualifications
- Bachelor’s or Master’s degree in Computer Science, Robotics, Electronics, Electrical Engineering, Artificial Intelligence, or a related field.
- 2 or more years of professional experience in robotics, computer vision, autonomous systems, or applied AI.
- Strong programming skills in Python and C++.
- Experience deploying and debugging camera systems on real hardware.
- Strong understanding of image processing, camera calibration, coordinate transformations, and computer-vision fundamentals.
- Hands-on experience with deep-learning frameworks such as PyTorch or TensorFlow.
- Experience with computer-vision models, including object detection, segmentation, classification, or tracking.
- Experience with ROS or ROS 2.
- Familiarity with Linux development and debugging.
- Experience deploying models on NVIDIA GPUs or edge-compute platforms.
- Understanding of model training, validation, benchmarking, and dataset management.
- Strong debugging and problem-solving skills.
Preferred Qualifications
- Experience with Vision-Language Models such as CLIP, LLaVA, Qwen-VL, InternVL, Florence, or similar multimodal models.
- Experience with Vision-Language-Action models, imitation learning, behaviour cloning, reinforcement learning, or robot foundation models.
- Experience fine-tuning multimodal models using LoRA, QLoRA, instruction tuning, or related techniques.
- Experience with NVIDIA Jetson platforms, CUDA, TensorRT, DeepStream, or JetPack.
- Experience with GStreamer and hardware-accelerated video pipelines.
- Experience with MIPI CSI, GMSL, FPD-Link, RealSense, stereo, or depth cameras.
- Experience with multimodal sensor fusion involving cameras, LiDAR, radar, IMU, or odometry.
- Familiarity with transformer architectures and large-language-model inference.
- Experience with Docker, Git, CI/CD, and cloud-based model-training workflows.
- Experience deploying autonomous robots in outdoor, industrial, warehouse, airport, campus, or public environments.
- Exposure to SLAM, localization, navigation, motion planning, or robot-control systems.
Technical Skills
- Languages: Python, C++
- Robotics: ROS, ROS 2, TF, sensor integration
- Computer Vision: OpenCV, camera calibration, detection, segmentation, tracking
- Deep Learning: PyTorch, TensorFlow, transformers
- Multimodal AI: VLMs, VLAs, vision transformers, language-conditioned policies
- Deployment: ONNX, TensorRT, CUDA, quantization
- Video Systems: GStreamer, DeepStream, V4L2
- Platforms: NVIDIA Jetson, embedded Linux, GPU-based edge systems
- Tools: Docker, Git, Linux debugging and profiling tools
What We Value
- A strong hands-on engineering mindset.
- Ability to debug across software, operating systems, drivers, cameras, and hardware.
- Interest in deploying AI on real robots rather than working only with offline datasets.
- Ability to balance model accuracy with latency, compute, reliability, and operational constraints.
- Ownership of projects from experimentation through field deployment.
- Comfort working in a fast-moving robotics environment where systems are tested under real-world conditions.
What You Will Work On
- Camera-based perception for autonomous mobile robots.
- Small-object and obstacle detection in complex environments.
- Multimodal scene understanding using VLMs.
- Language-conditioned robot behaviours and VLA-based research.
- Edge deployment and optimization of deep-learning models.
- Reliable perception pipelines for continuously operating robots.
- Tools for data collection, model evaluation, diagnostics, and field debugging.