Computer Vision Toolbox
R2026bComputer Vision Toolbox™ provides algorithms, apps, and AI models for designing, simulating, calibrating, and deploying computer vision systems. You can perform object detection and tracking, feature matching, and optical flow. You can automate camera calibration, including multi-sensor configurations. For 3D vision, the toolbox supports structure from motion, real-time SLAM, and novel view synthesis such as NeRF. Computer vision apps enable calibration and multi-user image and video labeling, including automation capabilities.
The toolbox provides AI techniques, including pretrained convolutional neural networks, vision transformers, and vision-language models. You can use these pre-trained models for tasks such as image classification, object detection, segmentation, pose estimation, image captioning, visual question answering, and OCR, or further customize them through transfer learning.
You can generate code in C/C++ and HDL, and for GPU or NPU hardware targets (with MATLAB® Coder™, HDL Coder™, GPU Coder™, and hardware support packages). You can also build custom apps (with MATLAB Compiler™)
Get Started
Learn the basics of Computer Vision Toolbox
Detect, Extract, and Match Features
Detect interest points, extract feature descriptors, match features, register and retrieve images
Ground Truth Images and Video
Interactively label images and videos using AI-assisted automation, create training data for AI models, and manage collaborative team labeling for large data sets
Detect and Segment Objects
Detect objects, recognize text (OCR), barcodes, and fiducial markers, perform semantic and instance segmentation using AI models
Classify Images and Videos
Classify images and videos and perform activity recognition using AI models
Vision-Language Models
Perform image classification, retrieval, captioning, and object detection tasks using vision-language models
Calibrate Cameras
Automate intrinsic and extrinsic parameter calibration for single, fisheye, and stereo camera configurations
Calibrate Multi-Sensor Systems
Estimate the position and orientation between cameras, lidars, IMUs, and perform robot hand-eye calibration
3-D Vision
Estimate camera poses, perform stereo vision, reconstruct 3-D scenes from stereo or structure from motion (SfM), implement real-time visual SLAM with inertial sensor fusion
Track Objects and Estimate Motion
Track multiple objects, track feature points, object re-identification (ReID), optical flow, and template matching