Skip to main content

SO Development

Ultralytics YOLO27: The Next Generation of Computer Vision

Introduction

Real-time computer vision is rapidly evolving. From autonomous vehicles and drones to robotics, industrial inspection, smart cameras, and medical imaging, AI systems increasingly need to understand visual information with both high accuracy and extremely low latency.

The YOLO family has played a major role in this evolution. After years of improvements in speed, accuracy, and deployment, Ultralytics is preparing its next-generation model family: YOLO27.

YOLO27 is designed to continue the YOLO philosophy of fast, practical computer vision while introducing new architectural ideas. The upcoming family combines a streamlined CNN architecture for its smaller models with a query-based, NMS-free detection architecture for its larger models.

According to the official Ultralytics documentation, YOLO27 is currently undergoing final research and development. The models are expected to launch later in 2026, but no official release date has been announced yet. The weights, configurations, and implementation code are not publicly available, and the current features and benchmark results may change before release.

For developers interested in the evolution of YOLO, our article on YOLO26: The Next Evolution of Real-Time Computer Vision provides useful background on the previous generation and the industry’s movement toward end-to-end, deployment-ready vision models.

What Is Ultralytics YOLO27?

Ultralytics YOLO27 is an upcoming family of real-time computer vision models designed for a broad range of vision tasks.

The family includes four planned model sizes:

  • YOLO27n – Nano
  • YOLO27s – Small
  • YOLO27m – Medium
  • YOLO27l – Large

The models are planned to support:

  • Object detection
  • Instance segmentation
  • Semantic segmentation
  • Depth estimation
  • Image classification
  • Pose estimation
  • Oriented object detection
  • Multi-object tracking for supported models

One of the most interesting aspects of YOLO27 is that its detection models use two different architectural approaches.

YOLO27n and YOLO27s use a streamlined CNN-based design optimized for efficient inference, while YOLO27m and YOLO27l use a query-based architecture that can generate detections without traditional Non-Maximum Suppression (NMS).

This makes YOLO27 less of a simple model-size upgrade and more of an architectural evolution of the YOLO family.

YOLO27 and the Evolution of YOLO Models

The YOLO architecture has evolved significantly since its introduction.

Earlier generations focused primarily on making object detection faster and more accurate. Later models introduced improvements in feature extraction, training, model efficiency, and deployment.

Developers who want to understand the practical evolution of YOLO can also explore our guide on How to Use YOLOv11 for Object Detection.

YOLO27 continues this progression by addressing several challenges that remain important in modern computer vision:

  • Small-object detection
  • Inference latency
  • Post-processing overhead
  • Deployment complexity
  • Accuracy versus model size
  • Efficient edge inference

The result is a model family designed around different deployment requirements rather than a single architecture optimized for every situation.

YOLO27 vs. YOLO26

YOLO27 is positioned as the successor to YOLO26.

However, the key difference is not simply that YOLO27 is larger or more accurate.

The new generation introduces different architectural strategies for different model sizes.

YOLO27n and YOLO27s

The smaller models focus on:

  • Efficient CNN architecture
  • Dual-scale detection
  • Small-object detection
  • Lower detection-head computation
  • Edge deployment
  • Real-time video

YOLO27m and YOLO27l

The larger models focus on:

  • Query-based detection
  • Transformer decoding
  • NMS-free inference
  • High accuracy
  • GPU deployment
  • Global visual context

The official documentation currently recommends continuing to use YOLO26 for production projects until YOLO27 is released.

For a detailed look at the previous generation, see our full guide to YOLO26: The Next Evolution of Real-Time Computer Vision.

YOLO26 vs YOLO27_ AI Detection Showdown

YOLO27’s Two Detection Architectures

One of the biggest innovations in YOLO27 is that the four detection models do not all use the same detection architecture.

Ultralytics has divided them into two groups.

YOLO27n and YOLO27s: Streamlined CNN Architecture

YOLO27n and YOLO27s are designed for applications where inference speed and computational efficiency are critical.

Traditional object detectors often predict objects using three feature maps:

  1. Fine-resolution features
  2. Medium-resolution features
  3. Coarse-resolution features

YOLO27n and YOLO27s remove the medium prediction scale.

Instead, they use:

  • A fine feature map for small objects
  • A coarse feature map for large objects

Removing the medium prediction map reduces a significant portion of detection-head computation while retaining the feature scales considered most important for small and large objects.

This design makes the smaller models particularly interesting for:

  • Edge AI
  • Drones
  • Robotics
  • Smart cameras
  • Real-time video
  • Embedded systems

Small-Object Detection In YOLO27

Small-object detection remains one of the hardest challenges in computer vision.

Objects that occupy only a small number of pixels contain less visual information, making their classification and localization more difficult.

This is especially important in:

  • Drone imagery
  • Traffic monitoring
  • Satellite imagery
  • Surveillance
  • Autonomous vehicles
  • Large-area video analytics

Our detailed guide on Small Object Detection in Computer Vision: Challenges, Techniques, and Future Trends explores why small targets are difficult to detect and the techniques used to improve their recognition.

YOLO27n and YOLO27s address this challenge by widening an early high-resolution feature stage so that the network can retain more fine-grained information.

Combined with the remaining fine-resolution prediction map, this is intended to improve small-object localization and regression.

Foreground Alignment Supervision

YOLO27n and YOLO27s also introduce foreground alignment supervision.

During training, a lightweight auxiliary branch learns to separate foreground objects from background regions. This bridges the gap between dense training supervision and the final one-to-one prediction head. 

According to Ultralytics, the accuracy gap is reduced from approximately:

  • 0.9 mAP to 0.4 mAP for YOLO26n → YOLO27n
  • 0.8 mAP to 0.4 mAP for YOLO26s → YOLO27s

The additional branch is used only during training and removed during inference and export, meaning it does not add deployment-time computation.

YOLO27m and YOLO27l: NMS-Free Detection

The larger YOLO27 models take a different approach.

Instead of generating many overlapping predictions and filtering them using Non-Maximum Suppression, YOLO27m and YOLO27l use a query-based transformer decoder.

The decoder works with a fixed set of object queries and progressively refines them to produce the final detections.

This means the final predictions can be generated directly without traditional NMS post-processing.

Why Does NMS-Free Detection Matter?

Traditional object detection systems often produce multiple candidate boxes for the same object.

NMS is then responsible for removing redundant predictions.

Although NMS works well, it adds an additional post-processing step to the inference pipeline.

Removing this step can potentially provide:

  • Lower latency
  • Simpler deployment
  • More predictable inference
  • Reduced post-processing
  • Better integration with end-to-end systems

NMS-free detection is an important trend in modern real-time computer vision, and YOLO27 continues this direction.

Our previous article on YOLO26 also discusses the transition toward NMS-free, end-to-end detection.

YOLO27m: The Accuracy-Speed Balance

YOLO27m is positioned as the middle ground between the lightweight N/S models and the accuracy-focused L model.

The model combines the query-based transformer decoder with a convolutional backbone based on the YOLO26 architecture.

According to the preliminary Ultralytics benchmark, YOLO27m achieves:

55.8 mAP at 640-pixel input resolution

with approximately:

1.39 ms inference latency on an NVIDIA RTX PRO 6000 using TensorRT 11.

This makes YOLO27m particularly interesting for GPU-based production applications where organizations need a strong balance between accuracy and latency.

YOLO27l and UltraViT

YOLO27l is the flagship model of the new family, combining query-based detection with an UltraViT backbone. The deepest stage utilizes self-attention mechanisms to capture global visual context across the entire frame.

This global receptive field enables the network to accurately recognize complex relationships between distant visual elements. According to Ultralytics, YOLO27l is the first model in their lineup to cross the 60 mAP threshold on COCO.

The preliminary results report:

  • 60.4 mAP at 640 pixels
  • 61.2 mAP at 800 pixels
  • Approximately 2.32 ms latency at 640 pixels on an NVIDIA RTX PRO 6000

These results are preliminary and may change before the final release.

YOLO27 Performance Benchmarks

Ultralytics has published preliminary detection benchmarks for the four YOLO27 models.

Model

Input

mAP50-95

CPU ONNX

RTX PRO 6000 TensorRT

Parameters

FLOPs

YOLO27n

640

42.3

16.1 ms

0.62 ms

3.0M

7.2B

YOLO27s

640

49.6

33.2 ms

0.79 ms

11.8M

28.2B

YOLO27m

640

55.8

67.2 ms

1.39 ms

22.8M

65.0B

YOLO27l

640

60.4

149.6 ms

2.32 ms

72.3M

165.3B

The official benchmarks were measured on an NVIDIA RTX PRO 6000 using TensorRT 11 and FP16 for GPU inference, while CPU measurements use an AMD EPYC 9655 with ONNX Runtime and FP32.

These figures should be treated as research benchmarks rather than guaranteed production performance.

Actual performance will depend on:

  • Hardware
  • Input resolution
  • Runtime
  • Precision
  • Dataset
  • Batch size
  • Optimization settings
  • Deployment environment

What Do the YOLO27 Numbers Mean?

The four model sizes target different requirements.

YOLO27n

YOLO27n has only 3 million parameters and is designed for highly efficient inference.

Potential applications include:

  • Edge devices
  • Embedded cameras
  • Drones
  • Robotics
  • IoT vision systems

YOLO27s

YOLO27s provides greater model capacity while maintaining very low reported latency.

It can be useful when an application needs more accuracy while still requiring real-time performance.

YOLO27m

YOLO27m is the middle-ground option.

Its 55.8 mAP preliminary result makes it attractive for GPU-based applications that need a balance between accuracy and inference speed.

YOLO27l

YOLO27l prioritizes accuracy.

With a preliminary 60.4 mAP at 640 pixels, it is aimed at applications where detection quality is more important than minimizing model size.

YOLO27 Supported Computer Vision Tasks

YOLO27 is engineered as a unified multi-task vision framework rather than a standalone object detector. Beyond bounding box detection, the ecosystem natively supports multiple downstream tasks within the same architecture:

Task

Core Capability

Key Applications

Object Detection & OBB

Standard & Oriented Bounding Boxes

Security, Aerial/Satellite, Retail Analytics

Segmentation

Instance & Semantic Masking

Medical Imaging, Autonomous Driving, Industrial Defect Detection

Pose & Tracking

Keypoint Estimation & Multi-Object Tracking

Sports Analytics, Workplace Safety, Video Surveillance

Classification & Depth

Scene Categorization & Distance Estimation

Product Categorization, Robotics, 3D Scene Understanding

For developers looking to understand the wider object-detection landscape, our Best Object Detection Models for Computer Vision in 2026 guide compares major approaches and models.

YOLO27 for Real-Time Video

Real-time video is one of the most important use cases for YOLO models.

A computer vision system processing video may need to analyze dozens of frames every second while maintaining accurate detections.

This creates a fundamental trade-off:

More accuracy often requires more computation.

More speed often requires a smaller model.

YOLO27 addresses this through multiple model sizes.

YOLO27n and YOLO27s target highly efficient deployments, while YOLO27m and YOLO27l provide higher model capacity for GPU-based applications.

This allows developers to select a model according to their hardware and accuracy requirements.

YOLO27 for Drones

Drones require computer vision systems that can operate with limited computational resources.

A drone may need to detect:

  • People
  • Vehicles
  • Buildings
  • Roads
  • Obstacles
  • Other aircraft
  • Objects of interest

At the same time, onboard computing hardware must operate within power and weight constraints.

YOLO27n and YOLO27s are specifically positioned for edge devices, drones, and real-time video.

The improved small-object detection of the smaller models could also be valuable for aerial applications where targets may occupy only a small part of each frame.

YOLO27 for Autonomous Vehicles

Autonomous vehicles require continuous perception of their surroundings.

A perception system may need to identify:

  • Cars
  • Trucks
  • Pedestrians
  • Cyclists
  • Traffic signs
  • Traffic lights
  • Road obstacles

Inference latency is critical because decisions may need to be made within milliseconds.

YOLO27’s combination of real-time inference and improved detection accuracy makes the model family potentially useful for automotive perception systems.

However, benchmark results on COCO should not be interpreted as proof of autonomous-driving performance. Real-world automotive systems require domain-specific datasets, extensive testing, safety validation, and hardware-specific optimization.

For organizations building autonomous-vehicle perception systems, high-quality training data is equally important. Our autonomous vehicle perception annotation case study explores how large-scale annotation can support computer vision development.

YOLO27 for Security and Surveillance

Security systems increasingly rely on computer vision to process large amounts of video.

Potential detection tasks include:

  • Person detection
  • Vehicle detection
  • Restricted-area monitoring
  • Crowd analysis
  • Abandoned-object detection
  • Safety monitoring

Lightweight models such as YOLO27n and YOLO27s could be deployed close to cameras, reducing the amount of raw video that needs to be transmitted to centralized servers.

Edge processing can potentially reduce:

  • Network bandwidth
  • Cloud processing costs
  • Response latency

It can also support privacy-oriented architectures where video is processed locally rather than continuously uploaded.

YOLO27 for Industrial Inspection

Manufacturers increasingly use computer vision to automate quality control.

A trained YOLO model can detect defects such as:

  • Scratches
  • Cracks
  • Missing components
  • Incorrect assembly
  • Surface damage
  • Manufacturing defects

For high-speed production lines, inference latency is especially important.

A model that takes too long to process each image can become a bottleneck in the manufacturing workflow.

YOLO27’s range of model sizes could allow manufacturers to balance detection accuracy and production-line speed.

YOLO27 for Medical AI

Computer vision is also increasingly important in medical AI.

Potential applications include:

  • Anatomical structure detection
  • Lesion detection
  • Organ segmentation
  • Medical image classification
  • Cell analysis
  • Medical image quality control

YOLO27’s planned support for detection, segmentation, classification, and depth estimation provides a broad technical foundation for specialized medical computer vision systems.

However, medical AI requires significantly more validation than ordinary computer vision applications.

A strong benchmark score does not automatically mean that a model is clinically safe, diagnostically accurate, or suitable for regulatory use.

Medical AI projects require carefully curated datasets, expert annotation, validation, and appropriate clinical oversight.

YOLO27 Applications Infographic

Training YOLO27 on Custom Datasets

One of the biggest advantages of the Ultralytics ecosystem is its focus on custom training.

Once YOLO27 is officially released, developers are expected to be able to use the familiar Ultralytics API to train models on custom datasets.

The official documentation currently provides a planned workflow similar to:

from ultralytics import YOLO

# Load a YOLO27 model

model = YOLO(“yolo27n.pt”)

# Run inference

results = model(“image.jpg”)

# Train on a dataset

results = model.train(

    data=”coco8.yaml”,

    epochs=100,

    imgsz=640

)

However, this code will not work with the current public package, because the YOLO27 weights and implementation have not yet been released.

Before training a production YOLO model, the quality of the dataset should be a major consideration.

Our guide Fine-Tuning YOLO Models with an Automated Data-Labeling Pipeline explores how automated labeling, human verification, and iterative model improvement can be combined to create efficient YOLO training workflows.

Why Data Quality Matters for YOLO27

A more advanced model does not eliminate the need for high-quality training data.

If training data contains:

  • Incorrect bounding boxes
  • Missing objects
  • Inconsistent labels
  • Poor image quality
  • Class imbalance
  • Incorrect segmentation masks

then even a powerful model can produce unreliable results.

For production computer vision systems, the data pipeline can be just as important as the model architecture.

A typical workflow may include:

  1. Data collection
  2. Data cleaning
  3. Annotation
  4. Quality assurance
  5. Model training
  6. Validation
  7. Error analysis
  8. Additional data collection
  9. Retraining
  10. Production deployment

This is particularly important for specialized applications such as autonomous driving, medical imaging, industrial inspection, and geospatial AI.

YOLO27 and Data Annotation

YOLO27’s potential performance ultimately depends on the quality and diversity of the data used to train and evaluate it.

For custom computer vision applications, organizations may need:

  • Bounding-box annotation
  • Polygon segmentation
  • Semantic segmentation
  • Keypoint annotation
  • Oriented bounding boxes
  • LiDAR annotation
  • Video annotation
  • Image classification

SO Development provides image annotation services for computer vision projects requiring large-scale, production-ready training data.

Combining high-quality annotation with automated pre-labeling and human quality control can help teams create datasets suitable for fine-tuning specialized vision models.

One Interface for Different YOLO27 Architectures

Despite the architectural differences between YOLO27n/s and YOLO27m/l, developers are expected to interact with the models through the same YOLO Python interface.

This is an important usability advantage.

For example, after release, a developer could load a Nano model with:

model = YOLO(“yolo27n.pt”)

and later switch to:

model = YOLO(“yolo27m.pt”)

without redesigning the entire application.

Ultralytics states that the correct training, validation, prediction, and export pipeline is selected automatically based on the model architecture.

This should make it easier to benchmark different model sizes against the same application and dataset.

Which YOLO27 Model Should You Choose?

The right model depends on the deployment environment.

Application

Potential Choice

Edge AI

YOLO27n

Drones

YOLO27n / YOLO27s

Smart cameras

YOLO27n / YOLO27s

Real-time video

YOLO27n / YOLO27s

GPU production

YOLO27m

Accuracy-speed balance

YOLO27m

Accuracy-critical applications

YOLO27l

High-resolution detection

YOLO27l

According to Ultralytics’ current guidance, YOLO27n and YOLO27s are intended for edge devices, drones, and real-time video; YOLO27m is positioned as the accuracy-speed option for GPU deployment; and YOLO27l targets accuracy-critical applications.

YOLO27 vs. Other Object Detection Models

YOLO27 enters an increasingly competitive computer vision ecosystem.

Developers can choose from:

  • YOLO models
  • RT-DETR
  • Transformer-based detectors
  • Open-vocabulary models
  • Vision foundation models
  • Specialized edge models

Choosing the best architecture depends on the requirements of the project.

Factors to consider include:

  • Accuracy
  • Latency
  • Hardware
  • Dataset size
  • Training cost
  • Deployment environment
  • Open-vocabulary requirements
  • Segmentation requirements
  • Tracking requirements

Our Best Object Detection Models for Computer Vision in 2026 guide provides a broader comparison of the major computer vision models and approaches.

Should You Upgrade From YOLO26 to YOLO27?

Not yet. 

YOLO27 is still unreleased.

Ultralytics currently recommends using YOLO26 for existing projects and evaluating YOLO27 after the final release.

Once YOLO27 becomes publicly available, teams should benchmark it on their own workloads rather than relying exclusively on COCO results.

Important metrics include:

  • mAP
  • Precision
  • Recall
  • Inference latency
  • Memory usage
  • CPU performance
  • GPU performance
  • Export compatibility
  • Power consumption
  • Real-world error rates

A model that performs better on COCO may not necessarily perform better on a specialized production dataset.

When Will YOLO27 Be Released?

According to the current Ultralytics documentation, YOLO27 is undergoing final research and development, with a launch anticipated later in 2026.

However, no official launch date has been announced.

The model weights, configuration files, and implementation code are not currently available. Ultralytics also warns that the features and benchmarks shown in the preview may change before release.

Developers interested in the release can follow the official YOLO27 documentation and wait for the public package and pretrained weights.

The Future of Real-Time Computer Vision

YOLO27 highlights an important direction in computer vision: the convergence of efficient CNN architectures, transformer-based detection, and end-to-end inference.

Instead of using one architecture for every model size, Ultralytics is using different approaches depending on the requirements.

Smaller models

YOLO27n and YOLO27s focus on:

  • Low latency
  • Efficient computation
  • Small-object detection
  • Edge deployment

Larger models

YOLO27m and YOLO27l focus on:

  • Query-based detection
  • NMS-free inference
  • Transformer decoding
  • Higher accuracy
  • GPU deployment

This approach reflects a broader shift in computer vision toward models that are optimized not only for benchmark accuracy but also for practical deployment.

Final Thoughts

Ultralytics YOLO27 represents an interesting next step in the evolution of real-time computer vision.

Rather than simply increasing model size, the upcoming YOLO27 family introduces different architectures for different deployment requirements.

YOLO27n and YOLO27s are designed around efficient CNN-based detection, dual-scale prediction, and improved small-object representation.

YOLO27m and YOLO27l introduce query-based, NMS-free detection using transformer decoding.

At the top end, YOLO27l combines this approach with an UltraViT backbone and currently reports a preliminary 60.4 mAP at 640 pixels and 61.2 mAP at 800 pixels on COCO.

However, it is important to remember that YOLO27 is not yet publicly released.

Its current benchmarks are preliminary, its implementation is unavailable, and the final specifications may change.

For production projects today, YOLO26 remains the recommended choice. Once YOLO27 is released, developers should benchmark the new models on their own datasets and hardware before migrating.

The most important lesson is that model architecture is only one part of a successful computer vision system. High-quality training data, accurate annotation, rigorous quality assurance, appropriate hardware, and continuous evaluation remain essential.

For organizations building specialized vision systems, the combination of advanced models + high-quality data + human-verified annotation will continue to be a major factor in achieving reliable production performance.

As YOLO evolves toward faster, more accurate, and increasingly end-to-end architectures, YOLO27 could become an important step toward the next generation of real-time AI perception.

Frequently Asked Questions

What is YOLO27?

YOLO27 is an upcoming family of real-time computer vision models from Ultralytics. It includes Nano, Small, Medium, and Large variants and is planned to support detection, segmentation, classification, pose, depth estimation, and oriented object detection.

Is YOLO27 available now?

No. YOLO27 is currently undergoing final R&D. The official documentation states that a launch is anticipated later in 2026, but no fixed release date has been announced.

What is the main improvement in YOLO27?

The major improvements include dual-scale detection for YOLO27n/s, stronger small-object detection, foreground alignment supervision, and query-based NMS-free detection for YOLO27m/l.

Is YOLO27 better than YOLO26?

Preliminary benchmarks show improved accuracy across the YOLO27 detection family, but YOLO27 is not yet released. Ultralytics currently recommends using YOLO26 for production projects.

Which YOLO27 model is best for edge devices?

YOLO27n and YOLO27s are the models currently positioned for edge devices, drones, and real-time video.

Which YOLO27 model has the highest accuracy?

Among the preliminary 640-pixel detection benchmarks, YOLO27l has the highest reported accuracy at 60.4 mAP, increasing to 61.2 mAP at 800 pixels.

Can YOLO27 be trained on custom datasets?

The official documentation provides training examples intended for use after YOLO27 is released and package support becomes available. The current public package does not yet support YOLO27.

Why is training data important for YOLO27?

Even advanced computer vision models depend on accurate and representative training data. Incorrect annotations, missing objects, inconsistent labels, and poor-quality images can significantly reduce model performance.

Related SO Development Articles

If you want to learn more about YOLO and computer vision, explore:

Official source: Ultralytics YOLO27 documentation

Visit Our Data Annotation Service