v0.30.8 — Sharper video, cleaner labels
Video, YOLO labels, masks, VLM parsing and mAP all get more accurate.
VideoSinkkeeps OpenCV's video quality when OpenCV isn't installed.from_yoloreads pose labels as boxes instead of polygons.MeanAveragePrecisionscores class-agnostic runs right when only one side has class IDs.from_vlmreturns one Florence-2 detection per object, not one per polygon.from_inferencemasks no longer drift up to a pixel up and left.
Drop-in upgrade. Without OpenCV, videos get larger; YOLO labels with a negative width or height now raise ValueError.
✨ Spotlights / highlights
sv.VideoSink and sv.process_video keep quality without OpenCV
The PyAV fallback left the encoder bit rate unset, so mp4v and MJPG files came out at under half of what cv2.VideoWriter writes. It now uses OpenCV's rate settings for every codec except H.264. Files get larger and vp09 may encode more slowly; codec="avc1" keeps files small where an H.264 encoder is available. (#2661)
import supervision as sv
video_info = sv.VideoInfo.from_video_path("in.mp4")
with sv.VideoSink("out.mp4", video_info) as sink: # OpenCV not installed
for frame in sv.get_video_frames_generator("in.mp4"):
sink.write_frame(frame)
# before: mp4v written at under half OpenCV's bit rate, visibly softer
# now: same bit rate OpenCV's writer usesYOLO pose labels load as boxes
A pose row is a box followed by keypoints. from_yolo used to parse the whole row as a polygon, giving wrong boxes and masks nobody asked for. It now reads the box and skips the keypoints that kpt_shape declares. (#2655)
ds = sv.DetectionDataset.from_yolo(
images_directory_path="pose/images",
annotations_directory_path="pose/labels",
data_yaml_path="pose/data.yaml", # kpt_shape: [17, 3]
)
# before: polygon-parsed boxes and masks
# now: one box per rowClass-agnostic mAP with one-sided class IDs
A perfect match scored zero when only one side carried class IDs, such as SAM proposals checked against labeled ground truth. With class_agnostic=True, both sides now count as one class.
One Florence-2 detection per object (#2648)
Florence-2 returns a segmented object as a list of polygons, one per connected region. An object split in two used to come back as two detections; the polygons now merge into one mask with one box around all of them.
Roboflow masks sit on the right pixels (#2649)
from_inference truncated sub-pixel polygon vertices, shifting each mask up and left by up to a pixel. Vertices are now rounded, the way the COCO, YOLO, LabelMe and Pascal VOC loaders already do.
🔄 Migration guide
No migration required for this release.
📝 Notable changes
🌱 Changed
sv.DetectionDataset.from_yoloraisesValueErrornaming the annotation file when a label has a negative width or height; it used to load a box withx_minpastx_max, which madeDetections.areanegative and skewed IoU and NMS.as_yolonow orders the corners of a reversed box before measuring, so it no longer writes a file the loader refuses. (#2663)
🔧 Fixed
sv.VideoSinkandsv.process_videowritemp4v,MJPGand other non-H.264 video at OpenCV's bit rate when OpenCV isn't installed. A frame rate of zero or less now raisesRuntimeErrorinsv.VideoSink, as it does with OpenCV. (#2661)sv.DetectionDataset.from_yoloreads the box of Ultralytics pose labels and skips their keypoints; akpt_shapeother than[K, 2]or[K, 3]raisesValueError. (#2655)sv.DetectionDataset.from_yolonames a malformed annotation line, one with too few values or, for OBB, not nine, in aValueErrorinstead of failing on an array shape. (#2665)sv.metrics.MeanAveragePrecision(class_agnostic=True)treats detections without class IDs as the same class as labeled ones, and unsigned class ID arrays no longer raiseOverflowErroron NumPy 2. (#2650)sv.Detections.from_vlmwithsv.VLM.FLORENCE_2merges the polygons of one instance into one detection for<REFERRING_EXPRESSION_SEGMENTATION>and<REGION_TO_SEGMENTATION>, and skips instances with no usable polygon. (#2648)sv.Detections.from_vlmwithsv.VLM.QWEN_2_5_VLorsv.VLM.QWEN_3_VLrecovers complete detections from a response cut off inside abbox_2darray or right after a complete object. (#2666)sv.Detections.from_inferencerounds polygon vertices to the nearest pixel before rasterising masks, and raisesValueErrorfor NaN or infinite vertices. (#2649)sv.xyxy_to_maskreturns an empty mask for a box entirely left of or above the image when its maximum coordinate is a negative fraction. (#2646)sv.Detections.get_anchors_coordinatescomputes axis-aligned midpoint anchors without integer overflow. (#2660)sv.LineZone.triggerages crossing history on frames whose detections lacktracker_id, so a reused track ID no longer creates a false crossing after the track expired. (#2644)sv.LineZoneAnnotator(text_orient_to_line=True)no longer raisesTypeErrorwithout OpenCV for lines drawn right to left. (#2659)- The Ultralytics, Inference and YOLO-NAS speed estimation examples measure elapsed time from frame indices; a vehicle missed in one frame of three was reported about 44% too fast. (#2654)
examples/speed_estimation/rfdetr_example.pyno longer raisesAttributeError:supervision._cv2now providesgetPerspectiveTransformandperspectiveTransform. (#2652)
🏆 Contributors
- Mohammad Hijjawi (@MohammadHijjawi97, LinkedIn) — fixed video quality without OpenCV, Florence-2 instance merging and Roboflow mask rounding.
- Kari Pikkarainen (@kari-pikkarainen, LinkedIn) — fixed speed estimation timing, the NumPy
flipfallback and the perspective-transform fallbacks. - Miral Amin (@aminmiral) — made YOLO loading reject negative extents and name malformed lines.
- NIKHIL (@Nikhi00718) — fixed
LineZonehistory expiry and anchor overflow. - kevin (@kevin9327) — made Qwen parsing recover from cut-off responses.
- A Aswanth Raj (@aswanth-07, LinkedIn) — fixed class-agnostic mAP.
- Devulapalli Naga Sri Vaishnavi (@Vaishnavi220506) — fixed masks for off-frame fractional boxes.
- JANG BYUNGKUN (@8rulerstar) — fixed YOLO pose label loading.
Automated contributions: @dependabot
Full changelog: 0.30.7...0.30.8