github meta-pytorch/torchcodec v0.17.0
TorchCodec 0.17 - Low-level decoding APIs

2 hours ago

TorchCodec 0.17 is out! It is compatible with torch >= 2.11 and Python >= 3.10. This is one of our biggest releases so far: we're releasing a new set of low-level decoding APIs that expose the decoding stages separately (demuxing, decoding, and conversion). This enables a lot of new capabilities that the existing VideoDecoder and AudioDecoder couldn't do:

  • Multi-threaded pipelines: unlock performance gains by running demuxing, decoding and color-conversion concurrently on several threads, and choose where to split them. Each stage releases the GIL. Tutorial
  • Access Raw YUV data, for SDR and HDR sources: you can skip the conversion stage and use the decoder's own YUV planes directly, at the source's own precision, with no conversion and no copy. A 10-bit HDR video gives uint16 planes with all 10 bits intact. Tutorial
  • Custom transformations of YUV data: apply a custom colorspace conversion, train a model directly in YUV space, or write a kernel fused with the first layer of your model! Tutorial
  • Multi-stream decoding: follow several streams of a container at once: multiple video streams, multiple audio streams, or a mix of both, in a single pass over the input. Tutorial
  • Endless streams: decode a source with no duration, no frame count, and no end. Tutorial
  • Retrieving key frames: a scan gives you every key frame position, and reaching a key frame costs exactly one decoded frame. That's ideal for thumbnails, coarse previews, or a cheap sampler for training. Tutorial

VideoDecoder and AudioDecoder are unchanged, and they remain the easiest way to decode video and audio.

The low-level APIs are Beta: their signatures and semantics may change slightly based on user feedback.

Low-level decoding APIs

These APIs expose the three stages of a decoding pipeline in different objects, for video and audio:

Demuxer  ->  VideoPacketDecoder  ->  ColorConverter
 Packet         RawFrame               RGB Frame

Demuxer  ->  AudioPacketDecoder  ->  AudioConverter
 Packet        RawAudioSamples          AudioSamples

Below is an example of multi-stream decoding: decode audio and video in a single pass, with notes on where the other features plug in. Refer to the tutorials linked below for more details!

from torchcodec.decoders import AudioConverter, ColorConverter, Demuxer

# Also works with file-like objects, and with endless sources like pipes.
demuxer = Demuxer("video.mp4", streams=("video", "audio"))
video_stream, audio_stream = demuxer.streams
# Optional: video_stream.scan() returns a FrameIndex with exact timestamps and key frames.

decoders = {
    video_stream.index: video_stream.make_decoder(device="cuda"),
    audio_stream.index: audio_stream.make_decoder(),
}
color_converter = ColorConverter(device="cuda")
audio_converter = AudioConverter(sample_rate=16_000, num_channels=1)

for packet in demuxer:  # The demuxer could run on its own thread.
    for raw in decoders[packet.stream_index].decode(packet):
        if packet.stream_index == video_stream.index:
            # Or skip the ColorConverter and use the raw YUV planes directly: raw.planes
            frame = color_converter.convert(raw)
        else:
            # Or use the source's own samples directly: raw.data
            samples = audio_converter.convert(raw)

# Once the demuxer is exhausted, call drain() on each decoder and on the audio converter.

Refer to our tutorials to learn more about the different capabilities of these
APIs:

The full API is in the API reference.

Improvements

  • Wider native NVDEC coverage: HEVC 4:4:4 (8, 10 and 12-bit) and 10-bit AV1 now decode natively on the GPU instead of falling back to the CPU (#1635).
  • Free-threaded Python: importing TorchCodec no longer re-enables the GIL on free-threaded CPython builds (#1740).
  • Clearer error message when trying to seek in unseekable formats (#1647).

Bug Fixes

  • swscale over-read: fixed a potential over-read with libswscale when resizing happens (#1750).
  • HEIC: fixed a potential out-of-bounds read in decode_heic (#1751).
  • >10-bit video on CPU: the input color range is now passed explicitly to libswscale, fixing the colors of high-bit-depth videos (#1708).
  • Video encoder color range: the RGB-to-YUV conversion now respects the output color range and color space, so full-range encodes are no longer converted as limited range (#1707).
  • PNG: fixed alpha stripping for palette PNGs with tRNS transparency (#1732).
  • CUDA decoding:
    • Fixed VP9 decoding by working around incorrect pts reported by NVCUVID (#1743).
    • Fixed an NVDEC cache collision bug (#1706).
  • CUDA JPEG decoding: nvJPEG hardware decoding now waits for the caller's stream before writing its output, working around an existing nvJPEG bug experienced on A100 (#1634).
  • macOS wheels: restored the Homebrew rpath, so a Homebrew-installed FFmpeg is found at runtime again (#1632).
  • File-like objects: the FFmpeg protocol allow-list now matches the one used for local files, so nested protocols are forbidden for file-like inputs too (#1738).
  • AVIF: added the missing num_threads parameter to decode_avif in builds without AVIF support (#1671).

Don't miss a new torchcodec release

NewReleases is sending notifications on new releases.