TorchCodec 0.17 is out! It is compatible with torch >= 2.11 and Python >= 3.10. This is one of our biggest releases so far: we're releasing a new set of low-level decoding APIs that expose the decoding stages separately (demuxing, decoding, and conversion). This enables a lot of new capabilities that the existing VideoDecoder and AudioDecoder couldn't do:
- Multi-threaded pipelines: unlock performance gains by running demuxing, decoding and color-conversion concurrently on several threads, and choose where to split them. Each stage releases the GIL. Tutorial
- Access Raw YUV data, for SDR and HDR sources: you can skip the conversion stage and use the decoder's own YUV planes directly, at the source's own precision, with no conversion and no copy. A 10-bit HDR video gives
uint16planes with all 10 bits intact. Tutorial - Custom transformations of YUV data: apply a custom colorspace conversion, train a model directly in YUV space, or write a kernel fused with the first layer of your model! Tutorial
- Multi-stream decoding: follow several streams of a container at once: multiple video streams, multiple audio streams, or a mix of both, in a single pass over the input. Tutorial
- Endless streams: decode a source with no duration, no frame count, and no end. Tutorial
- Retrieving key frames: a scan gives you every key frame position, and reaching a key frame costs exactly one decoded frame. That's ideal for thumbnails, coarse previews, or a cheap sampler for training. Tutorial
VideoDecoder and AudioDecoder are unchanged, and they remain the easiest way to decode video and audio.
The low-level APIs are Beta: their signatures and semantics may change slightly based on user feedback.
Low-level decoding APIs
These APIs expose the three stages of a decoding pipeline in different objects, for video and audio:
Demuxer -> VideoPacketDecoder -> ColorConverter
Packet RawFrame RGB Frame
Demuxer -> AudioPacketDecoder -> AudioConverter
Packet RawAudioSamples AudioSamples
Below is an example of multi-stream decoding: decode audio and video in a single pass, with notes on where the other features plug in. Refer to the tutorials linked below for more details!
from torchcodec.decoders import AudioConverter, ColorConverter, Demuxer
# Also works with file-like objects, and with endless sources like pipes.
demuxer = Demuxer("video.mp4", streams=("video", "audio"))
video_stream, audio_stream = demuxer.streams
# Optional: video_stream.scan() returns a FrameIndex with exact timestamps and key frames.
decoders = {
video_stream.index: video_stream.make_decoder(device="cuda"),
audio_stream.index: audio_stream.make_decoder(),
}
color_converter = ColorConverter(device="cuda")
audio_converter = AudioConverter(sample_rate=16_000, num_channels=1)
for packet in demuxer: # The demuxer could run on its own thread.
for raw in decoders[packet.stream_index].decode(packet):
if packet.stream_index == video_stream.index:
# Or skip the ColorConverter and use the raw YUV planes directly: raw.planes
frame = color_converter.convert(raw)
else:
# Or use the source's own samples directly: raw.data
samples = audio_converter.convert(raw)
# Once the demuxer is exhausted, call drain() on each decoder and on the audio converter.Refer to our tutorials to learn more about the different capabilities of these
APIs:
- Build your own decoding pipeline: the three stages, multi-stream decoding, seeking, scanning, key frames, metadata, and endless streams.
- Multi-threaded decoding pipelines: where to split a pipeline across threads, on CPU and CUDA.
- Raw frames and raw audio samples: YUV planes, HDR, custom color conversion, and raw audio samples.
- CUDA streams: what you need to take care of when you consume frames on a different CUDA stream than the one they were decoded on.
The full API is in the API reference.
Improvements
- Wider native NVDEC coverage: HEVC 4:4:4 (8, 10 and 12-bit) and 10-bit AV1 now decode natively on the GPU instead of falling back to the CPU (#1635).
- Free-threaded Python: importing TorchCodec no longer re-enables the GIL on free-threaded CPython builds (#1740).
- Clearer error message when trying to seek in unseekable formats (#1647).
Bug Fixes
- swscale over-read: fixed a potential over-read with libswscale when resizing happens (#1750).
- HEIC: fixed a potential out-of-bounds read in
decode_heic(#1751). - >10-bit video on CPU: the input color range is now passed explicitly to libswscale, fixing the colors of high-bit-depth videos (#1708).
- Video encoder color range: the RGB-to-YUV conversion now respects the output color range and color space, so full-range encodes are no longer converted as limited range (#1707).
- PNG: fixed alpha stripping for palette PNGs with tRNS transparency (#1732).
- CUDA decoding:
- CUDA JPEG decoding: nvJPEG hardware decoding now waits for the caller's stream before writing its output, working around an existing nvJPEG bug experienced on A100 (#1634).
- macOS wheels: restored the Homebrew rpath, so a Homebrew-installed FFmpeg is found at runtime again (#1632).
- File-like objects: the FFmpeg protocol allow-list now matches the one used for local files, so nested protocols are forbidden for file-like inputs too (#1738).
- AVIF: added the missing
num_threadsparameter todecode_avifin builds without AVIF support (#1671).