iterframes

Rust decoder for Python · v0.5.0

Your loop never waits
for the decoder.

Video frames reach your Python code as NumPy arrays, decoded by a Rust thread that has already moved on to the next ones. That thread never takes the GIL, so decoding overlaps your work instead of adding to it.

Documentation Source
pip install
FFmpeg comes in the wheel
CPython 3.11+
Linux, macOS, Windows
No GIL
held by the decoder thread

Twelve frames through the same loop

Without iterframes

Your loop decodes a frame, works on it, then decodes the next one. The two costs are paid one after the other.

Your loop
Frames out

With iterframes

The decoder thread runs ahead of you. By the time your loop asks for the next frame, that frame is already decoded and waiting in memory.

Your loop
Decoder thread
Frames out

Both runs do the same work on the same twelve frames. Only the waiting is different: one loop waits for the decoder, the other finds every frame already there.

Both timelines are drawn from the overlap benchmark in tests/test_benchmark.py, where per-frame work as expensive as decoding made the whole run 16% longer instead of twice as long. 1× is the time the same video takes to decode with nothing else happening.

NumPy arrays, no copy batches in one block random access by frame NVDEC and VideoToolbox FFmpeg in the wheel

The loop you already write

read returns an iterator. Nothing to open, nothing to close, no callback to register: the background thread starts with the first next and stops when you stop.

Every frame as an array

Each frame is a (height, width, 3) block of RGB uint8, read straight out of FFmpeg's buffer.

for frame in iterframes.read(video):
    model(frame)

Resized while it decodes

Ask for a size and FFmpeg scales the frame on its way out, instead of a cv2.resize in your loop.

iterframes.read(video,
                 height=224,
                 width=224)

Batches in one block

read_batches writes a whole batch into a single allocation, ready to hand to a model without stacking anything.

iterframes.read_batches(video, 12)
# (12, 224, 224, 3)

Stop whenever you like

Break out of the loop and the decoder thread notices on its next send and returns. Nothing is decoded that you never asked for.

for index, frame in enumerate(reader):
    if index == 100:
        break

The wheels carry FFmpeg with them. No system packages, no ffmpeg binary to call, nothing to match versions with.

Read three frames, not fourteen thousand

Give read a list of frame numbers, or a start, stop and step, and iterframes reads only what it has to. Here is what it does with frames=[934, 4522, 11711] on a ten minute video, exactly or approximately.

3 Frames requested
0 Frames decoded
14,316 Frames never touched

Nothing decoded yet.

Key frame Decoded, then thrown away Kept
  1. Index the file once. Every packet is demuxed and none is decoded, which is enough to learn where each frame sits and which frames are key frames.
  2. Jump back to a key frame. Frame 934 cannot be decoded on its own, so the decoder seeks to frame 900, the key frame in front of it.
  3. Decode forward and drop. The frames in between are decoded so the next one can be, then thrown away without ever being converted to RGB.
  4. Hand over three frames. 70 frames decoded, 14,246 never read at all.

iterframes.read("bunny.mp4", frames=[934, 4522, 11711])

Negative numbers count from the end, frame n is the one a plain read yields nth, and frames asked for in order cost no more than reading the video straight through.

approximate is the switch above: each number is read as the key frame nearest to it, as long as that key frame is within the frames given, so nothing in between is decoded. A frame further away than that is still read exactly, which keeps the error inside the number you chose.

Decode on the GPU

Pass device and the decoding leaves the CPU, which then stays free for your model. The device names are PyTorch's, and both decoders are already inside the wheel.

Three paths out of the same file. On the CPU, packets are decoded and scaled by swscale and land in host memory as a NumPy array. On NVIDIA, NVDEC decodes and resizes on the GPU, and the frame either reaches PyTorch through DLPack without leaving the GPU or is copied to the host as a NumPy array. On Apple silicon, VideoToolbox decodes and the frame arrives in host memory as a NumPy array. clip.mp4 H.264 packets device="cpu" CPU decoder Swscale resizes and converts Host memory RGB24, rows unpadded NumPy array No copy device="cuda" NVDEC Decodes and resizes on the GPU GPU memory NV12, Y and UV planes Torch tensor DLPack, stays on the GPU NumPy array Copied to the host device="mps" VideoToolbox Apple's media engine decodes Host memory RGB24 NumPy array No copy

NVIDIA, through NVDEC

Linux and Windows. The driver is loaded at run time, and the GPU resizes the frame as well when you ask for a size. With on_device=True the planes never leave the GPU: PyTorch takes them through DLPack.

for frame in iterframes.read(video, device="cuda",
                              on_device=True):
    y = torch.from_dlpack(frame.y)
    uv = torch.from_dlpack(frame.uv)

Apple silicon, through VideoToolbox

macOS on Apple silicon, on the media engine rather than the CPU cores. Frames arrive in host memory as ordinary NumPy arrays, exactly as they do on the CPU path.

for frame in iterframes.read(video, device="mps"):
    model(frame)

device="auto" takes whatever the machine has and falls back to the CPU, and iterframes.DEVICES says what that turned out to be. A GPU saves CPU time but is not always faster than the CPU decoder, so measure both.

The array is FFmpeg's buffer

RGB frames are allocated with an alignment of one, so their rows carry no padding and Python can read them as contiguous memory. The frame object implements the buffer protocol: NumPy wraps the same bytes FFmpeg wrote, and nothing is copied between the decoder and your model.

With read_batches, swscale writes every frame of a batch into its slice of one allocation, so a batch reaches you as a single four-dimensional array rather than a list to stack.

One allocation holds twelve frames back to back; the NumPy array is an outline drawn around the same bytes. One allocation, 12 × h × w × 3 bytes Numpy.asarray(batch) wraps the same bytes, shape (12, h, w, 3) Nothing is copied between the decoder and your model

Where iterframes fits

iterframesOpenCV VideoCapturedecordPyAV
Decoding runsAhead of your code, on a background threadWhen you call read()When you index the readerWhen you ask for the next frame
PixelsRGBBGRRGBAny format FFmpeg supports
Resize while decodingYesNo, cv2.resize afterYesYes, with reformat
Batches as one arrayYesNoYesNo
Random accessYes, by frame number, indexedYes, by frame numberYes, by frame number, indexedYes, by timestamp
Audio, encoding, muxingNoEncoding with VideoWriterAudio readingYes
Hardware decodingNVIDIA and Apple silicon, in the wheelsDepends on the buildNVIDIA, built from sourceDepends on the build
PlatformsLinux x86_64 and aarch64, macOS on Apple silicon, Windows x86_64Linux, macOS and Windows, x86_64 and arm64Linux, macOS on Intel and Windows, x86_64 onlyLinux, macOS and Windows, x86_64 and arm64

Reach for PyAV when you need the rest of FFmpeg: audio, encoding, streams, precise control over the decoder. OpenCV is the natural choice when the rest of the pipeline already uses it.

Start reading frames

Wheels for Linux (x86_64, aarch64), macOS on Apple silicon and Windows, for every CPython from 3.11 on. Other platforms build from source, FFmpeg included.