Every frame as an array
Each frame is a (height, width, 3) block of RGB
uint8, read straight out of FFmpeg's buffer.
for frame in iterframes.read(video):
model(frame)
Rust decoder for Python · v0.5.0
Video frames reach your Python code as NumPy arrays, decoded by a Rust thread that has already moved on to the next ones. That thread never takes the GIL, so decoding overlaps your work instead of adding to it.
Your loop decodes a frame, works on it, then decodes the next one. The two costs are paid one after the other.
The decoder thread runs ahead of you. By the time your loop asks for the next frame, that frame is already decoded and waiting in memory.
Both runs do the same work on the same twelve frames. Only the waiting is different: one loop waits for the decoder, the other finds every frame already there.
Both timelines are drawn from the overlap benchmark in tests/test_benchmark.py, where per-frame work as expensive as decoding made the whole run 16% longer instead of twice as long. 1× is the time the same video takes to decode with nothing else happening.
read returns an iterator. Nothing to open, nothing to close,
no callback to register: the background thread starts with the first
next and stops when you stop.
Each frame is a (height, width, 3) block of RGB
uint8, read straight out of FFmpeg's buffer.
for frame in iterframes.read(video):
model(frame)
Ask for a size and FFmpeg scales the frame on its way out, instead of a
cv2.resize in your loop.
iterframes.read(video,
height=224,
width=224)
read_batches writes a whole batch into a single allocation,
ready to hand to a model without stacking anything.
iterframes.read_batches(video, 12)
# (12, 224, 224, 3)
Break out of the loop and the decoder thread notices on its next send and returns. Nothing is decoded that you never asked for.
for index, frame in enumerate(reader):
if index == 100:
break
The wheels carry FFmpeg with them. No system packages, no
ffmpeg binary to call, nothing to match versions with.
Give read a list of frame numbers, or a start,
stop and step, and iterframes reads only what it
has to. Here is what it does with
frames=[934, 4522, 11711] on a ten minute video,
exactly or approximately.
Nothing decoded yet.
iterframes.read("bunny.mp4", frames=[934, 4522, 11711])
Negative numbers count from the end, frame n is the one a plain
read yields nth, and frames asked for in order cost
no more than reading the video straight through.
approximate is the switch above: each number is read as the key
frame nearest to it, as long as that key frame is within the frames given,
so nothing in between is decoded. A frame further away than that is still
read exactly, which keeps the error inside the number you chose.
Pass device and the decoding leaves the CPU, which then stays
free for your model. The device names are PyTorch's, and both decoders are
already inside the wheel.
Linux and Windows. The driver is loaded at run time, and the GPU resizes
the frame as well when you ask for a size. With
on_device=True the planes never leave the GPU: PyTorch takes
them through DLPack.
for frame in iterframes.read(video, device="cuda",
on_device=True):
y = torch.from_dlpack(frame.y)
uv = torch.from_dlpack(frame.uv)
macOS on Apple silicon, on the media engine rather than the CPU cores. Frames arrive in host memory as ordinary NumPy arrays, exactly as they do on the CPU path.
for frame in iterframes.read(video, device="mps"):
model(frame)
device="auto" takes whatever the machine has and falls back to
the CPU, and iterframes.DEVICES says what that turned out to be.
A GPU saves CPU time but is not always faster than the CPU decoder, so
measure both.
RGB frames are allocated with an alignment of one, so their rows carry no padding and Python can read them as contiguous memory. The frame object implements the buffer protocol: NumPy wraps the same bytes FFmpeg wrote, and nothing is copied between the decoder and your model.
With read_batches, swscale writes every frame of a batch into
its slice of one allocation, so a batch reaches you as a single
four-dimensional array rather than a list to stack.
| iterframes | OpenCV VideoCapture | decord | PyAV | |
|---|---|---|---|---|
| Decoding runs | Ahead of your code, on a background thread | When you call read() | When you index the reader | When you ask for the next frame |
| Pixels | RGB | BGR | RGB | Any format FFmpeg supports |
| Resize while decoding | Yes | No, cv2.resize after | Yes | Yes, with reformat |
| Batches as one array | Yes | No | Yes | No |
| Random access | Yes, by frame number, indexed | Yes, by frame number | Yes, by frame number, indexed | Yes, by timestamp |
| Audio, encoding, muxing | No | Encoding with VideoWriter | Audio reading | Yes |
| Hardware decoding | NVIDIA and Apple silicon, in the wheels | Depends on the build | NVIDIA, built from source | Depends on the build |
| Platforms | Linux x86_64 and aarch64, macOS on Apple silicon, Windows x86_64 | Linux, macOS and Windows, x86_64 and arm64 | Linux, macOS on Intel and Windows, x86_64 only | Linux, macOS and Windows, x86_64 and arm64 |
Reach for PyAV when you need the rest of FFmpeg: audio, encoding, streams, precise control over the decoder. OpenCV is the natural choice when the rest of the pipeline already uses it.
Wheels for Linux (x86_64, aarch64), macOS on Apple silicon and Windows, for every CPython from 3.11 on. Other platforms build from source, FFmpeg included.