Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 8 additions & 5 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -147,11 +147,14 @@ locally.

## Docs

Three pages: `index.md` (overview), `reference.md` (API and errors),
`development.md`. Every example must match the behavior of a freshly
built module; check them instead of writing them from memory. Keep the
text short, in American English, with no performance claims that have not
been measured.
Four pages: `index.md` (overview and quick start), `guides.md` (what the
arguments are for, and what they cost), `api.md` (signatures, arguments,
and errors), `development.md`. `scripts/docs_redirects.py` sends links to
the old `reference/` page, anchors included, to the page that now holds
each section; add to it when a section moves. Every example must match
the behavior of a freshly built module; check them instead of writing
them from memory. Keep the text short, in American English, with no
performance claims that have not been measured.

The landing page in `web/` is part of the docs: whenever an argument, a
default, or the behavior of a feature changes, update it in the same
Expand Down
12 changes: 6 additions & 6 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,8 +24,8 @@ The landing page at
[alesanfra.github.io/iterframes](https://alesanfra.github.io/iterframes/)
shows what the decoder does while your loop runs. The documentation is at
[iterframes.readthedocs.io](https://iterframes.readthedocs.io): the
[reference](https://iterframes.readthedocs.io/en/latest/reference/)
documents every argument and error, and the
[guides](https://iterframes.readthedocs.io/en/latest/guides/) explain the features, the
[API](https://iterframes.readthedocs.io/en/latest/api/) documents every argument and error, and the
[development guide](https://iterframes.readthedocs.io/en/latest/development/)
covers building from source.

Expand Down Expand Up @@ -90,7 +90,7 @@ frame `n` is the one `read` yields `n`-th. iterframes indexes the file
once, without decoding it, then decodes each frame from the key frame
before it; frames asked for in order cost no more than reading the video
straight through. See
[Reading frames by number](https://iterframes.readthedocs.io/en/latest/reference/#reading-frames-by-number).
[Reading frames by number](https://iterframes.readthedocs.io/en/latest/guides/#reading-frames-by-number).

When a frame nearby will do, `approximate` reads the key frame closest to
each of `frames`, which costs one decoded frame instead of the frames from
Expand All @@ -101,7 +101,7 @@ the key frame on:
frames = list(iterframes.read("video.mp4", frames=[100, 200, 300], approximate=5))
```

See [Approximate frames](https://iterframes.readthedocs.io/en/latest/reference/#approximate-frames).
See [Approximate frames](https://iterframes.readthedocs.io/en/latest/guides/#approximate-frames).

## Hardware decoding

Expand All @@ -126,7 +126,7 @@ print(iterframes.DEVICES) # ('cpu', 'mps') on a Mac

The frames still arrive as NumPy arrays in memory. A GPU saves CPU time
but is not always faster than the CPU decoder, so measure both; see
[Hardware decoding](https://iterframes.readthedocs.io/en/latest/reference/#hardware-decoding).
[Hardware decoding](https://iterframes.readthedocs.io/en/latest/guides/#hardware-decoding).

With an NVIDIA GPU, `on_device=True` keeps the frames on it, in NV12, for
PyTorch and other libraries to take without a copy:
Expand All @@ -139,7 +139,7 @@ for frame in iterframes.read("video.mp4", device="cuda", on_device=True):
uv = torch.from_dlpack(frame.uv) # (height / 2, width / 2, 2)
```

[Frames on the GPU](https://iterframes.readthedocs.io/en/latest/reference/#frames-on-the-gpu) shows how to
[Frames on the GPU](https://iterframes.readthedocs.io/en/latest/guides/#frames-on-the-gpu) shows how to
convert them to RGB there.

## Compared with OpenCV, decord, and PyAV
Expand Down
137 changes: 137 additions & 0 deletions docs/api.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,137 @@
# API

Everything lives in the top-level `iterframes` module.

## read

```python
read(path, height=None, width=None, prefetch_frames=1, device="cpu", on_device=False, frames=None, start=0, stop=None, step=1, approximate=None) -> Iterator[numpy.ndarray]
```

Yields the frames of the video at `path`, in order, or with `frames`, or
`start`, `stop`, and `step`, only those frames, in the order asked for.
Each frame is a C-contiguous, writable `numpy.ndarray` of shape
`(height, width, 3)` and dtype `uint8`, holding RGB pixels.

| Argument | Description |
| --- | --- |
| `path` | Path of the video, as a `str` or `os.PathLike` |
| `height` | Height of the frames. Defaults to the height of the video |
| `width` | Width of the frames. Defaults to the width of the video |
| `prefetch_frames` | How many decoded frames may wait for your code. Defaults to 1 |
| `device` | Where to decode: `"cpu"`, `"auto"`, or a name from `DEVICES`. Defaults to `"cpu"`. See [Hardware decoding](guides.md#hardware-decoding) |
| `on_device` | With `device="cuda"`, yield `CudaFrame` objects left on the GPU instead of arrays. See [Frames on the GPU](guides.md#frames-on-the-gpu) |
| `frames` | The numbers of the frames to read, in the order to read them. See [Reading frames by number](guides.md#reading-frames-by-number) |
| `start`, `stop`, `step` | The frames to read, as a slice of the video |
| `approximate` | With `frames`, read the key frame nearest to each of them instead, when it is within this many frames. `True` is any distance. See [Approximate frames](guides.md#approximate-frames) |

Frames are resized with bilinear interpolation. When only one of `height`
and `width` is given, the other keeps the size of the video, so the aspect
ratio changes.

Decoding starts when the first frame is requested, on a background thread
that runs ahead of your code by up to `prefetch_frames` frames. The thread
never takes the GIL, so it decodes the next frames while your code
processes the current one, even when that code holds the GIL. A larger
`prefetch_frames` smooths out frames that take longer to decode, at the
cost of memory: one 1080p frame takes about 6 MB.

The decoder stops when the iterator is exhausted or garbage-collected, for
example after a `break`.

```python
import iterframes

for frame in iterframes.read("video.mp4", height=270, width=480):
print(frame.shape) # (270, 480, 3)
```

## read_batches

```python
read_batches(path, batch_size, height=None, width=None, prefetch_frames=1, device="cpu", drop_last=False, frames=None, start=0, stop=None, step=1, approximate=None) -> Iterator[numpy.ndarray]
```

Yields the frames of the video in batches, in order, for models that take
several frames at once. Each batch is a C-contiguous, writable
`numpy.ndarray` of shape `(batch_size, height, width, 3)` and dtype
`uint8`. The frames are decoded straight into it, so unlike `numpy.stack`
over the frames of `read`, making a batch costs your code no copy.

| Argument | Description |
| --- | --- |
| `batch_size` | Frames per batch, at least 1 |
| `prefetch_frames` | How many decoded frames may wait for your code, rounded up to whole batches. Defaults to 1, that is one batch |
| `drop_last` | Drop the last batch when the video ends before it is full. By default it is yielded with the frames left over |

The other arguments are those of [`read`](#read); `on_device` is not
supported. Every frame of a batch has the size of the first frame of the
video, or `height` and `width`.

```python
for batch in iterframes.read_batches("video.mp4", 12, height=224, width=224):
print(batch.shape) # (12, 224, 224, 3), except maybe the last one
```

## Errors

Errors are raised by the first `next()` on the iterator, not by the call
to `read`, because decoding starts only then.

| Exception | When |
| --- | --- |
| `FileNotFoundError`, `PermissionError`, `OSError` | The file cannot be opened. The exception carries `errno` and `filename` |
| `ValueError` | The file is not a video FFmpeg can read, or has no video stream |
| `RuntimeError` | Decoding fails after the video has been opened, for instance on a codec that the build does not include |
| `IndexError` | `frames` asks for a frame the video does not have |

```python
try:
frames = list(iterframes.read("missing.mp4"))
except FileNotFoundError as error:
print(error.filename) # missing.mp4
```

## FrameReader

```python
FrameReader(path, height=None, width=None, prefetch_frames=1, device="cpu", on_device=False, batch_size=None, drop_last=False, frames=None, start=0, stop=None, step=1, approximate=None)
```

The iterator behind `read` and `read_batches`. It yields `Frame` objects,
`CudaFrame` objects with `on_device=True`, or `Batch` objects with
`batch_size`.

## Frame

A decoded frame. It holds the pixels and exposes them through the buffer
protocol as a writable, C-contiguous `(height, width, 3)` block of
unsigned bytes, so that other libraries can read them without a copy:

```python
import numpy as np
from iterframes import FrameReader

for frame in FrameReader("video.mp4"):
array = np.asarray(frame) # what read() yields
view = memoryview(frame) # no NumPy needed
```

The pixels stay alive as long as the frame or any array or view on it.
See [Frames without a copy](guides.md#frames-without-a-copy).

## Batch

Frames decoded into a single block of memory, which the buffer protocol
exposes as a writable, C-contiguous `(frames, height, width, 3)` block of
unsigned bytes. `read_batches` wraps it with `numpy.asarray`, which copies
nothing; its pixels stay alive as long as the batch or any array or view on
it.

## Constants

| Name | Value |
| --- | --- |
| `__version__` | Version of iterframes, such as `"0.4.0"` |
| `FFMPEG_VERSION` | Version of the FFmpeg that iterframes is linked to, such as `"9.0.2"` |
| `DEVICES` | Names accepted by `device` besides `"auto"`, such as `("cpu", "mps")` |
52 changes: 38 additions & 14 deletions docs/development.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,21 +10,39 @@ sudo apt install build-essential curl python3-venv \
pkg-config libclang-dev nasm # Debian, Ubuntu
```

On Windows, FFmpeg is built with MSVC from an
[MSYS2](https://www.msys2.org/) shell, as CI does in
`.github/workflows/ci.yaml`. Install Visual Studio's C++ build tools and
LLVM (for libclang), then, in MSYS2:
macOS and Linux need nothing else: the build fetches nasm, meson, and
ninja when they are missing.

```console
pacman -S make diffutils curl tar xz \
mingw-w64-ucrt-x86_64-pkgconf mingw-w64-ucrt-x86_64-nasm
rm /usr/bin/link.exe # it shadows MSVC's link.exe
uv tool install meson && uv tool install ninja
```
### Windows

The wheels target MSVC, so FFmpeg is compiled with MSVC too. Its build
script is a shell script, so it runs in an [MSYS2](https://www.msys2.org/)
shell, as in the `windows` job of `.github/workflows/ci.yaml`.

1. Install Visual Studio's C++ build tools, [LLVM](https://llvm.org/)
(for libclang), MSYS2, rustup, and uv.
2. Open the *x64 Native Tools Command Prompt* and start MSYS2's UCRT64
shell from it, keeping its `PATH` so that `cl`, `cargo`, and `uv` stay
on it:

Start the MSYS2 UCRT64 shell from a Visual Studio developer prompt with
`msys2_shell.cmd -ucrt64 -use-full-path`, so that `cl`, `cargo`, and `uv`
stay on `PATH`, and run every command below from it.
```console
C:\msys64\msys2_shell.cmd -ucrt64 -use-full-path
```

3. In that shell, install the build tools. `pacman` is MSYS2's package
manager; meson and ninja come from uv instead, because they must be
Windows programs to drive MSVC:

```console
pacman -S make diffutils curl tar xz \
mingw-w64-ucrt-x86_64-pkgconf mingw-w64-ucrt-x86_64-nasm
mv /usr/bin/link.exe /usr/bin/link-msys2.exe # hides MSVC's link.exe
uv tool install meson
uv tool install ninja
export PATH="$PATH:$(cygpath -u "$(uv tool dir --bin)")"
```

Run every command below from that shell.

## FFmpeg

Expand Down Expand Up @@ -63,10 +81,13 @@ environment. Run it again after every change to a `.rs` file.
| `src/lib.rs` | Python module: `Frame`, `Batch`, `FrameReader`, and the error mapping |
| `src/decoder.rs` | Decoding thread: demux, decode, convert to RGB |
| `src/ffmpeg.rs` | Safe wrappers over the FFmpeg calls the crate needs |
| `src/dlpack.rs` | DLPack capsules for the planes of `CudaFrame` |
| `iterframes/__init__.py` | `read` and `read_batches`, which wrap `FrameReader` |
| `scripts/build-ffmpeg.sh` | Static FFmpeg build for the wheels |
| `build.rs` | Builds and links FFmpeg, generates its bindings |
| `scripts/build-ffmpeg.sh` | Static FFmpeg build, run by `build.rs` |
| `tests/` | pytest suite, which checks the frames against PyAV |
| `docs/` | This site |
| `web/` | Landing page, published on GitHub Pages |

## Tests

Expand Down Expand Up @@ -106,6 +127,9 @@ uv run --no-sync mkdocs serve
Read the Docs builds the site with `uv sync` from `uv.lock`, without
compiling the extension.

The landing page in `web/` has no build step: open `web/index.html`, or
run `python3 -m http.server --directory web`.

## Wheels

The wheels use the stable ABI of CPython (abi3), so one wheel per platform
Expand Down
Loading