From 0d344a9575134fc83514b45094f94e070bca36c2 Mon Sep 17 00:00:00 2001 From: Alessio Sanfratello Date: Mon, 28 Sep 2026 11:25:41 +0200 Subject: [PATCH 1/3] docs: split the reference into Guides and API, and fix the Windows setup The reference mixed the API with explanations of when to use each argument. It is now two pages: guides.md (frames by number, approximate frames, threads, hardware decoding, frames on the GPU) and api.md (signatures, arguments, errors, types, constants). A MkDocs hook writes reference/index.html, which sends each old anchor to the page that now holds it, so the links in the README of released versions keep working. The quick start gains random access in any order, device="auto", reading videos in parallel, and when errors are raised. The Windows setup in development.md explains that pacman is MSYS2's, adds uv's tool directory to PATH for meson and ninja, and renames MSYS2's link.exe instead of deleting it. The layout table lists build.rs, src/dlpack.rs, and web/. Co-Authored-By: Claude Opus 5.5 --- AGENTS.md | 7 +- README.md | 12 +-- docs/api.md | 136 +++++++++++++++++++++++++++++ docs/development.md | 52 +++++++++--- docs/{reference.md => guides.md} | 141 +------------------------------ docs/index.md | 39 +++++++-- mkdocs.yml | 6 +- scripts/docs_redirects.py | 41 +++++++++ 8 files changed, 267 insertions(+), 167 deletions(-) create mode 100644 docs/api.md rename docs/{reference.md => guides.md} (54%) create mode 100644 scripts/docs_redirects.py diff --git a/AGENTS.md b/AGENTS.md index cb4335e..44c6575 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -147,8 +147,11 @@ locally. ## Docs -Three pages: `index.md` (overview), `reference.md` (API and errors), -`development.md`. Every example must match the behavior of a freshly +Four pages: `index.md` (overview and quick start), `guides.md` (what the +arguments are for, and what they cost), `api.md` (signatures, arguments, +and errors), `development.md`. `scripts/docs_redirects.py` sends links to +the old `reference/` page, anchors included, to the page that now holds +each section; add to it when a section moves. Every example must match the behavior of a freshly built module; check them instead of writing them from memory. Keep the text short, in American English, with no performance claims that have not been measured. diff --git a/README.md b/README.md index a42de69..2ffb92a 100644 --- a/README.md +++ b/README.md @@ -24,8 +24,8 @@ The landing page at [alesanfra.github.io/iterframes](https://alesanfra.github.io/iterframes/) shows what the decoder does while your loop runs. The documentation is at [iterframes.readthedocs.io](https://iterframes.readthedocs.io): the -[reference](https://iterframes.readthedocs.io/en/latest/reference/) -documents every argument and error, and the +[guides](https://iterframes.readthedocs.io/en/latest/guides/) explain the features, the +[API](https://iterframes.readthedocs.io/en/latest/api/) documents every argument and error, and the [development guide](https://iterframes.readthedocs.io/en/latest/development/) covers building from source. @@ -90,7 +90,7 @@ frame `n` is the one `read` yields `n`-th. iterframes indexes the file once, without decoding it, then decodes each frame from the key frame before it; frames asked for in order cost no more than reading the video straight through. See -[Reading frames by number](https://iterframes.readthedocs.io/en/latest/reference/#reading-frames-by-number). +[Reading frames by number](https://iterframes.readthedocs.io/en/latest/guides/#reading-frames-by-number). When a frame nearby will do, `approximate` reads the key frame closest to each of `frames`, which costs one decoded frame instead of the frames from @@ -101,7 +101,7 @@ the key frame on: frames = list(iterframes.read("video.mp4", frames=[100, 200, 300], approximate=5)) ``` -See [Approximate frames](https://iterframes.readthedocs.io/en/latest/reference/#approximate-frames). +See [Approximate frames](https://iterframes.readthedocs.io/en/latest/guides/#approximate-frames). ## Hardware decoding @@ -126,7 +126,7 @@ print(iterframes.DEVICES) # ('cpu', 'mps') on a Mac The frames still arrive as NumPy arrays in memory. A GPU saves CPU time but is not always faster than the CPU decoder, so measure both; see -[Hardware decoding](https://iterframes.readthedocs.io/en/latest/reference/#hardware-decoding). +[Hardware decoding](https://iterframes.readthedocs.io/en/latest/guides/#hardware-decoding). With an NVIDIA GPU, `on_device=True` keeps the frames on it, in NV12, for PyTorch and other libraries to take without a copy: @@ -139,7 +139,7 @@ for frame in iterframes.read("video.mp4", device="cuda", on_device=True): uv = torch.from_dlpack(frame.uv) # (height / 2, width / 2, 2) ``` -[Frames on the GPU](https://iterframes.readthedocs.io/en/latest/reference/#frames-on-the-gpu) shows how to +[Frames on the GPU](https://iterframes.readthedocs.io/en/latest/guides/#frames-on-the-gpu) shows how to convert them to RGB there. ## Compared with OpenCV, decord, and PyAV diff --git a/docs/api.md b/docs/api.md new file mode 100644 index 0000000..90d30a5 --- /dev/null +++ b/docs/api.md @@ -0,0 +1,136 @@ +# API + +Everything lives in the top-level `iterframes` module. + +## read + +```python +read(path, height=None, width=None, prefetch_frames=1, device="cpu", on_device=False, frames=None, start=0, stop=None, step=1, approximate=None) -> Iterator[numpy.ndarray] +``` + +Yields the frames of the video at `path`, in order, or with `frames`, or +`start`, `stop`, and `step`, only those frames, in the order asked for. +Each frame is a C-contiguous, writable `numpy.ndarray` of shape +`(height, width, 3)` and dtype `uint8`, holding RGB pixels. + +| Argument | Description | +| --- | --- | +| `path` | Path of the video, as a `str` or `os.PathLike` | +| `height` | Height of the frames. Defaults to the height of the video | +| `width` | Width of the frames. Defaults to the width of the video | +| `prefetch_frames` | How many decoded frames may wait for your code. Defaults to 1 | +| `device` | Where to decode: `"cpu"`, `"auto"`, or a name from `DEVICES`. Defaults to `"cpu"`. See [Hardware decoding](guides.md#hardware-decoding) | +| `on_device` | With `device="cuda"`, yield `CudaFrame` objects left on the GPU instead of arrays. See [Frames on the GPU](guides.md#frames-on-the-gpu) | +| `frames` | The numbers of the frames to read, in the order to read them. See [Reading frames by number](guides.md#reading-frames-by-number) | +| `start`, `stop`, `step` | The frames to read, as a slice of the video | +| `approximate` | With `frames`, read the key frame nearest to each of them instead, when it is within this many frames. `True` is any distance. See [Approximate frames](guides.md#approximate-frames) | + +Frames are resized with bilinear interpolation. When only one of `height` +and `width` is given, the other keeps the size of the video, so the aspect +ratio changes. + +Decoding starts when the first frame is requested, on a background thread +that runs ahead of your code by up to `prefetch_frames` frames. The thread +never takes the GIL, so it decodes the next frames while your code +processes the current one, even when that code holds the GIL. A larger +`prefetch_frames` smooths out frames that take longer to decode, at the +cost of memory: one 1080p frame takes about 6 MB. + +The decoder stops when the iterator is exhausted or garbage-collected, for +example after a `break`. + +```python +import iterframes + +for frame in iterframes.read("video.mp4", height=270, width=480): + print(frame.shape) # (270, 480, 3) +``` + +## read_batches + +```python +read_batches(path, batch_size, height=None, width=None, prefetch_frames=1, device="cpu", drop_last=False, frames=None, start=0, stop=None, step=1, approximate=None) -> Iterator[numpy.ndarray] +``` + +Yields the frames of the video in batches, in order, for models that take +several frames at once. Each batch is a C-contiguous, writable +`numpy.ndarray` of shape `(batch_size, height, width, 3)` and dtype +`uint8`. The frames are decoded straight into it, so unlike `numpy.stack` +over the frames of `read`, making a batch costs your code no copy. + +| Argument | Description | +| --- | --- | +| `batch_size` | Frames per batch, at least 1 | +| `prefetch_frames` | How many decoded frames may wait for your code, rounded up to whole batches. Defaults to 1, that is one batch | +| `drop_last` | Drop the last batch when the video ends before it is full. By default it is yielded with the frames left over | + +The other arguments are those of [`read`](#read); `on_device` is not +supported. Every frame of a batch has the size of the first frame of the +video, or `height` and `width`. + +```python +for batch in iterframes.read_batches("video.mp4", 12, height=224, width=224): + print(batch.shape) # (12, 224, 224, 3), except maybe the last one +``` + +## Errors + +Errors are raised by the first `next()` on the iterator, not by the call +to `read`, because decoding starts only then. + +| Exception | When | +| --- | --- | +| `FileNotFoundError`, `PermissionError`, `OSError` | The file cannot be opened. The exception carries `errno` and `filename` | +| `ValueError` | The file is not a video FFmpeg can read, or has no video stream | +| `RuntimeError` | Decoding fails after the video has been opened, for instance on a codec that the build does not include | +| `IndexError` | `frames` asks for a frame the video does not have | + +```python +try: + frames = list(iterframes.read("missing.mp4")) +except FileNotFoundError as error: + print(error.filename) # missing.mp4 +``` + +## FrameReader + +```python +FrameReader(path, height=None, width=None, prefetch_frames=1, device="cpu", on_device=False, batch_size=None, drop_last=False, frames=None, start=0, stop=None, step=1, approximate=None) +``` + +The iterator behind `read` and `read_batches`. It yields `Frame` objects, +`CudaFrame` objects with `on_device=True`, or `Batch` objects with +`batch_size`. + +## Frame + +A decoded frame. It holds the pixels and exposes them through the buffer +protocol as a writable, C-contiguous `(height, width, 3)` block of +unsigned bytes, so that other libraries can read them without a copy: + +```python +import numpy as np +from iterframes import FrameReader + +for frame in FrameReader("video.mp4"): + array = np.asarray(frame) # what read() yields + view = memoryview(frame) # no NumPy needed +``` + +The pixels stay alive as long as the frame or any array or view on it. + +## Batch + +Frames decoded into a single block of memory, which the buffer protocol +exposes as a writable, C-contiguous `(frames, height, width, 3)` block of +unsigned bytes. `read_batches` wraps it with `numpy.asarray`, which copies +nothing; its pixels stay alive as long as the batch or any array or view on +it. + +## Constants + +| Name | Value | +| --- | --- | +| `__version__` | Version of iterframes, such as `"0.4.0"` | +| `FFMPEG_VERSION` | Version of the FFmpeg that iterframes is linked to, such as `"9.0.2"` | +| `DEVICES` | Names accepted by `device` besides `"auto"`, such as `("cpu", "mps")` | diff --git a/docs/development.md b/docs/development.md index 049bb13..36d0e81 100644 --- a/docs/development.md +++ b/docs/development.md @@ -10,21 +10,39 @@ sudo apt install build-essential curl python3-venv \ pkg-config libclang-dev nasm # Debian, Ubuntu ``` -On Windows, FFmpeg is built with MSVC from an -[MSYS2](https://www.msys2.org/) shell, as CI does in -`.github/workflows/ci.yaml`. Install Visual Studio's C++ build tools and -LLVM (for libclang), then, in MSYS2: +macOS and Linux need nothing else: the build fetches nasm, meson, and +ninja when they are missing. -```console -pacman -S make diffutils curl tar xz \ - mingw-w64-ucrt-x86_64-pkgconf mingw-w64-ucrt-x86_64-nasm -rm /usr/bin/link.exe # it shadows MSVC's link.exe -uv tool install meson && uv tool install ninja -``` +### Windows + +The wheels target MSVC, so FFmpeg is compiled with MSVC too. Its build +script is a shell script, so it runs in an [MSYS2](https://www.msys2.org/) +shell, as in the `windows` job of `.github/workflows/ci.yaml`. + +1. Install Visual Studio's C++ build tools, [LLVM](https://llvm.org/) + (for libclang), MSYS2, rustup, and uv. +2. Open the *x64 Native Tools Command Prompt* and start MSYS2's UCRT64 + shell from it, keeping its `PATH` so that `cl`, `cargo`, and `uv` stay + on it: -Start the MSYS2 UCRT64 shell from a Visual Studio developer prompt with -`msys2_shell.cmd -ucrt64 -use-full-path`, so that `cl`, `cargo`, and `uv` -stay on `PATH`, and run every command below from it. + ```console + C:\msys64\msys2_shell.cmd -ucrt64 -use-full-path + ``` + +3. In that shell, install the build tools. `pacman` is MSYS2's package + manager; meson and ninja come from uv instead, because they must be + Windows programs to drive MSVC: + + ```console + pacman -S make diffutils curl tar xz \ + mingw-w64-ucrt-x86_64-pkgconf mingw-w64-ucrt-x86_64-nasm + mv /usr/bin/link.exe /usr/bin/link-msys2.exe # hides MSVC's link.exe + uv tool install meson + uv tool install ninja + export PATH="$PATH:$(cygpath -u "$(uv tool dir --bin)")" + ``` + +Run every command below from that shell. ## FFmpeg @@ -63,10 +81,13 @@ environment. Run it again after every change to a `.rs` file. | `src/lib.rs` | Python module: `Frame`, `Batch`, `FrameReader`, and the error mapping | | `src/decoder.rs` | Decoding thread: demux, decode, convert to RGB | | `src/ffmpeg.rs` | Safe wrappers over the FFmpeg calls the crate needs | +| `src/dlpack.rs` | DLPack capsules for the planes of `CudaFrame` | | `iterframes/__init__.py` | `read` and `read_batches`, which wrap `FrameReader` | -| `scripts/build-ffmpeg.sh` | Static FFmpeg build for the wheels | +| `build.rs` | Builds and links FFmpeg, generates its bindings | +| `scripts/build-ffmpeg.sh` | Static FFmpeg build, run by `build.rs` | | `tests/` | pytest suite, which checks the frames against PyAV | | `docs/` | This site | +| `web/` | Landing page, published on GitHub Pages | ## Tests @@ -106,6 +127,9 @@ uv run --no-sync mkdocs serve Read the Docs builds the site with `uv sync` from `uv.lock`, without compiling the extension. +The landing page in `web/` has no build step: open `web/index.html`, or +run `python3 -m http.server --directory web`. + ## Wheels The wheels use the stable ABI of CPython (abi3), so one wheel per platform diff --git a/docs/reference.md b/docs/guides.md similarity index 54% rename from docs/reference.md rename to docs/guides.md index 3068324..e6cf94a 100644 --- a/docs/reference.md +++ b/docs/guides.md @@ -1,77 +1,6 @@ -# Reference +# Guides -Everything lives in the top-level `iterframes` module. - -## read - -```python -read(path, height=None, width=None, prefetch_frames=1, device="cpu", on_device=False, frames=None, start=0, stop=None, step=1, approximate=None) -> Iterator[numpy.ndarray] -``` - -Yields the frames of the video at `path`, in order, or with `frames`, or -`start`, `stop`, and `step`, only those frames, in the order asked for. -Each frame is a C-contiguous, writable `numpy.ndarray` of shape -`(height, width, 3)` and dtype `uint8`, holding RGB pixels. - -| Argument | Description | -| --- | --- | -| `path` | Path of the video, as a `str` or `os.PathLike` | -| `height` | Height of the frames. Defaults to the height of the video | -| `width` | Width of the frames. Defaults to the width of the video | -| `prefetch_frames` | How many decoded frames may wait for your code. Defaults to 1 | -| `device` | Where to decode: `"cpu"`, `"auto"`, or a name from `DEVICES`. Defaults to `"cpu"`. See [Hardware decoding](#hardware-decoding) | -| `on_device` | With `device="cuda"`, yield `CudaFrame` objects left on the GPU instead of arrays. See [Frames on the GPU](#frames-on-the-gpu) | -| `frames` | The numbers of the frames to read, in the order to read them. See [Reading frames by number](#reading-frames-by-number) | -| `start`, `stop`, `step` | The frames to read, as a slice of the video | -| `approximate` | With `frames`, read the key frame nearest to each of them instead, when it is within this many frames. `True` is any distance. See [Approximate frames](#approximate-frames) | - -Frames are resized with bilinear interpolation. When only one of `height` -and `width` is given, the other keeps the size of the video, so the aspect -ratio changes. - -Decoding starts when the first frame is requested, on a background thread -that runs ahead of your code by up to `prefetch_frames` frames. The thread -never takes the GIL, so it decodes the next frames while your code -processes the current one, even when that code holds the GIL. A larger -`prefetch_frames` smooths out frames that take longer to decode, at the -cost of memory: one 1080p frame takes about 6 MB. - -The decoder stops when the iterator is exhausted or garbage-collected, for -example after a `break`. - -```python -import iterframes - -for frame in iterframes.read("video.mp4", height=270, width=480): - print(frame.shape) # (270, 480, 3) -``` - -## read_batches - -```python -read_batches(path, batch_size, height=None, width=None, prefetch_frames=1, device="cpu", drop_last=False, frames=None, start=0, stop=None, step=1, approximate=None) -> Iterator[numpy.ndarray] -``` - -Yields the frames of the video in batches, in order, for models that take -several frames at once. Each batch is a C-contiguous, writable -`numpy.ndarray` of shape `(batch_size, height, width, 3)` and dtype -`uint8`. The frames are decoded straight into it, so unlike `numpy.stack` -over the frames of `read`, making a batch costs your code no copy. - -| Argument | Description | -| --- | --- | -| `batch_size` | Frames per batch, at least 1 | -| `prefetch_frames` | How many decoded frames may wait for your code, rounded up to whole batches. Defaults to 1, that is one batch | -| `drop_last` | Drop the last batch when the video ends before it is full. By default it is yielded with the frames left over | - -The other arguments are those of [`read`](#read); `on_device` is not -supported. Every frame of a batch has the size of the first frame of the -video, or `height` and `width`. - -```python -for batch in iterframes.read_batches("video.mp4", 12, height=224, width=224): - print(batch.shape) # (12, 224, 224, 3), except maybe the last one -``` +What the arguments of [`read`](api.md#read) are for, and what they cost. ## Reading frames by number @@ -86,7 +15,7 @@ every_fifth = iterframes.read("video.mp4", start=100, stop=200, step=5) thumbnail = next(iterframes.read("video.mp4", frames=[-1])) ``` -Both work with [`read_batches`](#read_batches), which decodes the frames +Both work with [`read_batches`](api.md#read_batches), which decodes the frames straight into the batch: ```python @@ -105,7 +34,7 @@ RGB. Frames asked for in order cost no more than reading the video straight through, since the decoder goes on from the frame it decoded last whenever that is closer than the key frame. -Frame numbers are those of [`read`](#read): frame `n` is the one `read` +Frame numbers are those of [`read`](api.md#read): frame `n` is the one `read` yields `n`-th. A number the video does not have raises `IndexError`, while a slice past the end stops at the last frame, as a list does. @@ -135,25 +64,6 @@ saving is in decoding alone. Two frames near the same key frame become the same frame, which is decoded once for each of them: the iterator yields as many frames as `frames` asks for, in the same order. -## Errors - -Errors are raised by the first `next()` on the iterator, not by the call -to `read`, because decoding starts only then. - -| Exception | When | -| --- | --- | -| `FileNotFoundError`, `PermissionError`, `OSError` | The file cannot be opened. The exception carries `errno` and `filename` | -| `ValueError` | The file is not a video FFmpeg can read, or has no video stream | -| `RuntimeError` | Decoding fails after the video has been opened, for instance on a codec that the build does not include | -| `IndexError` | `frames` asks for a frame the video does not have | - -```python -try: - frames = list(iterframes.read("missing.mp4")) -except FileNotFoundError as error: - print(error.filename) # missing.mp4 -``` - ## Threads Waiting for the next frame releases the GIL, and each call to `read` has @@ -171,41 +81,6 @@ with ThreadPoolExecutor() as pool: FFmpeg also spreads the decoding of each video over several threads. -## FrameReader - -```python -FrameReader(path, height=None, width=None, prefetch_frames=1, device="cpu", on_device=False, batch_size=None, drop_last=False, frames=None, start=0, stop=None, step=1, approximate=None) -``` - -The iterator behind `read` and `read_batches`. It yields `Frame` objects, -`CudaFrame` objects with `on_device=True`, or `Batch` objects with -`batch_size`. - -## Frame - -A decoded frame. It holds the pixels and exposes them through the buffer -protocol as a writable, C-contiguous `(height, width, 3)` block of -unsigned bytes, so that other libraries can read them without a copy: - -```python -import numpy as np -from iterframes import FrameReader - -for frame in FrameReader("video.mp4"): - array = np.asarray(frame) # what read() yields - view = memoryview(frame) # no NumPy needed -``` - -The pixels stay alive as long as the frame or any array or view on it. - -## Batch - -Frames decoded into a single block of memory, which the buffer protocol -exposes as a writable, C-contiguous `(frames, height, width, 3)` block of -unsigned bytes. `read_batches` wraps it with `numpy.asarray`, which copies -nothing; its pixels stay alive as long as the batch or any array or view on -it. - ## Hardware decoding `device` moves decoding to a hardware device, which leaves more CPU to @@ -303,11 +178,3 @@ for frame in frames: This has not run on an NVIDIA GPU yet; please report how it works for you. - -## Constants - -| Name | Value | -| --- | --- | -| `__version__` | Version of iterframes, such as `"0.4.0"` | -| `FFMPEG_VERSION` | Version of the FFmpeg that iterframes is linked to, such as `"9.0.2"` | -| `DEVICES` | Names accepted by `device` besides `"auto"`, such as `("cpu", "mps")` | diff --git a/docs/index.md b/docs/index.md index 52369f4..8896172 100644 --- a/docs/index.md +++ b/docs/index.md @@ -57,17 +57,19 @@ for batch in iterframes.read_batches("video.mp4", 12, height=224, width=224): assert batch.shape[1:] == (224, 224, 3) ``` -Read frames at given positions, without decoding the whole video: +Read frames by number, in any order, without decoding the whole video: ```python -clip = list(iterframes.read("video.mp4", frames=[0, 30, 60])) +clip = list(iterframes.read("video.mp4", frames=[120, 0, 60])) +last = next(iterframes.read("video.mp4", frames=[-1])) every_fifth = iterframes.read("video.mp4", start=100, stop=200, step=5) ``` -Or, when a frame nearby will do, read the nearest key frame to each of -them and skip the decoding in between: +When a frame nearby will do, `approximate` reads the nearest key frame +instead, which skips the decoding in between: ```python +# The nearest key frame within 5 frames, else the frame itself. clip = list(iterframes.read("video.mp4", frames=[0, 30, 60], approximate=5)) ``` @@ -79,8 +81,31 @@ for index, frame in enumerate(iterframes.read("video.mp4")): break ``` -The [reference](reference.md) describes every argument and the errors -raised. +Errors, such as a missing file, are raised by the first `next()`, not by +`read`, because decoding starts only then. + +Decode on a GPU when the machine has one, else on the CPU: + +```python +for frame in iterframes.read("video.mp4", device="auto"): + ... +``` + +Read several videos in parallel, one thread each; waiting for a frame +releases the GIL: + +```python +from concurrent.futures import ThreadPoolExecutor + +def count_frames(path): + return sum(1 for _ in iterframes.read(path)) + +with ThreadPoolExecutor() as pool: + counts = list(pool.map(count_frames, paths)) +``` + +The [guides](guides.md) explain random access, hardware decoding, and +frames on the GPU; the [API](api.md) lists every argument and error. ## Formats @@ -90,4 +115,4 @@ MPEG-4, ProRes, and MJPEG. AV1 is decoded by [dav1d](https://code.videolan.org/videolan/dav1d). Only local files are read: network protocols are left out of the build. Decoding runs on the CPU unless you ask for a hardware device; see -[Hardware decoding](reference.md#hardware-decoding). +[Hardware decoding](guides.md#hardware-decoding). diff --git a/mkdocs.yml b/mkdocs.yml index a57e39b..d66e7fa 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -52,9 +52,13 @@ markdown_extensions: plugins: - search +hooks: + - scripts/docs_redirects.py + nav: - Home: index.md - - Reference: reference.md + - Guides: guides.md + - API: api.md - Development: development.md extra: diff --git a/scripts/docs_redirects.py b/scripts/docs_redirects.py new file mode 100644 index 0000000..5a7d001 --- /dev/null +++ b/scripts/docs_redirects.py @@ -0,0 +1,41 @@ +"""Redirect the old reference page, which became the API and Guides pages. + +mkdocs-redirects keeps the anchor but not its page, and the sections of +the old page went to two pages, so each anchor is mapped here. +""" + +import json +import os + +OLD = "reference" +GUIDES = [ + "reading-frames-by-number", + "approximate-frames", + "threads", + "hardware-decoding", + "frames-on-the-gpu", +] + +PAGE = """ + + + +Redirecting + + + +API ยท Guides + +""" + + +def on_post_build(config, **kwargs): + path = os.path.join(config["site_dir"], OLD, "index.html") + os.makedirs(os.path.dirname(path), exist_ok=True) + with open(path, "w") as file: + file.write(PAGE.format(guides=json.dumps(GUIDES))) From 5b425b27a51dab72111f2b07387915de0364659d Mon Sep 17 00:00:00 2001 From: Alessio Sanfratello Date: Mon, 28 Sep 2026 15:04:03 +0200 Subject: [PATCH 2/3] docs: explain how the buffer protocol spares a copy of each frame Co-Authored-By: Claude Opus 5.5 --- docs/api.md | 1 + docs/guides.md | 28 ++++++++++++++++++++++++++++ docs/index.md | 3 ++- 3 files changed, 31 insertions(+), 1 deletion(-) diff --git a/docs/api.md b/docs/api.md index 90d30a5..5f76d40 100644 --- a/docs/api.md +++ b/docs/api.md @@ -118,6 +118,7 @@ for frame in FrameReader("video.mp4"): ``` The pixels stay alive as long as the frame or any array or view on it. +See [Frames without a copy](guides.md#frames-without-a-copy). ## Batch diff --git a/docs/guides.md b/docs/guides.md index e6cf94a..fcec267 100644 --- a/docs/guides.md +++ b/docs/guides.md @@ -2,6 +2,34 @@ What the arguments of [`read`](api.md#read) are for, and what they cost. +## Frames without a copy + +Each frame is decoded into memory that iterframes allocates, and +swscale writes its RGB pixels there. `read` does not copy them into a new +NumPy array. The `Frame` that owns the memory implements Python's +[buffer protocol](https://docs.python.org/3/c-api/buffer.html): it tells +any library that asks where the pixels are, their shape +`(height, width, 3)`, their strides, and their type, `uint8`. +`numpy.asarray` builds the array on those same bytes. + +What this changes for your code: + +- **No copy per frame.** A copy would read and write every pixel once + more, about 6 MB for a 1080p frame. Wrapping costs the same at any + size. +- **Contiguous arrays.** The rows have no padding, so the array is + C-contiguous. Libraries that need contiguous memory take it as it is: + `torch.from_numpy(frame)` shares the memory too. +- **Batches in one block.** With `read_batches`, swscale writes each + frame into its slice of one allocation. The batch arrives as one + `(batch, height, width, 3)` array, with no `numpy.stack`. +- **Memory you own.** The decoder never writes to a frame again once it + hands it over, so the array is writable and you can change it in place. + The memory is freed when the last array or view on it goes away. + +Other libraries can read a `Frame` without NumPy; see +[`Frame`](api.md#frame). + ## Reading frames by number `frames` reads the frames with those numbers, in the order given, instead diff --git a/docs/index.md b/docs/index.md index 8896172..cc25f9e 100644 --- a/docs/index.md +++ b/docs/index.md @@ -14,7 +14,8 @@ While `model` runs on one frame, FFmpeg decodes the next ones on a background thread, written in Rust with [PyO3](https://pyo3.rs/). The thread never takes the GIL, so it keeps decoding even while your code holds it, and decoding overlaps with your work instead of adding to it. -The frames reach NumPy without a copy. +The frames reach NumPy without a copy, through the +[buffer protocol](guides.md#frames-without-a-copy). ## Installation From 48603783febb81f93b0bd74c3d4ab296686a2525 Mon Sep 17 00:00:00 2001 From: Alessio Sanfratello Date: Mon, 28 Sep 2026 22:20:09 +0200 Subject: [PATCH 3/3] docs: point to the new pages and rewrap AGENTS.md [skip ci] Co-Authored-By: Claude Opus 5.5 --- AGENTS.md | 8 ++++---- tests/test_read.py | 2 +- 2 files changed, 5 insertions(+), 5 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 44c6575..de74f1a 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -151,10 +151,10 @@ Four pages: `index.md` (overview and quick start), `guides.md` (what the arguments are for, and what they cost), `api.md` (signatures, arguments, and errors), `development.md`. `scripts/docs_redirects.py` sends links to the old `reference/` page, anchors included, to the page that now holds -each section; add to it when a section moves. Every example must match the behavior of a freshly -built module; check them instead of writing them from memory. Keep the -text short, in American English, with no performance claims that have not -been measured. +each section; add to it when a section moves. Every example must match +the behavior of a freshly built module; check them instead of writing +them from memory. Keep the text short, in American English, with no +performance claims that have not been measured. The landing page in `web/` is part of the docs: whenever an argument, a default, or the behavior of a feature changes, update it in the same diff --git a/tests/test_read.py b/tests/test_read.py index 41c1e74..c9eb5df 100644 --- a/tests/test_read.py +++ b/tests/test_read.py @@ -310,7 +310,7 @@ def test_on_device_matches_cpu(video_path, pyav_frames, assert_close): def nv12_to_rgb(torch, frame): - """The conversion shown in docs/reference.md.""" + """The conversion shown in docs/guides.md.""" y = torch.from_dlpack(frame.y).float() uv = torch.from_dlpack(frame.uv).float() uv = uv.repeat_interleave(2, 0).repeat_interleave(2, 1)