Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 21 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,6 +54,13 @@ as `iterframes.iterframes`, and `iterframes/__init__.py` wraps its
frame, which decoding cannot return either, so frame numbers are those of
a plain `read`. A stream that cannot seek, such as raw H.264, is opened
again and read from the start (`Source::open`).
- `approximate` goes with `frames` only, and carries a tolerance in frames
(`True` is `usize::MAX`), in `Selection::Frames`. `selected` then maps
each index through `Seeker::nearest_key`, which snaps it to the key frame
closest to it when that is within the tolerance, and leaves it alone
otherwise. Nothing else changes: the index pass still runs, and two
indices that snap to the same key frame decode it once each, so the caller
gets one frame per number asked for.
- Errors travel through the channel and become Python exceptions in
`impl From<Error> for PyErr`. A closed channel means the end of the video.
- Dropping the reader closes the channel; the thread notices on its next
Expand Down Expand Up @@ -146,6 +153,20 @@ built module; check them instead of writing them from memory. Keep the
text short, in American English, with no performance claims that have not
been measured.

The landing page in `web/` is part of the docs: whenever an argument, a
default, or the behavior of a feature changes, update it in the same
change as the code, and say so in the pull request. Its diagrams are
drawn from the numbers declared in `web/main.js`, so change the number
and let the drawing and the captions follow; never write a figure into
the HTML that the script also computes. `web/README.md` lists those
numbers and the files. Check the result in a browser before proposing the
change: open `web/index.html`, or serve the directory with
`python3 -m http.server --directory web`. The page has no build step and no
third-party request: its two typefaces are self-hosted woff2 files in
`web/assets/fonts/`, and everything else is plain HTML, CSS, and JavaScript.
`web/README.md` describes the look and the colour rules; keep them rather
than restyling one component on its own.

## Landing page

`web/` is the page at
Expand Down
11 changes: 11 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -92,6 +92,17 @@ before it; frames asked for in order cost no more than reading the video
straight through. See
[Reading frames by number](https://iterframes.readthedocs.io/en/latest/reference/#reading-frames-by-number).

When a frame nearby will do, `approximate` reads the key frame closest to
each of `frames`, which costs one decoded frame instead of the frames from
the key frame on:

```python
# The nearest key frame within 5 frames, else the frame itself.
frames = list(iterframes.read("video.mp4", frames=[100, 200, 300], approximate=5))
```

See [Approximate frames](https://iterframes.readthedocs.io/en/latest/reference/#approximate-frames).

## Hardware decoding

Pass `device` to decode on a GPU instead of the CPU, which then stays free
Expand Down
7 changes: 7 additions & 0 deletions docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -64,6 +64,13 @@ clip = list(iterframes.read("video.mp4", frames=[0, 30, 60]))
every_fifth = iterframes.read("video.mp4", start=100, stop=200, step=5)
```

Or, when a frame nearby will do, read the nearest key frame to each of
them and skip the decoding in between:

```python
clip = list(iterframes.read("video.mp4", frames=[0, 30, 60], approximate=5))
```

Stop whenever you like; the decoder stops with the loop:

```python
Expand Down
33 changes: 30 additions & 3 deletions docs/reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ Everything lives in the top-level `iterframes` module.
## read

```python
read(path, height=None, width=None, prefetch_frames=1, device="cpu", on_device=False, frames=None, start=0, stop=None, step=1) -> Iterator[numpy.ndarray]
read(path, height=None, width=None, prefetch_frames=1, device="cpu", on_device=False, frames=None, start=0, stop=None, step=1, approximate=None) -> Iterator[numpy.ndarray]
```

Yields the frames of the video at `path`, in order, or with `frames`, or
Expand All @@ -23,6 +23,7 @@ Each frame is a C-contiguous, writable `numpy.ndarray` of shape
| `on_device` | With `device="cuda"`, yield `CudaFrame` objects left on the GPU instead of arrays. See [Frames on the GPU](#frames-on-the-gpu) |
| `frames` | The numbers of the frames to read, in the order to read them. See [Reading frames by number](#reading-frames-by-number) |
| `start`, `stop`, `step` | The frames to read, as a slice of the video |
| `approximate` | With `frames`, read the key frame nearest to each of them instead, when it is within this many frames. `True` is any distance. See [Approximate frames](#approximate-frames) |

Frames are resized with bilinear interpolation. When only one of `height`
and `width` is given, the other keeps the size of the video, so the aspect
Expand All @@ -48,7 +49,7 @@ for frame in iterframes.read("video.mp4", height=270, width=480):
## read_batches

```python
read_batches(path, batch_size, height=None, width=None, prefetch_frames=1, device="cpu", drop_last=False, frames=None, start=0, stop=None, step=1) -> Iterator[numpy.ndarray]
read_batches(path, batch_size, height=None, width=None, prefetch_frames=1, device="cpu", drop_last=False, frames=None, start=0, stop=None, step=1, approximate=None) -> Iterator[numpy.ndarray]
```

Yields the frames of the video in batches, in order, for models that take
Expand Down Expand Up @@ -108,6 +109,32 @@ Frame numbers are those of [`read`](#read): frame `n` is the one `read`
yields `n`-th. A number the video does not have raises `IndexError`,
while a slice past the end stops at the last frame, as a list does.

## Approximate frames

Decoding a frame in the middle of a group of pictures costs every frame
from the key frame on. `approximate` spends one decoded frame instead, by
reading the key frame nearest to each frame asked for. It goes with
`frames` only; with `start`, `stop`, and `step` it raises `ValueError`.

```python
# The nearest key frame to each of these, however far away it is.
frames = list(iterframes.read("video.mp4", frames=[100, 200, 300], approximate=True))

# The nearest key frame within 5 frames, else the frame itself.
frames = list(iterframes.read("video.mp4", frames=[100, 200, 300], approximate=5))
```

A number bounds the error: a frame moves by at most that many frames, and
one with no key frame that close is decoded exactly. `True` accepts any
distance, which in a video with a key frame every 10 seconds means a frame
up to 5 seconds away from the one asked for. `False`, `0`, and `None`
read every frame exactly.

The video is still indexed, so its packets are read either way, and the
saving is in decoding alone. Two frames near the same key frame become
the same frame, which is decoded once for each of them: the iterator
yields as many frames as `frames` asks for, in the same order.

## Errors

Errors are raised by the first `next()` on the iterator, not by the call
Expand Down Expand Up @@ -147,7 +174,7 @@ FFmpeg also spreads the decoding of each video over several threads.
## FrameReader

```python
FrameReader(path, height=None, width=None, prefetch_frames=1, device="cpu", on_device=False, batch_size=None, drop_last=False, frames=None, start=0, stop=None, step=1)
FrameReader(path, height=None, width=None, prefetch_frames=1, device="cpu", on_device=False, batch_size=None, drop_last=False, frames=None, start=0, stop=None, step=1, approximate=None)
```

The iterator behind `read` and `read_batches`. It yields `Frame` objects,
Expand Down
16 changes: 15 additions & 1 deletion iterframes/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,7 @@ def read(
start: int = 0,
stop: Optional[int] = None,
step: int = 1,
approximate: Union[bool, int, None] = None,
) -> Iterator[Union[np.ndarray, CudaFrame]]:
"""Yield the frames of the video at ``path``, in order.

Expand Down Expand Up @@ -81,6 +82,15 @@ def read(
the file first, by reading its packets without decoding them, and then
decodes each frame from the key frame before it, so frames asked for
in order cost no more than reading the video straight through.

``approximate`` trades accuracy for speed, and goes with ``frames``
only: each frame asked for is read as the key frame nearest to it, so
it costs one decoded frame instead of the frames from the key frame
on. ``approximate=n`` moves a frame by at most ``n`` frames and reads
the rest exactly, which bounds the error; ``approximate=True`` moves
it by any distance. The video still has to be indexed, so the packets
are read either way. Two frames near the same key frame are then the
same frame, and are decoded once each.
"""
reader = FrameReader(
path,
Expand All @@ -93,6 +103,7 @@ def read(
start=start,
stop=stop,
step=step,
approximate=approximate,
)
if on_device:
yield from reader
Expand All @@ -114,6 +125,7 @@ def read_batches(
start: int = 0,
stop: Optional[int] = None,
step: int = 1,
approximate: Union[bool, int, None] = None,
) -> Iterator[np.ndarray]:
"""Yield the frames of the video at ``path`` in batches, in order.

Expand All @@ -125,7 +137,8 @@ def read_batches(

Takes the same arguments as :func:`read`, except ``on_device``, so
``frames``, or ``start``, ``stop``, and ``step``, batch the frames
with those numbers instead of the whole video. The background thread
with those numbers instead of the whole video, and ``approximate``
batches the key frames nearest to ``frames``. The background thread
decodes up to ``prefetch_frames`` frames ahead, rounded up to whole
batches.
"""
Expand All @@ -143,6 +156,7 @@ def read_batches(
start=start,
stop=stop,
step=step,
approximate=approximate,
)
# The arrays share memory with the batches, which they keep alive.
for batch in reader:
Expand Down
42 changes: 38 additions & 4 deletions src/decoder.rs
Original file line number Diff line number Diff line change
Expand Up @@ -38,8 +38,13 @@ pub enum Selection {
/// straight through as well, and stops early.
First(usize),
/// These frames, in this order; a negative number counts from the end
/// of the video.
Frames(Vec<isize>),
/// of the video. With `approximate`, each of them is read as the
/// nearest key frame within that many frames of it, which costs one
/// decoded frame instead of the ones from the key frame on.
Frames {
wanted: Vec<isize>,
approximate: Option<usize>,
},
/// The frames from `start` to `stop`, every `step` of them, as a
/// Python slice: negative bounds count from the end, and `stop` of
/// `None` is the end of the video.
Expand Down Expand Up @@ -67,7 +72,7 @@ impl Selection {
match self {
Selection::All => Ok((0..frames).collect()),
Selection::First(count) => Ok((0..frames.min(*count)).collect()),
Selection::Frames(wanted) => wanted.iter().copied().map(resolve).collect(),
Selection::Frames { wanted, .. } => wanted.iter().copied().map(resolve).collect(),
// Out of range bounds clamp, as a slice of a list does.
Selection::Range { start, stop, step } => {
let clamp = |index: isize| {
Expand Down Expand Up @@ -250,8 +255,18 @@ fn selected(
tx: &Sender<Message>,
) -> Result<bool, Error> {
let mut seeker = Seeker::new(source, input, decoder)?;
let mut indices = selection.indices(seeker.len())?;
if let Selection::Frames {
approximate: Some(tolerance),
..
} = selection
{
for index in &mut indices {
*index = seeker.nearest_key(*index, tolerance);
}
}
let mut frame = Frame::new();
for index in selection.indices(seeker.len())? {
for index in indices {
seeker.frame(index, &mut frame)?;
if !converter.send(&mut frame, tx)? {
return Ok(false);
Expand Down Expand Up @@ -518,6 +533,25 @@ impl Seeker<'_> {
}
}

/// The key frame nearest to `index`, when no more than `tolerance`
/// frames away from it, else `index` itself. A tie goes to the key
/// frame before, which decoding is more likely to reach by going on.
fn nearest_key(&self, index: usize, tolerance: usize) -> usize {
let before = self.key_frame_before(index);
let nearest = match self
.keys
.get(self.keys.partition_point(|key| *key <= index))
{
Some(after) if after - index < index - before => *after,
_ => before,
};
if nearest.abs_diff(index) <= tolerance {
nearest
} else {
index
}
}

/// The last key frame at or before `index`. The numbers rise, so a
/// long video costs a binary search rather than a scan.
fn key_frame_before(&self, index: usize) -> usize {
Expand Down
39 changes: 34 additions & 5 deletions src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -332,7 +332,8 @@ impl Plane {
/// `batch_size`, a `Batch` of that many frames (fewer in the last one,
/// unless `drop_last`), whose `prefetch_frames` rounds up to whole batches.
/// `frames`, or `start`, `stop`, and `step`, pick the frames to decode by
/// number instead of reading the whole video. Use `iterframes.read` and
/// number instead of reading the whole video, and `approximate` reads each
/// of `frames` as the key frame nearest to it. Use `iterframes.read` and
/// `iterframes.read_batches`, which wrap them in NumPy arrays.
#[pyclass(module = "iterframes")]
struct FrameReader {
Expand All @@ -344,7 +345,8 @@ impl FrameReader {
#[new]
#[pyo3(signature = (
path, height=None, width=None, prefetch_frames=1, device="cpu", on_device=false,
batch_size=None, drop_last=false, frames=None, start=0, stop=None, step=1
batch_size=None, drop_last=false, frames=None, start=0, stop=None, step=1,
approximate=None
))]
#[allow(clippy::too_many_arguments)]
fn new(
Expand All @@ -360,8 +362,9 @@ impl FrameReader {
start: isize,
stop: Option<isize>,
step: isize,
approximate: Option<&Bound<'_, PyAny>>,
) -> PyResult<Self> {
let selection = selection(frames, start, stop, step)?;
let selection = selection(frames, start, stop, step, approximate)?;
let batching = match batch_size {
None => None,
Some(0) => return Err(PyValueError::new_err("batch_size must be at least 1")),
Expand Down Expand Up @@ -432,19 +435,28 @@ impl FrameReader {

/// Which frames to decode, from the arguments Python passes: `frames`
/// lists them, while `start`, `stop`, and `step` slice them as a list is
/// sliced. The two ways cannot be mixed.
/// sliced. The two ways cannot be mixed, and only `frames` can be
/// approximate.
fn selection(
frames: Option<Vec<isize>>,
start: isize,
stop: Option<isize>,
step: isize,
approximate: Option<&Bound<'_, PyAny>>,
) -> PyResult<Selection> {
let sliced = start != 0 || stop.is_some() || step != 1;
let approximate = tolerance(approximate)?;
match frames {
Some(_) if sliced => Err(PyValueError::new_err(
"frames does not go with start, stop, or step",
)),
Some(frames) => Ok(Selection::Frames(frames)),
Some(wanted) => Ok(Selection::Frames {
wanted,
approximate,
}),
None if approximate.is_some() => Err(PyValueError::new_err(
"approximate needs frames: a slice cannot be approximate",
)),
None if step < 1 => Err(PyValueError::new_err("step must be at least 1")),
// Frames counted from the first are read straight through, with
// no index and no seeking.
Expand All @@ -461,6 +473,23 @@ fn selection(
}
}

/// How far a frame may move to the nearest key frame, in frames, from
/// what Python passes as `approximate`: `True` is any distance, a number
/// is at most that many frames, and `False` or `None` is no move at all.
fn tolerance(approximate: Option<&Bound<'_, PyAny>>) -> PyResult<Option<usize>> {
let Some(approximate) = approximate else {
return Ok(None);
};
// A bool is an int in Python, so it has to be read first.
if let Ok(any_distance) = approximate.extract::<bool>() {
return Ok(any_distance.then_some(usize::MAX));
}
let frames = approximate.extract::<usize>().map_err(|_| {
PyValueError::new_err("approximate must be True, False, or a number of frames")
})?;
Ok(Some(frames))
}

/// The hardware devices of this build, by the names PyTorch gives them,
/// which Python users know better than FFmpeg's.
fn hardware_devices() -> impl Iterator<Item = (&'static str, ffmpeg::DeviceType)> {
Expand Down
26 changes: 26 additions & 0 deletions tests/test_benchmark.py
Original file line number Diff line number Diff line change
Expand Up @@ -108,3 +108,29 @@ def decode():
whole, second_half = read(), read(start=450)
print(f"\nwhole {whole:.3f}s, start=450 {second_half:.3f}s")
assert second_half < 0.75 * whole


def test_approximate_frames_decode_less(video_path):
"""Approximate frames must cost one decoded frame each.

Reading frames in the middle of their group of pictures decodes every
frame from the key frame on, while the nearest key frames decode one
frame each. Both pay for the pass that indexes the file.
"""
with av.open(str(video_path)) as container:
frames = list(container.decode(video=0))
keys = [index for index, frame in enumerate(frames) if frame.key_frame]
# Frames far enough into their group of pictures to cost several
# decoded frames each, which approximate reading skips.
wanted = [key + 20 for key in keys if key + 20 < len(frames)][:20]

def read(**arguments):
def decode():
for _ in iterframes.read(video_path, frames=wanted, **arguments):
pass

return best_of(decode)

exact, approximate = read(), read(approximate=True)
print(f"\nexact {exact:.3f}s, approximate {approximate:.3f}s")
assert approximate < exact
Loading
Loading