Skip to content

Latest commit

 

History

History
305 lines (218 loc) · 25.8 KB

File metadata and controls

305 lines (218 loc) · 25.8 KB

Streaming configuration

Everything between the input URI and the output URL: real sources and sinks, the encoder presets, GPU offload, audio, and the full flag reference.

← Back to the README · AI workers · Operations

Real-world pipelines

SRT in, SRT out — the common broadcast shape:

cargo run -p braidpipe --release -- \
  --uri 'srt://0.0.0.0:9000?mode=listener' \
  --sink 'videoconvert ! video/x-raw,format=I420 \
          ! x264enc tune=zerolatency bitrate=4000 speed-preset=veryfast key-int-max=60 \
          ! h264parse config-interval=-1 ! mpegtsmux \
          ! srtsink sync=false uri="srt://0.0.0.0:8891?mode=listener&latency=200"'

sync=false is not incidental — it is worth about 48 ms, for the reasons in Measuring latency.

NDI in, via the ndi:// pseudo-scheme:

cargo run -p braidpipe --release -- --uri 'ndi://STUDIO-PC%20(Camera%201)' --sink 'videoconvert ! autovideosink'

ndisrc registers no GStreamer URI handler, so the daemon maps this scheme itself onto ndisrc ndi-name="STUDIO-PC (Camera 1)" ! ndisrcdemux name=decoder. The name is the full one receivers list — MACHINE (source), percent-encoded — and it must match exactly: the machine part is whatever the sender's NDI runtime advertises, which on macOS includes the .LOCAL suffix (EMDIPLES-MACBOOK-PRO.LOCAL (File Feed)). A wrong name does not fail fast — ndisrc waits out its connect timeout and the pipeline dies with EOS without available srcpad(s) — so check the advertised name first:

dns-sd -B _ndi._tcp local        # macOS; avahi-browse -r _ndi._tcp on Linux

An optional ?url=host:port skips discovery and connects straight to the sender (ndisrc url-address), which helps when mDNS does not cross the network. --audio taps the demuxer like any other source and needs no extra setup. Needs the ndi plugin from gst-plugins-rs and the NDI runtime.

The daemon waits for the source rather than timing out. ndisrc on its own gives up after 10 s if the source is not there yet (connect-timeout) and after 5 s without a frame once it is (timeout), and either one ends the pipeline with the EOS without available srcpad(s) error above. The daemon sets both to 0, which the element treats as "never": start braidpipe before the camera is on and it sits waiting, and if the source drops out for a while the video simply resumes when it comes back, with the output staying up throughout. To get the fail-fast behaviour back, pass the timeouts in milliseconds on the URI, e.g. ndi://STUDIO-PC%20(Camera%201)?connect-timeout=10000&timeout=5000.

A Blackmagic DeckLink capture card (SDI/HDMI), via the decklink:// pseudo-scheme:

cargo run -p braidpipe --release -- \
  --uri 'decklink://0?mode=1080p50&connection=sdi' \
  --width 1920 --height 1080 --fps 50 \
  --preset lowlatency --output 'srt://0.0.0.0:8891?mode=listener'

decklinkvideosrc registers no GStreamer URI handler, so the daemon maps this scheme itself: the number after // is the card (decklink:// alone means the first one), and the optional mode and connection query parameters go straight to the element — gst-inspect-1.0 decklinkvideosrc lists the valid values, and mode=auto follows whatever the deck delivers on cards that support format detection. Needs the decklink plugin from gst-plugins-bad and Blackmagic's Desktop Video drivers installed. --audio works here too: a capture card has no demuxer to tap, so the audio branch is sourced from decklinkaudiosrc on the same device number, picking up the embedded SDI/HDMI audio.

Raw RTP in, which needs explicit caps and a depayloader — that's what --source is for:

cargo run -p braidpipe --release -- \
  --source 'udpsrc port=5000 caps="application/x-rtp,media=video,encoding-name=H264,payload=96" \
            ! rtph264depay ! h264parse ! decodebin3 ! videoconvert ! videoscale' \
  --sink 'videoconvert ! autovideosink'

A webcam, on macOS (avfvideosrc) or Linux (v4l2src device=/dev/video0):

cargo run -p braidpipe --release -- \
  --source 'avfvideosrc device-index=0 ! videoconvert ! videoscale ! video/x-raw,width=1280,height=720,framerate=30/1 ! videoconvert' \
  --sink 'videoconvert ! autovideosink'

The scaling matters: --width/--height fix the shared-memory slot geometry, and a camera negotiates whatever resolution it likes unless you tell it otherwise. List your devices with gst-device-monitor-1.0 Video/Source — the index is not always the one you expect, and on macOS index 0 is often an iPhone offering itself as a Continuity Camera. A camera that is present but not actually streaming sits in PLAYING and delivers nothing, which looks exactly like a braidpipe stall; check it in isolation first:

gst-launch-1.0 avfvideosrc device-index=0 num-buffers=10 ! fakesink

That should reach end-of-stream in a couple of seconds. If it hangs, the problem is the camera, not this project.

1080p60:

cargo run -p braidpipe --release -- --width 1920 --height 1080 --fps 60 \
  --source 'videotestsrc is-live=true pattern=ball ! video/x-raw,width=1920,height=1080,framerate=60/1 ! videoconvert' \
  --sink 'videoconvert ! autovideosink'

--uri and --source are mutually exclusive. Use --uri when GStreamer can figure out the source on its own (it picks srtsrc ! decodebin3 for srt:// and uridecodebin3 for everything else); use --source when you need to spell out elements yourself.

Output presets

Writing the sink by hand, as above, gives full control — but most deployments want one of a few well-understood points on the latency/bandwidth curve. --output plus --preset builds the whole encoder + mux + sink chain for you, the way ffmpeg's -preset expands into a bag of x264 options:

cargo run -p braidpipe --release -- \
  --uri 'srt://0.0.0.0:9000?mode=listener' \
  --preset lowlatency --output 'srt://0.0.0.0:8891?mode=listener'

--output understands rtmp://, srt://, udp://host:port and ndi://<name>, and picks the right mux for each (FLV for RTMP, MPEG-TS for SRT/UDP, the NDI sink combiner for NDI). The daemon logs the sink it built at startup, so you can copy it out and use it as a --sink starting point.

NDI is the odd one out: it carries raw frames and the NDI SDK compresses them itself, so none of the encoder settings below apply to an ndi:// output — only the sink sync flag does. See the NDI recipe.

Preset Encoder settings GOP VBV Sink sync SRT latency Intent
zerolatency ultrafast + zerolatency, 6000 kbps 1 s 100 ms false 50 ms Every latency lever pulled; bandwidth pays for it
lowlatency (default) veryfast + zerolatency, 4500 kbps 2 s 200 ms false 125 ms The measured sweet spot — same ~40 ms p50 as the tuned harness sink
balanced medium + zerolatency, 3000 kbps 2 s 500 ms false 250 ms Better compression, still no B-frame delay
bandwidth slow, B-frames + lookahead, 1800 kbps 4 s 1000 ms true 500 ms Minimum bits for the quality; adds several frames of encoder delay by design

The VBV column is what makes the bitrate column mean something on the wire: it bounds how far above the target the encoder may burst, and how much encoded data a receiver has to be ready to buffer — so it is simultaneously a bandwidth cap and hidden latency. x264's own default (600 ms) would allow bursts more than half a second long.

Bitrates assume 720p30 — scale them for other formats. A preset only decides defaults; every parameter yields to an environment variable, so you can start from a profile and turn one knob:

Variable Overrides Values
BRAIDPIPE_ENCODER encoder auto (default — best GPU encoder, else x264), x264, vtenc, nvenc, va, vaapi, qsv, mf, amf — see GPU acceleration
BRAIDPIPE_HW GPU decode/encode detection off forces software both ways
BRAIDPIPE_BITRATE_KBPS target bitrate kbps
BRAIDPIPE_SPEED_PRESET x264 speed preset ultrafast … placebo
BRAIDPIPE_ZEROLATENCY zero-latency tuning 1/0 — x264 tune=zerolatency, vtenc realtime
BRAIDPIPE_GOP_SECONDS keyframe interval seconds
BRAIDPIPE_VBV_BUF_MS x264 VBV buffer (burst bound) milliseconds
BRAIDPIPE_SINK_SYNC sink clock sync 1/0 — see Measuring latency for why 0 is worth ~48 ms
BRAIDPIPE_SRT_LATENCY_MS srtsink latency budget milliseconds
BRAIDPIPE_SRT_WAIT_FOR_CONNECTION srtsink blocking on a missing viewer 1/0 — default 0: the pipeline runs and drops output until a viewer connects, so the input is consumed from startup
# lowlatency profile, but cap the bandwidth
BRAIDPIPE_BITRATE_KBPS=2500 cargo run -p braidpipe --release -- \
  --preset lowlatency --output srt://127.0.0.1:8888

Verified end-to-end with the latency harness: --preset lowlatency --output rtmp://… measured 39.8 ms p50 / 45.6 ms p99 worker→receiver at 720p30, identical to the hand-tuned sink.

Recipe: NDI 1080p50, low latency, high quality

When bandwidth is not a constraint and the goal is the lowest latency at the best picture, stay on NDI end to end — the processed feed goes back onto the network as a new NDI source, and no H.264 encoder ever touches it:

cargo run -p braidpipe --release -- \
  --uri 'ndi://STUDIO-PC%20(Studio%20Camera)' \
  --width 1920 --height 1080 --fps 50 \
  --audio --preset lowlatency \
  --output 'ndi://Studio%20Camera%20AI'

The name after ndi:// is what receivers (vMix, OBS, TriCaster, NDI Studio Monitor) see in their source list, percent-encoded the same way the input side addresses a camera. Under the hood this expands to videoconvert ! video/x-raw,format=UYVY ! ndisinkcombiner name=mux ! ndisink sync=false ndi-name="Studio Camera AI", with audio joining the combiner as PCM. Why this is the right shape:

  • No encoder, no encoder delay. NDI takes raw UYVY frames and the SDK applies its own intra-frame SpeedHQ compression, so the bitrate, speed-preset, GOP and VBV knobs simply do not exist here. --encoder and BRAIDPIPE_BITRATE_KBPS are ignored for an ndi:// output.
  • --preset lowlatency still matters for one thing: it sets sync=false on the sink, so frames leave the moment the relay hands them back rather than being held to the pipeline clock.
  • --audio stays uncompressed too — interleaved F32 PCM into the combiner's audio pad, which is the only layout it accepts. The AAC variables in Audio passthrough do not apply.
  • Bandwidth is the price. Full-bandwidth NDI at 1080p50 runs around 150–200 Mbps per stream, so this recipe wants a wired gigabit LAN, ideally on its own VLAN. That is the trade the title promises: latency and quality, paid for in network.

It needs the ndi plugin from gst-plugins-rs and the NDI runtime installed on the host — the Docker images do not ship either, so this recipe runs natively.

To check the feed from GStreamer, receive it by its advertised name — the sending machine's NDI name plus the --output name in parentheses:

gst-launch-1.0 ndisrc ndi-name="STUDIO-PC.LOCAL (Studio Camera AI)" ! ndisrcdemux name=d \
  d.video ! queue ! videoconvert ! autovideosink sync=false \
  d.audio ! queue ! audioconvert ! autoaudiosink

The name must match what dns-sd -B _ndi._tcp local (or avahi-browse -r _ndi._tcp) lists, .LOCAL suffix included where the runtime adds one. A near-miss does not fail fast: ndisrc waits out its connect timeout and then the demuxer errors with EOS without available srcpad(s). If discovery does not reach the receiving machine at all, url-address=host:port in place of ndi-name connects directly (lsof -nP -iTCP -sTCP:LISTEN -a -p <daemon pid> shows the port the SDK picked).

If the output has to leave the LAN instead, keep the same source and swap the output for SRT — then the encoder is back in the path, and it is worth starting from lowlatency and turning exactly two knobs:

BRAIDPIPE_BITRATE_KBPS=20000 BRAIDPIPE_SPEED_PRESET=fast \
cargo run -p braidpipe --release -- \
  --uri 'ndi://STUDIO-PC%20(Studio%20Camera)' \
  --width 1920 --height 1080 --fps 50 \
  --preset lowlatency --encoder auto \
  --output 'srt://0.0.0.0:8891?mode=listener'

Why these values and nothing else:

  • --preset lowlatency already pulls every latency lever that matters — tune=zerolatency (no B-frames, no lookahead), a 200 ms VBV burst bound, sync=false on the sink, a 2 s GOP (key-int-max=100 at 50 fps). None of those need restating.
  • BRAIDPIPE_BITRATE_KBPS=20000 — the preset's 4500 kbps assumes 720p30; 1080p50 is ~3.7× the pixel rate, and for H.264 at this format quality saturates around 20–25 Mbps. Past that you are spending bits without seeing them.
  • BRAIDPIPE_SPEED_PRESET=fast — the speed preset costs CPU, not latency, as long as the encoder keeps real time. fast buys compression efficiency over the preset's veryfast; go slower only if CPU headroom says so, and never shrink the VBV to compensate.
  • --encoder auto — at 20 Mbps the quality gap between hardware encoders and x264 collapses, so let a GPU take the job and keep the CPU for the worker (which has a 20 ms/frame budget at 50 fps).

Do not reach for the bandwidth preset here: its lookahead and B-frames add several frames of encoder delay by design.

Measured bandwidth

Bitrate targets are promises until you look at the wire, so scripts/preset-bandwidth.sh runs each preset through the full daemon + worker path, captures the RTMP output, and reports what actually left the encoder. Content is a moving scene blended with 30% white noise — a stand-in for camera footage, hard enough to push rate control against its cap. 20-second runs at 720p30, startup excluded:

Preset Target Mean on wire Worst 1 s Worst 250 ms burst Output fps
zerolatency 6000 kbps 5123 5332 6481 30.0
lowlatency 4500 kbps 4501 4525 4941 30.0
balanced 3000 kbps 3000 3033 3291 30.0
bandwidth 1800 kbps 64 142 439 30.0

Three things worth reading out of that table:

  • lowlatency and balanced hold their targets to within 1%, and no preset's worst 250 ms burst exceeds its target by more than 10% — that is the VBV bound doing its job. Provision the link for the target plus ~10% and it will not be surprised.
  • zerolatency runs ~15% under target on hard content. The 100 ms VBV is tight enough to constrain ABR itself, trading a little quality for the strictest burst bound. That is the correct trade for its use case.
  • bandwidth spends almost nothing on this content, and that is by design, not a bug. Without tune=zerolatency, x264's mbtree lookahead rates every block by how much future frames can predict from it — and noise predicts nothing, so mbtree declines to encode it. The synthetic content is 30% noise; real footage has structure everywhere and will sit far closer to target. Either way the target is a hard ceiling, never a floor: this preset buys quality-per-bit, not constant bandwidth. (On pure noise — BRAIDPIPE_BW_PATTERN=snow — the effect is even starker, while the three zerolatency presets still hold their caps, since zerolatency disables mbtree.)

GPU acceleration

On a machine with a capable GPU, both halves of the codec work move off the CPU automatically. The mechanism differs per platform because every OS has its own video API:

Platform Decode Encode (H.264, first available wins)
macOS VideoToolbox (vtdec_hw, vtdec) VideoToolbox (vtenc_h264)
Linux NVDEC → VA-API (va plugin, then legacy vaapi) → QuickSync nvh264enc → vah264enc → qsvh264enc → vaapih264enc
Windows Direct3D 12 → Direct3D 11 → NVDEC → QuickSync nvh264enc → qsvh264enc → amfh264enc → mfh264enc

Decoding needs no pipeline changes at all. decodebin3 picks decoders by element rank, so at startup the daemon promotes the platform's hardware decoders above the software avdec_* family and autoplugging does the rest — for any codec the source carries, on both --uri and custom --source pipelines that use a decodebin. The Hardware decoders promoted for autoplugging log line lists what was found. Frames still land in system memory for the AI branch, which needs the raw pixels anyway: the win is the decode itself, not zero-copy.

Either way, the daemon logs every codec element the pipeline actually ends up using, tagged hardware or software — the encoder at startup, the decoder as soon as the input's caps are known:

INFO braidpipe_engine::pipeline: Video encoder in use (hardware) element=vtenc_h264
INFO braidpipe_engine::pipeline: Video decoder in use (hardware) element=vtdec_hw

Whether the GPU is then actually doing work shows up in Monitoring: braidpipe_gpu_utilization_percent samples machine-wide GPU load every 5 seconds (via ioreg on macOS, nvidia-smi or the amdgpu sysfs on Linux), NVIDIA additionally breaks out the dedicated _encoder_/_decoder_ block utilization, and the Grafana dashboard has a GPU row for all of them.

Encoding is decided when --output builds the sink from a preset: the best hardware encoder present replaces the preset's x264 default, and the Built sink from preset log shows which one won. Detection is reliable because the NVIDIA/VA/QSV/AMF/MediaFoundation plugins only register their elements when the device probe succeeds — if nvh264enc exists in the registry, there is an NVENC-capable GPU behind it. The preset's parameters map onto each encoder's own vocabulary: bitrate and GOP always, the zero-latency switch per encoder (realtime, low-latency, ultra-low-latency usage), and the VBV burst bound where one is exposed (NVENC vbv-buffer-size, VA cpb-size).

Two knobs control this, each a CLI flag with an environment-variable twin (the flag wins when both are given):

  • --hw off (or BRAIDPIPE_HW=off) — software everywhere: no decoder promotion, no encoder auto-pick.
  • --encoder <name> (or BRAIDPIPE_ENCODER=<name>) — pin the encoder regardless of detection; --encoder auto explicitly re-enables detection. Hardware encoders trade some quality-per-bit for speed, so the bandwidth preset's intent is best served by pinning x264. Pinning the encoder leaves GPU decoding on — use --hw off to force software both ways.

A hand-written --sink bypasses encoder selection entirely — you name the encoder yourself.

In docker

The plain compose stack is software-only by design — it must run anywhere. On a Linux host with an NVIDIA GPU, add the GPU overlay:

docker compose -f docker-compose.yml -f docker-compose.gpu.yml up --build -d

The overlay hands the container every GPU; it does not change the image, because the braidpipe image is already the GPU image. NVENC/NVDEC come from the host's nvidia-container-toolkit injecting libnvidia-encode/libnvcuvid at run time, which it does on any base, and gstreamer1.0-plugins-bad in the image dlopens them. Success is the same log line as everywhere else, Video encoder in use (hardware) element=nvh264enc.

--build — on this command and on the plain docker compose up alike — is only needed the first time and after something the image is built from changes — Rust source, Cargo.toml/Cargo.lock, the Dockerfiles, the worker's Python files, or build args. Compose never detects source changes on its own: without the flag it silently runs the stale image. Edits to the compose files themselves (flags, environment, ports) need no rebuild — plain up -d recreates the container with the new settings. Rebuilds are incremental thanks to the Dockerfile's cargo cache mount, so a habitual --build costs seconds, not the full compile. Both invocations build the same braidpipe image, so a rebuild under either one serves the other. The host needs the NVIDIA driver working (nvidia-smi) and the nvidia-container-toolkit registered with docker — the overlay's header comments carry the check and install commands. For Intel/AMD VA-API instead, the header shows how to swap the GPU reservation for a /dev/dri device mount; the image already ships va-driver-all.

On macOS there is no GPU path in docker at all — Docker Desktop cannot pass the GPU through, so a containerized daemon always encodes with x264. Run the daemon natively (cargo run) to get VideoToolbox.

Audio passthrough

Real sources carry audio, and the output should too. --audio routes the source's audio around the AI branch — decoded, re-encoded to AAC, and joined back at the output muxer:

cargo run -p braidpipe --release -- \
  --uri 'srt://0.0.0.0:9000?mode=listener' \
  --audio --preset lowlatency --output 'srt://0.0.0.0:8891?mode=listener'

How sync works. There is no dedicated sync machinery, because none is needed: the relay pushes every video frame back into the pipeline with its original PTS — whether the worker processed it or the deadline passed and it went through unchanged — and audio keeps the PTS the source gave it. The muxer pairs the two streams by timestamp, exactly as it would in a plain GStreamer pipeline. That also means failover cannot desynchronize anything: the input-selector switches video branches while audio never stops flowing.

Measured on a live SRT source (video + audio) relayed to RTMP at 720p30: steady-state packet cadence was exactly 33.33 ms for video and 21.33 ms for audio (1024 samples at 48 kHz), with under one video frame of relative drift over a 19 s run — both with the worker healthy and with a worker that missed every deadline.

--audio needs to know where audio comes from and where it goes:

  • Source — with --uri this is automatic. With a custom --source, name your demuxer or decodebin decoder so the audio branch can tap it.
  • Sink — with --output this is automatic (preset muxers are named). With a custom --sink, name your muxer mux.

An ndi:// output takes audio as PCM rather than AAC, because NDI carries it uncompressed: the branch becomes … ! audioconvert ! audioresample ! audio/x-raw,format=F32LE,layout=interleaved ! queue ! mux. and the encoder variables below are ignored.

Variable Overrides Default
BRAIDPIPE_AUDIO_ENCODER AAC encoder element avenc_aac (in gst-libav; fdkaacenc, faac also work)
BRAIDPIPE_AUDIO_BITRATE_KBPS audio bitrate 128
BRAIDPIPE_AUDIO_BRANCH the entire generated branch, verbatim decoder. ! queue ! audio/x-raw ! audioconvert ! audioresample ! avenc_aac bitrate=128000 ! aacparse ! queue ! mux.

If the source has no audio stream, don't pass --audio — the audio branch would wait forever for a pad that never appears and GStreamer fails the pipeline with a delayed-linking error.

Command-line reference

Flag Default Purpose
-i, --source <PIPELINE> test pattern Explicit GStreamer source fragment
--uri <URI> — Input URI decoded by GStreamer (srt://, udp://, rtp://, file://), an NDI source (ndi://<MACHINE%20(name)>[?url=host:port&connect-timeout=ms&timeout=ms], timeouts default to never), or a DeckLink capture card (decklink://<device>?mode=…&connection=…)
-o, --sink <PIPELINE> videoconvert ! autovideosink Output fragment appended after the selector
--output <URL> — Publish target (rtmp://, srt://, udp://host:port, ndi://<name>); builds the sink from --preset
--preset <NAME> lowlatency Latency/bandwidth profile for --output, see Output presets
--hw <auto|off> auto GPU mode, see GPU acceleration; off forces software decode and encode
--encoder <NAME> auto Encoder for --output: auto picks the best hardware encoder, or pin x264, vtenc, nvenc, va, vaapi, qsv, mf, amf
--audio off Carry source audio to the output muxer, see Audio passthrough
-p, --python-script <PATH> python/braidpipe/worker.py Worker to launch
--external-worker off Don't spawn or supervise a worker; connect to an externally managed AI process, see External worker mode
-f, --fps <N> 30 Frame rate; sets the relay deadline and watchdog tick
--width <N> / --height <N> 1280 / 720 Shared-memory slot geometry
--rust-sock <PATH> /tmp/braidpipe_rust.sock Where the daemon listens for acks and hellos
--python-sock <PATH> /tmp/braidpipe_python.sock Where the worker listens for notifications
--worker-listen <IP:PORT> — Also accept workers from other machines (tcp-raw), see Workers on another machine
--passthrough-only off Media path only; no worker, no shared memory
--metrics-port <N> 9184 Prometheus endpoint on 127.0.0.1, see Monitoring; 0 disables
--metrics-drain-ms <N> 2000 How long to keep serving metrics after a shutdown signal, so the down state gets scraped

--width/--height must match the frames your source actually produces after videoscale, because they define the slot size that both sides index into.

Set RUST_LOG=debug to see per-frame relay activity, including dropped and stale acks.