Skip to content

Latest commit

 

History

25 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

VidObjectifier — camera to noise box

A tiny GPU watches video and scribbles a score; SuperCollider turns the scribble into a noisy choir.

This repo is a log, and it is a toolkit. It's for anyone who wants to poke at the idea of objects in a camera frame becoming musical voices. The code is kept small and loud on purpose so you can read it like liner notes.


In plain language:

There are two main moving parts (plus a lo-fi Processing sketch if you want to stay in Java land):

  1. Analyzer (analyzer/vid2score.py)
    • Python script that runs on a Jetson Nano (or any CUDA Jetson).
    • Looks at a video file or live stream, spots objects with YOLO, and writes a timestamped score as CSV or JSONL.
  2. Renderer (renderer/render.scd or renderer/render_ring8.scd)
    • SuperCollider patch that reads the score and gives every object a voice.
    • Default budget is 20 voices, so things stay musical instead of mush.
  3. Processing sketch (processing/VidObjectifierProcessing.pde)
    • Java/Processing rewrite that spots moving blobs with OpenCV and hands each one a sine voice.

Picture a security camera feeding a garage band.

[video file / TouchDesigner] → [Jetson: detector + characterizer] → score.csv / score.jsonl
                                                     ↓
                                            [SuperCollider renderer]
                                                     ↓
                                     (stereo / 8‑channel ring / binaural)

Why bother?

  • Low cost, high grit. Everything here runs on free tools, save the Jetson Nano which can be had for ~$225 right now. Not my favorite situation, but when life gives you beautiful Jetsons, you run things like: SuperCollider, JACK, and a Jetson that you probably already cooked noodles on.
  • Every object gets a personality. The analyzer measures position, speed, color, shape, and even a janky "glitch" metric. The renderer maps those numbers to timbre.
  • Live or offline. Feed it prerecorded footage, or sling live frames from TouchDesigner over RTSP/NDI.

What's in the box?

.
├── analyzer/
│   └── vid2score.py           # video/stream → score.csv
├── renderer/
│   ├── render.scd             # SuperCollider renderer (stereo)
│   ├── render_ring8.scd       # 8‑channel ring renderer
│   ├── mapping.scd            # generated class→timbre map (see config/timbre_map.yaml)
│   ├── mapping_steel_mill.scd # preset 1: steel, presses, belts
│   ├── mapping_neon_glass.scd # preset 2: glass, hiss, sheen
│   └── mapping_rust_choir.scd # preset 3: drones, rust, vocals-of-metal
├── config/
│   └── timbre_map.yaml        # the real mapping source; generator writes mapping.scd
├── examples/
│   ├── input.mp4              # drop your video here
│   ├── score_example.csv      # small real-ish example (CSV)
│   └── score_example.jsonl    # tiny JSONL example (one row)
│   └── score_example_template.csv # header-only CSV template
├── processing/
│   ├── VidObjectifierProcessing.pde # webcam → blob tracker → sine choir
│   └── README.md                    # how to run and hack the gremlin
└── README.md                  # you are here

Clone the repo, throw your own media into examples/, and hack away. The *_template files are there to teach the shape of the data — swap in real material before you run anything loud.


Getting set up

On the Jetson (Nano is fine)

sudo apt update && sudo apt install -y ffmpeg python3-pip
pip3 install --upgrade pip
pip3 install ultralytics opencv-python numpy

# optional but makes the Nano run spicy hot
sudo nvpmodel -m 0 && sudo jetson_clocks

On the render machine (Linux or macOS)

  • SuperCollider — free, friendly synth language.
  • JACK2 + QjackCtl — for routing audio devices.
  • Optional: Ardour (open-source) or REAPER (generous demo) to record multichannel stems.

Running the analyzer

The script writes a newline‑timed score with spatial, color, and shape features. The score is just text; open it in a spreadsheet if that makes you smile.

cd analyzer
# Replace examples/input_template.mp4 with your own footage (or point at a
# different path) before you run this.
python3 vid2score.py ../examples/input.mp4 --out ../examples/score_example.csv --stream_id camA

# JSONL version (one JSON object per line, no header row)
python3 vid2score.py ../examples/input.mp4 --out ../examples/score_example.jsonl --stream_id camA --format jsonl

Want live video from TouchDesigner? Add a Stream Out TOP (RTSP) or NDI Out TOP in TD and point the script at the URL:

python3 vid2score.py "rtsp://<ip>:<port>/<name>" --out score_camA.csv --stream_id camA

Best practice is to analyze each pre‑mix stream for object‑level timbres, then analyze the post‑mix once to pull out macro "mood" controls.


Score Schema

Same data, two formats. CSV is a header row plus values. JSONL is one JSON object per line with the exact same keys. Pick your poison; both are loud and legible. This is the schema-as-contract — tweak it if you must, but don’t be surprised if your synth complains.

Column Type Units Range Notes
t float seconds >= 0 Timestamp since start (rounded to 0.001).
stream string n/a any Source ID you passed in (camA, camB, etc.).
oid int n/a >= 0 Tracker object ID.
cls int n/a >= 0 YOLO class index.
az float degrees -180..180 Azimuth around the listener.
el float degrees -30..30 Elevation (ish).
dist float normalized 0..1 Fake distance derived from box area.
spd float normalized/sec 0..~2 Speed in normalized screen units per second.
conf float probability 0..1 Detection confidence.
glitch float normalized 0..1 Horizontal-edge chaos meter.
hue float degrees 0..360 Average hue.
sat float normalized 0..1 Average saturation.
val float normalized 0..1 Average brightness/value.
edge float normalized 0..1 Edge density.
shape float normalized 0..1 Compactness-ish shape score.

JSONL example

{"t":0.033,"stream":"camA","oid":12,"cls":0,"az":14.2,"el":-3.1,"dist":0.742,"spd":0.02,"conf":0.91,"glitch":0.11,"hue":210.4,"sat":0.48,"val":0.62,"edge":0.33,"shape":0.29}

See examples/score_example.jsonl if you want a real file to poke at.


Running the renderer

Open SuperCollider, start JACK, and load one of:

  • renderer/render.scd — stereo, Pan2 based.
  • renderer/render_ring8.scd — 8‑channel ring without needing VBAP.

Each mapping file in renderer/ tweaks the personality of the voices. Swap in mapping_steel_mill.scd, mapping_neon_glass.scd, or mapping_rust_choir.scd for different vibes. They all respect the same 20‑voice budget.

If you want to tweak the default mapping, edit config/timbre_map.yaml and run:

python3 renderer/generate_mapping.py

That writes a fresh renderer/mapping.scd and keeps the YAML as the single source of truth. It's like tuning your synth with a wrench instead of random vibes.

Once a renderer is running, call ~playScore with the path to your CSV:

~playScore.("/full/path/to/examples/score_run.csv");

The renderer keeps a voice alive a few seconds after the last update so tails breathe instead of choking.


Processing quickie

If you're allergic to command lines or want everything in one Java file, peek at processing/VidObjectifierProcessing.pde. It's an alternate learning path, not the primary pipeline. Open it in the Processing IDE, install the Minim and OpenCV for Processing libraries, and run. It tracks moving blobs, spits out color/motion/shape stats, and turns each blob into a sine voice. Cheap, loud, educational.


Voice policy (musical & safe)

  • Global cap: ~MAX_VOICES = 20.
  • Per stream soft cap: ~PER_STREAM = 4.
  • Priority: new voices beat loud voices which beat fast voices. Something has to win.
  • Hysteresis: letting notes hang for a moment keeps things human.

TouchDesigner quick tips

  • On each TD source, add Stream Out TOP (RTSP) and give it a unique name (camA, camB, ...).
  • Run one analyzer per stream on the Jetson and set --stream_id to match.
  • To sniff the final mix, add another Stream Out TOP and analyze it into a second score (e.g. macro.csv).

Roadmap

  • Optional DeepStream backend.
  • VBAP/FOA decoders for 8‑out and binaural prints.
  • One‑shot event bus on object birth/death.
  • JSON mapping loaded at runtime instead of hard‑coding mapping.scd.
  • Simple OSC live mode (Jetson → renderer in real time).

License

MIT, because art should travel.


TL;DR quickstart

# 1) Analyze a file
cd analyzer
python3 vid2score.py ../examples/input_template.mp4 --out ../examples/score_run.csv --stream_id camA

# 2) Render it (in SuperCollider)
# open renderer/render.scd OR renderer/render_ring8.scd and run; adjust path in ~playScore

About

Video frame analysis - complex synthesis

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages