A tiny GPU watches video and scribbles a score; SuperCollider turns the scribble into a noisy choir.
This repo is a log, and it is a toolkit. It's for anyone who wants to poke at the idea of objects in a camera frame becoming musical voices. The code is kept small and loud on purpose so you can read it like liner notes.
There are two main moving parts (plus a lo-fi Processing sketch if you want to stay in Java land):
- Analyzer (
analyzer/vid2score.py)- Python script that runs on a Jetson Nano (or any CUDA Jetson).
- Looks at a video file or live stream, spots objects with YOLO, and writes a
timestamped score as
CSVorJSONL.
- Renderer (
renderer/render.scdorrenderer/render_ring8.scd)- SuperCollider patch that reads the score and gives every object a voice.
- Default budget is 20 voices, so things stay musical instead of mush.
- Processing sketch (
processing/VidObjectifierProcessing.pde)- Java/Processing rewrite that spots moving blobs with OpenCV and hands each one a sine voice.
Picture a security camera feeding a garage band.
[video file / TouchDesigner] → [Jetson: detector + characterizer] → score.csv / score.jsonl
↓
[SuperCollider renderer]
↓
(stereo / 8‑channel ring / binaural)
- Low cost, high grit. Everything here runs on free tools, save the Jetson Nano which can be had for ~$225 right now. Not my favorite situation, but when life gives you beautiful Jetsons, you run things like: SuperCollider, JACK, and a Jetson that you probably already cooked noodles on.
- Every object gets a personality. The analyzer measures position, speed, color, shape, and even a janky "glitch" metric. The renderer maps those numbers to timbre.
- Live or offline. Feed it prerecorded footage, or sling live frames from TouchDesigner over RTSP/NDI.
.
├── analyzer/
│ └── vid2score.py # video/stream → score.csv
├── renderer/
│ ├── render.scd # SuperCollider renderer (stereo)
│ ├── render_ring8.scd # 8‑channel ring renderer
│ ├── mapping.scd # generated class→timbre map (see config/timbre_map.yaml)
│ ├── mapping_steel_mill.scd # preset 1: steel, presses, belts
│ ├── mapping_neon_glass.scd # preset 2: glass, hiss, sheen
│ └── mapping_rust_choir.scd # preset 3: drones, rust, vocals-of-metal
├── config/
│ └── timbre_map.yaml # the real mapping source; generator writes mapping.scd
├── examples/
│ ├── input.mp4 # drop your video here
│ ├── score_example.csv # small real-ish example (CSV)
│ └── score_example.jsonl # tiny JSONL example (one row)
│ └── score_example_template.csv # header-only CSV template
├── processing/
│ ├── VidObjectifierProcessing.pde # webcam → blob tracker → sine choir
│ └── README.md # how to run and hack the gremlin
└── README.md # you are here
Clone the repo, throw your own media into examples/, and hack away. The
*_template files are there to teach the shape of the data — swap in real
material before you run anything loud.
sudo apt update && sudo apt install -y ffmpeg python3-pip
pip3 install --upgrade pip
pip3 install ultralytics opencv-python numpy
# optional but makes the Nano run spicy hot
sudo nvpmodel -m 0 && sudo jetson_clocks- SuperCollider — free, friendly synth language.
- JACK2 + QjackCtl — for routing audio devices.
- Optional: Ardour (open-source) or REAPER (generous demo) to record multichannel stems.
The script writes a newline‑timed score with spatial, color, and shape features. The score is just text; open it in a spreadsheet if that makes you smile.
cd analyzer
# Replace examples/input_template.mp4 with your own footage (or point at a
# different path) before you run this.
python3 vid2score.py ../examples/input.mp4 --out ../examples/score_example.csv --stream_id camA
# JSONL version (one JSON object per line, no header row)
python3 vid2score.py ../examples/input.mp4 --out ../examples/score_example.jsonl --stream_id camA --format jsonlWant live video from TouchDesigner? Add a Stream Out TOP (RTSP) or NDI Out TOP in TD and point the script at the URL:
python3 vid2score.py "rtsp://<ip>:<port>/<name>" --out score_camA.csv --stream_id camABest practice is to analyze each pre‑mix stream for object‑level timbres, then analyze the post‑mix once to pull out macro "mood" controls.
Same data, two formats. CSV is a header row plus values. JSONL is one JSON object per line with the exact same keys. Pick your poison; both are loud and legible. This is the schema-as-contract — tweak it if you must, but don’t be surprised if your synth complains.
| Column | Type | Units | Range | Notes |
|---|---|---|---|---|
t |
float | seconds | >= 0 |
Timestamp since start (rounded to 0.001). |
stream |
string | n/a | any | Source ID you passed in (camA, camB, etc.). |
oid |
int | n/a | >= 0 |
Tracker object ID. |
cls |
int | n/a | >= 0 |
YOLO class index. |
az |
float | degrees | -180..180 |
Azimuth around the listener. |
el |
float | degrees | -30..30 |
Elevation (ish). |
dist |
float | normalized | 0..1 |
Fake distance derived from box area. |
spd |
float | normalized/sec | 0..~2 |
Speed in normalized screen units per second. |
conf |
float | probability | 0..1 |
Detection confidence. |
glitch |
float | normalized | 0..1 |
Horizontal-edge chaos meter. |
hue |
float | degrees | 0..360 |
Average hue. |
sat |
float | normalized | 0..1 |
Average saturation. |
val |
float | normalized | 0..1 |
Average brightness/value. |
edge |
float | normalized | 0..1 |
Edge density. |
shape |
float | normalized | 0..1 |
Compactness-ish shape score. |
JSONL example
{"t":0.033,"stream":"camA","oid":12,"cls":0,"az":14.2,"el":-3.1,"dist":0.742,"spd":0.02,"conf":0.91,"glitch":0.11,"hue":210.4,"sat":0.48,"val":0.62,"edge":0.33,"shape":0.29}See examples/score_example.jsonl if you want a real file to poke at.
Open SuperCollider, start JACK, and load one of:
renderer/render.scd— stereo, Pan2 based.renderer/render_ring8.scd— 8‑channel ring without needing VBAP.
Each mapping file in renderer/ tweaks the personality of the voices. Swap in
mapping_steel_mill.scd, mapping_neon_glass.scd, or mapping_rust_choir.scd
for different vibes. They all respect the same 20‑voice budget.
If you want to tweak the default mapping, edit config/timbre_map.yaml and run:
python3 renderer/generate_mapping.pyThat writes a fresh renderer/mapping.scd and keeps the YAML as the single source
of truth. It's like tuning your synth with a wrench instead of random vibes.
Once a renderer is running, call ~playScore with the path to your CSV:
~playScore.("/full/path/to/examples/score_run.csv");The renderer keeps a voice alive a few seconds after the last update so tails breathe instead of choking.
If you're allergic to command lines or want everything in one Java file, peek at processing/VidObjectifierProcessing.pde. It's an alternate learning path, not the primary pipeline. Open it in the Processing IDE, install the Minim and OpenCV for Processing libraries, and run. It tracks moving blobs, spits out color/motion/shape stats, and turns each blob into a sine voice. Cheap, loud, educational.
- Global cap:
~MAX_VOICES = 20. - Per stream soft cap:
~PER_STREAM = 4. - Priority: new voices beat loud voices which beat fast voices. Something has to win.
- Hysteresis: letting notes hang for a moment keeps things human.
- On each TD source, add Stream Out TOP (RTSP) and give it a unique name
(
camA,camB, ...). - Run one analyzer per stream on the Jetson and set
--stream_idto match. - To sniff the final mix, add another Stream Out TOP and analyze it into a
second score (e.g.
macro.csv).
- Optional DeepStream backend.
- VBAP/FOA decoders for 8‑out and binaural prints.
- One‑shot event bus on object birth/death.
- JSON mapping loaded at runtime instead of hard‑coding
mapping.scd. - Simple OSC live mode (Jetson → renderer in real time).
MIT, because art should travel.
# 1) Analyze a file
cd analyzer
python3 vid2score.py ../examples/input_template.mp4 --out ../examples/score_run.csv --stream_id camA
# 2) Render it (in SuperCollider)
# open renderer/render.scd OR renderer/render_ring8.scd and run; adjust path in ~playScore