MeetIngest is a headless meeting bot and data capture layer that joins live Google Meet calls, taps each participant's WebRTC audio and video streams, identifies who every stream belongs to, and produces an isolated, audio-video synced recording per participant, along with a timeline of when each participant joined and left.
It is built as the ingestion layer beneath an orchestration system. Mixed meeting audio makes it impossible to tell who said what and where it was said. MeetIngest instead delivers attributed, per-participant data so that the layer above can extract accurate, individual-level insights and locate exactly where each piece of information was shared.
Meeting Replay & Final Output
After a meeting ends, MeetIngest automatically publishes the recordings and opens a replay page. All participant videos play together on one master clock in a Meet-style grid, and tiles animate in and out exactly when participants joined and left. The complete replay can be exported as a single combined .mp4 as proof of the captured meeting.
Per-Participant Recordings
Every participant gets their own .mp4 containing only their video and only their voice, with a black frame and name label whenever their camera is off, so each file stays continuous and correctly attributed from join to leave.
Meeting Timeline
A timeline.json and manifest.json record each participant's identity, recording start offset and output file, giving downstream systems the exact position of every participant's data within the meeting.
- Features
- How It Works
- System Design
- Project Structure
- Tech Stack
- Participant Identification
- Recording Pipeline
- Audio-Video Bonding
- Mid-Meeting Re-Negotiation
- Traffic Light Braking
- Stitching & Publishing
- Storage Architecture
- Edge Cases & Reliability
- Setup
- Current Status
- Future Improvements
- Design Principles
- Author
- Joins Google Meet calls through a Playwright-controlled Chromium browser
- Persistent browser profile for signed-in sessions
- Hooks WebRTC peer connections to access live media tracks
- Runs without manual interaction during the meeting
- Video streams mapped to participants through DOM SSRC attributes
- Audio streams mapped through active-speaker correlation
- Participant names resolved and kept updated from the Meet UI
- Screen-share presentations excluded from participant recordings
- One dedicated recorder per participant
- Audio and video recorded together in one container through MediaRecorder
- Canvas compositor keeps video continuous when the camera is off
- Recording chunks streamed to disk over a local WebSocket
- 4-strategy bonding filter to accept correct matches and reject noise
- Gate windows for overlapping speech and ambiguous tile claims
- Strength test, dual confirmation and exclusive track bonding
- Two-signal detection of mid-meeting WebRTC re-negotiation
- Dual-recorder hand-off with no recording gap
- Segment-based recording stitched into one final file per participant
- Fault-isolated stitching: one participant's failure never affects another
- 3-signal traffic light braking system (RED → YELLOW → GREEN)
- All files finalized before the bot leaves the call
- Shutdown in 5-25 seconds regardless of meeting load
- Automatic publishing to Cloudinary after the meeting
- Meet-style synced replay grid with animated join/leave tiles
- Master-clock playback with drift correction
- One-click export of the full replay as a combined
.mp4
Bot Joins Meeting
↓
WebRTC Track Hooking
↓
Participant Identification
↓
Video-Only Recorder Starts
↓
Audio-Video Bonding
↓
Dual-Recorder Hand-Off (Audio + Video)
↓
Chunk Streaming over WebSocket
↓
Re-Negotiation Handling
↓
Traffic Light Braking
↓
FFmpeg Normalize + Stitch
↓
Per-Participant .mp4 + Timeline
↓
Publish + Replay + Combined Export
MeetIngest is split into two cooperating runtimes. Inside the browser, injected scripts hook Google Meet's WebRTC connections, identify participants and run one recorder per participant. In Node.js, a local WebSocket server receives the recording chunks, writes them to disk, coordinates shutdown and hands the finished segments to FFmpeg.
Heavy post-processing (upload, publishing, replay) runs in a separate process after the bot exits, so it never adds load to live recording.
┌──────────────────────┐
│ Google Meet Call │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Playwright Chromium │
│ (WebRTC Hooks) │
└──────────┬───────────┘
│
┌─────────────┴─────────────┐
│ │
▼ ▼
┌─────────────────┐ ┌─────────────────┐
│ Video Mapping │ │ Audio Mapping │
│ (DOM SSRC) │ │ (Active Speaker)│
└────────┬────────┘ └────────┬────────┘
│ │
└────────────┬─────────────┘
│
▼
┌──────────────────────┐
│ 4-Strategy A/V Bond │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Per-Participant │
│ MediaRecorders │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Local WebSocket │
│ Server (Node.js) │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ .webm Segments │
│ + .meta.json │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ FFmpeg Stitching │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Per-Participant .mp4 │
│ + timeline.json │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Publisher + Replay │
└──────────────────────┘
MeetIngest/
│
├── src/
│ ├── bot.js # Meeting join, WebSocket server, orchestration, braking
│ ├── recorder.js # In-browser per-participant recorders
│ ├── mapper.js # Participant ↔ stream identity mapping
│ ├── stitch.js # FFmpeg segment normalization and stitching
│ ├── timeline.js # Participant start offsets and output mapping
│ ├── publish.js # Upload, manifest, local server, recording save
│ │
│ ├── player/
│ │ └── player.html # Meet-style synced replay and combined export
│ │
│ ├── recordings/ # .webm segments + .meta.json (runtime)
│ ├── finalRecorded/ # Per-participant .mp4 + timeline/manifest (runtime)
│ └── finalOutput/ # Combined meeting .mp4 exports (runtime)
│
├── start-bot.ps1 # Launcher with Ctrl+C stop signal
├── package.json
├── .env # Cloudinary credentials (not committed)
└── README.md
Note: Runtime-generated directories — such as the browser profile, recordings, stitched outputs, combined exports and
node_modules— are intentionally excluded from version control.
| Category | Technologies |
|---|---|
| Runtime | Node.js |
| Browser Automation | Playwright, playwright-extra, Stealth plugin |
| Real-Time Media | WebRTC, MediaRecorder API, Canvas API, Web Audio API |
| Transport | WebSockets (ws) |
| Media Processing | FFmpeg (ffmpeg-static), H.264, AAC |
| Cloud Storage | Cloudinary |
| Replay Frontend | HTML, CSS, Vanilla JavaScript |
| Launcher | PowerShell |
Google Meet delivers every stream as an anonymous WebRTC track. MeetIngest resolves identity separately for video and audio.
Each participant tile in the Meet UI carries a data-ssrc attribute matching the SSRC of that participant's incoming video stream. The bot reads it and binds the video track to the participant's name.
Participant Tile (data-ssrc)
↓
Matching Video Track SSRC
↓
Video ↔ Participant Bound
Audio tracks carry no DOM identifier. When someone speaks, Meet lights up the speaking indicator on their tile, and at the same moment an audio track's level spikes. The bot correlates the two events within a tight time window to link the audio track to that participant.
Speaking Indicator Change + Audio Level Spike
↓
Time-Window Correlation
↓
Audio ↔ Participant Candidate
As soon as a participant's video is identified, a video-only recorder starts, so their footage is captured from the moment they join.
Each participant's video is drawn onto a dedicated canvas. When frames arrive, the live camera is shown; when the camera is off, a black frame with the participant's name is drawn instead. The video stream therefore never has gaps.
Once the participant's audio is bonded, recording switches to a combined audio + video MediaRecorder. Audio and video share one container and one clock, so they are synced at capture time instead of being aligned later.
Recorders emit chunks that are streamed to Node.js over a local WebSocket (127.0.0.1) and appended to segment files on disk, each with a .meta.json describing the segment.
Participant Media
↓
Canvas Compositor
↓
MediaRecorder
↓
WebSocket Chunks
↓
<participant>__segN.webm + .meta.json
Overlapping speech, background noise and flickering speaking indicators can cause one participant's audio to be matched to another participant's video. Every candidate match passes through a 4-strategy filter.
| Strategy | Rule | Purpose |
|---|---|---|
| 1. A/V Gate Window — Case 1 | More than one audio track spiking within 300 ms → skip | Rejects overlapping speech |
| 1. A/V Gate Window — Case 2 | More than one tile claiming the same track within 150 ms → skip | Rejects ambiguous claims |
| 2. Strength Test | Correlation gap ≤ 120 ms and audio level ≥ 0.1 | Accepts only tight, clear matches |
| 3. Dual Confirmation | Same pairing repeats 200 ms later with no rival claims | Rejects one-off coincidences |
| 4. A/V Bond | Confirmed track is locked to the participant | Prevents re-claiming by others |
Candidate Match
↓
Gate Window (Case 1 & 2)
↓
Strength Test
↓
Dual Confirmation
↓
A/V Bond
In a significant share of meetings (~40% in testing), Google Meet silently re-negotiates the connection mid-call and delivers participants' media on new tracks, which would otherwise freeze the recording.
- Signal 1: a new peer connection appears and is registered
- Signal 2: that connection delivers a track for a participant already being recorded on an older connection, confirming a real re-negotiation (not a new joiner or screen share)
The new recorder starts before the old one stops, producing a new segment with no gap in coverage. The same hand-off is used when switching from the video-only recorder to the audio + video recorder.
Old Recorder (segN) ──────────┐
│ overlap
New Recorder (segN+1) ──────┴──────────▶
Under heavy meeting load, the browser and Node.js event loop can become saturated, and leaving the call too early destroys recordings that are not yet saved. Shutdown is therefore split into three independent phases.
| Signal | Duration | Action |
|---|---|---|
| 🔴 RED | ~4 s | Accept in-flight chunks (Node.js only, no browser dependency) |
| 🟡 YELLOW | ~2 s | Force-close sockets and finalize every file on disk |
| 🟢 GREEN | — | Leave the call only after all files are safe, then stitch |
Shutdown dropped from up to ~8 minutes to 5-25 seconds, with no lost recordings.
Each .webm segment is normalized to H.264 video (640×360, 15 fps) and AAC audio. Segments without audio receive a silent track so every segment has the same structure. Normalization runs in parallel using up to CPU cores − 1 FFmpeg processes.
Normalized segments are concatenated in order into one continuous .mp4 per participant in finalRecorded/. Every task is individually fault-isolated, so a failed segment or participant never blocks the others.
timeline.json records the meeting start and each participant's recording start and output file, giving every participant's exact offset within the meeting.
After the bot has fully exited, the publisher runs in its own window. It uploads the recordings to Cloudinary, writes manifest.json, starts a local server and opens the replay page. If an upload fails, that file is served locally instead.
The replay page plays all participants on a single master clock with drift correction, in a Meet-style grid where tiles animate in and out at each participant's join and leave times. The Record button replays the meeting and saves a combined 720p / 30 fps .mp4 to finalOutput/.
.webm Segments
↓
Normalize (parallel)
↓
Concatenate
↓
Per-Participant .mp4
↓
Cloudinary Upload
↓
Replay Grid
↓
Combined .mp4 Export
- Raw
.webmsegments per participant .meta.jsonper segment (e.g. audio presence)
- Stitched per-participant
.mp4files timeline.json— participant start offsets and output filesmanifest.json— published video URLs and offsets
- Combined meeting recordings, timestamped per export
- Hosted per-participant videos used by the replay page
MeetIngest is designed to handle:
- Participants joining mid-meeting
- Participants leaving early
- Camera turned off (continuous black frame with name)
- Muted or silent participants (silent audio track)
- Overlapping speakers and background noise
- Mid-meeting WebRTC re-negotiation
- Screen-share presentations (excluded from participant files)
- Failed segment normalization or concatenation
- Unresponsive browser during shutdown
- Failed cloud uploads (local playback fallback)
- Browser autoplay restrictions on the replay page
git clone <your-repo>
cd MeetIngestnpm installCreate a .env file in the project root:
CLOUDINARY_CLOUD_NAME=your_cloud_name
CLOUDINARY_API_KEY=your_api_key
CLOUDINARY_API_SECRET=your_api_secretWindows (PowerShell)
.\start-bot.ps1Press Ctrl+C. The launcher creates a stop signal, the bot runs the traffic light shutdown, stitches the recordings and opens the replay page automatically.
- ✅ Automated Google Meet joining
- ✅ WebRTC track hooking
- ✅ Video identification via DOM SSRC
- ✅ Audio identification via active-speaker correlation
- ✅ Per-participant canvas + MediaRecorder recording
- ✅ Audio-video sync at capture time
- ✅ WebSocket chunk streaming
- ✅ 4-strategy audio-video bonding
- ✅ Two-signal re-negotiation detection
- ✅ Dual-recorder hand-off
- ✅ Traffic light braking
- ✅ Parallel FFmpeg stitching
- ✅ Timeline and manifest generation
- ✅ Cloudinary publishing
- ✅ Meet-style synced replay
- ✅ Combined meeting
.mp4export - ✅ Validated in live 4-participant meetings
- 🚧 SFU audio-slot shuffling handling (switchboard technique, in progress)
Potential areas for further development include:
- Switchboard-based routing for SFU audio-slot reassignment
- Screen-share capture as a separate attributed stream
- Per-participant transcription with speaker attribution
- Direct integration with the orchestration layer
- Support for additional meeting platforms
- Automated session folder management
- Containerized headless deployment
- Horizontal scaling for concurrent meetings
- Improved observability and recording health metrics
- Attribution over aggregation
- Sync at capture, not after
- Zero-gap recording
- Accept correct, reject ambiguous
- Fault isolation per participant
- Data safety before exit
- No extra load on the live bot
- Modular components
- Separation of capture and post-processing
- Graceful fallbacks
Ravi Sharma Full Stack Developer / AI Engineer


