Skip to content

Latest commit

 

History

15 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

MeetIngest

🎙️ Attributed Meeting Data Capture Bot

MeetIngest is a headless meeting bot and data capture layer that joins live Google Meet calls, taps each participant's WebRTC audio and video streams, identifies who every stream belongs to, and produces an isolated, audio-video synced recording per participant, along with a timeline of when each participant joined and left.

It is built as the ingestion layer beneath an orchestration system. Mixed meeting audio makes it impossible to tell who said what and where it was said. MeetIngest instead delivers attributed, per-participant data so that the layer above can extract accurate, individual-level insights and locate exactly where each piece of information was shared.


Meeting Replay & Final Output

After a meeting ends, MeetIngest automatically publishes the recordings and opens a replay page. All participant videos play together on one master clock in a Meet-style grid, and tiles animate in and out exactly when participants joined and left. The complete replay can be exported as a single combined .mp4 as proof of the captured meeting.



    


Per-Participant Recordings

Every participant gets their own .mp4 containing only their video and only their voice, with a black frame and name label whenever their camera is off, so each file stays continuous and correctly attributed from join to leave.

Meeting Timeline

A timeline.json and manifest.json record each participant's identity, recording start offset and output file, giving downstream systems the exact position of every participant's data within the meeting.


📑 Table of Contents


✨ Features

🤖 Automated Meeting Bot

  • Joins Google Meet calls through a Playwright-controlled Chromium browser
  • Persistent browser profile for signed-in sessions
  • Hooks WebRTC peer connections to access live media tracks
  • Runs without manual interaction during the meeting

🧭 Participant Identification

  • Video streams mapped to participants through DOM SSRC attributes
  • Audio streams mapped through active-speaker correlation
  • Participant names resolved and kept updated from the Meet UI
  • Screen-share presentations excluded from participant recordings

🎥 Per-Participant Recording

  • One dedicated recorder per participant
  • Audio and video recorded together in one container through MediaRecorder
  • Canvas compositor keeps video continuous when the camera is off
  • Recording chunks streamed to disk over a local WebSocket

🔒 Correct Audio-Video Attribution

  • 4-strategy bonding filter to accept correct matches and reject noise
  • Gate windows for overlapping speech and ambiguous tile claims
  • Strength test, dual confirmation and exclusive track bonding

🔁 Reliability

  • Two-signal detection of mid-meeting WebRTC re-negotiation
  • Dual-recorder hand-off with no recording gap
  • Segment-based recording stitched into one final file per participant
  • Fault-isolated stitching: one participant's failure never affects another

🛑 Safe Shutdown

  • 3-signal traffic light braking system (RED → YELLOW → GREEN)
  • All files finalized before the bot leaves the call
  • Shutdown in 5-25 seconds regardless of meeting load

📺 Replay & Export

  • Automatic publishing to Cloudinary after the meeting
  • Meet-style synced replay grid with animated join/leave tiles
  • Master-clock playback with drift correction
  • One-click export of the full replay as a combined .mp4

🧠 How It Works

Bot Joins Meeting
       ↓
WebRTC Track Hooking
       ↓
Participant Identification
       ↓
Video-Only Recorder Starts
       ↓
Audio-Video Bonding
       ↓
Dual-Recorder Hand-Off (Audio + Video)
       ↓
Chunk Streaming over WebSocket
       ↓
Re-Negotiation Handling
       ↓
Traffic Light Braking
       ↓
FFmpeg Normalize + Stitch
       ↓
Per-Participant .mp4 + Timeline
       ↓
Publish + Replay + Combined Export

🏗️ System Design

MeetIngest is split into two cooperating runtimes. Inside the browser, injected scripts hook Google Meet's WebRTC connections, identify participants and run one recorder per participant. In Node.js, a local WebSocket server receives the recording chunks, writes them to disk, coordinates shutdown and hands the finished segments to FFmpeg.

Heavy post-processing (upload, publishing, replay) runs in a separate process after the bot exits, so it never adds load to live recording.

                      ┌──────────────────────┐
                      │   Google Meet Call   │
                      └──────────┬───────────┘
                                 │
                                 ▼
                      ┌──────────────────────┐
                      │ Playwright Chromium  │
                      │  (WebRTC Hooks)      │
                      └──────────┬───────────┘
                                 │
                   ┌─────────────┴─────────────┐
                   │                           │
                   ▼                           ▼
          ┌─────────────────┐        ┌─────────────────┐
          │ Video Mapping   │        │ Audio Mapping   │
          │ (DOM SSRC)      │        │ (Active Speaker)│
          └────────┬────────┘        └────────┬────────┘
                   │                          │
                   └────────────┬─────────────┘
                                │
                                ▼
                      ┌──────────────────────┐
                      │ 4-Strategy A/V Bond  │
                      └──────────┬───────────┘
                                 │
                                 ▼
                      ┌──────────────────────┐
                      │ Per-Participant      │
                      │ MediaRecorders       │
                      └──────────┬───────────┘
                                 │
                                 ▼
                      ┌──────────────────────┐
                      │ Local WebSocket      │
                      │ Server (Node.js)     │
                      └──────────┬───────────┘
                                 │
                                 ▼
                      ┌──────────────────────┐
                      │ .webm Segments       │
                      │ + .meta.json         │
                      └──────────┬───────────┘
                                 │
                                 ▼
                      ┌──────────────────────┐
                      │ FFmpeg Stitching     │
                      └──────────┬───────────┘
                                 │
                                 ▼
                      ┌──────────────────────┐
                      │ Per-Participant .mp4 │
                      │ + timeline.json      │
                      └──────────┬───────────┘
                                 │
                                 ▼
                      ┌──────────────────────┐
                      │ Publisher + Replay   │
                      └──────────────────────┘

📁 Project Structure

MeetIngest/
│
├── src/
│   ├── bot.js              # Meeting join, WebSocket server, orchestration, braking
│   ├── recorder.js         # In-browser per-participant recorders
│   ├── mapper.js           # Participant ↔ stream identity mapping
│   ├── stitch.js           # FFmpeg segment normalization and stitching
│   ├── timeline.js         # Participant start offsets and output mapping
│   ├── publish.js          # Upload, manifest, local server, recording save
│   │
│   ├── player/
│   │   └── player.html     # Meet-style synced replay and combined export
│   │
│   ├── recordings/         # .webm segments + .meta.json (runtime)
│   ├── finalRecorded/      # Per-participant .mp4 + timeline/manifest (runtime)
│   └── finalOutput/        # Combined meeting .mp4 exports (runtime)
│
├── start-bot.ps1           # Launcher with Ctrl+C stop signal
├── package.json
├── .env                    # Cloudinary credentials (not committed)
└── README.md

Note: Runtime-generated directories — such as the browser profile, recordings, stitched outputs, combined exports and node_modules — are intentionally excluded from version control.


🛠️ Tech Stack

Category Technologies
Runtime Node.js
Browser Automation Playwright, playwright-extra, Stealth plugin
Real-Time Media WebRTC, MediaRecorder API, Canvas API, Web Audio API
Transport WebSockets (ws)
Media Processing FFmpeg (ffmpeg-static), H.264, AAC
Cloud Storage Cloudinary
Replay Frontend HTML, CSS, Vanilla JavaScript
Launcher PowerShell

🧭 Participant Identification

Google Meet delivers every stream as an anonymous WebRTC track. MeetIngest resolves identity separately for video and audio.

1. 🎥 Active Video Correlation (SSRC)

Each participant tile in the Meet UI carries a data-ssrc attribute matching the SSRC of that participant's incoming video stream. The bot reads it and binds the video track to the participant's name.

Participant Tile (data-ssrc)
          ↓
Matching Video Track SSRC
          ↓
Video ↔ Participant Bound

2. 🔊 Active Speaker Correlation

Audio tracks carry no DOM identifier. When someone speaks, Meet lights up the speaking indicator on their tile, and at the same moment an audio track's level spikes. The bot correlates the two events within a tight time window to link the audio track to that participant.

Speaking Indicator Change   +   Audio Level Spike
                  ↓
        Time-Window Correlation
                  ↓
        Audio ↔ Participant Candidate

⚙️ Recording Pipeline

1. 🎬 Video-Only Recorder

As soon as a participant's video is identified, a video-only recorder starts, so their footage is captured from the moment they join.

2. 🖼️ Canvas Compositor

Each participant's video is drawn onto a dedicated canvas. When frames arrive, the live camera is shown; when the camera is off, a black frame with the participant's name is drawn instead. The video stream therefore never has gaps.

3. 🎙️ Audio-Video Recorder

Once the participant's audio is bonded, recording switches to a combined audio + video MediaRecorder. Audio and video share one container and one clock, so they are synced at capture time instead of being aligned later.

4. 📡 Chunk Streaming

Recorders emit chunks that are streamed to Node.js over a local WebSocket (127.0.0.1) and appended to segment files on disk, each with a .meta.json describing the segment.

Participant Media
       ↓
Canvas Compositor
       ↓
MediaRecorder
       ↓
WebSocket Chunks
       ↓
<participant>__segN.webm + .meta.json

🔒 Audio-Video Bonding

Overlapping speech, background noise and flickering speaking indicators can cause one participant's audio to be matched to another participant's video. Every candidate match passes through a 4-strategy filter.

Strategy Rule Purpose
1. A/V Gate Window — Case 1 More than one audio track spiking within 300 ms → skip Rejects overlapping speech
1. A/V Gate Window — Case 2 More than one tile claiming the same track within 150 ms → skip Rejects ambiguous claims
2. Strength Test Correlation gap ≤ 120 ms and audio level ≥ 0.1 Accepts only tight, clear matches
3. Dual Confirmation Same pairing repeats 200 ms later with no rival claims Rejects one-off coincidences
4. A/V Bond Confirmed track is locked to the participant Prevents re-claiming by others
Candidate Match
      ↓
Gate Window (Case 1 & 2)
      ↓
Strength Test
      ↓
Dual Confirmation
      ↓
A/V Bond

🔁 Mid-Meeting Re-Negotiation

In a significant share of meetings (~40% in testing), Google Meet silently re-negotiates the connection mid-call and delivers participants' media on new tracks, which would otherwise freeze the recording.

Two-Signal Detection

  • Signal 1: a new peer connection appears and is registered
  • Signal 2: that connection delivers a track for a participant already being recorded on an older connection, confirming a real re-negotiation (not a new joiner or screen share)

Dual-Recorder Hand-Off

The new recorder starts before the old one stops, producing a new segment with no gap in coverage. The same hand-off is used when switching from the video-only recorder to the audio + video recorder.

Old Recorder (segN)  ──────────┐
                               │  overlap
New Recorder (segN+1)    ──────┴──────────▶

🚦 Traffic Light Braking

Under heavy meeting load, the browser and Node.js event loop can become saturated, and leaving the call too early destroys recordings that are not yet saved. Shutdown is therefore split into three independent phases.

Signal Duration Action
🔴 RED ~4 s Accept in-flight chunks (Node.js only, no browser dependency)
🟡 YELLOW ~2 s Force-close sockets and finalize every file on disk
🟢 GREEN — Leave the call only after all files are safe, then stitch

Shutdown dropped from up to ~8 minutes to 5-25 seconds, with no lost recordings.


🎞️ Stitching & Publishing

1. 🧩 Normalize

Each .webm segment is normalized to H.264 video (640×360, 15 fps) and AAC audio. Segments without audio receive a silent track so every segment has the same structure. Normalization runs in parallel using up to CPU cores − 1 FFmpeg processes.

2. 🔗 Concatenate

Normalized segments are concatenated in order into one continuous .mp4 per participant in finalRecorded/. Every task is individually fault-isolated, so a failed segment or participant never blocks the others.

3. 🗓️ Timeline

timeline.json records the meeting start and each participant's recording start and output file, giving every participant's exact offset within the meeting.

4. ☁️ Publish

After the bot has fully exited, the publisher runs in its own window. It uploads the recordings to Cloudinary, writes manifest.json, starts a local server and opens the replay page. If an upload fails, that file is served locally instead.

5. 📺 Replay & Export

The replay page plays all participants on a single master clock with drift correction, in a Meet-style grid where tiles animate in and out at each participant's join and leave times. The Record button replays the meeting and saves a combined 720p / 30 fps .mp4 to finalOutput/.

.webm Segments
      ↓
Normalize (parallel)
      ↓
Concatenate
      ↓
Per-Participant .mp4
      ↓
Cloudinary Upload
      ↓
Replay Grid
      ↓
Combined .mp4 Export

🗃️ Storage Architecture

src/recordings/

  • Raw .webm segments per participant
  • .meta.json per segment (e.g. audio presence)

src/finalRecorded/

  • Stitched per-participant .mp4 files
  • timeline.json — participant start offsets and output files
  • manifest.json — published video URLs and offsets

src/finalOutput/

  • Combined meeting recordings, timestamped per export

Cloudinary

  • Hosted per-participant videos used by the replay page

⚠️ Edge Cases & Reliability

MeetIngest is designed to handle:

  • Participants joining mid-meeting
  • Participants leaving early
  • Camera turned off (continuous black frame with name)
  • Muted or silent participants (silent audio track)
  • Overlapping speakers and background noise
  • Mid-meeting WebRTC re-negotiation
  • Screen-share presentations (excluded from participant files)
  • Failed segment normalization or concatenation
  • Unresponsive browser during shutdown
  • Failed cloud uploads (local playback fallback)
  • Browser autoplay restrictions on the replay page

🚀 Setup

1. Clone the Repository

git clone <your-repo>
cd MeetIngest

2. Install Dependencies

npm install

3. Configure Environment

Create a .env file in the project root:

CLOUDINARY_CLOUD_NAME=your_cloud_name
CLOUDINARY_API_KEY=your_api_key
CLOUDINARY_API_SECRET=your_api_secret

4. Start the Bot

Windows (PowerShell)

.\start-bot.ps1

5. Stop the Bot

Press Ctrl+C. The launcher creates a stop signal, the bot runs the traffic light shutdown, stitches the recordings and opens the replay page automatically.


🎯 Current Status

  • ✅ Automated Google Meet joining
  • ✅ WebRTC track hooking
  • ✅ Video identification via DOM SSRC
  • ✅ Audio identification via active-speaker correlation
  • ✅ Per-participant canvas + MediaRecorder recording
  • ✅ Audio-video sync at capture time
  • ✅ WebSocket chunk streaming
  • ✅ 4-strategy audio-video bonding
  • ✅ Two-signal re-negotiation detection
  • ✅ Dual-recorder hand-off
  • ✅ Traffic light braking
  • ✅ Parallel FFmpeg stitching
  • ✅ Timeline and manifest generation
  • ✅ Cloudinary publishing
  • ✅ Meet-style synced replay
  • ✅ Combined meeting .mp4 export
  • ✅ Validated in live 4-participant meetings
  • 🚧 SFU audio-slot shuffling handling (switchboard technique, in progress)

🔮 Future Improvements

Potential areas for further development include:

  • Switchboard-based routing for SFU audio-slot reassignment
  • Screen-share capture as a separate attributed stream
  • Per-participant transcription with speaker attribution
  • Direct integration with the orchestration layer
  • Support for additional meeting platforms
  • Automated session folder management
  • Containerized headless deployment
  • Horizontal scaling for concurrent meetings
  • Improved observability and recording health metrics

🧠 Design Principles

  • Attribution over aggregation
  • Sync at capture, not after
  • Zero-gap recording
  • Accept correct, reject ambiguous
  • Fault isolation per participant
  • Data safety before exit
  • No extra load on the live bot
  • Modular components
  • Separation of capture and post-processing
  • Graceful fallbacks

👤 Author

Ravi Sharma Full Stack Developer / AI Engineer

About

MeetIngest is a headless meeting bot and data capture layer that joins live Google Meet calls, taps each participant's WebRTC audio and video streams, identifies who every stream belongs to, and produces an isolated, audio-video synced recording per participant, along with a timeline of when each participant joined and left.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages