Skip to content
@MM-Speech

MM-Speech

Welcome to MM-Speech 👋

Our lab is committed to cutting-edge research in speech generation, spoken dialogue systems, and spatial audio generation. We strive to develop intelligent, natural, and immersive audio technologies that advance human–machine interaction and multimedia experiences.

Popular repositories Loading

  1. SwanSphere SwanSphere Public

    [ICML 2026] Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer

    Python 120

  2. DuplexSurvey DuplexSurvey Public

    [EMNLP 2026 Main] Speaking While Listening: a survey and empirical audit of full-duplex spoken dialogue systems — L0–L3 architectural hierarchy, T×I×R interaction ontology, and a five-state decisio…

    73 1

  3. VoxMind VoxMind Public

    [ACL 2026] VoxMind: An End-to-End Agentic Spoken Dialogue System

    Python 40 4

  4. DualAxisRM DualAxisRM Public

    [ACL 2026] Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models

    Python 22 1

  5. EMO-TTS EMO-TTS Public

    [ACL 2026] Rectifying the Emotional Flow: Aligning Priors and Dynamic Guidance for High-Arousal Text-to-Speech

    Python 15

  6. DiTReducio DiTReducio Public

    [ACL 2026] DiTReducio: A Training-Free Acceleration for DiT-Based TTS viaProgressive Calibration

    Python 14 2

Repositories

Showing 10 of 19 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…