AuK: An Open-Source Foundational Model for Speech Generation and Editing
-
Updated
Sep 25, 2026 - Python
AuK: An Open-Source Foundational Model for Speech Generation and Editing
Native AuK speech generation and editing nodes for ComfyUI V3
Code and samples for the paper "Acoustic token admixture for joint speaker and content anonymization"
Phone-level speech inpainting: swap or repair individual phones in speech audio
Orchestration layer for AuK speech editing: multi-intent planning, context-budget-aware chunking with seamless stitching, and BS.1770 quality gates that verify each edit. Pure-numpy core, no GPU needed to run the tests.
Zero-shot speech editing and TTS using neural codec language models
Run Tencent AuK (1.5B speech generation & editing foundation model) with the Qwen2.5-Omni encoder on Kaggle's free 2xT4 GPUs - 24 playable before/after audio examples, weights served from public Kaggle datasets
wavepainter: multimodal LLM-guided diffusion for text-based speech editing — edit a word in the transcript and only that span of the spectrogram is re-predicted
Tencent Just Built a 1.5B AI That Can Generate AND Edit Audio! (AuK-Flash Setup) - Open-source 1.5B audio generation and editing foundation model setup, PyTorch pipeline, and 4-step fast distillation guide.
[剪辑工具·达芬奇] 达芬奇AI初剪工具|本地检测静音和换气并生成口播粗剪时间线(MIT 协议)
To associate your repository with the speech-editing topic, visit your repo's landing page and select "manage topics."