Unified-Modal Speech-Text Pre-Training for Spoken Language Processing
-
Updated
Apr 24, 2024 - Python
Unified-Modal Speech-Text Pre-Training for Spoken Language Processing
Real time audio to audio translation over sockets. With virtual microphones, you can use this in any video conferencing software you'd like!
🔊😊 A fastapi voice-assistant framework to quickly prototype LLM-powered voice assistants in <5 minutes.
Fine-tune SpeechT5 for non-English text-to-speech task, implemented in PyTorch.
Text-to-speech, expressive speech synthesis, and text-to-music generation using SpeechT5, Bark, and MusicGen.
This repository contains any code related to audio generated with artificial intelligence, such as voice cloning, text-to-speech, audio classification, speech recognition, etc.
Zero-shot voice cloning and voice conversion - SpeechBrain x-vector speaker embeddings conditioning microsoft/speecht5_tts, with Whisper transcription. Gradio apps included.
Different Task Guides for Audio Data
To associate your repository with the speecht5 topic, visit your repo's landing page and select "manage topics."