PL BERT
yl4579
Phoneme-Level BERT for Enhanced Prosody of Text-to-Speech with Grapheme Predictions
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
Page 40 · 16 shown · 743 public entries
Results: 743
yl4579
Phoneme-Level BERT for Enhanced Prosody of Text-to-Speech with Grapheme Predictions
zenstory-ai
Direct game art and creative vision. Turn GAME_DESIGN into a production-level ART_DIRECTION defining a recognizable visual style, camera and composition, world and chara…
eosphoros-ai
johnGettings
Long-Inference, High Quality Synthetic Speaker (AI avatar/ AI presenter)
davidmartinrius
🔊 Create labeled datasets, enhance audio quality, identify speakers, support diverse dataset types. 🎧👥📊 Advanced audio processing.
NTT123
Vietnamese Text to Speech library
keonlee9420
Official repository of DailyTalk: Spoken Dialogue Dataset for Conversational Text-to-Speech, ICASSP 2023
ammaarreshi
Open-source clone of the MidJourney web interface featuring real AI image and video generation powered by Google's Gemini SDK. Use Imagen 4 to generate images and Veo 2…
KevinMIN95
Official implementation of Meta-StyleSpeech and StyleSpeech
yl4579
HiFTNet: A Fast High-Quality Neural Vocoder with Harmonic-plus-Noise Filter and Inverse Short Time Fourier Transform
jxzhanggg
Implementation code of non-parallel sequence-to-sequence VC
keonlee9420
PyTorch implementation of DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism (focused on DiffSpeech)
robinhad
Ukrainian TTS (text-to-speech) using ESPNET
Hagsten
Javascript Text to speech library
rishikksh20
PyTorch Implementation of FastSpeech 2 : Fast and High-Quality End-to-End Text to Speech
calesthio
Provider-independent production workflow for AI avatar spokesperson videos, including presenter briefs, consent and likeness rights, disclosure and platform labeling, ca…