WavTokenizer
jishengpeng
[ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language modeling
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
241–256 / 478
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 478
jishengpeng
[ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language modeling
Uminosachi
Inpaint Anything extension performs stable diffusion inpainting on a browser UI using masks from Segment Anything.
snap-research
Code for Motion Representations for Articulated Animation paper
ictnlp
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
huanngzh
[ICCV 2025] Official impl. of "MV-Adapter: Multi-view Consistent Image Generation Made Easy"
sdkcarlos
A voice control - voice commands - speech recognition and speech synthesis javascript library. Create your own siri,google now or cortana with Google Chrome within your…
alphacep
WebSocket, gRPC and WebRTC speech recognition server based on Vosk and Kaldi libraries
gitmylo
A webui for different audio related Neural Networks
mravanelli
SincNet is a neural architecture for efficiently processing raw audio samples.
ali-vilab
Code for SCIS-2025 Paper "UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation".
cure-lab
[ICLR24] Official implementation of the paper “MagicDrive: Street View Generation with Diverse 3D Geometry Control”
ermongroup
PyTorch implementation for SDEdit: Image Synthesis and Editing with Stochastic Differential Equations
shadow2496
Official PyTorch implementation of "VITON-HD: High-Resolution Virtual Try-On via Misalignment-Aware Normalization" (CVPR 2021)
zai-org
CogView4, CogView3-Plus and CogView3(ECCV 2024)
alumae
Real-time full-duplex speech recognition server, based on the Kaldi toolkit and the GStreamer framwork.