Sana
NVlabs
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–7 / 7
Results: 7
NVlabs
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
fudan-generative-vision
[ECCV 2024] Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance
fikrikarim
On-device, real-time multimodal AI. Have natural voice and vision conversations with an AI that runs entirely on your machine. Powered by Gemma 4 E2B and Kokoro.
ARahim3
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
byjlw
Analyze videos using LLMs, Computer Vision and Automatic Speech Recognition
Francis-Rings
We present StableAvatar, the first end-to-end video diffusion transformer, which synthesizes infinite-length high-quality audio-driven avatar videos without any post-pro…
sooftware
[Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition" (INTERSPEECH 2020)
Owner-curated external sources. Not filtered by the scores or compatibility controls above; excluded from GitHub rankings and automatic installation.
No external entries match this query.