DreamOmni2
JIA-Lab-research
This project is the official implementation of 'DreamOmni2: Multimodal Instruction-based Editing and Generation (CVPR2026 Highlight)''
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Search retrieves candidates across the registry and ranks a bounded shortlist by task fit. This count is matching candidates, not the registry total. No suitable match? Try a specific tool or task.
17–32 / 61
Results: 61
JIA-Lab-research
This project is the official implementation of 'DreamOmni2: Multimodal Instruction-based Editing and Generation (CVPR2026 Highlight)''
bytedance
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
JimLiu
Open-source Agent Skill that drives the BaoCut macOS app CLI (transcribe · subtitle · translate · cut) from Claude Code, Codex, and other agents
Pluviobyte
Build a reproducible local rough-cut workflow for talking-head or narrated screen-recording videos with transcript review gates.
Pluviobyte
Generate a controlled local narration workflow with auditions, version tracking, and subtitle-ready final audio.
FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
nexu-io
Audio generation skill — jingles, beds, voiceover, and sound effects. Routes music requests to Suno V5 / Udio / Lyria, speech to MiniMax TTS / FishAudio / ElevenLabs V3,…
huggingface
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
apple
Core ML tools contain supporting tools for Core ML model conversion, editing, and validation.
Eventual-Inc
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
whitphx
Real-time video and audio processing on Streamlit
Yuan-ManX
Your AI Game Dev Hub. The ultimate resource hub for AI-powered game development tools. Discover cutting-edge LLMs, World Model, Agent, Code, Image, Texture, Shader, 3D M…
ThioJoe
Automatically translates the text of a video based on a subtitle file, and then uses AI voice services to create a new dubbed & translated audio track where the speech i…
calesthio
Comprehensive guide for BFL FLUX image generation models. Covers prompting, T2I, I2I, structured JSON, hex colors, typography, multi-reference editing, and model-specifi…
calesthio
Video and audio processing with FFmpeg. Use for format conversion, resizing, compression, audio extraction, and preparing assets for Remotion. Triggers include convertin…
Owner-curated external sources. Not filtered by the scores or compatibility controls above; excluded from GitHub rankings and automatic installation.
No external entries match this query.