StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models
$ npx skills add yl4579/StyleTTS2Scenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets