End-to-end speech recognition large model: 31 languages, dialects, accents, lyrics, hotwords, timestamps, speaker diarization. Trained on tens of millions of hours.
$ npx skills add FunAudioLLM/Fun-ASRScenario Multimodal media · I need my agent to process images, video, or audio and extract useful information.
CLI + Codex · 4 targets