No visual example yet
Explore the skillDoctr
mindee
docTR (Document Text Recognition) - a seamless, high-performing & accessible library for OCR-related tasks powered by Deep Learning.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
17–32 / 63
Results: 63
No visual example yet
Explore the skillmindee
docTR (Document Text Recognition) - a seamless, high-performing & accessible library for OCR-related tasks powered by Deep Learning.
No visual example yet
Explore the skillwenet-e2e
Production First and Production Ready End-to-End Speech Recognition Toolkit
No visual example yet
Explore the skillMahmoudAshraf97
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
No visual example yet
Explore the skillflashlight
Facebook AI Research's Automatic Speech Recognition Toolkit
No visual example yet
Explore the skillgautamkrishnar
Show your latest blog posts from any sources or StackOverflow activity or Youtube Videos on your GitHub profile/project readme automatically using the RSS feed
No visual example yet
Explore the skillantgroup
[CVPR 2025] EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation
No visual example yet
Explore the skilljianchang512
Voice Recognition to Text Tool / 一个离线运行的本地音视频转字幕工具,输出json、srt字幕、纯文字格式
No visual example yet
Explore the skillLynpoint
Self hosted, real-time digital human agent platform. Build voice-first AI agents with WebRTC, persona memory, tools, RAG, and optional digital-human video.
No visual example yet
Explore the skillFireRedTeam
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering ou…
No visual example yet
Explore the skillFunAudioLLM
End-to-end speech recognition large model: 31 languages, dialects, accents, lyrics, hotwords, timestamps, speaker diarization. Trained on tens of millions of hours.
No visual example yet
Explore the skillroyshil
OBS plugin for local speech recognition and captioning using AI

yanliudesign
Generate original one-ink or controlled two-ink editorial images from any theme, sentence, article idea, object, or reference photo. Always use this skill when the user…
View examples · 1No visual example yet
Explore the skillyeyupiaoling
Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inf…
No visual example yet
Explore the skilldeclare-lab
MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversation
No visual example yet
Explore the skillgoogleworkspace
Google Workspace CLI — one command-line tool for Drive, Gmail, Calendar, Sheets, Docs, Chat, Admin, and more. Dynamically built from Google Discovery Service. Includes A…
No visual example yet
Explore the skillakfamily
AKShare is an elegant and simple financial data interface library for Python, built for human beings! 开源财经数据接口库