Daft
Eventual-Inc
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
Page 1 · 16 shown · 16 public entries
Results: 16
Eventual-Inc
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
bghira
A general fine-tuning kit geared toward image/video/audio diffusion models.
linto-ai
Multilingual Automatic Speech Recognition with word-level timestamps and confidence
AliAkhtari78
Extract public Spotify data — tracks, albums, artists, playlists, podcasts & lyrics — without the official API. Sync + async, typed models, one dependency.
Xewdy444
A Python library for solving reCAPTCHA v2 and v3 with Playwright
welovemedia
FFmate is a modern and powerful automation layer built on top of FFmpeg - designed to make video and audio transcoding simpler, smarter, and easier to integrate
WofWca
⏩ Fast-forwards long pauses between sentences — watch lectures ~1.5x faster (browser extension)
cosmicstack-labs
Automated daily tech briefing — multi-source collection → knowledge-base deduplication → AI summarization → TTS speech synthesis, generating MP3 audio briefings
ccoreilly
A speech recognition library running in the browser thanks to a WebAssembly build of Vosk
KoljaB
Local AI talk with a custom voice based on Zephyr 7B model. Uses RealtimeSTT with faster_whisper for transcription and RealtimeTTS with Coqui XTTS for synthesis.
MartinDelophy
Use native WebMCP tools to edit the project open in Timeline Studio. Inspect clips and available assets, preview and apply multi-track, caption style/position and projec…
richmondu
libfaceid is a research framework for prototyping of face recognition solutions. It seamlessly integrates multiple detection, recognition and liveness models w/ speech s…
ictnlp
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across various modality combinations.
seanwood
Real-time GCC-NMF Blind Speech Separation and Enhancement
eddieoz
MARCELO: an AI powered bot to automate the editing and thumbnail creation for your Youtube clips channel
liarjsdev
Audit a browser fingerprint for internal contradictions with the liarjs CLI - canvas, WebGL, WebGL2, WebGPU, audio, 220 fonts, WebRTC and timezone probes, scored against…