Transformers
huggingface
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and traini…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
Page 1 · 16 shown · 20 public entries
Results: 20
huggingface
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and traini…
screenpipe
YC (S26) | AI that knows what you've seen, said, or heard. Records everything you do, say, hear 24/7, local, private, secure
2noise
A generative speech model for daily dialogue.
xorbitsai
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all throug…
yanshengjia
Machine Learning and Agentic AI Resources, Practice and Research
ARahim3
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
instillai
:speech_balloon: Machine Learning Course with Python:
apocas
RESTai is an AIaaS (AI as a Service) open-source platform. Supports many public and local LLM suported by Ollama/vLLM/etc. Precise embeddings usage, tuning, analytics et…
alexpinel
Text-To-Speech, RAG, and LLMs. All local!
mgonzs13
llama.cpp (GGUF LLMs) and llava.cpp (GGUF VLMs) for ROS 2
vndee
A talking LLM that runs on your own computer without needing the internet.
ferranpons
True on-device AI for Kotlin Multiplatform (Android, iOS, Desktop, JVM, WASM). LLM, Speech-to-Text and Image Generation — powered by llama.cpp, whisper.cpp and stable-di…
cerul-ai
The video search layer for AI agents. Search video by meaning — across speech, visuals, and on-screen text.
microsoft
Secure AI conversations with documents, video, audio, and more. Personal workspaces for focused context, group spaces for shared insight. Classify docs, reuse prompts, a…
Aratako
Multilingual TTS model with voice cloning and duration control, based on T5Gemma encoder-decoder LLM
code-100-precent
LingEcho is an intelligent voice interaction platform that provides a comprehensive AI voice interaction solution. It integrates advanced speech recognition (ASR), text-…