Fara
microsoft
Fara-7B: An Efficient Agentic Model for Computer Use
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
65–80 / 97
Results: 97
microsoft
Fara-7B: An Efficient Agentic Model for Computer Use
NirantK
Curated list of Machine Learning, NLP, Vision, Recommender Systems Project Ideas
icereed
Use LLMs and LLM Vision (OCR) to handle paperless-ngx - Document Digitalization powered by AI
magnitudedev
Open-source, vision-first browser agent
fikrikarim
On-device, real-time multimodal AI. Have natural voice and vision conversations with an AI that runs entirely on your machine. Powered by Gemma 4 E2B and Kokoro.
NVIDIA-AI-Blueprints
The NVIDIA VSS Blueprint is a suite of reference architectures for building GPU-accelerated vision agents and AI-powered video analytics applications.
A comprehensive list of Deep Learning / Artificial Intelligence and Machine Learning tutorials - rapidly expanding into areas of AI/Deep Learning / Machine Vision / NLP…
ARahim3
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
siddsachar
Row-Bot - Personal AI Sovereignty. A local-first AI assistant with integrated tools, a personal knowledge graph, voice, vision, shell, browser automation, scheduled task…
Turbo1123
Android Automation Tool Based on Vision-Language Models
emcf
Get clean data from tricky documents, powered by vision-language models ⚡
sparklabx
Teach your AI to draw correct, beautiful draw.io diagrams — declarative layout engine, ground-truth stencils, structural validator, vision self-check. AWS · Azure · GCP…
tanelpoder
0x.Tools: X-Ray vision for Linux systems
trycua
A curated list of resources about AI agents for Computer Use, including research papers, projects, frameworks, and tools.
kairyou
Reusable Agent Skills, plus integrations (statusline, provider usage, vision) that install into Codex, Claude Code, and opencode.
TIGER-AI-Lab
This repo contains the code for "VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks" [ICLR 2025]