No visual example yet
Explore the skillagent-evaluation
NeoLabHQ
Evaluate and improve Claude Code commands, skills, and agents. Use when testing prompt effectiveness, validating context engineering choices, or measuring improvement qu…
OPENAGENTSKILL / DIRECTORY
Encuentra una habilidad para tu próxima tarea con Codex, Claude Code, Cursor y más.
48 Skills
Resultados: 48
No visual example yet
Explore the skillNeoLabHQ
Evaluate and improve Claude Code commands, skills, and agents. Use when testing prompt effectiveness, validating context engineering choices, or measuring improvement qu…
No visual example yet
Explore the skillhuggingface
Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!
No visual example yet
Explore the skillGiskard-AI
🐢 Open-Source Evaluation & Testing library for LLM Agents
No visual example yet
Explore the skillcoze-dev
Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from developmen…
No visual example yet
Explore the skilltruera
Evaluation and Tracking for LLM Experiments and AI Agents
No visual example yet
Explore the skillopen-compass
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
No visual example yet
Explore the skillraga-ai-hub
Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm and tools tracing, debugging multi-agentic system, self-hos…
No visual example yet
Explore the skillAgenta-AI
The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.
No visual example yet
Explore the skillMarker-Inc-Korea
AutoRAG: An Open-Source Framework for Retrieval-Augmented Generation (RAG) Evaluation & Optimization with AutoML-Style Automation
No visual example yet
Explore the skillmodelscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
No visual example yet
Explore the skilljuanjuandog
AI equity research agent with resilient workflows, Redis Lua single-flight, pgvector RAG, versioned reports, evidence tracing, and RAG evaluation.
No visual example yet
Explore the skillxinshuoweng
(IROS 2020, ECCVW 2020) Official Python Implementation for "3D Multi-Object Tracking: A Baseline and New Evaluation Metrics"
No visual example yet
Explore the skilllangfuse
Agent Skills for Langfuse, the open source LLM engineering platform for tracing, prompt management, and evaluation
No visual example yet
Explore the skillyaojingang
YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.
No visual example yet
Explore the skillagentscope-ai
OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
No visual example yet
Explore the skillsangrokjung
Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities…