OPENAGENTSKILL / DIRECTORY

AI Agent Skills

Encuentra una habilidad para tu próxima tarea con Codex, Claude Code, Cursor y más.

Resultados · “evaluation”

50 Skills

Resultados: 50

IA y conocimiento

No visual example yet

Explore the skill

Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!

Precio sin confirmarIA y conocimientoRevisar antes de usar
2,1 milGitHub
Hardware e IoT

No visual example yet

Explore the skill

VLMEvalKit

open-compass

Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

Precio sin confirmarHardware e IoTClaude CodeOpenAI AgentsRevisar antes de usar
4,2 milGitHub
Desarrollo

No visual example yet

Explore the skill

RagaAI Catalyst

raga-ai-hub

Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm and tools tracing, debugging multi-agentic system, self-hos…

Precio sin confirmarDesarrolloRevisar antes de usar
16,1 milGitHub
IA y conocimiento

No visual example yet

Explore the skill

Agenta

Agenta-AI

The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.

Precio sin confirmarIA y conocimientoRevisar antes de usar
4,5 milGitHub
Desarrollo

No visual example yet

Explore the skill

Coze Loop

coze-dev

Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from developmen…

Precio sin confirmarDesarrolloRevisar antes de usar
5,6 milGitHub
Navegador y automatización

No visual example yet

Explore the skill

AB3DMOT

xinshuoweng

(IROS 2020, ECCVW 2020) Official Python Implementation for "3D Multi-Object Tracking: A Baseline and New Evaluation Metrics"

Precio sin confirmarNavegador y automatizaciónRevisar antes de usar
1,8 milGitHub
IA y conocimiento

No visual example yet

Explore the skill

Skills

langfuse

Agent Skills for Langfuse, the open source LLM engineering platform for tracing, prompt management, and evaluation

Precio sin confirmarIA y conocimientoClaude CodeRevisar antes de usar
214GitHub
IA y conocimiento

No visual example yet

Explore the skill

🦄 Unitxt is a Python library for enterprise-grade evaluation of AI performance, offering the world's largest catalog of tools and data for end-to-end AI benchmarking

Precio sin confirmarIA y conocimientoRevisar antes de usar
214GitHub

Guías y comparativas

Para desarrolladores