OPENAGENTSKILL / DIRECTORY

AI Agent Skills

为下一项任务找到合适的技能。探索适用于 Codex、Claude Code、Cursor 等 Agent 的工具。

搜索结果 · “evaluation”

49 Skills

搜索结果: 49

AI 与知识库

暂未收录效果图

查看技能说明

Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!

价格未确认AI 与知识库使用前请审核
2124GitHub
硬件与物联网

暂未收录效果图

查看技能说明

VLMEvalKit

open-compass

Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

价格未确认硬件与物联网Claude CodeOpenAI Agents使用前请审核
4229GitHub
开发与测试

暂未收录效果图

查看技能说明

RagaAI Catalyst

raga-ai-hub

Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm and tools tracing, debugging multi-agentic system, self-hos…

价格未确认开发与测试使用前请审核
1.6万GitHub
AI 与知识库

暂未收录效果图

查看技能说明

Agenta

Agenta-AI

The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.

价格未确认AI 与知识库使用前请审核
4450GitHub
开发与测试

暂未收录效果图

查看技能说明

Coze Loop

coze-dev

Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from developmen…

价格未确认开发与测试使用前请审核
5606GitHub
浏览器与自动化

暂未收录效果图

查看技能说明

AB3DMOT

xinshuoweng

(IROS 2020, ECCVW 2020) Official Python Implementation for "3D Multi-Object Tracking: A Baseline and New Evaluation Metrics"

价格未确认浏览器与自动化使用前请审核
1840GitHub
AI 与知识库

暂未收录效果图

查看技能说明

Skills

langfuse

Agent Skills for Langfuse, the open source LLM engineering platform for tracing, prompt management, and evaluation

价格未确认AI 与知识库Claude Code使用前请审核
214GitHub
AI 与知识库

暂未收录效果图

查看技能说明

🦄 Unitxt is a Python library for enterprise-grade evaluation of AI performance, offering the world's largest catalog of tools and data for end-to-end AI benchmarking

价格未确认AI 与知识库使用前请审核
214GitHub

指南与对比

开发者入口