暂未收录效果图
查看技能说明agent-evaluation
NeoLabHQ
Evaluate and improve Claude Code commands, skills, and agents. Use when testing prompt effectiveness, validating context engineering choices, or measuring improvement qu…
OPENAGENTSKILL / DIRECTORY
为下一项任务找到合适的技能。探索适用于 Codex、Claude Code、Cursor 等 Agent 的工具。
48 Skills
搜索结果: 48
暂未收录效果图
查看技能说明NeoLabHQ
Evaluate and improve Claude Code commands, skills, and agents. Use when testing prompt effectiveness, validating context engineering choices, or measuring improvement qu…
暂未收录效果图
查看技能说明huggingface
Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!
暂未收录效果图
查看技能说明Giskard-AI
暂未收录效果图
查看技能说明coze-dev
Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from developmen…
暂未收录效果图
查看技能说明truera
Evaluation and Tracking for LLM Experiments and AI Agents
暂未收录效果图
查看技能说明open-compass
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
暂未收录效果图
查看技能说明raga-ai-hub
Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm and tools tracing, debugging multi-agentic system, self-hos…
暂未收录效果图
查看技能说明Agenta-AI
The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.
暂未收录效果图
查看技能说明Marker-Inc-Korea
AutoRAG: An Open-Source Framework for Retrieval-Augmented Generation (RAG) Evaluation & Optimization with AutoML-Style Automation
暂未收录效果图
查看技能说明modelscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
暂未收录效果图
查看技能说明juanjuandog
AI equity research agent with resilient workflows, Redis Lua single-flight, pgvector RAG, versioned reports, evidence tracing, and RAG evaluation.
暂未收录效果图
查看技能说明xinshuoweng
(IROS 2020, ECCVW 2020) Official Python Implementation for "3D Multi-Object Tracking: A Baseline and New Evaluation Metrics"
暂未收录效果图
查看技能说明langfuse
Agent Skills for Langfuse, the open source LLM engineering platform for tracing, prompt management, and evaluation
暂未收录效果图
查看技能说明yaojingang
YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.
暂未收录效果图
查看技能说明agentscope-ai
OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
暂未收录效果图
查看技能说明sangrokjung
Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities…