No visual example yet
Explore the skillEvaluation Guidebook
huggingface
Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!
OPENAGENTSKILL / DIRECTORY
次のタスクに合うスキルを。Codex、Claude Code、Cursor などのツールを探せます。
12 Skills
検索結果: 12
No visual example yet
Explore the skillhuggingface
Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!
No visual example yet
Explore the skillAgenta-AI
The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.
No visual example yet
Explore the skillMarker-Inc-Korea
AutoRAG: An Open-Source Framework for Retrieval-Augmented Generation (RAG) Evaluation & Optimization with AutoML-Style Automation
No visual example yet
Explore the skilltruera
Evaluation and Tracking for LLM Experiments and AI Agents
No visual example yet
Explore the skillmodelscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
No visual example yet
Explore the skilljuanjuandog
AI equity research agent with resilient workflows, Redis Lua single-flight, pgvector RAG, versioned reports, evidence tracing, and RAG evaluation.
No visual example yet
Explore the skilllangfuse
Agent Skills for Langfuse, the open source LLM engineering platform for tracing, prompt management, and evaluation
No visual example yet
Explore the skillagentscope-ai
OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
No visual example yet
Explore the skillIBM
🦄 Unitxt is a Python library for enterprise-grade evaluation of AI performance, offering the world's largest catalog of tools and data for end-to-end AI benchmarking
No visual example yet
Explore the skillAgentEvalHQ
AgentEval is the comprehensive .NET toolkit for AI agent evaluation—tool usage validation, RAG quality metrics, stochastic evaluation, and model comparison—built first f…
No visual example yet
Explore the skillgithub
Patterns and techniques for evaluating and improving AI agent outputs. Use this skill when: - Implementing self-critique and reflection loops - Building evaluator-optimi…
No visual example yet
Explore the skillRaudaschl
RAG-Fusion: multi-query generation + Reciprocal Rank Fusion for better retrieval-augmented generation. Includes evaluation harness with NFCorpus/BEIR.