OPENAGENTSKILL / DIRECTORY

AI Agent Skills

다음 작업에 맞는 스킬을 찾아보세요. Codex, Claude Code, Cursor 등을 지원합니다.

검색 결과 · “evaluation”

12 Skills

검색 결과: 12

AI 및 지식

No visual example yet

Explore the skill

Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!

가격 미확인AI 및 지식사용 전 검토
2.1천GitHub
AI 및 지식

No visual example yet

Explore the skill

Agenta

Agenta-AI

The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.

가격 미확인AI 및 지식사용 전 검토
4.5천GitHub
AI 및 지식

No visual example yet

Explore the skill

Skills

langfuse

Agent Skills for Langfuse, the open source LLM engineering platform for tracing, prompt management, and evaluation

가격 미확인AI 및 지식Claude Code사용 전 검토
214GitHub
AI 및 지식

No visual example yet

Explore the skill

🦄 Unitxt is a Python library for enterprise-grade evaluation of AI performance, offering the world's largest catalog of tools and data for end-to-end AI benchmarking

가격 미확인AI 및 지식사용 전 검토
214GitHub
AI 및 지식

No visual example yet

Explore the skill

Patterns and techniques for evaluating and improving AI agent outputs. Use this skill when: - Implementing self-critique and reflection loops - Building evaluator-optimi…

가격 미확인AI 및 지식Claude Code
3.9만GitHub
AI 및 지식

No visual example yet

Explore the skill

RAG Fusion

Raudaschl

RAG-Fusion: multi-query generation + Reciprocal Rank Fusion for better retrieval-augmented generation. Includes evaluation harness with NFCorpus/BEIR.

가격 미확인AI 및 지식OpenAI Agents사용 전 검토
946GitHub

가이드 및 비교

개발자용