No visual example yet
Explore the skillEvaluation Guidebook
huggingface
Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!
OPENAGENTSKILL / DIRECTORY
Temukan skill untuk tugas berikutnya dengan Codex, Claude Code, Cursor, dan lainnya.
50 Skills
Hasil: 50
No visual example yet
Explore the skillhuggingface
Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!
No visual example yet
Explore the skillopen-compass
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
No visual example yet
Explore the skillraga-ai-hub
Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm and tools tracing, debugging multi-agentic system, self-hos…
No visual example yet
Explore the skillAgenta-AI
The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.
No visual example yet
Explore the skillGiskard-AI
🐢 Open-Source Evaluation & Testing library for LLM Agents
No visual example yet
Explore the skillcoze-dev
Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from developmen…
No visual example yet
Explore the skillMarker-Inc-Korea
AutoRAG: An Open-Source Framework for Retrieval-Augmented Generation (RAG) Evaluation & Optimization with AutoML-Style Automation
No visual example yet
Explore the skilltruera
Evaluation and Tracking for LLM Experiments and AI Agents
No visual example yet
Explore the skillmodelscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
No visual example yet
Explore the skilljuanjuandog
AI equity research agent with resilient workflows, Redis Lua single-flight, pgvector RAG, versioned reports, evidence tracing, and RAG evaluation.
No visual example yet
Explore the skillxinshuoweng
(IROS 2020, ECCVW 2020) Official Python Implementation for "3D Multi-Object Tracking: A Baseline and New Evaluation Metrics"
No visual example yet
Explore the skilllangfuse
Agent Skills for Langfuse, the open source LLM engineering platform for tracing, prompt management, and evaluation
No visual example yet
Explore the skillyaojingang
YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.
No visual example yet
Explore the skillagentscope-ai
OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
No visual example yet
Explore the skillablab
Genome assembly evaluation tool
No visual example yet
Explore the skillIBM
🦄 Unitxt is a Python library for enterprise-grade evaluation of AI performance, offering the world's largest catalog of tools and data for end-to-end AI benchmarking