OPENAGENTSKILL / DIRECTORY

AI Agent Skills

Temukan skill untuk tugas berikutnya dengan Codex, Claude Code, Cursor, dan lainnya.

Hasil · “evaluation”

49 Skills

Hasil: 49

AI & pengetahuan

No visual example yet

Explore the skill

Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!

Harga belum dikonfirmasiAI & pengetahuanTinjau sebelum digunakan
2,1 rbGitHub
Lihat skill
Perangkat keras & IoT

No visual example yet

Explore the skill

VLMEvalKit

open-compass

Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

Harga belum dikonfirmasiPerangkat keras & IoTClaude CodeOpenAI AgentsTinjau sebelum digunakan
4,2 rbGitHub
Lihat skill
Pengembangan

No visual example yet

Explore the skill

RagaAI Catalyst

raga-ai-hub

Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm and tools tracing, debugging multi-agentic system, self-hos…

Harga belum dikonfirmasiPengembanganTinjau sebelum digunakan
16,1 rbGitHub
Lihat skill
AI & pengetahuan

No visual example yet

Explore the skill

Agenta

Agenta-AI

The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.

Harga belum dikonfirmasiAI & pengetahuanTinjau sebelum digunakan
4,5 rbGitHub
Lihat skill
Pengembangan

No visual example yet

Explore the skill

Coze Loop

coze-dev

Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from developmen…

Harga belum dikonfirmasiPengembanganTinjau sebelum digunakan
5,6 rbGitHub
Lihat skill
AI & pengetahuan

No visual example yet

Explore the skill

AutoRAG

Marker-Inc-Korea

AutoRAG: An Open-Source Framework for Retrieval-Augmented Generation (RAG) Evaluation & Optimization with AutoML-Style Automation

Harga belum dikonfirmasiAI & pengetahuanTinjau sebelum digunakan
4,8 rbGitHub
Lihat skill
AI & pengetahuan

No visual example yet

Explore the skill

FinSight AI

juanjuandog

AI equity research agent with resilient workflows, Redis Lua single-flight, pgvector RAG, versioned reports, evidence tracing, and RAG evaluation.

Harga belum dikonfirmasiAI & pengetahuanTinjau sebelum digunakan
1,2 rbGitHub
Lihat skill
Browser & otomatisasi

No visual example yet

Explore the skill

AB3DMOT

xinshuoweng

(IROS 2020, ECCVW 2020) Official Python Implementation for "3D Multi-Object Tracking: A Baseline and New Evaluation Metrics"

Harga belum dikonfirmasiBrowser & otomatisasiTinjau sebelum digunakan
1,8 rbGitHub
Lihat skill
AI & pengetahuan

No visual example yet

Explore the skill

Skills

langfuse

Agent Skills for Langfuse, the open source LLM engineering platform for tracing, prompt management, and evaluation

Harga belum dikonfirmasiAI & pengetahuanClaude CodeTinjau sebelum digunakan
214GitHub
Lihat skill
AI & pengetahuan

No visual example yet

Explore the skill

🦄 Unitxt is a Python library for enterprise-grade evaluation of AI performance, offering the world's largest catalog of tools and data for end-to-end AI benchmarking

Harga belum dikonfirmasiAI & pengetahuanTinjau sebelum digunakan
214GitHub
Lihat skill

Panduan & perbandingan

Untuk pengembang