No visual example yet
Explore the skillevaluating-code-models
sangrokjung
Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities…
OPENAGENTSKILL / DIRECTORY
다음 작업에 맞는 스킬을 찾아보세요. Codex, Claude Code, Cursor 등을 지원합니다.
7 Skills
검색 결과: 7
No visual example yet
Explore the skillsangrokjung
Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities…
No visual example yet
Explore the skillmicrosoft
Windows Agent Arena (WAA) 🪟 is a scalable OS platform for testing and benchmarking of multi-modal AI agents.
No visual example yet
Explore the skillagentscope-ai
Use when the user wants to build an evaluation system for an LLM/agent application but doesn't know where to start — they have traces, prompts, RAG pipelines, or nothing…
No visual example yet
Explore the skillMozerWang
[EMNLP 2024 (Oral)] Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
No visual example yet
Explore the skillscdenney
Reviews an existing conjoint study for threats to inference and returns prioritized findings across five areas — design integrity (attributes, profile restrictions, task…
No visual example yet
Explore the skillrealjaymes
Quickly creates new Claude Code skills or translates ChatGPT projects into Claude Code skills. Handles skill scaffolding, frontmatter, directory structure, and ChatGPT-t…
No visual example yet
Explore the skillrustyrazorblade
Helps the user design a lab workflow by asking questions, working through steps together, and writing the result to a plan.md file. Use when planning a new lab run, benc…