暂未收录效果图
查看技能说明evaluating-code-models
sangrokjung
Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities…
OPENAGENTSKILL / DIRECTORY
为下一项任务找到合适的技能。探索适用于 Codex、Claude Code、Cursor 等 Agent 的工具。
7 Skills
搜索结果: 7
暂未收录效果图
查看技能说明sangrokjung
Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities…
暂未收录效果图
查看技能说明microsoft
Windows Agent Arena (WAA) 🪟 is a scalable OS platform for testing and benchmarking of multi-modal AI agents.
暂未收录效果图
查看技能说明agentscope-ai
Use when the user wants to build an evaluation system for an LLM/agent application but doesn't know where to start — they have traces, prompts, RAG pipelines, or nothing…
暂未收录效果图
查看技能说明MozerWang
[EMNLP 2024 (Oral)] Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
暂未收录效果图
查看技能说明scdenney
Reviews an existing conjoint study for threats to inference and returns prioritized findings across five areas — design integrity (attributes, profile restrictions, task…
暂未收录效果图
查看技能说明realjaymes
Quickly creates new Claude Code skills or translates ChatGPT projects into Claude Code skills. Handles skill scaffolding, frontmatter, directory structure, and ChatGPT-t…
暂未收录效果图
查看技能说明rustyrazorblade
Helps the user design a lab workflow by asking questions, working through steps together, and writing the result to a plan.md file. Use when planning a new lab run, benc…