OPENAGENTSKILL / DIRECTORY

AI Agent Skills

Finde den passenden Skill für deine nächste Aufgabe mit Codex, Claude Code, Cursor und mehr.

Ergebnisse · “benchmark”

6 Skills

Ergebnisse: 6

Entwicklung

No visual example yet

Explore the skill

Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities…

Preis unbestätigtEntwicklungClaude Code
825GitHub
Skill ansehen
Daten & Analyse

No visual example yet

Explore the skill

arbor

K-Dense-AI

Autonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinemen…

Preis unbestätigtDaten & AnalyseClaude Code
33.974GitHub
Skill ansehen
Entwicklung

No visual example yet

Explore the skill

run-load-test

omnigent-ai

Run the Omnigent load test and produce a results file explaining the latencies. Load when the user wants to load-test / stress-test / benchmark Omnigent under concurrenc…

Preis unbestätigtEntwicklungClaude Code
9609GitHub
Skill ansehen
Entwicklung

No visual example yet

Explore the skill

run-deep-swe

davidondrej

Score any AI model on the DeepSWE coding-agent benchmark via the OpenRouter API. Use when the user wants an independent, reproducible coding-agent eval — "run DeepSWE",…

Preis unbestätigtEntwicklungClaude Code
3850GitHub
Skill ansehen

Anleitungen & Vergleiche

Für Entwickler