Skill 디렉토리

AI Agent를 위한 재사용 가능한 Skill을 찾으세요.

작업으로 실제 GitHub Skill을 검색하고 사용 전에 Stars, 신뢰, 감사, 카테고리, 설치 경로를 확인하세요.

모든 추천은 리포지토리, 감사, 설치 경로와 명확하게 연결됩니다.

검색 결과: swe-bench

영문 디렉토리

SWE-agent takes a GitHub issue and tries to automatically fix it, using your LM of choice. It can also be employed for offensive cybersecurity or competitive coding challenges. [NeurIPS 2024]

20K
Stars
87/100
신뢰
카테고리: coding-agents감사

An Open-Source Asynchronous Coding Agent

10.0K
Stars
81/100
신뢰
카테고리: coding-agents감사

Checks whether Kubernetes is deployed according to security best practices as defined in the CIS Kubernetes Benchmark

8.1K
Stars
86/100
신뢰
카테고리: devops감사

A simple SWE style browser agent framework that achieves SOTA results on long horizon web tasks.

5.5K
Stars
84/100
신뢰
카테고리: agent-frameworks감사

The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—but scores >74% on SWE-bench verified!

5.3K
Stars
80/100
신뢰
카테고리: agent-frameworks감사

A FREE pragmatic DevOps learning to kickstart your DevOps career and knowledge in the Cloud Native era following the Agile MVP style! ⭐ (2026 plans for DevOps, Cloud, Platform, SRE, SWE)

2.4K
Stars
83/100
신뢰
카테고리: devops감사

A self-learning skill layer for Claude Code that automatically distills, merges, updates, and prunes skills from real sessions.

413
Stars
75/100
신뢰
카테고리: coding-agents감사

Kodezi Chronos is a debugging-first language model that achieves state-of-the-art results on SWE-bench Lite (80.33%) and 67% real-world fix accuracy, over six times better than GPT-4. Built with Adaptive Graph-Guided Retrieval and Persistent Debug Memory. Model available Q1 2026 via Kodezi OS.

4.9K
Stars
73/100
신뢰
카테고리: ml-automation감사

A Claude Code plugin that automates a multi-agent software development pipeline from feature spec to reviewed PR.

136
Stars
75/100
신뢰
카테고리: coding-agents감사

Autonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinement (HTR) from the Arbor paper. Use this whenever someone wants to iteratively optimize something over many experiments without overfitting — e.g. "get my model's eval score up", "improve this agent/harness", "tune this pipeline", "beat the baseline on this benchmark", "run a search over approaches and keep the best", "do an MLE-bench / Kaggle-style optimization", or any long-horizon "make this artifact better and don't just memorize the dev set" task. Trigger it even when the user doesn't say "Arbor" or "hypothesis tree" but describes repeated experiment-and-evaluate loops, branching exploration of competing ideas, or worries about a dev/test gap. Runs Claude itself as the coordinator with subagent executors in isolated git worktrees; for the standalone `arbor` CLI tool see references/arbor-upstream.md.

34K
Stars
77/100
신뢰
카테고리: research감사

Autonomous software engineering fleet of AI agents for production-grade PRs on AgentField: plan, code, test, and ship.

969
Stars
74/100
신뢰
카테고리: agent-frameworks감사

Measuring frontier coding agents on original, long-horizon engineering tasks

944
Stars
67/100
신뢰
카테고리: coding-agents감사