Skill 디렉토리

AI Agent를 위한 재사용 가능한 Skill을 찾으세요.

작업으로 실제 GitHub Skill을 검색하고 사용 전에 Stars, 신뢰, 감사, 카테고리, 설치 경로를 확인하세요.

모든 추천은 리포지토리, 감사, 설치 경로와 명확하게 연결됩니다.

검색 결과: eval

영문 디렉토리

28 eval-informed mental models and critical-thinking skills for Claude Code, GitHub Copilot, Codex, Cursor, and other Agent Skills-compatible tools

941
Stars
84/100
신뢰
카테고리: utility감사

A powerful tool for creating datasets for LLM fine-tuning 、RAG and Eval

14K
Stars
75/100
신뢰
카테고리: data감사

A Claude skill that removes 54 neural network fingerprints from Russian text to bypass AI detectors like GPTZero and RuBERT.

223
Stars
76/100
신뢰
카테고리: utility감사

Awesome QA Skills — a bilingual (zh/en) AI testing Agent Skills library for Codex, Cursor, Claude Code, Kiro, OpenCode, and Trae. Ships 4 testing workflows and 25 testing-type skills (58 skill folders with language parity): independently installable, composable, and eval-ready with skill-up. Covers requirements, strategy, cases, API/performance/sec

151
Stars
79/100
신뢰
카테고리: utility감사

A self-learning skill layer for Claude Code that automatically distills, merges, updates, and prunes skills from real sessions.

413
Stars
75/100
신뢰
카테고리: coding-agents감사

A meta-skill that creates, evaluates, and improves other AI agent skills with multiple modes and evidence-based validation.

133
Stars
77/100
신뢰
카테고리: utility감사

A modular agent skill package for directing Seedance 2.0 filmmaking workflows across text, image, video, audio, references, safety rewrites, and production handoff.

796
Stars
67/100
신뢰
카테고리: Creative감사

Autonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinement (HTR) from the Arbor paper. Use this whenever someone wants to iteratively optimize something over many experiments without overfitting — e.g. "get my model's eval score up", "improve this agent/harness", "tune this pipeline", "beat the baseline on this benchmark", "run a search over approaches and keep the best", "do an MLE-bench / Kaggle-style optimization", or any long-horizon "make this artifact better and don't just memorize the dev set" task. Trigger it even when the user doesn't say "Arbor" or "hypothesis tree" but describes repeated experiment-and-evaluate loops, branching exploration of competing ideas, or worries about a dev/test gap. Runs Claude itself as the coordinator with subagent executors in isolated git worktrees; for the standalone `arbor` CLI tool see references/arbor-upstream.md.

34K
Stars
77/100
신뢰
카테고리: research감사

Turn any domain folder of skills into a bounded agentic loop: compile a goal into a verifiable task plan, execute tasks with the domain's own tools, verify every task with machine-run checks, retry with caps, escalate to a human when budgets exhaust, and refuse to close until everything is verified or explicitly waived. Use when you want an agent or subagent to pick up a goal and drive it to a verified close across one of this repo's 18 domains ('run this goal through the engineering harness', 'set up an agentic loop for marketing work', 'make the finance domain self-verifying'). NOT for authoring Claude Code Workflow-tool .js scripts (workflow-builder), N-agent tournaments on one task (agenthub), single-file metric optimization (autoresearch-agent), or discovering published loop recipes (loop-library).

25K
Stars
72/100
신뢰
카테고리: research감사

Use when the user asks to design a multi-agent system, pick an orchestration pattern (supervisor/swarm/pipeline), generate tool schemas for agents, or evaluate agent execution logs for cost, latency, and failure bottlenecks. Examples: 'design an agent architecture for research automation', 'generate Anthropic tool schemas from these tool descriptions', 'analyze these agent run logs for bottlenecks'. NOT for Claude Code workflow files (use workflow-builder) or single-agent prompt design (use agent-workflow-designer).

25K
Stars
83/100
신뢰
카테고리: research감사

NEO Emacs (WIP): GPU powered Emacs written in Rust with a modern display engine. Aiming for modern design & multi-threaded Elisp, 10x performance, zero-pause GC and 100% Emacs compatibility.

879
Stars
67/100
신뢰
카테고리: productivity-automation감사

A test runner for agentskills.io-style AI agent skills

596
Stars
65/100
신뢰
카테고리: agent-frameworks감사