No visual example yet
Explore the skillDarwin Skill
alchaincyf
达尔文.skill —— 一个让你的Skill无限进化的系统:评估→改进→测试→保留或回滚 | Autoresearch-inspired autonomous skill optimization for Claude Code. Evaluate, improve, test, keep or revert.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
17–32 / 40
Results: 40
No visual example yet
Explore the skillalchaincyf
达尔文.skill —— 一个让你的Skill无限进化的系统:评估→改进→测试→保留或回滚 | Autoresearch-inspired autonomous skill optimization for Claude Code. Evaluate, improve, test, keep or revert.
No visual example yet
Explore the skillHelicone
🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓
No visual example yet
Explore the skillevidentlyai
Evidently is an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ m…
No visual example yet
Explore the skillWindy3f3f3f3f
Deep dive into Claude Code internals — architecture, agent loop, context engineering, and more. / 深入解析 Claude Code 源码:架构、Agent 循环、上下文工程、工具系统等
No visual example yet
Explore the skillsmixs
Architecture-first skill lifecycle for AI agents. 6 modes: CREATE / IMPROVE / VALIDATE / REVIEW / OPTIMIZE / PACKAGE. BinEval binary scoring with threshold-blind, cross-…
No visual example yet
Explore the skillTHUDM
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
No visual example yet
Explore the skillAnandChowdhary
🔂 Ralph loop with PRs: Run Claude Code in a continuous loop, autonomously creating PRs, waiting for checks, and merging
No visual example yet
Explore the skillWeizhena
Structured deep research skill for Claude Code/Open Code/Codex with human-in-the-loop control
No visual example yet
Explore the skillopen-edge-platform
Train, Evaluate, Optimize, Deploy Computer Vision Models via OpenVINO™
No visual example yet
Explore the skillevo-hq
turns your codebase into an autoresearch loop — discovers what to measure, instruments the benchmark, then runs tree search with parallel subagents.
No visual example yet
Explore the skillVchitect
[CVPR2024 Highlight] VBench - We Evaluate Video Generation
No visual example yet
Explore the skillzhaopeiym
This is an IoT device communication protocol implementation client, which will include common industrial communication protocols such as mainstream PLC communication rea…
No visual example yet
Explore the skilladdyosmani
Drives development with tests using the red-green-refactor loop. Use when implementing any logic, fixing any bug, or changing any behavior. Use when you need to prove th…
No visual example yet
Explore the skillbeir-cellar
A Heterogeneous Benchmark for Information Retrieval. Easy to use, evaluate your models across 15+ diverse IR datasets.
No visual example yet
Explore the skillEmbeddedLLM
The collaborative spreadsheet for AI. Chain cells into powerful pipelines, experiment with prompts and models, and evaluate LLM responses in real-time. Work together sea…
No visual example yet
Explore the skillzjunlp
Create, Evaluate, and Connect AI Skills