Agent outcome loop

Let rankings learn from real agent runs.

Every resolved skill should report what happened after one narrow task: success, failure, setup friction, or risk block. Those aggregate signals feed Trust Score, rankings, skill pages, and install recommendations.
653
Outcomes
87%
Success
634
Installs
1
Risk blocks
0
Production
43/100
Avg proven

Resolve

The agent gets one recommended skill.

Resolve returns the selected skill, alternatives, install plan, Trust Score, safety policy, and a unique feedback event id.

Run

The agent tries one narrow task.

Use a sandbox workflow first. Record whether install was used, whether setup was required, and whether risk blocked execution.

Learn

The registry updates trust signals.

Aggregate outcomes improve rankings without exposing raw agent notes or per-user identifiers publicly.

Outcome leaderboard

Skills with the strongest adoption evidence

228 skills currently have reported agent outcomes.

01

Frontend Design

Promising agent evidence

Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.

Promising agent evidence: 56 outcomes, 93% success, 56 install attempts, 0 risk blocks, 0 setup-required reports, Agent Proven Score 69/100.

93%
Success
69/100
Proven
56
Outcomes
7%
Recent fail
—
Quality
95/100
Trust
02

mono-color

Promising agent evidence

Generate original one-ink or controlled two-ink editorial images from any theme, sentence, article idea, object, or reference photo. Always use this skill when the user asks for 单色海报、双色印刷、单色调视觉、蓝色/绿色孔版印刷、risograph、网点照片、复古或当代编辑排版、zine poster, monochrome editorial poster, duotone print, or asks to use the mono-color style. It uses an adaptive white, gray, or pale-beige substrate, no more than two printing inks, active negative space, terse human language, and strong serif/grotesk/mono typography without making retro styling the default or copying a source composition, wording, logo, or artwork. Produce both the final generation prompt and the generated raster image unless the user explicitly asks for prompt only.

Promising agent evidence: 45 outcomes, 100% success, 45 install attempts, 0 risk blocks, 0 setup-required reports, Agent Proven Score 80/100.

100%
Success
80/100
Proven
45
Outcomes
0%
Recent fail
—
Quality
89/100
Trust
03

Vox Director

Promising agent evidence

Turn one topic into a narrated Vox-style paper-collage explainer or ad video, from script through captions.

Promising agent evidence: 38 outcomes, 100% success, 38 install attempts, 0 risk blocks, 0 setup-required reports, Agent Proven Score 80/100.

100%
Success
80/100
Proven
38
Outcomes
0%
Recent fail
—
Quality
91/100
Trust
04

Taste Skill: Anti-Slop Frontend

Early agent signal

Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.

Early agent signal: 32 outcomes, 88% success, 32 install attempts, 0 risk blocks, 0 setup-required reports, Agent Proven Score 66/100.

88%
Success
66/100
Proven
32
Outcomes
12%
Recent fail
—
Quality
98/100
Trust
05

Guizang Ppt Skill

Promising agent evidence

AI-agent Skill for generating polished HTML slide decks: editorial magazine and Swiss layouts, image prompts, social covers, and a WebGL/low-power presentation runtime.

Promising agent evidence: 25 outcomes, 100% success, 25 install attempts, 0 risk blocks, 0 setup-required reports, Agent Proven Score 74/100.

100%
Success
74/100
Proven
25
Outcomes
0%
Recent fail
—
Quality
97/100
Trust
06

Archify

Promising agent evidence

Any agent Skill: generate beautiful architecture diagrams with dark/light theme toggle and PNG/JPEG/WebP/SVG export

Promising agent evidence: 17 outcomes, 100% success, 17 install attempts, 0 risk blocks, 0 setup-required reports, Agent Proven Score 75/100.

100%
Success
75/100
Proven
17
Outcomes
0%
Recent fail
—
Quality
92/100
Trust
07

Canvas Design

Promising agent evidence

Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.

Promising agent evidence: 13 outcomes, 100% success, 13 install attempts, 0 risk blocks, 0 setup-required reports, Agent Proven Score 77/100.

100%
Success
77/100
Proven
13
Outcomes
0%
Recent fail
—
Quality
95/100
Trust
08

Web Design Guidelines

Promising agent evidence

Review UI code for web interface guidelines, UX quality, accessibility, and interaction design best practices.

Promising agent evidence: 12 outcomes, 100% success, 12 install attempts, 0 risk blocks, 0 setup-required reports, Agent Proven Score 77/100.

100%
Success
77/100
Proven
12
Outcomes
0%
Recent fail
—
Quality
94/100
Trust
09

Superpowers

Promising agent evidence

An agentic skills framework & software development methodology that works.

Promising agent evidence: 13 outcomes, 92% success, 13 install attempts, 0 risk blocks, 0 setup-required reports, Agent Proven Score 69/100.

92%
Success
69/100
Proven
13
Outcomes
8%
Recent fail
—
Quality
92/100
Trust
10

unslop

Promising agent evidence

Applies Cursor's prose-discipline rules to remove filler and AI writing tells from agent-facing and user-facing text.

Promising agent evidence: 13 outcomes, 100% success, 13 install attempts, 0 risk blocks, 0 setup-required reports, Agent Proven Score 75/100.

100%
Success
75/100
Proven
13
Outcomes
0%
Recent fail
—
Quality
88/100
Trust
11

Frontend Slides

Early agent signal

Create beautiful slides on the web using a coding agent's frontend skills

Early agent signal: 12 outcomes, 92% success, 11 install attempts, 0 risk blocks, 0 setup-required reports, Agent Proven Score 67/100.

92%
Success
67/100
Proven
12
Outcomes
11%
Recent fail
—
Quality
93/100
Trust
12

Webapp Testing

Promising agent evidence

Use Playwright to interact with and test local web applications, capture screenshots, debug UI behavior, and inspect browser logs.

Promising agent evidence: 8 outcomes, 100% success, 8 install attempts, 0 risk blocks, 0 setup-required reports, Agent Proven Score 73/100.

100%
Success
73/100
Proven
8
Outcomes
0%
Recent fail
—
Quality
93/100
Trust

POST /api/agent/outcome

Report after one narrow run.

Resolve responses include a unique feedback.event_id. Agents should reuse it when reporting the result so retries stay idempotent.

{
  "event_id": "resolve_...",
  "skill_slug": "crawl4ai",
  "task": "scrape pricing pages",
  "agent": "codex",
  "outcome": "success",
  "install_used": true,
  "time_to_useful_ms": 120000
}

Outcome meanings

Use the smallest honest label.

successThe skill helped complete the task.
failedThe skill was attempted but did not work.
not_relevantThe selected skill did not fit the task.
blocked_by_riskAudit, license, token, shell, or network risk stopped execution.
setup_requiredThe skill looked relevant but needed missing keys, data, or configuration.