Agent outcome loop

Let rankings learn from real agent runs.

Every resolved skill should report what happened after one narrow task: success, failure, setup friction, or risk block. Those aggregate signals feed Trust Score, rankings, skill pages, and install recommendations.
0
Outcomes
No data
Success
0
Installs
0
Risk blocks
0
Production
No data
Avg proven

Resolve

The agent gets one recommended skill.

Resolve returns the selected skill, alternatives, install plan, Trust Score, safety policy, and a unique feedback event id.

Run

The agent tries one narrow task.

Use a sandbox workflow first. Record whether install was used, whether setup was required, and whether risk blocked execution.

Learn

The registry updates trust signals.

Aggregate outcomes improve rankings without exposing raw agent notes or per-user identifiers publicly.

Outcome leaderboard

Skills with the strongest adoption evidence

No public outcome reports yet; this page is ready for the first Resolve-powered runs.

01

Taste Skill: Anti-Slop Frontend

Needs first agent run

Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.

No agent outcome reports yet. The first resolved run should report success, failed, not_relevant, blocked_by_risk, or setup_required.

No data
Success
0/100
Proven
0
Outcomes
—
Recent fail
—
Quality
93/100
Trust
02

Frontend Design

Needs first agent run

Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.

No agent outcome reports yet. The first resolved run should report success, failed, not_relevant, blocked_by_risk, or setup_required.

No data
Success
0/100
Proven
0
Outcomes
—
Recent fail
—
Quality
90/100
Trust
03

Anthropic Brand Guidelines

Needs first agent run

Apply Anthropic official brand colors, typography, and visual standards to appropriate Anthropic-related artifacts.

No agent outcome reports yet. The first resolved run should report success, failed, not_relevant, blocked_by_risk, or setup_required.

No data
Success
0/100
Proven
0
Outcomes
—
Recent fail
—
Quality
90/100
Trust
04

Canvas Design

Needs first agent run

Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.

No agent outcome reports yet. The first resolved run should report success, failed, not_relevant, blocked_by_risk, or setup_required.

No data
Success
0/100
Proven
0
Outcomes
—
Recent fail
—
Quality
89/100
Trust
05

Webapp Testing

Needs first agent run

Use Playwright to interact with and test local web applications, capture screenshots, debug UI behavior, and inspect browser logs.

No agent outcome reports yet. The first resolved run should report success, failed, not_relevant, blocked_by_risk, or setup_required.

No data
Success
0/100
Proven
0
Outcomes
—
Recent fail
—
Quality
88/100
Trust
06

Crawl4AI

Needs first agent run

Open-source LLM-friendly web crawler and scraper for agent workflows.

No agent outcome reports yet. The first resolved run should report success, failed, not_relevant, blocked_by_risk, or setup_required.

No data
Success
0/100
Proven
0
Outcomes
—
Recent fail
—
Quality
80/100
Trust
07

MarkItDown

Needs first agent run

Convert PDFs, Office documents, and web files into clean markdown for agents.

No agent outcome reports yet. The first resolved run should report success, failed, not_relevant, blocked_by_risk, or setup_required.

No data
Success
0/100
Proven
0
Outcomes
—
Recent fail
—
Quality
79/100
Trust
08

Playwright

Needs first agent run

Reliable browser automation and testing engine for web agent tasks.

No agent outcome reports yet. The first resolved run should report success, failed, not_relevant, blocked_by_risk, or setup_required.

No data
Success
0/100
Proven
0
Outcomes
—
Recent fail
—
Quality
79/100
Trust
09

Agent Skills

Needs first agent run

Production-grade engineering skills and quality gates for AI coding agents.

No agent outcome reports yet. The first resolved run should report success, failed, not_relevant, blocked_by_risk, or setup_required.

No data
Success
0/100
Proven
0
Outcomes
—
Recent fail
—
Quality
77/100
Trust
10

Browser Use

Needs first agent run

Browser automation layer for agents that need to interact with websites.

No agent outcome reports yet. The first resolved run should report success, failed, not_relevant, blocked_by_risk, or setup_required.

No data
Success
0/100
Proven
0
Outcomes
—
Recent fail
—
Quality
75/100
Trust
11

n8n

Needs first agent run

Workflow automation for connecting agents to repeated operational tasks.

No agent outcome reports yet. The first resolved run should report success, failed, not_relevant, blocked_by_risk, or setup_required.

No data
Success
0/100
Proven
0
Outcomes
—
Recent fail
—
Quality
72/100
Trust
12

OpenBB

Needs first agent run

Open-source investment research platform for financial analysis agents.

No agent outcome reports yet. The first resolved run should report success, failed, not_relevant, blocked_by_risk, or setup_required.

No data
Success
0/100
Proven
0
Outcomes
—
Recent fail
—
Quality
77/100
Trust

POST /api/agent/outcome

Report after one narrow run.

Resolve responses include a unique feedback.event_id. Agents should reuse it when reporting the result so retries stay idempotent.

{
  "event_id": "resolve_...",
  "skill_slug": "crawl4ai",
  "task": "scrape pricing pages",
  "agent": "codex",
  "outcome": "success",
  "install_used": true,
  "time_to_useful_ms": 120000
}

Outcome meanings

Use the smallest honest label.

successThe skill helped complete the task.
failedThe skill was attempted but did not work.
not_relevantThe selected skill did not fit the task.
blocked_by_riskAudit, license, token, shell, or network risk stopped execution.
setup_requiredThe skill looked relevant but needed missing keys, data, or configuration.