Creator · agentscope-ai
Last updated · Sep 5, 2026
Use when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch. Also use when the user says "I need to start evaluating my app but don't know where to begin," "I want to set up eval for a new product," or has just identified fa
Sandbox only
Install targets
Codex install prompt
Install the "bootstrap" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/08-bootstrap. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Use when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch. Also use when the user says "I need to start evaluating my app but don't know where to begin," "I want to set up eval for a new product," or has just identified failure modes and needs to turn them into principles. Outputs a v0 grader in 30 minutes using OpenJudge SimpleRubricsGenerator, plus a roadmap to reach calibrated evaluation. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"agentscope-ai-bootstrap","task":"Install bootstrap","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Design assets, images, video, audio, multimodal media, presentation, and creative production skills.
Scenario
Design and creative
I need my agent to produce design assets, UI directions, presentations, or creative media workflows.
Agent fit
Claude Code + OpenAI Agents + CLI
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add agentscope-ai/OpenJudge --skill bootstrap
Maintenance
active
1mo since push
Risk
Needs review
The SKILL.md content appears truncated at the end (Step 5: Output Roadmap is incomplete). The provided excerpt ends abruptly with 'The v0' and lacks the full roadmap details.
GitHub quality
816
70/100 Quality · 76/100 Trust
Coverage tags
Review notes
The SKILL.md content appears truncated at the end (Step 5: Output Roadmap is incomplete). The provided excerpt ends abruptly with 'The v0' and lacks the full roadmap details. · Quality score needs review
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
StrongSolid option that is likely worth shortlisting for production workflows.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
816 GitHub stars
Repo activity
816 stars, 65 forks
Maintenance
1mo since push
License
Apache-2.0
Install
npx skills add agentscope-ai/OpenJudge --skill bootstrap
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add agentscope-ai/OpenJudge --skill bootstrapDo not use when
Alternative
175.1K Stars
npx skills add anthropics/skills --skill frontend-design
Alternative
85.2K Stars
npx skills add Leonxlnx/taste-skill --skill design-taste-frontend
Alternative
1.8K Stars
npx skills add Alisa0808/vox-director --skill vox-director
Alternative
175.1K Stars
npx skills add anthropics/skills --skill canvas-design
Agent safety v2
Usable candidate, but the agent should surface permission and audit notes before installation.
Require human approval before installing into a real workspace.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
medium
Skill may inspect schemas, query databases, or work with persistent stores.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20bootstrap%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20bootstrap%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/agentscope-ai-bootstrap/install
Agent should check
Copy prompt
Task: Use bootstrap in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20bootstrap%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/agentscope-ai-bootstrap/install
Install command: npx skills add agentscope-ai/OpenJudge --skill bootstrap
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/agentscope-ai-bootstrap/install
LLM text format
/api/skills/agentscope-ai-bootstrap/install?format=text
Find alternatives
/api/skills/search?q=bootstrap&limit=3
Agent prompt
Use bootstrap for this task. Review https://www.openagentskill.com/api/skills/agentscope-ai-bootstrap/install, then install with: npx skills add agentscope-ai/OpenJudge --skill bootstrapRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/agentscope-ai-bootstrap
LLM text
/api/registry/manifest/agentscope-ai-bootstrap?format=text
Install alias
/api/registry/install/agentscope-ai-bootstrap
Recommend
/api/registry/recommend?task=Use%20bootstrap%20in%20an%20agent%20workflow&limit=3
Agent fit
Browser automation
Use-case tags
Platforms
Claude Code, OpenAI Agents
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Shortlist this skill and compare it with close alternatives before production adoption.
Role in stack
Companion skill
Primary fit
Browser automation
Trust label
Strong shortlist
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
INFO816 GitHub stars
Stars/forks activity
INFO816 stars, 65 forks; issue activity unavailable in current metadata
Recent maintenance
PASS1mo since push
License clarity
PASSApache-2.0
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Solid option that is likely worth shortlisting for production workflows.
Workflow fit
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Investigate faster
I need my agent to research a topic, compare sources, and produce a concise report.
Verify behavior
I need my agent to test a web app, reproduce bugs, and verify fixes.
Workflow fit
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Design, build, test, and ship interfaces
A practical workflow for agents that turn product briefs or Figma designs into polished frontend code, review the result, test it in a browser, and prepare a safe deployment.
Find, compare, and synthesize
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Alternative shortlist
Similar skills that may fit this task.
Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.
Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.
Turn one topic into a narrated Vox-style paper-collage explainer or ad video, from script through captions.
Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.
--- name: bootstrap description: > Use when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch. Also use when the user says "I need to start evaluating my app but don't know where to begin," "I want to set up eval for a new product," or has just identified failure modes and needs to turn them into principles. Outputs a v0 grader in 30 minutes using OpenJudge SimpleRubricsGenerator, plus a roadmap to reach calibrated evaluation. ---
<HARD-GATE> NO v0 grader deployed WITHOUT explicitly marking it as uncalibrated. NO synthetic labels — LLM can generate eval inputs, but labels MUST come from real system output + human judgment. NO principle without a source label documenting where it came from. </HARD-GATE>
# Bootstrap
Cold-start an evaluation system when you have nothing. In 30 minutes you get a working v0 grader and a clear path to a calibrated, trustworthy evaluation.
> **Requires OpenJudge** (`pip install py-openjudge`) for the grader generators > (`SimpleRubricsGenerator` / `IterativeRubricsGenerator`). The interview, stratification, > and calibration-roadmap methodology is SDK-independent.
## Checklist
You MUST create a task for each item and complete them in order:
1. **Understand the product** — one-shot interview, not question-by-question 2. **Generate v0 grader** — use OpenJudge SimpleRubricsGenerator 3. **Synthesize eval inputs** — 30 inputs with 60/30/10 stratification 4. **Run v0 evaluation** — GradingRunner with the generated grader 5. **Output roadmap** — exactly how to reach 50 labels → calibrate
## Step 1: Product Interview (One Shot)
Ask the user to describe their system in one go:
``` To bootstrap your evaluation, I need to understand what you're building. Please describe (all at once):
- What does your system do? Who uses it? - What are 3 examples of perfect outputs? - What are 3 things the system must never do? - What failures worry you most? ```
Don't drip-feed these questions. One prompt, one answer. If the user provides a spec doc or design document instead, read that directly.
## Step 2: Generate v0 Grader
Use OpenJudge's `SimpleRubricsGenerator` to create a zero-shot grader from the product description:
```python import asyncio from openjudge.models.openai_chat_model import OpenAIChatModel from openjudge.generator.simple_rubric.generator import ( SimpleRubricsGenerator, SimpleRubricsGeneratorConfig, ) from openjudge.runner.grading_runner import GradingRunner
# OpenAIChatModel reads OPENAI_API_KEY / OPENAI_BASE_URL from the environment. # For Aliyun DashScope (Bailian): set OPENAI_BASE_URL to # https://dashscope.aliyuncs.com/compatible-mode/v1 and OPENAI_API_KEY to your key. model = OpenAIChatModel(model="qwen-plus") # or "gpt-4o", etc.
config = SimpleRubricsGeneratorConfig( grader_name="Initial Quality Grader", model=model, task_description="<summarize from the interview>", scenario="<usage context from interview>", min_score=0, max_score=1, )
generator = SimpleRubricsGenerator(config) grader = await generator.generate( dataset=[], sample_queries=[ "<example query 1 from interview>", "<example query 2 from interview>", "<example query 3 from interview>", ], ) ```
Why zero-shot instead of asking the user to write criteria? At this stage, the user doesn't know what "good" means operationally. The generator produces a reasonable starting point. The user refines it after seeing v0 results.
## Step 3: Synthesize Eval Inputs
Generate 30 test inputs with stratification. Use 3 different prompt templates for diversity:
``` Template 1: "Generate a typical {domain} query for a {user_type}" Template 2: "Create an ambiguous {domain} query where intent is unclear" Template 3: "Generate an edge-case {domain} query that's unusual but realistic" ```
Target distribution: - 60% common/typical queries (18 inputs) - 30% boundary/ambiguous queries (9 inputs) - 10% edge-case/unusual queries (3 inputs)
**Critical**: Generate inputs ONLY. Never generate labels. The labels come from running the actual system and getting human judgments.
```python # The dataset format for GradingRunner dataset = [ { "query": "What's the status of my order #12345?", "response": "<will be filled by running the system>", }, # ... 30 inputs ] ```
## Step 4: Run v0 Evaluation
Plug the generated grader into GradingRunner:
```python from openjudge.runner.grading_runner import GradingRunner from openjudge.graders.schema import GraderScore, GraderError
runner = GradingRunner( grader_configs={"v0_quality": grader}, max_concurrency=8, )
results = await runner.arun(dataset)
scores = [r.score for r in results["v0_quality"] if isinstance(r, GraderScore)] errors = [r for r in results["v0_quality"] if isinstance(r, GraderError)] print(f"V0 Results: avg={sum(scores)/len(scores):.2f}, errors={len(errors)}") ```
## Step 5: Output Roadmap
The v0 grader is uncalibrated — you don't know its TPR/TNR yet. Give the user an exact path to trustworthiness:
``` Your v0 evaluation is ready. Here's the path to a calibrated system:
Phase 1 (now): Run the v0 grader on 30 inputs to get a baseline. → The grader is UNCALIBRATED. Treat scores as directional, not definitive.
Phase 2 (1-2 weeks): Collect 50 human-labeled examples (25 pass + 25 fail). → For each system output, have a human mark pass/fail against the criterion. → Store labels in labels/<grader_name>.jsonl
Phase 3: When you have 50 labels, run 03-align-human to: → Measure TPR/TNR of the v0 grader → Detect biases (position, verbosity, self-enhancement) → Get a human-reduction roadmap
Phase 4: When TPR >= 0.8 and TNR >= 0.8: → The grader is calibrated and can be used as a production gate ```
## Quick Mode vs Deep Mode
- **Quick mode (default)**: Steps 1-5 above. 30 minutes to v0. Use when stakes=low or when exploring. - **Deep mode**: If the user has 20+ labeled examples, use `IterativeRubricsGenerator` instead of `SimpleRubricsGenerator` for data-driven grader creation:
```python from openjudge.generator.iterative_rubric.generator import ( IterativeRubricsGenerator, IterativePointwiseRubricsGeneratorConfig, )
config = IterativePointwiseRubricsGeneratorConfig( grader_name="Data-Driven Grader", model=model, task_description="<from interview>", min_score=0, max_score=1, max_epochs=3, batch_size=10, ) generator = IterativeRubricsGenerator(config) grader = await generator.generate(dataset=labeled_data) # 20+ labeled examples ```
## Red Flags — STOP and Re-evaluate
- "I'll generate both inputs and labels with the LLM to save time" → STOP. LLM-generating labels creates a self-consistency loop. TPR will look great until you test on real data, then it collapses. - "The v0 grader looks good, let's deploy it as a gate" → STOP. Uncalibrated graders have unknown TPR/TNR. They might pass everything or fail everything. - "I'll skip the roadmap, the user knows what to do next" → STOP. The roadmap IS the deliverable. Without it, bootstrap just produces an untrustworthy grader.
## Common Mistakes
- **Over-interviewing**. One prompt with 4 questions. Don't ask follow-ups unless the answers are genuinely unclear. - **Too many principles in v0**. SimpleRubricsGenerator works best with a focused task description. Don't try to evaluate 10 dimensions in v0 — start with the 2-3 most important ones. - **Skipping stratification in synthetic inputs**. If all 30 inputs are typical queries, you'll never see how the system handles edge cases. - **Presenting v0 scores as truth**. Always prefix v0 results with "UNVERIFIED — these scores are directional only."
## Next Skills
After `08-bootstrap`: - **`03-align-human`**: Once 50 human labels are collected, calibrate the grader. - **`01-eval-design`**: If you want a properly stratified dataset beyond the v0 30 inputs. - **`02-metric-design`**: If you need multiple graders for different dimensions.
Source provenance
Decision snapshot
816 GitHub stars
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for bootstrap, ready for a manual X post.
bootstrap: Use when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 e... 816 stars https://www.openagentskill.com/skills/agentscope-ai-bootstrap?ref=x
Listing + install path for bootstrap: https://www.openagentskill.com/skills/agentscope-ai-bootstrap?ref=x Install: npx skills add agentscope-ai/OpenJudge --skill bootstrap
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to agentscope-ai but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/agentscope-ai-bootstrap?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/agentscope-ai-bootstrap?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/agentscope-ai-bootstrap/audit)
[](https://www.openagentskill.com/skills/agentscope-ai-bootstrap?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)agentscope-ai
@agentscope-ai
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
Frontend Design
Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.
175.1K StarsTaste Skill: Anti-Slop Frontend
Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.
85.2K StarsVox Director
Turn one topic into a narrated Vox-style paper-collage explainer or ad video, from script through captions.
1.8K StarsCanvas Design
Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.
175.1K StarsPermission surface
filesystem or document access, database access
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness