ab-test-setup

Kuat · 80
Diindeks di Registry

When the user wants to plan, design, or implement an A/B test or experiment. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "conversion experiment," "statistical significance," or "tes

Verified installs0
Star24.8K
Versi1.0.0
Kualitas91/100 · Sangat baik
Kepercayaan80/100 · Tinjau sebelum memasang
Audit89/100 · Aman untuk dicoba

Profil aset

Riset dan pekerjaan pengetahuan

Deep research, source comparison, literature review, RAG, knowledge search, and reports.

Lihat kategori

Skenario

Agent riset

I need my agent to research a topic, compare sources, and produce a concise report.

Kecocokan Agent

Claude Code + CLI + Codex

Cocok untuk Codex, Claude Code, Cursor, CLI, atau Agent khusus.

Pasang

Siap

npx skills add alirezarezvani/claude-skills --skill ab-test-setup

Pemeliharaan

Terkini

Diperbarui hari ini

Risiko

Aman untuk dicoba

Quality score needs review

Kualitas GitHub

25K

91/100 Kualitas · 85/100 Kepercayaan

Tag cakupan

RisetAgent risetDesain dan kreatifagent-skill

Catatan ulasan

Quality score needs review

Kartu adopsi Agent

Kepercayaan, audit, dan kesiapan pemasangan dalam sekali lihat

Skor ini menggabungkan metadata repositori publik, sinyal ulasan OpenAgentSkill, kebaruan pemeliharaan, dan kesiapan pemasangan. Ini adalah sinyal shortlist, bukan pengganti peninjauan manusia.

Kualitas

Sangat baik
91

High-confidence pick with strong adoption and healthy maintenance signals.

Kepercayaan

Tinjau sebelum memasang
80

Sinyal shortlist yang baik, tetapi Agent harus meninjau catatan audit, kebijakan pemasangan, dan bukti hasil sebelum menjalankannya.

Audit

Aman untuk dicoba
89

Tinjauan yang dapat dibaca mesin tentang kesiapan pemasangan, metadata keamanan, pemeliharaan, dan risiko adopsi.

Trust Score OpenAgentSkill v5

Tinjauan manusia sebelum pemasangan

Gunakan sebagai kandidat utama setelah tinjauan manusia atau sandbox.

CodexClaude CodeCursorOpenAgentSkill CLI

Star

25K star GitHub

Aktivitas repositori

25K star dan 3.5K fork

Pemeliharaan

Diperbarui hari ini

Lisensi

MIT

Pasang

npx skills add alirezarezvani/claude-skills --skill ab-test-setup

Keamanan pemasangan

Jalur pemasangan paket atau runtime standar

Cakupan izin

shell or command execution, filesystem or document access

Hasil Agent

Belum ada data hasil Agent

Dokumentasi

Konteks README/SKILL.md kuat

Ringkasan risiko

Risiko metadata rendah

  • Quality score needs review

Kesiapan pemasangan

Jalur pemasangan tersedia

  • Jalur pemasangan tersedia
  • Bukti repositori tersedia
  • Lisensi dinyatakan
  • Belum ada bukti hasil Agent-Proven

Metadata yang dapat dibaca Agent

Data keputusan yang dapat dibaca mesin untuk skill ini.

Gunakan blok ini atau JSON tersemat untuk memutuskan apakah Agent perlu memasang skill ini, memilih alternatif, atau meminta tinjauan manusia terlebih dahulu.

Buka JSON

Tugas yang sesuai

  • Alur kerja Agent riset
  • Tim Claude Code
  • Tim yang menghargai sinyal adopsi GitHub
  • Sumber pencarian

Agent yang sesuai

CodexClaude CodeCursorOpenAgentSkill CLICLI

Keputusan pemasangan

Perintah
npx skills add alirezarezvani/claude-skills --skill ab-test-setup
Kebijakan
Tinjau
Tinjauan manusia
Ya

Kepercayaan dan risiko

Kepercayaan
80/100
Audit
89/100
Tingkat risiko
Aman untuk dicoba

Lingkar hasil

Endpoint
/api/agent/outcome
ID event
resolve
Hasil
5

Perintah pemasangan

npx skills add alirezarezvani/claude-skills --skill ab-test-setup

Jangan gunakan ketika

  • Tim yang membutuhkan SLA dengan dukungan vendor
  • Lingkungan berkompliansi tinggi tanpa tinjauan keamanan internal
  • No OpenAgentSkill engagement data yet
  • Petunjuk izin berisiko tinggi: eksekusi shell atau perintah
  • Quality score needs review

Keamanan Agent v2

61/100 · Tinjau sebelum memasang

Ditinjau dengan catatan izinTinjau

Kandidat yang dapat digunakan, tetapi Agent harus menampilkan catatan izin dan audit sebelum memasang.

Memerlukan persetujuan manusia sebelum memasang ke workspace nyata.

Selesaikan via API

Tinggi

Eksekusi shell atau perintah

Metadata skill merujuk terminal, CLI, shell, subprocess, atau alur kerja eksekusi perintah.

Sedang

Akses jaringan

Skill kemungkinan mengambil halaman jarak jauh, API, repositori, atau layanan eksternal.

Sedang

Akses sistem file

Skill dapat membaca atau menulis file proyek, dokumen, artefak yang dihasilkan, atau status workspace lokal.

  • Petunjuk izin berisiko tinggi: eksekusi shell atau perintah
  • Quality score needs review

Target pemasangan

Pasang skill ini di alur Agent Anda

Gunakan endpoint publik untuk mengambil perintah, checklist keamanan, prompt target, dan tautan kanonis.

skill install

OpenAgentSkill CLI

Resolve policy, run the source installer safely, and report a verified install receipt.

$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install alirezarezvani-ab-test-setup

Rencana resolusi Agent

Biarkan Agent memverifikasi kecocokan sebelum memasang.

API Resolve mengembalikan skill utama, alternatif, kebijakan keamanan, catatan audit, target pemasangan, dan prompt siap pakai.

Buka rencana teks

Agent harus memeriksa

  • Task fit and alternatives from Resolve API.
  • Audit score, trust score, and safety policy warnings.
  • Install target compatibility for Codex, Claude Code, Cursor, or CLI.

Salin prompt

Task: Use ab-test-setup in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20ab-test-setup%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/alirezarezvani-ab-test-setup/install
Install command: npx skills add alirezarezvani/claude-skills --skill ab-test-setup
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.

Serah-terima Agent

Berikan jalur pemasangan kepada Agent, bukan direktori lain.

Gunakan endpoint publik untuk mengambil perintah, checklist keamanan, prompt target, dan tautan kanonis.

Buka API pemasangan

Prompt Agent

Use ab-test-setup for this task. Review https://www.openagentskill.com/api/skills/alirezarezvani-ab-test-setup/install, then install with: npx skills add alirezarezvani/claude-skills --skill ab-test-setup

Metadata Registry

Profil yang dapat dibaca Agent untuk pemilihan skill otomatis.

API Registry menyediakan sinyal keputusan, kepercayaan, audit, use case, dan pemasangan tanpa mengikis UI.

Buka Manifest

Kecocokan Agent

100/100

Agent riset

Platform

Claude Code

Laporan audit

Aman untuk dicoba · 89/100

Tinjauan yang dapat dibaca mesin tentang kesiapan pemasangan, metadata keamanan, pemeliharaan, dan risiko adopsi.

Lihat laporan auditLihat laporan evaluasi

Panel keputusan Agent

Pilihan utama untuk Agent riset

Use this as a leading candidate, then validate the README and install path in your own agent stack.

100
Kesiapan
Adopsi
Tahap

Peran di stack

Pilihan utama

Kecocokan utama

Agent riset

Label kepercayaan

Siap produksi

Jalur pemasangan

Perintah siap

Gunakan saat

  • Alur kerja Agent riset
  • Tim Claude Code
  • Tim yang menghargai sinyal adopsi GitHub

Bukti

  • 24,795 star GitHub
  • recent repository activity
  • install command or GitHub repo available
  • profil kualitas 91/100

tinjau dulu

  • No OpenAgentSkill engagement data yet

Jalur implementasi

  1. 1Pasang di Agent sandbox dan jalankan satu tugas Agent riset dari awal hingga akhir.
  2. 2Compare output quality, latency, and failure behavior against at least one alternative.
  3. 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.

Profil kepercayaan

Tinjau sebelum memasang

Sinyal shortlist yang baik, tetapi Agent harus meninjau catatan audit, kebijakan pemasangan, dan bukti hasil sebelum menjalankannya.

80
Trust Score OpenAgentSkill

Adopsi GitHub

Lulus

25K star GitHub

Aktivitas star/fork

Lulus

25K star dan 3.5K fork; aktivitas issue tidak tersedia dalam metadata saat ini

Pemeliharaan terbaru

Lulus

Diperbarui hari ini

Kejelasan lisensi

Lulus

MIT

Sinyal positif

  • Tinjauan AI disetujui
  • Jalur pemasangan tersedia
  • Bukti repositori tersedia
  • Repositori yang baru dipelihara
  • Large GitHub adoption signal
  • Perintah pemasangan tidak memiliki pola berisiko tinggi yang jelas
  • Loop hasil siap tetapi membutuhkan eksekusi Agent nyata pertama

Tinjau sebelum memasang

  • Quality score needs review
  • Belum ada laporan hasil Agent nyata
  • Tinjauan manusia diperlukan sebelum pemasangan tanpa pengawasan

Tindakan yang disarankan

Gunakan sebagai kandidat utama setelah tinjauan manusia atau sandbox.

Profil kualitas

Sangat baik kandidat untuk alur kerja Agent

High-confidence pick with strong adoption and healthy maintenance signals.

91
Star GitHub
25K
Keterkinian
Hari ini
Siap dipasang
Ya
Lisensi
MIT

Kecocokan alur kerja

Gunakan skill ini pada skenario berikut

Kecocokan alur kerja

Tambahkan ke alur kerja lengkap

Daftar alternatif

Bandingkan sebelum memasang

Similar skills that may fit this task.

Bandingkan semua

Ringkasan

--- name: "ab-test-setup" description: When the user wants to plan, design, or implement an A/B test or experiment. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "conversion experiment," "statistical significance," or "test this." For tracking implementation, see analytics-tracking. license: MIT metadata: version: 1.0.0 author: Alireza Rezvani category: marketing updated: 2026-03-06 ---

# A/B Test Setup

You are an expert in experimentation and A/B testing. Your goal is to help design tests that produce statistically valid, actionable results.

## Initial Assessment

**Check for product marketing context first:** If `.claude/product-marketing-context.md` exists, read it before asking questions. Use that context and only ask for information not already covered or specific to this task.

Before designing a test, understand:

1. **Test Context** - What are you trying to improve? What change are you considering? 2. **Current State** - Baseline conversion rate? Current traffic volume? 3. **Constraints** - Technical complexity? Timeline? Tools available?

---

## Core Principles

### 1. Start with a Hypothesis - Not just "let's see what happens" - Specific prediction of outcome - Based on reasoning or data

### 2. Test One Thing - Single variable per test - Otherwise you don't know what worked

### 3. Statistical Rigor - Pre-determine sample size - Don't peek and stop early - Commit to the methodology

### 4. Measure What Matters - Primary metric tied to business value - Secondary metrics for context - Guardrail metrics to prevent harm

---

## Hypothesis Framework

### Structure

``` Because [observation/data], we believe [change] will cause [expected outcome] for [audience]. We'll know this is true when [metrics]. ```

### Example

**Weak**: "Changing the button color might increase clicks."

**Strong**: "Because users report difficulty finding the CTA (per heatmaps and feedback), we believe making the button larger and using contrasting color will increase CTA clicks by 15%+ for new visitors. We'll measure click-through rate from page view to signup start."

---

## Test Types

| Type | Description | Traffic Needed | |------|-------------|----------------| | A/B | Two versions, single change | Moderate | | A/B/n | Multiple variants | Higher | | MVT | Multiple changes in combinations | Very high | | Split URL | Different URLs for variants | Moderate |

---

## Sample Size

### Calculate It (bundled tool)

Use this skill's own calculator — don't eyeball it:

```bash python3 scripts/sample_size_calculator.py --baseline 0.05 --mde 0.20 # human-readable python3 scripts/sample_size_calculator.py --baseline 0.05 --mde 0.20 --json # for pipelines python3 scripts/sample_size_calculator.py --baseline 0.05 --mde 0.20 --daily-traffic 2000 # adds test-duration estimate ```

Paste `sample_size_per_variation` and the duration estimate directly into the test plan's "Sample size + duration" row before any test is approved to run.

### Quick Reference

Generated by `sample_size_calculator.py` (two-proportion z-test, α=0.05 two-tailed, 80% power; relative MDE):

| Baseline | 10% Lift | 20% Lift | 50% Lift | |----------|----------|----------|----------| | 1% | 163k/variant | 43k/variant | 7.7k/variant | | 3% | 53k/variant | 14k/variant | 2.5k/variant | | 5% | 31k/variant | 8.2k/variant | 1.5k/variant | | 10% | 15k/variant | 3.8k/variant | 683/variant |

**Cross-check calculators** (should agree with the script within rounding): - [Evan Miller's](https://www.evanmiller.org/ab-testing/sample-size.html) - [Optimizely's](https://www.optimizely.com/sample-size-calculator/)

**For detailed sample size tables and duration calculations**: See [references/sample-size-guide.md](references/sample-size-guide.md)

---

## Metrics Selection

### Primary Metric - Single metric that matters most - Directly tied to hypothesis - What you'll use to call the test

### Secondary Metrics - Support primary metric interpretation - Explain why/how the change worked

### Guardrail Metrics - Things that shouldn't get worse - Stop test if significantly negative

### Example: Pricing Page Test - **Primary**: Plan selection rate - **Secondary**: Time on page, plan distribution - **Guardrail**: Support tickets, refund rate

---

## Designing Variants

### What to Vary

| Category | Examples | |----------|----------| | Headlines/Copy | Message angle, value prop, specificity, tone | | Visual Design | Layout, color, images, hierarchy | | CTA | Button copy, size, placement, number | | Content | Information included, order, amount, social proof |

### Best Practices - Single, meaningful change - Bold enough to make a difference - True to the hypothesis

---

## Traffic Allocation

| Approach | Split | When to Use | |----------|-------|-------------| | Standard | 50/50 | Default for A/B | | Conservative | 90/10, 80/20 | Limit risk of bad variant | | Ramping | Start small, increase | Technical risk mitigation |

**Considerations:** - Consistency: Users see same variant on return - Balanced exposure across time of day/week

---

## Implementation

### Client-Side - JavaScript modifies page after load - Quick to implement, can cause flicker - Tools: PostHog, Optimizely, VWO

### Server-Side - Variant determined before render - No flicker, requires dev work - Tools: PostHog, LaunchDarkly, Split

---

## Running the Test

### Pre-Launch Checklist - [ ] Hypothesis documented - [ ] Primary metric defined - [ ] Sample size calculated - [ ] Variants implemented correctly - [ ] Tracking verified - [ ] QA completed on all variants

### During the Test

**DO:** - Monitor for technical issues - Check segment quality - Document external factors

**DON'T:** - Peek at results and stop early - Make changes to variants - Add traffic from new sources

### The Peeking Problem Looking at results before reaching sample size and stopping early leads to false positives and wrong decisions. Pre-commit to sample size and trust the process.

---

## Analyzing Results

### Statistical Significance - 95% confidence = p-value < 0.05 - Means <5% chance result is random - Not a guarantee—just a threshold

### Analysis Checklist

1. **Reach sample size?** If not, result is preliminary 2. **Statistically significant?** Check confidence intervals 3. **Effect size meaningful?** Compare to MDE, project impact 4. **Secondary metrics consistent?** Support the primary? 5. **Guardrail concerns?** Anything get worse? 6. **Segment differences?** Mobile vs. desktop? New vs. returning?

### Interpreting Results

| Result | Conclusion | |--------|------------| | Significant winner | Implement variant | | Significant loser | Keep control, learn why | | No significant difference | Need more traffic or bolder test | | Mixed signals | Dig deeper, maybe segment |

---

## Documentation

Document every test with: - Hypothesis - Variants (with screenshots) - Results (sample, metrics, significance) - Decision and learnings

**For templates**: See [references/test-templates.md](references/test-templates.md)

---

## Common Mistakes

### Test Design - Testing too small a change (undetectable) - Testing too many things (can't isolate) - No clear hypothesis

### Execution - Stopping early - Changing things mid-test - Not checking implementation

### Analysis - Ignoring confidence intervals - Cherry-picking segments - Over-interpreting inconclusive results

---

## Task-Specific Questions

1. What's your current conversion rate? 2. How much traffic does this page get? 3. What change are you considering and why? 4. What's the smallest improvement worth detecting? 5. What tools do you have for testing? 6. Have you tested this area before?

---

## Proactive Triggers

Proactively offer A/B test design when:

1. **Conversion rate mentioned** — User shares a conversion rate and asks how to improve it; suggest designing a test rather than guessing at solutions. 2. **Copy or design decision is unclear** — When two variants of a headline, CTA, or layout are being debated, propose testing instead of opinionating. 3. **Campaign underperformance** — User reports a landing page or email performing below expectations; offer a structured test plan. 4. **Pricing page discussion** — Any mention of pricing page changes should trigger an offer to design a pricing test with guardrail metrics. 5. **Post-launch review** — After a feature or campaign goes live, propose follow-up experiments to optimize the result.

---

## Output Artifacts

| Artifact | Format | Description | |----------|--------|-------------| | Experiment Brief | Markdown doc | Hypothesis, variants, metrics, sample size, duration, owner | | Sample Size Calculator Input | Table | Baseline rate, MDE, confidence level, power | | Pre-Launch QA Checklist | Checklist | Implementation, tracking, variant rendering verification | | Results Analysis Report | Markdown doc | Statistical significance, effect size, segment breakdown, decision | | Test Backlog | Prioritized list | Ranked experiments by expected impact and feasibility |

---

## Communication

All outputs should meet the quality standard: clear hypothesis, pre-registered metrics, and documented decisions. Avoid presenting inconclusive results as wins. Every test should produce a learning, even if the variant loses. Reference `marketing-context` for product and audience framing before designing experiments.

---

## Related Skills

- **page-cro** — USE when you need ideas for *what* to test; NOT when you already have a hypothesis and just need test design. - **analytics-tracking** — USE to set up measurement infrastructure before running tests; NOT as a substitute for defining primary metrics upfront. - **campaign-analytics** — USE after tests conclude to fold results into broader campaign attribution; NOT during the test itself. - **pricing-strategy** — USE when test results affect pricing decisions; NOT to replace a controlled test with pure strategic reasoning. - **marketing-context** — USE as foundation before any test design to ensure hypotheses align with ICP and positioning; always load first.

Detail teknis

Versi
1.0.0
Lisensi
MIT
Pembaruan terakhir
22 Agu 2026
Diterbitkan
22 Agu 2026

Ringkasan keputusan

Pilihan utama

100
Siap
Adopsi
Tahap

24,795 star GitHub

Audit

Tinjauan pemasangan

Tinjauan pemasangan dan adopsi

89
Aman untuk dicoba
Keamanan
84/100
Pemeliharaan
100/100
Pasang
92/100
Buka audit lengkapLihat laporan evaluasi

Bukti tervalidasi Agent

Bukti tervalidasi Agent

Laporan hasil setelah resolve, tinjau, pasang, dan satu eksekusi terbatas.

0
Terbukti
Needs first agent runPasang otomatis: tinjau duluTerakhir: Tidak diketahui
Tingkat sukses
Kegagalan terbaru
Hasil
0
Kualitas output
Gagal
0
Tidak relevan
0
Pemasangan
0
Diblokir risiko
0
Perlu penyiapan
0
Produksi
0

Belum ada data hasil Agent. Eksekusi pertama dapat melaporkan keberhasilan, kebutuhan setup, blok risiko, kegagalan, atau tidak relevan melalui /api/agent/outcome.

Pasang

Tambahkan ke alur Agent

Gratis dan sumber terbuka. Tinjau laporan sebelum memasang pada Agent produksi.

Siklus pertumbuhan

Kit berbagi

X

Draf berbasis skenario untuk ab-test-setup, siap untuk posting manual di X.

Catatan kurator
ab-test-setup: When the user wants to plan, design, or implement an A/B test or experiment. Also use when th...

24.8K stars

https://www.openagentskill.com/skills/alirezarezvani-ab-test-setup?ref=x
Buka draf X
Balasan opsional dengan perintah pemasangan
Listing + install path for ab-test-setup:
https://www.openagentskill.com/skills/alirezarezvani-ab-test-setup?ref=x

Install: npx skills add alirezarezvani/claude-skills --skill ab-test-setup
Buka draf balasan

Sumber listing

Diindeks Registry

Dapat diklaim

Listing ini diindeks dari sumber publik dan belum ditandai resmi hingga klaim pemelihara disetujui.

Diindeks oleh
Indeks komunitas OpenAgentSkill

Atribusi menautkan ke repositori publik atau profil kreator. Kreator dapat mengklaim listing untuk memperbarui sinyal kepemilikan.

Klaim skill ini

Klaim pemilik

Klaim listing skill ini

Listing Diindeks Registry ini dikaitkan dengan alirezarezvani, tetapi belum ditandai resmi. Klaim untuk menambahkan sinyal pemilik terverifikasi dan membuat pembaruan peluncuran, pemasangan, serta audit berikutnya lebih tepercaya.

Kit backlink kreator

Tambahkan badge bukti ke README Anda

Tampilkan listing kanonis, sinyal kepercayaan dan audit saat ini, serta bukti Agent-Proven nyata di tempat pengembang mengevaluasi repositori.

[![Listed on OpenAgentSkill](https://www.openagentskill.com/api/badge/alirezarezvani-ab-test-setup?metric=listed&label=Listed)](https://www.openagentskill.com/skills/alirezarezvani-ab-test-setup)
[![OpenAgentSkill Trust](https://www.openagentskill.com/api/badge/alirezarezvani-ab-test-setup?metric=trust&label=Trust)](https://www.openagentskill.com/skills/alirezarezvani-ab-test-setup)
[![OpenAgentSkill Audit](https://www.openagentskill.com/api/badge/alirezarezvani-ab-test-setup?metric=audit&label=Audit)](https://www.openagentskill.com/skills/alirezarezvani-ab-test-setup/audit)
[![Agent Proven](https://www.openagentskill.com/api/badge/alirezarezvani-ab-test-setup?metric=proven&label=Agent%20Proven)](https://www.openagentskill.com/skills/alirezarezvani-ab-test-setup)

Penulis

A

alirezarezvani

@alirezarezvani

Kecocokan platform

Sinyal kesehatan

Star GitHub
24.8K
Skor kualitas
54/100
Push GitHub terakhir
22 Agu 2026
Petunjuk framework
Tidak diketahui
Tampilan OpenAgentSkill
0
Salinan pemasangan
0
Klik keluar
0

Sinyal komunitas

Bagikan apakah skill ini bermanfaat untuk alur kerja Agent Anda. Masukan gabungan meningkatkan peringkat dari waktu ke waktu.

Kepercayaan & keamanan

Tinjau sebelum memasang

80
  • Adopsi GitHub25K star GitHubLulus
  • Aktivitas star/fork25K star dan 3.5K fork; aktivitas issue tidak tersedia dalam metadata saat iniLulus
  • Pemeliharaan terbaruDiperbarui hari iniLulus
  • Kejelasan lisensiMITLulus
  • Kelengkapan README/SKILL.mdMetadata memuat konteks penggunaan dan alur kerja yang cukupLulus
  • Risiko dependensi/runtimeCakupan eksekusi perintahInfo