ab-test-setup

강함 · 80
Registry 색인

When the user wants to plan, design, or implement an A/B test or experiment. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "conversion experiment," "statistical significance," or "tes

Verified installs0
스타24.8K
버전1.0.0
품질91/100 · 우수
신뢰80/100 · 검토 후 설치
감사89/100 · 안전하게 시도 가능

공급 자산 프로필

리서치 및 지식 작업

Deep research, source comparison, literature review, RAG, knowledge search, and reports.

트랙 보기

시나리오

리서치 Agent

I need my agent to research a topic, compare sources, and produce a concise report.

Agent 적합도

Claude Code + CLI + Codex

Codex, Claude Code, Cursor, CLI 또는 맞춤형 Agent에 적합합니다.

설치

준비됨

npx skills add alirezarezvani/claude-skills --skill ab-test-setup

유지보수

최신

오늘 푸시됨

위험

안전하게 시도 가능

Quality score needs review

GitHub 품질

25K

91/100 품질 · 85/100 신뢰

커버리지 태그

리서치리서치 Agent디자인 및 크리에이티브agent-skill

검토 메모

Quality score needs review

Agent 채택 스코어카드

신뢰, 감사, 설치 준비 상태를 한눈에 확인하세요

이 점수는 공개 저장소 메타데이터, OpenAgentSkill 검토 신호, 유지보수 최신성, 설치 준비 상태를 결합합니다. 후보 선정 신호일 뿐, 사람의 검토를 대체하지 않습니다.

품질

우수
91

강한 채택 및 유지보수 신호를 갖춘 신뢰도 높은 추천입니다.

신뢰

검토 후 설치
80

좋은 후보 신호이지만 Agent는 실행 전에 감사 메모, 설치 정책 및 결과 근거를 검토해야 합니다.

감사

안전하게 시도 가능
89

설치 준비 상태, 보안 메타데이터, 유지보수 및 채택 위험에 대한 기계 판독형 검토입니다.

OpenAgentSkill 신뢰 점수 v5

설치 전 사람 검토

사람 검토 또는 샌드박스 검증 후 우선 후보로 사용하세요.

CodexClaude CodeCursorOpenAgentSkill CLI

스타

GitHub 스타 25K

저장소 활동

스타 25K, 포크 3.5K

유지보수

오늘 푸시됨

라이선스

MIT

설치

npx skills add alirezarezvani/claude-skills --skill ab-test-setup

설치 안전성

표준 패키지 또는 런타임 설치 경로

권한 범위

shell or command execution, filesystem or document access

Agent 결과

아직 Agent 결과 데이터가 없습니다

문서

README/SKILL.md 맥락이 충분합니다

위험 요약

낮은 메타데이터 위험

  • Quality score needs review

설치 준비 상태

설치 경로 사용 가능

  • 설치 경로를 사용할 수 있습니다
  • 저장소 근거를 사용할 수 있습니다
  • 라이선스가 명시되었습니다
  • 아직 Agent 검증 결과 근거가 없습니다

Agent 읽기용 메타데이터

이 스킬의 기계 판독형 의사결정 데이터.

이 블록 또는 포함된 JSON을 사용해 Agent가 이 스킬을 설치할지, 대안을 고를지, 먼저 사람의 검토를 요청할지 판단할 수 있습니다.

JSON 열기

적합한 작업

  • 리서치 Agent 워크플로
  • Claude Code 팀
  • GitHub 채택 신호를 중시하는 팀
  • 검색 소스

적합한 Agent

CodexClaude CodeCursorOpenAgentSkill CLICLI

설치 결정

명령어
npx skills add alirezarezvani/claude-skills --skill ab-test-setup
정책
검토
사람 검토

신뢰와 위험

신뢰
80/100
감사
89/100
위험 수준
안전하게 시도 가능

결과 루프

엔드포인트
/api/agent/outcome
이벤트 ID
resolve
결과
5

설치 명령어

npx skills add alirezarezvani/claude-skills --skill ab-test-setup

사용하지 말아야 할 경우

  • 벤더 지원 SLA가 필요한 팀
  • 내부 보안 검토가 없는 고규정 준수 환경
  • 아직 OpenAgentSkill 사용 피드백 데이터가 없습니다
  • 고위험 권한 힌트: Shell 또는 명령 실행
  • Quality score needs review

Agent 안전 v2

61/100 · 설치 전 검토

권한 메모와 함께 검토됨검토

사용 가능한 후보이지만 Agent는 설치 전에 권한과 감사 메모를 표시해야 합니다.

실제 작업 공간에 설치하기 전에 사람의 승인이 필요합니다.

API로 해결

높음

Shell 또는 명령 실행

Skill 메타데이터가 터미널, CLI, Shell, 하위 프로세스 또는 명령 실행 워크플로를 참조합니다.

중간

네트워크 접근

Skill은 원격 페이지, API, 저장소 또는 외부 서비스에 접근할 수 있습니다.

중간

파일 시스템 접근

Skill은 프로젝트 파일, 문서, 생성 산출물 또는 로컬 작업 공간 상태를 읽거나 쓸 수 있습니다.

  • 고위험 권한 힌트: Shell 또는 명령 실행
  • Quality score needs review

설치 대상

Agent 워크플로에 이 스킬 설치

공개 설치 엔드포인트에서 명령어, 안전 체크리스트, 대상 프롬프트와 정규 링크를 가져옵니다.

skill install

OpenAgentSkill CLI

Resolve policy, run the source installer safely, and report a verified install receipt.

$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install alirezarezvani-ab-test-setup

Agent 해결 계획

설치 전에 Agent가 적합성을 검증하게 하세요.

Resolve API는 최우선 스킬, 대안, 안전 정책, 감사 메모, 설치 대상 및 Agent가 페이지를 스크래핑하지 않고 사용할 수 있는 프롬프트를 반환합니다.

텍스트 계획 열기

Agent가 확인할 항목

  • Resolve API에서 작업 적합도와 대안을 확인합니다.
  • 감사 점수, 신뢰 점수 및 안전 정책 경고를 확인합니다.
  • Codex, Claude Code, Cursor 또는 CLI의 설치 대상 호환성을 확인합니다.

프롬프트 복사

Task: Use ab-test-setup in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20ab-test-setup%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/alirezarezvani-ab-test-setup/install
Install command: npx skills add alirezarezvani/claude-skills --skill ab-test-setup
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.

Agent 핸드오프

또 다른 디렉터리 페이지 대신 설치 경로를 Agent에게 제공합니다.

공개 설치 엔드포인트에서 명령어, 안전 체크리스트, 대상 프롬프트와 정규 링크를 가져옵니다.

설치 API 열기

Agent 프롬프트

Use ab-test-setup for this task. Review https://www.openagentskill.com/api/skills/alirezarezvani-ab-test-setup/install, then install with: npx skills add alirezarezvani/claude-skills --skill ab-test-setup

Registry 메타데이터

자동 스킬 선택을 위한 Agent 읽기용 프로필.

Registry API를 통해 동일한 결정, 신뢰, 감사, 사용 사례, 설치 신호를 제공하므로 Agent가 UI를 스크래핑하지 않고도 순위를 매길 수 있습니다.

Manifest 열기

Agent 적합도

100/100

리서치 Agent

플랫폼

Claude Code

감사 보고서

안전하게 시도 가능 · 89/100

설치 준비 상태, 보안 메타데이터, 유지보수 및 채택 위험에 대한 기계 판독형 검토입니다.

감사 보고서 보기평가 보고서 보기

Agent 결정 패널

리서치 Agent용 우선 추천

우선 후보로 사용하되, 자신의 Agent 환경에서 README와 설치 경로를 검증하세요.

100
준비 상태
채택
단계

스택 내 역할

우선 추천

주요 적합도

리서치 Agent

신뢰 라벨

프로덕션 준비 완료

설치 경로

명령어 준비됨

사용 시점

  • 리서치 Agent 워크플로
  • Claude Code 팀
  • GitHub 채택 신호를 중시하는 팀

근거

  • GitHub 스타 24,795
  • 최근 저장소 활동
  • 설치 명령 또는 GitHub 저장소를 사용할 수 있습니다
  • 품질 프로필 91/100

먼저 검토

  • 아직 OpenAgentSkill 사용 피드백 데이터가 없습니다

구현 경로

  1. 1샌드박스 Agent에 설치하고 리서치 Agent 작업을 처음부터 끝까지 한 번 실행하세요.
  2. 2Compare output quality, latency, and failure behavior against at least one alternative.
  3. 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.

신뢰 프로필

검토 후 설치

좋은 후보 신호이지만 Agent는 실행 전에 감사 메모, 설치 정책 및 결과 근거를 검토해야 합니다.

80
OpenAgentSkill 신뢰 점수

GitHub 채택도

통과

GitHub 스타 25K

스타/포크 활동

통과

스타 25K, 포크 3.5K; 현재 메타데이터에서 이슈 활동을 확인할 수 없습니다

최근 유지보수

통과

오늘 푸시됨

라이선스 명확성

통과

MIT

긍정 신호

  • AI 검토 승인됨
  • 설치 경로를 사용할 수 있습니다
  • 저장소 근거를 사용할 수 있습니다
  • 최근 유지보수된 저장소
  • Large GitHub adoption signal
  • 설치 명령에서 뚜렷한 고위험 패턴이 발견되지 않았습니다
  • 결과 루프는 준비되었지만 첫 실제 Agent 실행이 필요합니다

설치 전 검토

  • Quality score needs review
  • 아직 실제 Agent 결과 보고서가 없습니다
  • 무인 설치 전에 사람 검토가 필요합니다

권장 작업

사람 검토 또는 샌드박스 검증 후 우선 후보로 사용하세요.

품질 프로필

우수 Agent 워크플로용 후보

강한 채택 및 유지보수 신호를 갖춘 신뢰도 높은 추천입니다.

91
GitHub 스타
25K
최신성
오늘
설치 준비됨
라이선스
MIT

워크플로 적합도

이 스킬을 사용할 시나리오

워크플로 적합도

완전한 워크플로에 추가

대안 후보

설치 전 비교

이 작업에 적합할 수 있는 유사 스킬입니다.

모두 비교

개요

--- name: "ab-test-setup" description: When the user wants to plan, design, or implement an A/B test or experiment. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "conversion experiment," "statistical significance," or "test this." For tracking implementation, see analytics-tracking. license: MIT metadata: version: 1.0.0 author: Alireza Rezvani category: marketing updated: 2026-03-06 ---

# A/B Test Setup

You are an expert in experimentation and A/B testing. Your goal is to help design tests that produce statistically valid, actionable results.

## Initial Assessment

**Check for product marketing context first:** If `.claude/product-marketing-context.md` exists, read it before asking questions. Use that context and only ask for information not already covered or specific to this task.

Before designing a test, understand:

1. **Test Context** - What are you trying to improve? What change are you considering? 2. **Current State** - Baseline conversion rate? Current traffic volume? 3. **Constraints** - Technical complexity? Timeline? Tools available?

---

## Core Principles

### 1. Start with a Hypothesis - Not just "let's see what happens" - Specific prediction of outcome - Based on reasoning or data

### 2. Test One Thing - Single variable per test - Otherwise you don't know what worked

### 3. Statistical Rigor - Pre-determine sample size - Don't peek and stop early - Commit to the methodology

### 4. Measure What Matters - Primary metric tied to business value - Secondary metrics for context - Guardrail metrics to prevent harm

---

## Hypothesis Framework

### Structure

``` Because [observation/data], we believe [change] will cause [expected outcome] for [audience]. We'll know this is true when [metrics]. ```

### Example

**Weak**: "Changing the button color might increase clicks."

**Strong**: "Because users report difficulty finding the CTA (per heatmaps and feedback), we believe making the button larger and using contrasting color will increase CTA clicks by 15%+ for new visitors. We'll measure click-through rate from page view to signup start."

---

## Test Types

| Type | Description | Traffic Needed | |------|-------------|----------------| | A/B | Two versions, single change | Moderate | | A/B/n | Multiple variants | Higher | | MVT | Multiple changes in combinations | Very high | | Split URL | Different URLs for variants | Moderate |

---

## Sample Size

### Calculate It (bundled tool)

Use this skill's own calculator — don't eyeball it:

```bash python3 scripts/sample_size_calculator.py --baseline 0.05 --mde 0.20 # human-readable python3 scripts/sample_size_calculator.py --baseline 0.05 --mde 0.20 --json # for pipelines python3 scripts/sample_size_calculator.py --baseline 0.05 --mde 0.20 --daily-traffic 2000 # adds test-duration estimate ```

Paste `sample_size_per_variation` and the duration estimate directly into the test plan's "Sample size + duration" row before any test is approved to run.

### Quick Reference

Generated by `sample_size_calculator.py` (two-proportion z-test, α=0.05 two-tailed, 80% power; relative MDE):

| Baseline | 10% Lift | 20% Lift | 50% Lift | |----------|----------|----------|----------| | 1% | 163k/variant | 43k/variant | 7.7k/variant | | 3% | 53k/variant | 14k/variant | 2.5k/variant | | 5% | 31k/variant | 8.2k/variant | 1.5k/variant | | 10% | 15k/variant | 3.8k/variant | 683/variant |

**Cross-check calculators** (should agree with the script within rounding): - [Evan Miller's](https://www.evanmiller.org/ab-testing/sample-size.html) - [Optimizely's](https://www.optimizely.com/sample-size-calculator/)

**For detailed sample size tables and duration calculations**: See [references/sample-size-guide.md](references/sample-size-guide.md)

---

## Metrics Selection

### Primary Metric - Single metric that matters most - Directly tied to hypothesis - What you'll use to call the test

### Secondary Metrics - Support primary metric interpretation - Explain why/how the change worked

### Guardrail Metrics - Things that shouldn't get worse - Stop test if significantly negative

### Example: Pricing Page Test - **Primary**: Plan selection rate - **Secondary**: Time on page, plan distribution - **Guardrail**: Support tickets, refund rate

---

## Designing Variants

### What to Vary

| Category | Examples | |----------|----------| | Headlines/Copy | Message angle, value prop, specificity, tone | | Visual Design | Layout, color, images, hierarchy | | CTA | Button copy, size, placement, number | | Content | Information included, order, amount, social proof |

### Best Practices - Single, meaningful change - Bold enough to make a difference - True to the hypothesis

---

## Traffic Allocation

| Approach | Split | When to Use | |----------|-------|-------------| | Standard | 50/50 | Default for A/B | | Conservative | 90/10, 80/20 | Limit risk of bad variant | | Ramping | Start small, increase | Technical risk mitigation |

**Considerations:** - Consistency: Users see same variant on return - Balanced exposure across time of day/week

---

## Implementation

### Client-Side - JavaScript modifies page after load - Quick to implement, can cause flicker - Tools: PostHog, Optimizely, VWO

### Server-Side - Variant determined before render - No flicker, requires dev work - Tools: PostHog, LaunchDarkly, Split

---

## Running the Test

### Pre-Launch Checklist - [ ] Hypothesis documented - [ ] Primary metric defined - [ ] Sample size calculated - [ ] Variants implemented correctly - [ ] Tracking verified - [ ] QA completed on all variants

### During the Test

**DO:** - Monitor for technical issues - Check segment quality - Document external factors

**DON'T:** - Peek at results and stop early - Make changes to variants - Add traffic from new sources

### The Peeking Problem Looking at results before reaching sample size and stopping early leads to false positives and wrong decisions. Pre-commit to sample size and trust the process.

---

## Analyzing Results

### Statistical Significance - 95% confidence = p-value < 0.05 - Means <5% chance result is random - Not a guarantee—just a threshold

### Analysis Checklist

1. **Reach sample size?** If not, result is preliminary 2. **Statistically significant?** Check confidence intervals 3. **Effect size meaningful?** Compare to MDE, project impact 4. **Secondary metrics consistent?** Support the primary? 5. **Guardrail concerns?** Anything get worse? 6. **Segment differences?** Mobile vs. desktop? New vs. returning?

### Interpreting Results

| Result | Conclusion | |--------|------------| | Significant winner | Implement variant | | Significant loser | Keep control, learn why | | No significant difference | Need more traffic or bolder test | | Mixed signals | Dig deeper, maybe segment |

---

## Documentation

Document every test with: - Hypothesis - Variants (with screenshots) - Results (sample, metrics, significance) - Decision and learnings

**For templates**: See [references/test-templates.md](references/test-templates.md)

---

## Common Mistakes

### Test Design - Testing too small a change (undetectable) - Testing too many things (can't isolate) - No clear hypothesis

### Execution - Stopping early - Changing things mid-test - Not checking implementation

### Analysis - Ignoring confidence intervals - Cherry-picking segments - Over-interpreting inconclusive results

---

## Task-Specific Questions

1. What's your current conversion rate? 2. How much traffic does this page get? 3. What change are you considering and why? 4. What's the smallest improvement worth detecting? 5. What tools do you have for testing? 6. Have you tested this area before?

---

## Proactive Triggers

Proactively offer A/B test design when:

1. **Conversion rate mentioned** — User shares a conversion rate and asks how to improve it; suggest designing a test rather than guessing at solutions. 2. **Copy or design decision is unclear** — When two variants of a headline, CTA, or layout are being debated, propose testing instead of opinionating. 3. **Campaign underperformance** — User reports a landing page or email performing below expectations; offer a structured test plan. 4. **Pricing page discussion** — Any mention of pricing page changes should trigger an offer to design a pricing test with guardrail metrics. 5. **Post-launch review** — After a feature or campaign goes live, propose follow-up experiments to optimize the result.

---

## Output Artifacts

| Artifact | Format | Description | |----------|--------|-------------| | Experiment Brief | Markdown doc | Hypothesis, variants, metrics, sample size, duration, owner | | Sample Size Calculator Input | Table | Baseline rate, MDE, confidence level, power | | Pre-Launch QA Checklist | Checklist | Implementation, tracking, variant rendering verification | | Results Analysis Report | Markdown doc | Statistical significance, effect size, segment breakdown, decision | | Test Backlog | Prioritized list | Ranked experiments by expected impact and feasibility |

---

## Communication

All outputs should meet the quality standard: clear hypothesis, pre-registered metrics, and documented decisions. Avoid presenting inconclusive results as wins. Every test should produce a learning, even if the variant loses. Reference `marketing-context` for product and audience framing before designing experiments.

---

## Related Skills

- **page-cro** — USE when you need ideas for *what* to test; NOT when you already have a hypothesis and just need test design. - **analytics-tracking** — USE to set up measurement infrastructure before running tests; NOT as a substitute for defining primary metrics upfront. - **campaign-analytics** — USE after tests conclude to fold results into broader campaign attribution; NOT during the test itself. - **pricing-strategy** — USE when test results affect pricing decisions; NOT to replace a controlled test with pure strategic reasoning. - **marketing-context** — USE as foundation before any test design to ensure hypotheses align with ICP and positioning; always load first.

기술 세부 사항

버전
1.0.0
라이선스
MIT
최근 업데이트
2026년 8월 22일
게시일
2026년 8월 22일

결정 스냅샷

우선 추천

100
준비됨
채택
단계

GitHub 스타 24,795

감사

설치 검토

설치 및 채택 검토

89
안전하게 시도 가능
보안
84/100
유지보수
100/100
설치
92/100
전체 감사 열기평가 보고서 보기

Agent 검증 증거

Agent 검증 증거

Resolve, 검토, 설치 및 한 번의 제한된 실행 후 결과 보고서입니다.

0
검증됨
Needs first agent run자동 설치: 먼저 검토최근: 알 수 없음
성공률
최근 실패
결과
0
출력 품질
실패
0
관련 없음
0
설치
0
위험 차단
0
설정 필요
0
프로덕션
0

아직 Agent 결과 데이터가 없습니다. 첫 실행은 /api/agent/outcome을 통해 성공, 설정 필요, 위험 차단, 실패 또는 비관련 결과를 보고할 수 있습니다.

설치

Agent 워크플로에 추가

무료 오픈 소스. 프로덕션 Agent에 설치하기 전에 보고서를 검토하세요.

성장 루프

공유 키트

X

ab-test-setup용 시나리오 기반 초안입니다. X에 수동으로 게시할 수 있습니다.

큐레이터 노트
ab-test-setup: When the user wants to plan, design, or implement an A/B test or experiment. Also use when th...

24.8K stars

https://www.openagentskill.com/skills/alirezarezvani-ab-test-setup?ref=x
X 초안 열기
선택 사항: 설치 명령이 포함된 답글
Listing + install path for ab-test-setup:
https://www.openagentskill.com/skills/alirezarezvani-ab-test-setup?ref=x

Install: npx skills add alirezarezvani/claude-skills --skill ab-test-setup

등록 출처

Registry 색인

소유권 주장 가능

이 등록은 공개 소스에서 색인되었으며 유지보수자 소유권 주장이 승인될 때까지 공식으로 표시되지 않습니다.

색인 주체
OpenAgentSkill 커뮤니티 인덱스

귀속은 공개 저장소 또는 제작자 프로필에 연결됩니다. 제작자는 등록을 주장하여 소유권 신호를 업데이트할 수 있습니다.

이 스킬 소유권 주장

소유자 소유권 주장

이 스킬 등록 소유권 주장

이 Registry 색인 등록은 alirezarezvani에게 귀속되어 있지만 아직 공식으로 표시되지 않았습니다. 소유권을 주장하면 확인된 소유자 신호가 추가되어 이후 출시, 설치 및 감사 업데이트를 더 신뢰할 수 있습니다.

크리에이터 백링크 키트

README에 증거 배지 추가

개발자가 저장소를 평가하는 위치에 정규 등록, 현재 신뢰 및 감사 신호, 실제 Agent-Proven 증거를 표시합니다.

[![Listed on OpenAgentSkill](https://www.openagentskill.com/api/badge/alirezarezvani-ab-test-setup?metric=listed&label=Listed)](https://www.openagentskill.com/skills/alirezarezvani-ab-test-setup)
[![OpenAgentSkill Trust](https://www.openagentskill.com/api/badge/alirezarezvani-ab-test-setup?metric=trust&label=Trust)](https://www.openagentskill.com/skills/alirezarezvani-ab-test-setup)
[![OpenAgentSkill Audit](https://www.openagentskill.com/api/badge/alirezarezvani-ab-test-setup?metric=audit&label=Audit)](https://www.openagentskill.com/skills/alirezarezvani-ab-test-setup/audit)
[![Agent Proven](https://www.openagentskill.com/api/badge/alirezarezvani-ab-test-setup?metric=proven&label=Agent%20Proven)](https://www.openagentskill.com/skills/alirezarezvani-ab-test-setup)

작성자

A

alirezarezvani

@alirezarezvani

플랫폼 적합도

상태 신호

GitHub 스타
24.8K
품질 점수
54/100
최근 GitHub 푸시
2026년 8월 22일
프레임워크 힌트
알 수 없음
OpenAgentSkill 조회수
0
설치 명령 복사
0
외부 클릭
0

커뮤니티 신호

이 스킬이 Agent 워크플로에 유용한지 알려 주세요. 집계된 피드백은 시간이 지날수록 순위를 개선합니다.

신뢰와 안전

검토 후 설치

80
  • GitHub 채택도GitHub 스타 25K통과
  • 스타/포크 활동스타 25K, 포크 3.5K; 현재 메타데이터에서 이슈 활동을 확인할 수 없습니다통과
  • 최근 유지보수오늘 푸시됨통과
  • 라이선스 명확성MIT통과
  • README/SKILL.md 완성도메타데이터에 충분한 사용 및 워크플로 맥락이 포함되어 있습니다통과
  • 의존성/런타임 위험명령 실행 범위정보