add-llm-evals
ContextJet-ai
Use this when adding evaluation to an LLM/agent app - measuring output quality (correctness, faithfulness, relevance, safety) rather than just watching traces. Trigger o…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
Page 101 · 16 shown · 1,806 public entries
Results: 1806
ContextJet-ai
Use this when adding evaluation to an LLM/agent app - measuring output quality (correctness, faithfulness, relevance, safety) rather than just watching traces. Trigger o…
ContextJet-ai
Use this to build a good evaluation dataset for an LLM app, the part everyone underestimates. Trigger on "make an eval set", "what should I test my LLM on", "I don't hav…
ContextJet-ai
Use this to detect when an LLM is making things up, so you can flag or block confident-but-wrong answers before users see them. Trigger on "detect hallucinations", "is t…
ContextJet-ai
Use this to build or change an LLM feature the reliable way, by writing evals first and iterating against them, instead of tweaking prompts by vibes. Trigger on "how do…
fishzjp
API testing from OpenAPI/Swagger or case schemas: parameters, boundaries, auth, idempotency, concurrency, error responses, data consistency; runnable scripts; k6 handoff…
fishzjp
对已确认的 Bug 做根因定位、影响分析、回归建议时使用——复现 → 读代码定位根因 → 影响五面分析 → 回归建议,条目(根因/影响/Severity 依据/修复建议/回归建议五个扩展字段)落盘为 Bug 条目。不用于:仅收集 Bug 证据(automated-e2e-testing / api-testing)、疑似未定性缺陷(te…
fishzjp
Exploratory testing sessions when requirements are vague or docs are missing — charter-driven; outputs system understanding, risks, test ideas. Not for: automation prep,…
fishzjp
End-to-end QA entry: \"test this feature fully\" orchestrates requirements, strategy, cases, review, execution, bugs, regression, report; resumable. Single-stage tasks u…
fishzjp
After a code change (diff/fix/requirement change), decide what to regression-test: changed files → functions → features → cases traceability; outputs a ranked list. Not…
fishzjp
Govern flaky tests and suite reliability: rerun-pass verdicts, root-cause classes, quarantine, retry semantics, health metrics. Not for: in-run failures (e2e/api), triag…
llopresto87
Stand up a brand-new project from an empty (or near-empty) repo through the nine-phase from-scratch protocol — brainstorm, skeleton, grill, research, architecture, verif…
llopresto87
The test-SHAPING technique applied inside the test-first protocol — pick the lowest test level that exercises the behavior (unit, integration, contract, e2e, golden, pro…
llopresto87
Prove that a knowledge base actually works before trusting it — after building or adopting docs, a knowledge graph, or a wiki, verify it can orient a fresh agent and res…
blundergoat
Use when evaluating test coverage gaps, planning test strategy, or assessing testing risk for code changes.
mingdui
QA 流程倒数第二步的独立审查环节。在 qa-test-runner 执行完成后,以独立视角审查代码本身——检查架构、安全、并发正确性、数据一致性、 回归风险、测试覆盖缺口——不受执行方结论的影响。产出 code-review.json(含 severity 分级发现和 readiness 约束), 和 assert-code-revi…
mingdui
QA 上游阶段——为仓库构建事实画像和环境证据。当你需要收集模块的代码上下文、索引已有用例和测试、 检查前后端服务可达性、准备环境快照时触发。只收集事实,不做风险分析、不写用例、不判断质量。 通常在 quality-assurance-agent 的调度下作为第一阶段执行。 不适用:风险分析、用例设计、测试执行、质量判定(由下游阶段负责…