instrument-llm-observability
ContextJet-ai
Use this when adding tracing/observability to an LLM or AI-agent application - capturing prompts, tool calls, token usage, latency, and cost per step. Trigger whenever s…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
Page 100 · 16 shown · 1,645 public entries
Results: 1645
ContextJet-ai
Use this when adding tracing/observability to an LLM or AI-agent application - capturing prompts, tool calls, token usage, latency, and cost per step. Trigger whenever s…
ContextJet-ai
Use this to measure whether an AI agent actually completed its task end to end, not just whether individual LLM calls looked fine. Trigger on "is my agent working", "mea…
ContextJet-ai
Use this to measure and monitor the quality of a RAG (retrieval-augmented generation) pipeline - whether it retrieves the right context and answers faithfully. Trigger o…
ContextJet-ai
Use this to cut the cost of an LLM app using observability data. Trigger on "my OpenAI/Anthropic bill is too high", "reduce token usage", "the app is expensive", "optimi…
llopresto87
Maintain a project-local, version-pinned wiki of every external dependency. Use whenever a new library is being added, an existing one is being upgraded, an idiom for us…
llopresto87
Prove that a knowledge base actually works before trusting it — after building or adopting docs, a knowledge graph, or a wiki, verify it can orient a fresh agent and res…
tudoumashu
Maintain bounded, durable AI project memory in a repository's docs/ai/ pack. Use when Codex or Claude Code needs to create, read, update, or audit project memory (projec…
tudoumashu
Query and maintain the user's local LLM Wiki from any Codex or Claude Code conversation or project. Use when the user asks to search personal knowledge, recall prior res…
xSAVIKx
Guidance for an AI agent to enrich an Open Knowledge Format (OKF) bundle with high-quality concept descriptions using its own LLM — grounded in the bundle's schema, data…
mingdui
在风险分析完成后、写测试代码之前,生成中文业务用例并等待用户确认。这是整个 QA 流程中唯一的强制人工门禁——确认后不可回头改用例。 从 context、risk-analysis、已有用例索引出发,生成只含业务行为的 test-cases.json(不含技术检查如 Maven/build/安装), 渲染 test-cases.html…
PolicyEngine
Sensitivity registry for PolicyEngine microsim results — maps {program x deviation signature} to the calibration target or imputed variable most likely driving a mismatc…
bjcoombs
Compare two versions of an LLM-directed document - an original (teacher) and a candidate (student) - across a transfer set and return a per-case behavioural-equivalence…
sabahattink
Comprehensive prompt engineering framework for designing, optimizing, and iterating LLM prompts. Use when creating prompts, optimizing existing prompts, or improving AI…
dhruvanbhalara
Use when building AI workflows, tool calling agents, structured outputs, or LLM pipelines using the Genkit Dart SDK.
MarcosNahuel
Use the notebook knowledge base (the local SQLite RAG /agy:notebook builds from a folder of documents) to do precise, grounded, cited work — total amounts by category, f…
kennethkhoocy
Build human adjudication / hand-labeling sheets from LLM-pipeline data without evidence truncation. Use when: (1) preparing a CSV/Excel sheet for a human to rule on case…