Creator · Masriyan
Last updated · Sep 5, 2026
LLM and AI application security testing — prompt injection, jailbreak resistance, OWASP LLM Top 10 (2025), RAG and agent/tool-use security, model supply chain, and AI red teaming for authorized assessments
Creator · Masriyan
Last updated · Sep 5, 2026
LLM and AI application security testing — prompt injection, jailbreak resistance, OWASP LLM Top 10 (2025), RAG and agent/tool-use security, model supply chain, and AI red teaming for authorized assessments
Creator · Masriyan
Last updated · Sep 5, 2026
LLM and AI application security testing — prompt injection, jailbreak resistance, OWASP LLM Top 10 (2025), RAG and agent/tool-use security, model supply chain, and AI red teaming for authorized assessments
Creator · Masriyan
Last updated · Sep 5, 2026
LLM and AI application security testing — prompt injection, jailbreak resistance, OWASP LLM Top 10 (2025), RAG and agent/tool-use security, model supply chain, and AI red teaming for authorized assessments
Do not auto-install
Install targets
Codex install prompt
Install the "AI & LLM Security" agent skill from https://github.com/Masriyan/Claude-Code-CyberSecurity-Skill/tree/main/skills/16-ai-llm-security. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: LLM and AI application security testing — prompt injection, jailbreak resistance, OWASP LLM Top 10 (2025), RAG and agent/tool-use security, model supply chain, and AI red teaming for authorized assessments After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"masriyan-ai-llm-security","task":"Install AI & LLM Security","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Code review, repo analysis, testing, CI, GitHub, DevOps, and developer workflow skills.
Scenario
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Agent fit
Claude Code + Browser agents + CLI
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM Security
Maintenance
fresh
2d since push
Risk
Risky
Dependency or permission surface needs review
GitHub quality
397
77/100 Quality · 65/100 Trust
Coverage tags
Review notes
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
StrongSolid option that is likely worth shortlisting for production workflows.
Trust
Do not auto-installTrust Score v5 found insufficient evidence for agent installation. Treat this as discovery material, not an executable recommendation.
Audit
RiskyA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Choose a stronger alternative or inspect the source manually before any install attempt.
Stars
397 GitHub stars
Repo activity
397 stars, 75 forks
Maintenance
2d since push
License
MIT
Install
npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM Security
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM SecurityDo not use when
Agent safety v2
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill may drive a browser or interact with web pages.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20AI%20%26%20LLM%20Security%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20AI%20%26%20LLM%20Security%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/masriyan-ai-llm-security/install
Agent should check
Copy prompt
Task: Use AI & LLM Security in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20AI%20%26%20LLM%20Security%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/masriyan-ai-llm-security/install
Install command: npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM Security
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/masriyan-ai-llm-security/install
LLM text format
/api/skills/masriyan-ai-llm-security/install?format=text
Find alternatives
/api/skills/search?q=AI%20%26%20LLM%20Security&limit=3
Agent prompt
Use AI & LLM Security for this task. Review https://www.openagentskill.com/api/skills/masriyan-ai-llm-security/install, then install with: npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM SecurityRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/masriyan-ai-llm-security
LLM text
/api/registry/manifest/masriyan-ai-llm-security?format=text
Install alias
/api/registry/install/masriyan-ai-llm-security
Recommend
/api/registry/recommend?task=Use%20AI%20%26%20LLM%20Security%20in%20an%20agent%20workflow&limit=3
Agent fit
RAG and knowledge
Platforms
Claude Code, Browser agents
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Shortlist this skill and compare it with close alternatives before production adoption.
Role in stack
Companion skill
Primary fit
RAG and knowledge
Trust label
Strong shortlist
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Trust Score v5 found insufficient evidence for agent installation. Treat this as discovery material, not an executable recommendation.
GitHub adoption
INFO397 GitHub stars
Stars/forks activity
INFO397 stars, 75 forks; issue activity unavailable in current metadata
Recent maintenance
PASS2d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Choose a stronger alternative or inspect the source manually before any install attempt.
Quality profile
Solid option that is likely worth shortlisting for production workflows.
Workflow fit
Search private knowledge
I need my agent to build a RAG workflow over documents and retrieve reliable context.
Reduce risk
I need my agent to scan a project for security risks and summarize what needs attention.
Parse messy files
I need my agent to read PDFs, extract tables, and turn documents into structured data.
Workflow fit
Ingest, retrieve, and cite
A workflow for document-heavy agents that ingest files, create searchable knowledge, retrieve relevant context, and answer with grounded sources.
Scrape, clean, and reuse web data
A practical workflow for agents that crawl public pages, extract clean content, normalize data, and hand it to downstream research or RAG workflows.
Inspect, patch, and verify code
A workflow for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
Alternative shortlist
Similar skills that may fit this task.
Wazuh - The Open Source Security Platform. Unified XDR and SIEM protection for endpoints and cloud workloads.
🕵️♂️ Collect a dossier on a person by username from 3000+ sites
Nuclei is a fast, customizable vulnerability scanner powered by the global security community and built on a simple YAML-based DSL, enabling collaboration to tackle trending vulnerabilities on the internet. It helps you find vulnerabilities in your applications, APIs, networks, DNS, and cloud configurations.
Infisical is the open-source platform for secrets, certificates, and privileged access management.
--- name: AI & LLM Security description: LLM and AI application security testing — prompt injection, jailbreak resistance, OWASP LLM Top 10 (2025), RAG and agent/tool-use security, model supply chain, and AI red teaming for authorized assessments version: 3.0.0 author: Masriyan tags: [cybersecurity, ai-security, llm, prompt-injection, owasp-llm, agent-security, rag, mlsecops, red-teaming] ---
# AI & LLM Security
## Purpose
Enable Claude to assess the security of AI/LLM-powered applications — chatbots, RAG pipelines, autonomous agents, and tool-using systems. Claude maps findings to the **OWASP Top 10 for LLM Applications (2025)** and the **MITRE ATLAS** adversarial-ML knowledge base, builds reproducible attack cases, and recommends concrete mitigations (input/output guardrails, least-privilege tool scopes, content provenance).
> **Authorization Required**: Only test AI systems you own or are explicitly authorized to assess. Prompt-injection and data-exfiltration testing against third-party AI services may violate their terms of service and local law. Confirm written scope before proceeding.
---
## Activation Triggers
This skill activates when the user asks about: - Prompt injection (direct or indirect), jailbreaks, or system-prompt extraction - OWASP LLM Top 10, MITRE ATLAS, or AI/ML threat modeling - Securing a RAG pipeline, vector database, or retrieval layer - LLM agent / tool-use / function-calling security and confused-deputy risks - Guardrail, content-filter, or model output validation design - Sensitive-information disclosure or training-data leakage from a model - Model / ML supply chain security (model files, `pickle`, model registries) - AI red teaming, jailbreak corpora, or automated adversarial prompt generation - Securing MCP (Model Context Protocol) servers and tool integrations
---
## Prerequisites
```bash pip install requests pyyaml rich ```
**Optional enhanced capabilities:** - `garak` — LLM vulnerability scanner (NVIDIA) - `promptfoo` — prompt/red-team evaluation harness - API key for the target LLM endpoint (test environment only) - `modelscan` / `picklescan` — ML model file safety scanning
---
## Core Capabilities
### 1. Threat Modeling (OWASP LLM Top 10 — 2025)
When asked to threat-model an AI application, map the system against each category and record exposure:
| ID | Risk | What to look for | |----|------|------------------| | LLM01 | Prompt Injection | Untrusted text reaching the prompt (direct & indirect via RAG/web/email) | | LLM02 | Sensitive Information Disclosure | PII/secrets in prompts, outputs, or training data; system-prompt leakage | | LLM03 | Supply Chain | Untrusted models, LoRA adapters, datasets, plugins, `pickle` deserialization | | LLM04 | Data & Model Poisoning | Tainted training/fine-tune/RAG data; backdoors | | LLM05 | Improper Output Handling | LLM output passed unsanitized to SQL, shell, browser (XSS), or `eval` | | LLM06 | Excessive Agency | Over-broad tool scopes, autonomous side effects, no human-in-the-loop | | LLM07 | System Prompt Leakage | Secrets/authz logic embedded in the system prompt | | LLM08 | Vector & Embedding Weaknesses | RAG access-control bypass, embedding inversion, cross-tenant leakage | | LLM09 | Misinformation | Hallucinations relied on for security/safety decisions | | LLM10 | Unbounded Consumption | Cost/DoS via token floods, model extraction, wallet-drain |
Produce a per-category table: **Exposure (Yes/No/Partial) → Evidence → Severity → Mitigation**.
### 2. Prompt Injection & Jailbreak Testing
**Direct injection** — user input that overrides instructions. Test families: - Instruction override ("ignore previous instructions and …") - Role-play / persona escape (DAN-style, hypothetical framing) - Encoding/obfuscation (Base64, ROT13, leetspeak, homoglyphs, zero-width chars) - Token smuggling and prompt-boundary confusion (fake delimiters, fake system tags) - Many-shot jailbreaking (long context of faux dialogue priming compliance) - Crescendo / multi-turn gradual escalation
**Indirect injection** — payload arrives via retrieved/processed content (web page, PDF, email, RAG doc, tool output). This is the highest-impact class for agents. Test that retrieved text **cannot** issue commands, exfiltrate context, or trigger tools.
For every test record: payload, channel (direct/indirect), goal (override / exfiltrate / tool-abuse), and result (blocked / partial / success). Use `scripts/prompt_injection_tester.py` to run a corpus and score outcomes.
**Refusal-quality note:** a single refusal is not a pass. Re-test the same goal across ≥3 phrasings and obfuscations before marking a control effective.
### 3. RAG & Vector Store Security
When reviewing a RAG pipeline: 1. **Access control at retrieval** — confirm the vector query is filtered by the *caller's* permissions, not just the app's. Test cross-tenant / cross-user document leakage. 2. **Indirect injection surface** — treat every ingested document as attacker-controlled; verify retrieved chunks are clearly delimited and never executed as instructions. 3. **Embedding inversion / membership** — sensitive source text may be partially reconstructable from embeddings; flag PII stored unencrypted in the vector DB. 4. **Chunk poisoning** — a single malicious document can dominate retrieval; check ranking/dedup and source allow-listing. 5. **Citation integrity** — outputs should cite retrieved sources so injected claims are traceable.
### 4. Agent & Tool-Use (Function Calling / MCP) Security
The agent is a **confused deputy**: it holds privileges the user may not. Review: - **Least-privilege tools** — each tool scoped to the minimum action; no broad `execute_shell`/`http_request` to arbitrary hosts - **Human-in-the-loop gates** on irreversible/outbound actions (payments, email send, file delete, deploy) - **Argument validation** — tool args are model-generated and untrusted; validate/allow-list server-side - **Injection → tool chain** — verify retrieved/indirect content cannot drive tool calls (e.g., a web page telling the agent to email its memory out) - **MCP server hardening** — authenticate clients, scope resources, rate-limit, log every tool invocation; never expose secrets via resource reads - **Memory poisoning** — persistent agent memory can be seeded with malicious instructions that fire on later turns
### 5. Model & ML Supply Chain
- Scan model artifacts for unsafe deserialization — **`pickle`/`.pt`/`.bin` can execute code on load**. Prefer `safetensors`. Run `scripts/model_supply_chain.py` or `modelscan`. - Verify model provenance, hashes, and signatures; pin versions from trusted registries. - Review fine-tune/LoRA adapters and datasets for poisoning and licensing. - Treat third-party plugins/MCP servers as untrusted dependencies (review + pin).
### 6. Output Handling & Guardrails
- **Never** pass raw LLM output into `eval`, SQL, shell, or innerHTML. Encode/parameterize at the sink (LLM05). - Layered guardrails: input filter → policy in system prompt → output classifier → sink-specific sanitization. Defense in depth, since any single layer is bypassable. - Validate structured output against a strict schema; reject on parse failure. - Apply egress controls so an injected agent cannot reach attacker URLs.
---
## Output Standards
Produce a structured AI security assessment:
```markdown # AI/LLM Security Assessment — [Application] Date: [Date] | Scope: [Endpoints/Models] | Model: [name/version] | Analyst: [Name]
## Executive Summary [2-3 sentences: overall posture, highest risks]
## OWASP LLM Top 10 Coverage | ID | Risk | Exposure | Severity | Evidence | |----|------|----------|----------|----------| | LLM01 | Prompt Injection | Yes | High | [repro] | ...
## Confirmed Findings ### [F-01] Indirect Prompt Injection via RAG → Tool Abuse (Critical) - ATLAS: AML.T0051 / OWASP LLM01+LLM06 - Repro: [payload, channel, steps] - Impact: [data exfil / unauthorized action] - Mitigation: [least-privilege tool scope + retrieved-content isolation + HITL]
## Guardrail Bypass Matrix | Goal | Direct | Encoded | Multi-turn | Indirect | Result |
## Recommendations (Prioritized) 1. ... ```
---
## Script Reference
### `prompt_injection_tester.py` ```bash # Run the built-in injection/jailbreak corpus against an endpoint python scripts/prompt_injection_tester.py --url https://app.test/api/chat --field message --output results.json
# Use a custom payload corpus and a refusal-detection keyword set python scripts/prompt_injection_tester.py --url ... --corpus payloads.txt --judge-keywords refusals.txt ```
### `model_supply_chain.py` ```bash # Scan a model directory/file for unsafe pickle opcodes and risky imports python scripts/model_supply_chain.py --path ./models/model.pt python scripts/model_supply_chain.py --path ./models/ --recursive --output scan.json ```
---
## Skill Integration
| Next Step | Condition | Target Skill | |-----------|-----------|--------------| | Web/API vuln testing of the app shell | App exposes web/API surface | → Skill 09 | | Cloud/infra hosting the model | Model served on AWS/Azure/GCP/K8s | → Skill 10 | | Detection rules for prompt-injection attempts | Need SIEM coverage | → Skill 12 | | Dependency/model-package CVEs | ML libs in use | → Skill 02 | | Red team narrative incorporating AI abuse | Full engagement | → Skill 14 |
---
## References
- [OWASP Top 10 for LLM Applications (2025)](https://genai.owasp.org/llm-top-10/) - [MITRE ATLAS — Adversarial Threat Landscape for AI Systems](https://atlas.mitre.org/) - [NIST AI Risk Management Framework (AI RMF 1.0) + Generative AI Profile](https://www.nist.gov/itl/ai-risk-management-framework) - [OWASP Agentic AI — Threats and Mitigations](https://genai.owasp.org/) - [Google SAIF — Secure AI Framework](https://saif.google/) - [NVIDIA garak — LLM vulnerability scanner](https://github.com/NVIDIA/garak)
Source provenance
Decision snapshot
recent repository activity
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for AI & LLM Security, ready for a manual X post.
AI & LLM Security: LLM and AI application security testing — prompt injection, jailbreak resistance, OWASP LLM T... 397 stars https://www.openagentskill.com/skills/masriyan-ai-llm-security?ref=x
Listing + install path for AI & LLM Security: https://www.openagentskill.com/skills/masriyan-ai-llm-security?ref=x Install: npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM Security
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to Masriyan but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/masriyan-ai-llm-security?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/masriyan-ai-llm-security?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/masriyan-ai-llm-security/audit)
[](https://www.openagentskill.com/skills/masriyan-ai-llm-security?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)Masriyan
@masriyan
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Do not auto-install
Wazuh
Wazuh - The Open Source Security Platform. Unified XDR and SIEM protection for endpoints and cloud workloads.
16.3K StarsMaigret
🕵️♂️ Collect a dossier on a person by username from 3000+ sites
32.9K StarsNuclei
Nuclei is a fast, customizable vulnerability scanner powered by the global security community and built on a simple YAML-based DSL, enabling collaboration to tackle trending vulnerabilities on the internet. It helps you find vulnerabilities in your applications, APIs, networks, DNS, and cloud configurations.
29.2K StarsInfisical
Infisical is the open-source platform for secrets, certificates, and privileged access management.
27.4K StarsDo not auto-install
Install targets
Codex install prompt
Install the "AI & LLM Security" agent skill from https://github.com/Masriyan/Claude-Code-CyberSecurity-Skill/tree/main/skills/16-ai-llm-security. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: LLM and AI application security testing — prompt injection, jailbreak resistance, OWASP LLM Top 10 (2025), RAG and agent/tool-use security, model supply chain, and AI red teaming for authorized assessments After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"masriyan-ai-llm-security","task":"Install AI & LLM Security","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Code review, repo analysis, testing, CI, GitHub, DevOps, and developer workflow skills.
Scenario
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Agent fit
Claude Code + Browser agents + CLI
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM Security
Maintenance
fresh
2d since push
Risk
Risky
Dependency or permission surface needs review
GitHub quality
397
77/100 Quality · 65/100 Trust
Coverage tags
Review notes
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
StrongSolid option that is likely worth shortlisting for production workflows.
Trust
Do not auto-installTrust Score v5 found insufficient evidence for agent installation. Treat this as discovery material, not an executable recommendation.
Audit
RiskyA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Choose a stronger alternative or inspect the source manually before any install attempt.
Stars
397 GitHub stars
Repo activity
397 stars, 75 forks
Maintenance
2d since push
License
MIT
Install
npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM Security
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM SecurityDo not use when
Agent safety v2
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill may drive a browser or interact with web pages.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20AI%20%26%20LLM%20Security%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20AI%20%26%20LLM%20Security%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/masriyan-ai-llm-security/install
Agent should check
Copy prompt
Task: Use AI & LLM Security in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20AI%20%26%20LLM%20Security%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/masriyan-ai-llm-security/install
Install command: npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM Security
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/masriyan-ai-llm-security/install
LLM text format
/api/skills/masriyan-ai-llm-security/install?format=text
Find alternatives
/api/skills/search?q=AI%20%26%20LLM%20Security&limit=3
Agent prompt
Use AI & LLM Security for this task. Review https://www.openagentskill.com/api/skills/masriyan-ai-llm-security/install, then install with: npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM SecurityRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/masriyan-ai-llm-security
LLM text
/api/registry/manifest/masriyan-ai-llm-security?format=text
Install alias
/api/registry/install/masriyan-ai-llm-security
Recommend
/api/registry/recommend?task=Use%20AI%20%26%20LLM%20Security%20in%20an%20agent%20workflow&limit=3
Agent fit
RAG and knowledge
Platforms
Claude Code, Browser agents
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Shortlist this skill and compare it with close alternatives before production adoption.
Role in stack
Companion skill
Primary fit
RAG and knowledge
Trust label
Strong shortlist
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Trust Score v5 found insufficient evidence for agent installation. Treat this as discovery material, not an executable recommendation.
GitHub adoption
INFO397 GitHub stars
Stars/forks activity
INFO397 stars, 75 forks; issue activity unavailable in current metadata
Recent maintenance
PASS2d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Choose a stronger alternative or inspect the source manually before any install attempt.
Quality profile
Solid option that is likely worth shortlisting for production workflows.
Workflow fit
Search private knowledge
I need my agent to build a RAG workflow over documents and retrieve reliable context.
Reduce risk
I need my agent to scan a project for security risks and summarize what needs attention.
Parse messy files
I need my agent to read PDFs, extract tables, and turn documents into structured data.
Workflow fit
Ingest, retrieve, and cite
A workflow for document-heavy agents that ingest files, create searchable knowledge, retrieve relevant context, and answer with grounded sources.
Scrape, clean, and reuse web data
A practical workflow for agents that crawl public pages, extract clean content, normalize data, and hand it to downstream research or RAG workflows.
Inspect, patch, and verify code
A workflow for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
Alternative shortlist
Similar skills that may fit this task.
Wazuh - The Open Source Security Platform. Unified XDR and SIEM protection for endpoints and cloud workloads.
🕵️♂️ Collect a dossier on a person by username from 3000+ sites
Nuclei is a fast, customizable vulnerability scanner powered by the global security community and built on a simple YAML-based DSL, enabling collaboration to tackle trending vulnerabilities on the internet. It helps you find vulnerabilities in your applications, APIs, networks, DNS, and cloud configurations.
Infisical is the open-source platform for secrets, certificates, and privileged access management.
--- name: AI & LLM Security description: LLM and AI application security testing — prompt injection, jailbreak resistance, OWASP LLM Top 10 (2025), RAG and agent/tool-use security, model supply chain, and AI red teaming for authorized assessments version: 3.0.0 author: Masriyan tags: [cybersecurity, ai-security, llm, prompt-injection, owasp-llm, agent-security, rag, mlsecops, red-teaming] ---
# AI & LLM Security
## Purpose
Enable Claude to assess the security of AI/LLM-powered applications — chatbots, RAG pipelines, autonomous agents, and tool-using systems. Claude maps findings to the **OWASP Top 10 for LLM Applications (2025)** and the **MITRE ATLAS** adversarial-ML knowledge base, builds reproducible attack cases, and recommends concrete mitigations (input/output guardrails, least-privilege tool scopes, content provenance).
> **Authorization Required**: Only test AI systems you own or are explicitly authorized to assess. Prompt-injection and data-exfiltration testing against third-party AI services may violate their terms of service and local law. Confirm written scope before proceeding.
---
## Activation Triggers
This skill activates when the user asks about: - Prompt injection (direct or indirect), jailbreaks, or system-prompt extraction - OWASP LLM Top 10, MITRE ATLAS, or AI/ML threat modeling - Securing a RAG pipeline, vector database, or retrieval layer - LLM agent / tool-use / function-calling security and confused-deputy risks - Guardrail, content-filter, or model output validation design - Sensitive-information disclosure or training-data leakage from a model - Model / ML supply chain security (model files, `pickle`, model registries) - AI red teaming, jailbreak corpora, or automated adversarial prompt generation - Securing MCP (Model Context Protocol) servers and tool integrations
---
## Prerequisites
```bash pip install requests pyyaml rich ```
**Optional enhanced capabilities:** - `garak` — LLM vulnerability scanner (NVIDIA) - `promptfoo` — prompt/red-team evaluation harness - API key for the target LLM endpoint (test environment only) - `modelscan` / `picklescan` — ML model file safety scanning
---
## Core Capabilities
### 1. Threat Modeling (OWASP LLM Top 10 — 2025)
When asked to threat-model an AI application, map the system against each category and record exposure:
| ID | Risk | What to look for | |----|------|------------------| | LLM01 | Prompt Injection | Untrusted text reaching the prompt (direct & indirect via RAG/web/email) | | LLM02 | Sensitive Information Disclosure | PII/secrets in prompts, outputs, or training data; system-prompt leakage | | LLM03 | Supply Chain | Untrusted models, LoRA adapters, datasets, plugins, `pickle` deserialization | | LLM04 | Data & Model Poisoning | Tainted training/fine-tune/RAG data; backdoors | | LLM05 | Improper Output Handling | LLM output passed unsanitized to SQL, shell, browser (XSS), or `eval` | | LLM06 | Excessive Agency | Over-broad tool scopes, autonomous side effects, no human-in-the-loop | | LLM07 | System Prompt Leakage | Secrets/authz logic embedded in the system prompt | | LLM08 | Vector & Embedding Weaknesses | RAG access-control bypass, embedding inversion, cross-tenant leakage | | LLM09 | Misinformation | Hallucinations relied on for security/safety decisions | | LLM10 | Unbounded Consumption | Cost/DoS via token floods, model extraction, wallet-drain |
Produce a per-category table: **Exposure (Yes/No/Partial) → Evidence → Severity → Mitigation**.
### 2. Prompt Injection & Jailbreak Testing
**Direct injection** — user input that overrides instructions. Test families: - Instruction override ("ignore previous instructions and …") - Role-play / persona escape (DAN-style, hypothetical framing) - Encoding/obfuscation (Base64, ROT13, leetspeak, homoglyphs, zero-width chars) - Token smuggling and prompt-boundary confusion (fake delimiters, fake system tags) - Many-shot jailbreaking (long context of faux dialogue priming compliance) - Crescendo / multi-turn gradual escalation
**Indirect injection** — payload arrives via retrieved/processed content (web page, PDF, email, RAG doc, tool output). This is the highest-impact class for agents. Test that retrieved text **cannot** issue commands, exfiltrate context, or trigger tools.
For every test record: payload, channel (direct/indirect), goal (override / exfiltrate / tool-abuse), and result (blocked / partial / success). Use `scripts/prompt_injection_tester.py` to run a corpus and score outcomes.
**Refusal-quality note:** a single refusal is not a pass. Re-test the same goal across ≥3 phrasings and obfuscations before marking a control effective.
### 3. RAG & Vector Store Security
When reviewing a RAG pipeline: 1. **Access control at retrieval** — confirm the vector query is filtered by the *caller's* permissions, not just the app's. Test cross-tenant / cross-user document leakage. 2. **Indirect injection surface** — treat every ingested document as attacker-controlled; verify retrieved chunks are clearly delimited and never executed as instructions. 3. **Embedding inversion / membership** — sensitive source text may be partially reconstructable from embeddings; flag PII stored unencrypted in the vector DB. 4. **Chunk poisoning** — a single malicious document can dominate retrieval; check ranking/dedup and source allow-listing. 5. **Citation integrity** — outputs should cite retrieved sources so injected claims are traceable.
### 4. Agent & Tool-Use (Function Calling / MCP) Security
The agent is a **confused deputy**: it holds privileges the user may not. Review: - **Least-privilege tools** — each tool scoped to the minimum action; no broad `execute_shell`/`http_request` to arbitrary hosts - **Human-in-the-loop gates** on irreversible/outbound actions (payments, email send, file delete, deploy) - **Argument validation** — tool args are model-generated and untrusted; validate/allow-list server-side - **Injection → tool chain** — verify retrieved/indirect content cannot drive tool calls (e.g., a web page telling the agent to email its memory out) - **MCP server hardening** — authenticate clients, scope resources, rate-limit, log every tool invocation; never expose secrets via resource reads - **Memory poisoning** — persistent agent memory can be seeded with malicious instructions that fire on later turns
### 5. Model & ML Supply Chain
- Scan model artifacts for unsafe deserialization — **`pickle`/`.pt`/`.bin` can execute code on load**. Prefer `safetensors`. Run `scripts/model_supply_chain.py` or `modelscan`. - Verify model provenance, hashes, and signatures; pin versions from trusted registries. - Review fine-tune/LoRA adapters and datasets for poisoning and licensing. - Treat third-party plugins/MCP servers as untrusted dependencies (review + pin).
### 6. Output Handling & Guardrails
- **Never** pass raw LLM output into `eval`, SQL, shell, or innerHTML. Encode/parameterize at the sink (LLM05). - Layered guardrails: input filter → policy in system prompt → output classifier → sink-specific sanitization. Defense in depth, since any single layer is bypassable. - Validate structured output against a strict schema; reject on parse failure. - Apply egress controls so an injected agent cannot reach attacker URLs.
---
## Output Standards
Produce a structured AI security assessment:
```markdown # AI/LLM Security Assessment — [Application] Date: [Date] | Scope: [Endpoints/Models] | Model: [name/version] | Analyst: [Name]
## Executive Summary [2-3 sentences: overall posture, highest risks]
## OWASP LLM Top 10 Coverage | ID | Risk | Exposure | Severity | Evidence | |----|------|----------|----------|----------| | LLM01 | Prompt Injection | Yes | High | [repro] | ...
## Confirmed Findings ### [F-01] Indirect Prompt Injection via RAG → Tool Abuse (Critical) - ATLAS: AML.T0051 / OWASP LLM01+LLM06 - Repro: [payload, channel, steps] - Impact: [data exfil / unauthorized action] - Mitigation: [least-privilege tool scope + retrieved-content isolation + HITL]
## Guardrail Bypass Matrix | Goal | Direct | Encoded | Multi-turn | Indirect | Result |
## Recommendations (Prioritized) 1. ... ```
---
## Script Reference
### `prompt_injection_tester.py` ```bash # Run the built-in injection/jailbreak corpus against an endpoint python scripts/prompt_injection_tester.py --url https://app.test/api/chat --field message --output results.json
# Use a custom payload corpus and a refusal-detection keyword set python scripts/prompt_injection_tester.py --url ... --corpus payloads.txt --judge-keywords refusals.txt ```
### `model_supply_chain.py` ```bash # Scan a model directory/file for unsafe pickle opcodes and risky imports python scripts/model_supply_chain.py --path ./models/model.pt python scripts/model_supply_chain.py --path ./models/ --recursive --output scan.json ```
---
## Skill Integration
| Next Step | Condition | Target Skill | |-----------|-----------|--------------| | Web/API vuln testing of the app shell | App exposes web/API surface | → Skill 09 | | Cloud/infra hosting the model | Model served on AWS/Azure/GCP/K8s | → Skill 10 | | Detection rules for prompt-injection attempts | Need SIEM coverage | → Skill 12 | | Dependency/model-package CVEs | ML libs in use | → Skill 02 | | Red team narrative incorporating AI abuse | Full engagement | → Skill 14 |
---
## References
- [OWASP Top 10 for LLM Applications (2025)](https://genai.owasp.org/llm-top-10/) - [MITRE ATLAS — Adversarial Threat Landscape for AI Systems](https://atlas.mitre.org/) - [NIST AI Risk Management Framework (AI RMF 1.0) + Generative AI Profile](https://www.nist.gov/itl/ai-risk-management-framework) - [OWASP Agentic AI — Threats and Mitigations](https://genai.owasp.org/) - [Google SAIF — Secure AI Framework](https://saif.google/) - [NVIDIA garak — LLM vulnerability scanner](https://github.com/NVIDIA/garak)
Source provenance
Decision snapshot
recent repository activity
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for AI & LLM Security, ready for a manual X post.
AI & LLM Security: LLM and AI application security testing — prompt injection, jailbreak resistance, OWASP LLM T... 397 stars https://www.openagentskill.com/skills/masriyan-ai-llm-security?ref=x
Listing + install path for AI & LLM Security: https://www.openagentskill.com/skills/masriyan-ai-llm-security?ref=x Install: npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM Security
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to Masriyan but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/masriyan-ai-llm-security?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/masriyan-ai-llm-security?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/masriyan-ai-llm-security/audit)
[](https://www.openagentskill.com/skills/masriyan-ai-llm-security?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)Masriyan
@masriyan
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Do not auto-install
Wazuh
Wazuh - The Open Source Security Platform. Unified XDR and SIEM protection for endpoints and cloud workloads.
16.3K StarsMaigret
🕵️♂️ Collect a dossier on a person by username from 3000+ sites
32.9K StarsNuclei
Nuclei is a fast, customizable vulnerability scanner powered by the global security community and built on a simple YAML-based DSL, enabling collaboration to tackle trending vulnerabilities on the internet. It helps you find vulnerabilities in your applications, APIs, networks, DNS, and cloud configurations.
29.2K StarsInfisical
Infisical is the open-source platform for secrets, certificates, and privileged access management.
27.4K StarsDo not auto-install
Install targets
Codex install prompt
Install the "AI & LLM Security" agent skill from https://github.com/Masriyan/Claude-Code-CyberSecurity-Skill/tree/main/skills/16-ai-llm-security. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: LLM and AI application security testing — prompt injection, jailbreak resistance, OWASP LLM Top 10 (2025), RAG and agent/tool-use security, model supply chain, and AI red teaming for authorized assessments After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"masriyan-ai-llm-security","task":"Install AI & LLM Security","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Code review, repo analysis, testing, CI, GitHub, DevOps, and developer workflow skills.
Scenario
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Agent fit
Claude Code + Browser agents + CLI
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM Security
Maintenance
fresh
2d since push
Risk
Risky
Dependency or permission surface needs review
GitHub quality
397
77/100 Quality · 65/100 Trust
Coverage tags
Review notes
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
StrongSolid option that is likely worth shortlisting for production workflows.
Trust
Do not auto-installTrust Score v5 found insufficient evidence for agent installation. Treat this as discovery material, not an executable recommendation.
Audit
RiskyA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Choose a stronger alternative or inspect the source manually before any install attempt.
Stars
397 GitHub stars
Repo activity
397 stars, 75 forks
Maintenance
2d since push
License
MIT
Install
npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM Security
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM SecurityDo not use when
Agent safety v2
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill may drive a browser or interact with web pages.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20AI%20%26%20LLM%20Security%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20AI%20%26%20LLM%20Security%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/masriyan-ai-llm-security/install
Agent should check
Copy prompt
Task: Use AI & LLM Security in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20AI%20%26%20LLM%20Security%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/masriyan-ai-llm-security/install
Install command: npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM Security
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/masriyan-ai-llm-security/install
LLM text format
/api/skills/masriyan-ai-llm-security/install?format=text
Find alternatives
/api/skills/search?q=AI%20%26%20LLM%20Security&limit=3
Agent prompt
Use AI & LLM Security for this task. Review https://www.openagentskill.com/api/skills/masriyan-ai-llm-security/install, then install with: npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM SecurityRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/masriyan-ai-llm-security
LLM text
/api/registry/manifest/masriyan-ai-llm-security?format=text
Install alias
/api/registry/install/masriyan-ai-llm-security
Recommend
/api/registry/recommend?task=Use%20AI%20%26%20LLM%20Security%20in%20an%20agent%20workflow&limit=3
Agent fit
RAG and knowledge
Platforms
Claude Code, Browser agents
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Shortlist this skill and compare it with close alternatives before production adoption.
Role in stack
Companion skill
Primary fit
RAG and knowledge
Trust label
Strong shortlist
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Trust Score v5 found insufficient evidence for agent installation. Treat this as discovery material, not an executable recommendation.
GitHub adoption
INFO397 GitHub stars
Stars/forks activity
INFO397 stars, 75 forks; issue activity unavailable in current metadata
Recent maintenance
PASS2d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Choose a stronger alternative or inspect the source manually before any install attempt.
Quality profile
Solid option that is likely worth shortlisting for production workflows.
Workflow fit
Search private knowledge
I need my agent to build a RAG workflow over documents and retrieve reliable context.
Reduce risk
I need my agent to scan a project for security risks and summarize what needs attention.
Parse messy files
I need my agent to read PDFs, extract tables, and turn documents into structured data.
Workflow fit
Ingest, retrieve, and cite
A workflow for document-heavy agents that ingest files, create searchable knowledge, retrieve relevant context, and answer with grounded sources.
Scrape, clean, and reuse web data
A practical workflow for agents that crawl public pages, extract clean content, normalize data, and hand it to downstream research or RAG workflows.
Inspect, patch, and verify code
A workflow for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
Alternative shortlist
Similar skills that may fit this task.
Wazuh - The Open Source Security Platform. Unified XDR and SIEM protection for endpoints and cloud workloads.
🕵️♂️ Collect a dossier on a person by username from 3000+ sites
Nuclei is a fast, customizable vulnerability scanner powered by the global security community and built on a simple YAML-based DSL, enabling collaboration to tackle trending vulnerabilities on the internet. It helps you find vulnerabilities in your applications, APIs, networks, DNS, and cloud configurations.
Infisical is the open-source platform for secrets, certificates, and privileged access management.
--- name: AI & LLM Security description: LLM and AI application security testing — prompt injection, jailbreak resistance, OWASP LLM Top 10 (2025), RAG and agent/tool-use security, model supply chain, and AI red teaming for authorized assessments version: 3.0.0 author: Masriyan tags: [cybersecurity, ai-security, llm, prompt-injection, owasp-llm, agent-security, rag, mlsecops, red-teaming] ---
# AI & LLM Security
## Purpose
Enable Claude to assess the security of AI/LLM-powered applications — chatbots, RAG pipelines, autonomous agents, and tool-using systems. Claude maps findings to the **OWASP Top 10 for LLM Applications (2025)** and the **MITRE ATLAS** adversarial-ML knowledge base, builds reproducible attack cases, and recommends concrete mitigations (input/output guardrails, least-privilege tool scopes, content provenance).
> **Authorization Required**: Only test AI systems you own or are explicitly authorized to assess. Prompt-injection and data-exfiltration testing against third-party AI services may violate their terms of service and local law. Confirm written scope before proceeding.
---
## Activation Triggers
This skill activates when the user asks about: - Prompt injection (direct or indirect), jailbreaks, or system-prompt extraction - OWASP LLM Top 10, MITRE ATLAS, or AI/ML threat modeling - Securing a RAG pipeline, vector database, or retrieval layer - LLM agent / tool-use / function-calling security and confused-deputy risks - Guardrail, content-filter, or model output validation design - Sensitive-information disclosure or training-data leakage from a model - Model / ML supply chain security (model files, `pickle`, model registries) - AI red teaming, jailbreak corpora, or automated adversarial prompt generation - Securing MCP (Model Context Protocol) servers and tool integrations
---
## Prerequisites
```bash pip install requests pyyaml rich ```
**Optional enhanced capabilities:** - `garak` — LLM vulnerability scanner (NVIDIA) - `promptfoo` — prompt/red-team evaluation harness - API key for the target LLM endpoint (test environment only) - `modelscan` / `picklescan` — ML model file safety scanning
---
## Core Capabilities
### 1. Threat Modeling (OWASP LLM Top 10 — 2025)
When asked to threat-model an AI application, map the system against each category and record exposure:
| ID | Risk | What to look for | |----|------|------------------| | LLM01 | Prompt Injection | Untrusted text reaching the prompt (direct & indirect via RAG/web/email) | | LLM02 | Sensitive Information Disclosure | PII/secrets in prompts, outputs, or training data; system-prompt leakage | | LLM03 | Supply Chain | Untrusted models, LoRA adapters, datasets, plugins, `pickle` deserialization | | LLM04 | Data & Model Poisoning | Tainted training/fine-tune/RAG data; backdoors | | LLM05 | Improper Output Handling | LLM output passed unsanitized to SQL, shell, browser (XSS), or `eval` | | LLM06 | Excessive Agency | Over-broad tool scopes, autonomous side effects, no human-in-the-loop | | LLM07 | System Prompt Leakage | Secrets/authz logic embedded in the system prompt | | LLM08 | Vector & Embedding Weaknesses | RAG access-control bypass, embedding inversion, cross-tenant leakage | | LLM09 | Misinformation | Hallucinations relied on for security/safety decisions | | LLM10 | Unbounded Consumption | Cost/DoS via token floods, model extraction, wallet-drain |
Produce a per-category table: **Exposure (Yes/No/Partial) → Evidence → Severity → Mitigation**.
### 2. Prompt Injection & Jailbreak Testing
**Direct injection** — user input that overrides instructions. Test families: - Instruction override ("ignore previous instructions and …") - Role-play / persona escape (DAN-style, hypothetical framing) - Encoding/obfuscation (Base64, ROT13, leetspeak, homoglyphs, zero-width chars) - Token smuggling and prompt-boundary confusion (fake delimiters, fake system tags) - Many-shot jailbreaking (long context of faux dialogue priming compliance) - Crescendo / multi-turn gradual escalation
**Indirect injection** — payload arrives via retrieved/processed content (web page, PDF, email, RAG doc, tool output). This is the highest-impact class for agents. Test that retrieved text **cannot** issue commands, exfiltrate context, or trigger tools.
For every test record: payload, channel (direct/indirect), goal (override / exfiltrate / tool-abuse), and result (blocked / partial / success). Use `scripts/prompt_injection_tester.py` to run a corpus and score outcomes.
**Refusal-quality note:** a single refusal is not a pass. Re-test the same goal across ≥3 phrasings and obfuscations before marking a control effective.
### 3. RAG & Vector Store Security
When reviewing a RAG pipeline: 1. **Access control at retrieval** — confirm the vector query is filtered by the *caller's* permissions, not just the app's. Test cross-tenant / cross-user document leakage. 2. **Indirect injection surface** — treat every ingested document as attacker-controlled; verify retrieved chunks are clearly delimited and never executed as instructions. 3. **Embedding inversion / membership** — sensitive source text may be partially reconstructable from embeddings; flag PII stored unencrypted in the vector DB. 4. **Chunk poisoning** — a single malicious document can dominate retrieval; check ranking/dedup and source allow-listing. 5. **Citation integrity** — outputs should cite retrieved sources so injected claims are traceable.
### 4. Agent & Tool-Use (Function Calling / MCP) Security
The agent is a **confused deputy**: it holds privileges the user may not. Review: - **Least-privilege tools** — each tool scoped to the minimum action; no broad `execute_shell`/`http_request` to arbitrary hosts - **Human-in-the-loop gates** on irreversible/outbound actions (payments, email send, file delete, deploy) - **Argument validation** — tool args are model-generated and untrusted; validate/allow-list server-side - **Injection → tool chain** — verify retrieved/indirect content cannot drive tool calls (e.g., a web page telling the agent to email its memory out) - **MCP server hardening** — authenticate clients, scope resources, rate-limit, log every tool invocation; never expose secrets via resource reads - **Memory poisoning** — persistent agent memory can be seeded with malicious instructions that fire on later turns
### 5. Model & ML Supply Chain
- Scan model artifacts for unsafe deserialization — **`pickle`/`.pt`/`.bin` can execute code on load**. Prefer `safetensors`. Run `scripts/model_supply_chain.py` or `modelscan`. - Verify model provenance, hashes, and signatures; pin versions from trusted registries. - Review fine-tune/LoRA adapters and datasets for poisoning and licensing. - Treat third-party plugins/MCP servers as untrusted dependencies (review + pin).
### 6. Output Handling & Guardrails
- **Never** pass raw LLM output into `eval`, SQL, shell, or innerHTML. Encode/parameterize at the sink (LLM05). - Layered guardrails: input filter → policy in system prompt → output classifier → sink-specific sanitization. Defense in depth, since any single layer is bypassable. - Validate structured output against a strict schema; reject on parse failure. - Apply egress controls so an injected agent cannot reach attacker URLs.
---
## Output Standards
Produce a structured AI security assessment:
```markdown # AI/LLM Security Assessment — [Application] Date: [Date] | Scope: [Endpoints/Models] | Model: [name/version] | Analyst: [Name]
## Executive Summary [2-3 sentences: overall posture, highest risks]
## OWASP LLM Top 10 Coverage | ID | Risk | Exposure | Severity | Evidence | |----|------|----------|----------|----------| | LLM01 | Prompt Injection | Yes | High | [repro] | ...
## Confirmed Findings ### [F-01] Indirect Prompt Injection via RAG → Tool Abuse (Critical) - ATLAS: AML.T0051 / OWASP LLM01+LLM06 - Repro: [payload, channel, steps] - Impact: [data exfil / unauthorized action] - Mitigation: [least-privilege tool scope + retrieved-content isolation + HITL]
## Guardrail Bypass Matrix | Goal | Direct | Encoded | Multi-turn | Indirect | Result |
## Recommendations (Prioritized) 1. ... ```
---
## Script Reference
### `prompt_injection_tester.py` ```bash # Run the built-in injection/jailbreak corpus against an endpoint python scripts/prompt_injection_tester.py --url https://app.test/api/chat --field message --output results.json
# Use a custom payload corpus and a refusal-detection keyword set python scripts/prompt_injection_tester.py --url ... --corpus payloads.txt --judge-keywords refusals.txt ```
### `model_supply_chain.py` ```bash # Scan a model directory/file for unsafe pickle opcodes and risky imports python scripts/model_supply_chain.py --path ./models/model.pt python scripts/model_supply_chain.py --path ./models/ --recursive --output scan.json ```
---
## Skill Integration
| Next Step | Condition | Target Skill | |-----------|-----------|--------------| | Web/API vuln testing of the app shell | App exposes web/API surface | → Skill 09 | | Cloud/infra hosting the model | Model served on AWS/Azure/GCP/K8s | → Skill 10 | | Detection rules for prompt-injection attempts | Need SIEM coverage | → Skill 12 | | Dependency/model-package CVEs | ML libs in use | → Skill 02 | | Red team narrative incorporating AI abuse | Full engagement | → Skill 14 |
---
## References
- [OWASP Top 10 for LLM Applications (2025)](https://genai.owasp.org/llm-top-10/) - [MITRE ATLAS — Adversarial Threat Landscape for AI Systems](https://atlas.mitre.org/) - [NIST AI Risk Management Framework (AI RMF 1.0) + Generative AI Profile](https://www.nist.gov/itl/ai-risk-management-framework) - [OWASP Agentic AI — Threats and Mitigations](https://genai.owasp.org/) - [Google SAIF — Secure AI Framework](https://saif.google/) - [NVIDIA garak — LLM vulnerability scanner](https://github.com/NVIDIA/garak)
Source provenance
Decision snapshot
recent repository activity
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for AI & LLM Security, ready for a manual X post.
AI & LLM Security: LLM and AI application security testing — prompt injection, jailbreak resistance, OWASP LLM T... 397 stars https://www.openagentskill.com/skills/masriyan-ai-llm-security?ref=x
Listing + install path for AI & LLM Security: https://www.openagentskill.com/skills/masriyan-ai-llm-security?ref=x Install: npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM Security
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to Masriyan but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/masriyan-ai-llm-security?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/masriyan-ai-llm-security?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/masriyan-ai-llm-security/audit)
[](https://www.openagentskill.com/skills/masriyan-ai-llm-security?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)Masriyan
@masriyan
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Do not auto-install
Wazuh
Wazuh - The Open Source Security Platform. Unified XDR and SIEM protection for endpoints and cloud workloads.
16.3K StarsMaigret
🕵️♂️ Collect a dossier on a person by username from 3000+ sites
32.9K StarsNuclei
Nuclei is a fast, customizable vulnerability scanner powered by the global security community and built on a simple YAML-based DSL, enabling collaboration to tackle trending vulnerabilities on the internet. It helps you find vulnerabilities in your applications, APIs, networks, DNS, and cloud configurations.
29.2K StarsInfisical
Infisical is the open-source platform for secrets, certificates, and privileged access management.
27.4K StarsDo not auto-install
Install targets
Codex install prompt
Install the "AI & LLM Security" agent skill from https://github.com/Masriyan/Claude-Code-CyberSecurity-Skill/tree/main/skills/16-ai-llm-security. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: LLM and AI application security testing — prompt injection, jailbreak resistance, OWASP LLM Top 10 (2025), RAG and agent/tool-use security, model supply chain, and AI red teaming for authorized assessments After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"masriyan-ai-llm-security","task":"Install AI & LLM Security","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Code review, repo analysis, testing, CI, GitHub, DevOps, and developer workflow skills.
Scenario
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Agent fit
Claude Code + Browser agents + CLI
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM Security
Maintenance
fresh
2d since push
Risk
Risky
Dependency or permission surface needs review
GitHub quality
397
77/100 Quality · 65/100 Trust
Coverage tags
Review notes
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
StrongSolid option that is likely worth shortlisting for production workflows.
Trust
Do not auto-installTrust Score v5 found insufficient evidence for agent installation. Treat this as discovery material, not an executable recommendation.
Audit
RiskyA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Choose a stronger alternative or inspect the source manually before any install attempt.
Stars
397 GitHub stars
Repo activity
397 stars, 75 forks
Maintenance
2d since push
License
MIT
Install
npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM Security
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM SecurityDo not use when
Agent safety v2
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill may drive a browser or interact with web pages.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20AI%20%26%20LLM%20Security%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20AI%20%26%20LLM%20Security%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/masriyan-ai-llm-security/install
Agent should check
Copy prompt
Task: Use AI & LLM Security in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20AI%20%26%20LLM%20Security%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/masriyan-ai-llm-security/install
Install command: npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM Security
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/masriyan-ai-llm-security/install
LLM text format
/api/skills/masriyan-ai-llm-security/install?format=text
Find alternatives
/api/skills/search?q=AI%20%26%20LLM%20Security&limit=3
Agent prompt
Use AI & LLM Security for this task. Review https://www.openagentskill.com/api/skills/masriyan-ai-llm-security/install, then install with: npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM SecurityRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/masriyan-ai-llm-security
LLM text
/api/registry/manifest/masriyan-ai-llm-security?format=text
Install alias
/api/registry/install/masriyan-ai-llm-security
Recommend
/api/registry/recommend?task=Use%20AI%20%26%20LLM%20Security%20in%20an%20agent%20workflow&limit=3
Agent fit
RAG and knowledge
Platforms
Claude Code, Browser agents
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Shortlist this skill and compare it with close alternatives before production adoption.
Role in stack
Companion skill
Primary fit
RAG and knowledge
Trust label
Strong shortlist
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Trust Score v5 found insufficient evidence for agent installation. Treat this as discovery material, not an executable recommendation.
GitHub adoption
INFO397 GitHub stars
Stars/forks activity
INFO397 stars, 75 forks; issue activity unavailable in current metadata
Recent maintenance
PASS2d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Choose a stronger alternative or inspect the source manually before any install attempt.
Quality profile
Solid option that is likely worth shortlisting for production workflows.
Workflow fit
Search private knowledge
I need my agent to build a RAG workflow over documents and retrieve reliable context.
Reduce risk
I need my agent to scan a project for security risks and summarize what needs attention.
Parse messy files
I need my agent to read PDFs, extract tables, and turn documents into structured data.
Workflow fit
Ingest, retrieve, and cite
A workflow for document-heavy agents that ingest files, create searchable knowledge, retrieve relevant context, and answer with grounded sources.
Scrape, clean, and reuse web data
A practical workflow for agents that crawl public pages, extract clean content, normalize data, and hand it to downstream research or RAG workflows.
Inspect, patch, and verify code
A workflow for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
Alternative shortlist
Similar skills that may fit this task.
Wazuh - The Open Source Security Platform. Unified XDR and SIEM protection for endpoints and cloud workloads.
🕵️♂️ Collect a dossier on a person by username from 3000+ sites
Nuclei is a fast, customizable vulnerability scanner powered by the global security community and built on a simple YAML-based DSL, enabling collaboration to tackle trending vulnerabilities on the internet. It helps you find vulnerabilities in your applications, APIs, networks, DNS, and cloud configurations.
Infisical is the open-source platform for secrets, certificates, and privileged access management.
--- name: AI & LLM Security description: LLM and AI application security testing — prompt injection, jailbreak resistance, OWASP LLM Top 10 (2025), RAG and agent/tool-use security, model supply chain, and AI red teaming for authorized assessments version: 3.0.0 author: Masriyan tags: [cybersecurity, ai-security, llm, prompt-injection, owasp-llm, agent-security, rag, mlsecops, red-teaming] ---
# AI & LLM Security
## Purpose
Enable Claude to assess the security of AI/LLM-powered applications — chatbots, RAG pipelines, autonomous agents, and tool-using systems. Claude maps findings to the **OWASP Top 10 for LLM Applications (2025)** and the **MITRE ATLAS** adversarial-ML knowledge base, builds reproducible attack cases, and recommends concrete mitigations (input/output guardrails, least-privilege tool scopes, content provenance).
> **Authorization Required**: Only test AI systems you own or are explicitly authorized to assess. Prompt-injection and data-exfiltration testing against third-party AI services may violate their terms of service and local law. Confirm written scope before proceeding.
---
## Activation Triggers
This skill activates when the user asks about: - Prompt injection (direct or indirect), jailbreaks, or system-prompt extraction - OWASP LLM Top 10, MITRE ATLAS, or AI/ML threat modeling - Securing a RAG pipeline, vector database, or retrieval layer - LLM agent / tool-use / function-calling security and confused-deputy risks - Guardrail, content-filter, or model output validation design - Sensitive-information disclosure or training-data leakage from a model - Model / ML supply chain security (model files, `pickle`, model registries) - AI red teaming, jailbreak corpora, or automated adversarial prompt generation - Securing MCP (Model Context Protocol) servers and tool integrations
---
## Prerequisites
```bash pip install requests pyyaml rich ```
**Optional enhanced capabilities:** - `garak` — LLM vulnerability scanner (NVIDIA) - `promptfoo` — prompt/red-team evaluation harness - API key for the target LLM endpoint (test environment only) - `modelscan` / `picklescan` — ML model file safety scanning
---
## Core Capabilities
### 1. Threat Modeling (OWASP LLM Top 10 — 2025)
When asked to threat-model an AI application, map the system against each category and record exposure:
| ID | Risk | What to look for | |----|------|------------------| | LLM01 | Prompt Injection | Untrusted text reaching the prompt (direct & indirect via RAG/web/email) | | LLM02 | Sensitive Information Disclosure | PII/secrets in prompts, outputs, or training data; system-prompt leakage | | LLM03 | Supply Chain | Untrusted models, LoRA adapters, datasets, plugins, `pickle` deserialization | | LLM04 | Data & Model Poisoning | Tainted training/fine-tune/RAG data; backdoors | | LLM05 | Improper Output Handling | LLM output passed unsanitized to SQL, shell, browser (XSS), or `eval` | | LLM06 | Excessive Agency | Over-broad tool scopes, autonomous side effects, no human-in-the-loop | | LLM07 | System Prompt Leakage | Secrets/authz logic embedded in the system prompt | | LLM08 | Vector & Embedding Weaknesses | RAG access-control bypass, embedding inversion, cross-tenant leakage | | LLM09 | Misinformation | Hallucinations relied on for security/safety decisions | | LLM10 | Unbounded Consumption | Cost/DoS via token floods, model extraction, wallet-drain |
Produce a per-category table: **Exposure (Yes/No/Partial) → Evidence → Severity → Mitigation**.
### 2. Prompt Injection & Jailbreak Testing
**Direct injection** — user input that overrides instructions. Test families: - Instruction override ("ignore previous instructions and …") - Role-play / persona escape (DAN-style, hypothetical framing) - Encoding/obfuscation (Base64, ROT13, leetspeak, homoglyphs, zero-width chars) - Token smuggling and prompt-boundary confusion (fake delimiters, fake system tags) - Many-shot jailbreaking (long context of faux dialogue priming compliance) - Crescendo / multi-turn gradual escalation
**Indirect injection** — payload arrives via retrieved/processed content (web page, PDF, email, RAG doc, tool output). This is the highest-impact class for agents. Test that retrieved text **cannot** issue commands, exfiltrate context, or trigger tools.
For every test record: payload, channel (direct/indirect), goal (override / exfiltrate / tool-abuse), and result (blocked / partial / success). Use `scripts/prompt_injection_tester.py` to run a corpus and score outcomes.
**Refusal-quality note:** a single refusal is not a pass. Re-test the same goal across ≥3 phrasings and obfuscations before marking a control effective.
### 3. RAG & Vector Store Security
When reviewing a RAG pipeline: 1. **Access control at retrieval** — confirm the vector query is filtered by the *caller's* permissions, not just the app's. Test cross-tenant / cross-user document leakage. 2. **Indirect injection surface** — treat every ingested document as attacker-controlled; verify retrieved chunks are clearly delimited and never executed as instructions. 3. **Embedding inversion / membership** — sensitive source text may be partially reconstructable from embeddings; flag PII stored unencrypted in the vector DB. 4. **Chunk poisoning** — a single malicious document can dominate retrieval; check ranking/dedup and source allow-listing. 5. **Citation integrity** — outputs should cite retrieved sources so injected claims are traceable.
### 4. Agent & Tool-Use (Function Calling / MCP) Security
The agent is a **confused deputy**: it holds privileges the user may not. Review: - **Least-privilege tools** — each tool scoped to the minimum action; no broad `execute_shell`/`http_request` to arbitrary hosts - **Human-in-the-loop gates** on irreversible/outbound actions (payments, email send, file delete, deploy) - **Argument validation** — tool args are model-generated and untrusted; validate/allow-list server-side - **Injection → tool chain** — verify retrieved/indirect content cannot drive tool calls (e.g., a web page telling the agent to email its memory out) - **MCP server hardening** — authenticate clients, scope resources, rate-limit, log every tool invocation; never expose secrets via resource reads - **Memory poisoning** — persistent agent memory can be seeded with malicious instructions that fire on later turns
### 5. Model & ML Supply Chain
- Scan model artifacts for unsafe deserialization — **`pickle`/`.pt`/`.bin` can execute code on load**. Prefer `safetensors`. Run `scripts/model_supply_chain.py` or `modelscan`. - Verify model provenance, hashes, and signatures; pin versions from trusted registries. - Review fine-tune/LoRA adapters and datasets for poisoning and licensing. - Treat third-party plugins/MCP servers as untrusted dependencies (review + pin).
### 6. Output Handling & Guardrails
- **Never** pass raw LLM output into `eval`, SQL, shell, or innerHTML. Encode/parameterize at the sink (LLM05). - Layered guardrails: input filter → policy in system prompt → output classifier → sink-specific sanitization. Defense in depth, since any single layer is bypassable. - Validate structured output against a strict schema; reject on parse failure. - Apply egress controls so an injected agent cannot reach attacker URLs.
---
## Output Standards
Produce a structured AI security assessment:
```markdown # AI/LLM Security Assessment — [Application] Date: [Date] | Scope: [Endpoints/Models] | Model: [name/version] | Analyst: [Name]
## Executive Summary [2-3 sentences: overall posture, highest risks]
## OWASP LLM Top 10 Coverage | ID | Risk | Exposure | Severity | Evidence | |----|------|----------|----------|----------| | LLM01 | Prompt Injection | Yes | High | [repro] | ...
## Confirmed Findings ### [F-01] Indirect Prompt Injection via RAG → Tool Abuse (Critical) - ATLAS: AML.T0051 / OWASP LLM01+LLM06 - Repro: [payload, channel, steps] - Impact: [data exfil / unauthorized action] - Mitigation: [least-privilege tool scope + retrieved-content isolation + HITL]
## Guardrail Bypass Matrix | Goal | Direct | Encoded | Multi-turn | Indirect | Result |
## Recommendations (Prioritized) 1. ... ```
---
## Script Reference
### `prompt_injection_tester.py` ```bash # Run the built-in injection/jailbreak corpus against an endpoint python scripts/prompt_injection_tester.py --url https://app.test/api/chat --field message --output results.json
# Use a custom payload corpus and a refusal-detection keyword set python scripts/prompt_injection_tester.py --url ... --corpus payloads.txt --judge-keywords refusals.txt ```
### `model_supply_chain.py` ```bash # Scan a model directory/file for unsafe pickle opcodes and risky imports python scripts/model_supply_chain.py --path ./models/model.pt python scripts/model_supply_chain.py --path ./models/ --recursive --output scan.json ```
---
## Skill Integration
| Next Step | Condition | Target Skill | |-----------|-----------|--------------| | Web/API vuln testing of the app shell | App exposes web/API surface | → Skill 09 | | Cloud/infra hosting the model | Model served on AWS/Azure/GCP/K8s | → Skill 10 | | Detection rules for prompt-injection attempts | Need SIEM coverage | → Skill 12 | | Dependency/model-package CVEs | ML libs in use | → Skill 02 | | Red team narrative incorporating AI abuse | Full engagement | → Skill 14 |
---
## References
- [OWASP Top 10 for LLM Applications (2025)](https://genai.owasp.org/llm-top-10/) - [MITRE ATLAS — Adversarial Threat Landscape for AI Systems](https://atlas.mitre.org/) - [NIST AI Risk Management Framework (AI RMF 1.0) + Generative AI Profile](https://www.nist.gov/itl/ai-risk-management-framework) - [OWASP Agentic AI — Threats and Mitigations](https://genai.owasp.org/) - [Google SAIF — Secure AI Framework](https://saif.google/) - [NVIDIA garak — LLM vulnerability scanner](https://github.com/NVIDIA/garak)
Source provenance
Decision snapshot
recent repository activity
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for AI & LLM Security, ready for a manual X post.
AI & LLM Security: LLM and AI application security testing — prompt injection, jailbreak resistance, OWASP LLM T... 397 stars https://www.openagentskill.com/skills/masriyan-ai-llm-security?ref=x
Listing + install path for AI & LLM Security: https://www.openagentskill.com/skills/masriyan-ai-llm-security?ref=x Install: npx skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill AI & LLM Security
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to Masriyan but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/masriyan-ai-llm-security?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/masriyan-ai-llm-security?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/masriyan-ai-llm-security/audit)
[](https://www.openagentskill.com/skills/masriyan-ai-llm-security?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)Masriyan
@masriyan
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Do not auto-install
Wazuh
Wazuh - The Open Source Security Platform. Unified XDR and SIEM protection for endpoints and cloud workloads.
16.3K StarsMaigret
🕵️♂️ Collect a dossier on a person by username from 3000+ sites
32.9K StarsNuclei
Nuclei is a fast, customizable vulnerability scanner powered by the global security community and built on a simple YAML-based DSL, enabling collaboration to tackle trending vulnerabilities on the internet. It helps you find vulnerabilities in your applications, APIs, networks, DNS, and cloud configurations.
29.2K StarsInfisical
Infisical is the open-source platform for secrets, certificates, and privileged access management.
27.4K StarsPermission surface
secrets or environment access, shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness
Permission surface
secrets or environment access, shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness
Permission surface
secrets or environment access, shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness
Permission surface
secrets or environment access, shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness