No visual example yet
Explore the skillVllm Ascend
vllm-project
Community maintained hardware plugin for vLLM on Ascend
OPENAGENTSKILL / DIRECTORY
次のタスクに合うスキルを。Codex、Claude Code、Cursor などのツールを探せます。
24 Skills
検索結果: 24
No visual example yet
Explore the skillvllm-project
Community maintained hardware plugin for vLLM on Ascend
No visual example yet
Explore the skillshen-shanshan
A curated collection of Claude Code agent skills that accelerate the entire vLLM development lifecycle.
No visual example yet
Explore the skillshen-shanshan
Design and implement vLLM features. Given user requirements (feature description, related PRs, reference materials), produces (1) core code implementation — NO test case…
No visual example yet
Explore the skillshen-shanshan
Generate comprehensive Chinese technical tutorial documents for specific vLLM models (e.g., Qwen3-VL, DeepSeek-V3, Llama 4, InternVL3, etc.). Produces deep-dive model wa…
No visual example yet
Explore the skillvllm-project
A framework for efficient model inference with omni-modality models
No visual example yet
Explore the skillbricks-cloud
🔒 Enterprise-grade API gateway that helps you monitor and impose cost or rate limits per API key. Get fine-grained access control and monitoring per user, application,…
No visual example yet
Explore the skilldstackai
Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven. Guides task-first p…
No visual example yet
Explore the skillECNU-ICALK
An open-source educational chat model from ICALK, East China Normal University. 开源中英教育对话大模型。(通用基座模型,GPU部署,数据清理) 致敬: LLaMA, MOSS, BELLE, Ziya, vLLM
No visual example yet
Explore the skillagentsope
Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanica
No visual example yet
Explore the skillagentsope
Neutral, framework-agnostic decision tree for project kickoff: "which agent / RAG / LLM framework should I reach for?" Synthesizes the ecosystem sections of 7 landmark-p…
No visual example yet
Explore the skillBenchmarks LLM inference and drives GPU kernel optimization with Magpie. Use when the user wants to benchmark vLLM, SGLang, or Atom; capture torch traces; post-process i…
No visual example yet
Explore the skillServes an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda. Use for "vLLM on CPU", "zentorch serving", or an EPYC CP
No visual example yet
Explore the skillmicrosoft
Set up AI Runway on AKS — from bare cluster to running model. Covers cluster verification, controller install, GPU assessment, provider setup, and first deployment. WHEN…
No visual example yet
Explore the skillAutonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer. Given a
No visual example yet
Explore the skillinterestingLSY
A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with only 2k lines of code (2% of vLLM).
No visual example yet
Explore the skillSentdex
A tiny single-file coding agent for self-hosted models (llama.cpp / vLLM / SGLang).