No visual example yet
Explore the skillTransformerEngine
NVIDIA
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, t…
OPENAGENTSKILL / DIRECTORY
Trouvez un skill pour votre prochaine tâche avec Codex, Claude Code, Cursor et plus encore.
4 Skills
Résultats: 4
No visual example yet
Explore the skillNVIDIA
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, t…
No visual example yet
Explore the skillOrchestra-Research
Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime. Use when training large MoE models with FP8/INT4, needing train-inferen…
No visual example yet
Explore the skillOpenRaiser
Simplest distributed training API. 4 lines to add distributed support to any PyTorch script. Unified API for DeepSpeed/FSDP/Megatron/DDP. Automatic device placement, mix…
No visual example yet
Explore the skillRuns an end-to-end AMD Quark post-training quantization workflow for PyTorch / Hugging Face LLMs: inspect a Hub or local model, choose a quantization plan, create reprod…