SageAttention
thu-ml
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, i…
Media AutomationSkill source unconfirmed
3.4KGitHub
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–3 / 3
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 3
thu-ml
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, i…
thu-ml
[ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.