Convolution Vision Transformers
rishikksh20
PyTorch Implementation of CvT: Introducing Convolutions to Vision Transformers
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–16 / 16
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 16
rishikksh20
PyTorch Implementation of CvT: Introducing Convolutions to Vision Transformers
cdpierse
Model explainability that works seamlessly with 🤗 transformers. Explain your transformers model in just 2 lines of code.
huggingface
This repo is the homebase of a community driven course on Computer Vision with Neural Networks. Feel free to join us on the Hugging Face discord: hf.co/join/discord
google-research
Scenic: A Jax Library for Computer Vision Research and Beyond
huggingface
🚪✊Knock Knock: Get notified when your training ends with only two additional lines of code
huggingface
🦋A PyTorch implementation of BigGAN with pretrained weights and conversion scripts.
yuxumin
[ICCV 2021 Oral] PoinTr: Diverse Point Cloud Completion with Geometry-Aware Transformers
raoyongming
[NeurIPS 2021] [T-PAMI] DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification
huggingface
Minimal sharded dataset loaders, decoders, and utils for multi-modal document, image, and text datasets.
EvelynFan
[CVPR 2022] FaceFormer: Speech-Driven 3D Facial Animation with Transformers
microsoft
This is an official implementation of CvT: Introducing Convolutions to Vision Transformers.
arpitg1304
Robotics Data Toolkit | Convert between robotics dataset formats (RLDS, LeRobot v2/v3, Zarr, HDF5, Rosbag). Inspect, visualize, and analyze datasets. Works with HuggingF…
DerrickXuNu
[CoRL2022] CoBEVT: Cooperative Bird's Eye View Semantic Segmentation with Sparse Transformers
emla2805
Tensorflow implementation of the Vision Transformer (An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale)
htdt
Hyperbolic Vision Transformers: Combining Improvements in Metric Learning | Official repository
catherinesyeh
Visualizing query-key interactions in language + vision transformers (VIS 2023)