Vision Transformers Explained Series
Since their introduction in 2017 with Attention is All You Need¹, transformers have established themselves as the state of the art for natural language processing (NLP). In 2021, An Image is Worth 16x16 Words successfully adapted transformers for computer vision tasks. Since then, numerous transformer-based architectures have been proposed for computer vision. This article walks through the Vision Transformer (ViT) as laid out in An Image is Worth 16x16 Words.
97 MATHEMATICS AND COMPUTING↗