TensorTonicTensorTonic
Problems
Study PlansProjectsInterviewPricingFeedback

Research Papers

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale badge
1

Patch Embedding

Split image into patches and project them to embeddings

2

Position Embedding

Add learnable position information to patches

3

Class Token [CLS]

Prepend learnable classification token

4

ViT Encoder Block

Pre-LN Transformer block with GELU MLP

5

Classification Head

Extract [CLS] and project to classes

6

Complete ViT

Full Vision Transformer forward pass

Loading architecture visualization...