Split image into patches and project them to embeddings
Add learnable position information to patches
Prepend learnable classification token
Pre-LN Transformer block with GELU MLP
Extract [CLS] and project to classes
Full Vision Transformer forward pass