Large 11×11 and 5×5 kernels in early layers
6× faster training than tanh
Prevent overfitting with p=0.5
Cross-channel normalization
3×3 pooling with stride 2
Random crops, flips, and PCA color