AlexNet, Visually
Krizhevsky, Sutskever & Hinton · NIPS 2012

Don't tell the computer what to look for. Give it 1.2 million labeled photos and let it learn the features itself.

This page walks through the whole paper: the idea, the network, the tricks that made it train, the results, and what has held up since.

GPU 1 · 48 filtersmostly shape, little color
GPU 2 · 48 filtersmostly color blobs
Illustration of the paper's Figure 3: the 96 first-layer filters, each 11×11 pixels. Nobody designed these. The split between the two GPUs appeared on its own in every run.
1.2Mlabeled training images
1,000categories to choose from
60Mlearned parameters
5–6 dayson two 3 GB GTX 580 GPUs
The problem

From hand-written features to learned ones

ImageNet asks a model to label a photo with one of 1,000 classes, from mushrooms to container ships. Before 2012, the best systems used features that people designed, then fed them to a simple classifier. They had hit a wall around 45–47% top-1 error.

Before: features designed by people
Image
→
SIFT, Fisher vectorswritten by hand
→
Linear classifierlearned

Only the last step learns. If the handmade features miss something, the classifier can never see it.

AlexNet: everything is learned
Imageraw pixels
→
8 layersall learned
→
1,000 scoressoftmax

Every layer is trained together from the labels. The network decides what features matter.

The network

Follow one image through the eight layers

Five convolutional layers find patterns. Three fully connected layers combine them into a decision. Pick a layer to see what it does. Block height shows the spatial size and width shows the number of channels.

A detail most summaries skip

The weights live in one place, the work happens in another

The convolution layers do almost all the arithmetic but hold only 4% of the weights. The fully connected layers hold 96% of the weights. That is why the paper needs dropout on the FC layers, and why later networks such as GoogLeNet and ResNet cut those layers down.

Parameters60.97M total
96.2%
Compute per image≈724M multiply-adds
91.9%
Convolution layers 1–5Fully connected layers 6–8
Show per-layer numbers
LayerOutputParametersMultiply-adds

Computed from the layer sizes in Section 3.5, counting conv2, conv4 and conv5 as split across the two GPUs. The total matches the paper's "60 million parameters".

Why it trained

Four ideas that made it work

None of these was brand new in 2012. The contribution was showing that together they make a network this large trainable, without it simply memorizing the training set.

ReLU activation

6× faster training
ReLU = max(0, x) tanh flat at the ends: gradient ≈ 0

tanh flattens out for large inputs, so its gradient shrinks toward zero and learning stalls. ReLU has a slope of exactly 1 for any positive input. A 4-layer net on CIFAR-10 reached 25% training error six times faster.

Dropout

p = 0.5 on FC6, FC7

Each training step switches off a random half of the neurons, so no neuron can rely on a particular partner. At test time all neurons are on and outputs are halved. Without it the paper reports substantial overfitting. It costs about twice as many iterations.

Data augmentation

2,048 versions per image
photo: 256 × 256 dashed box = 224 × 224 crop offset 8, 8 not flipped 32 × 32 positions × 2 flips = 2,048
also shifts color slightly

Random crops and mirror flips cost nothing and give the model many slightly different views. A second trick adds small color shifts along the main color directions of ImageNet (PCA), teaching the model that lighting changes don't change the object. That alone cut top-1 error by over 1%.

Two GPUs

−1.7% top-1 error

One 3 GB card couldn't hold the network, so each GPU got half the filters. They only talk to each other at layer 3 and in the FC layers (solid crossing lines). The paper compares against a one-GPU net with half the filters, which is smaller, so the 1.7% gain is not a like-for-like test.

Local response normalization

Dropped later

A strongly active filter dampens its neighbors at the same pixel, borrowed from biology. Reported −1.4% top-1. Batch normalization replaced it in 2015.

Overlapping pooling

Small effect

3×3 pooling windows placed 2 pixels apart, so they overlap. Reported −0.4% top-1. Small enough that it could be run-to-run noise.

Training setup

The recipe, exactly as written

Optimizer
SGD + momentum 0.9
Batch size
128 images
Weight decay
0.0005
Learning rate
0.01, ÷10 ×3
Weight init
Gaussian, σ = 0.01
Bias init
1 in conv2, 4, 5, FC
Length
≈90 epochs
Loss
softmax cross-entropy

The learning rate was divided by 10 by hand whenever validation error stopped improving. Bias set to 1 gives ReLUs positive inputs early so they start learning. The paper notes weight decay lowered training error too, not only test error.

Results

Ten points better than anything before it

Top-1 error means the first guess is wrong. Top-5 error means the right label is not in the five best guesses. Lower is better. Hover a bar for details.

What this proves: an end-to-end CNN beats hand-built pipelines at this scale, by a margin far larger than noise. What it doesn't prove: that each individual trick is needed. The headline 15.3% also uses 7 models, 10 crops per test image, and extra pre-training data, so a single model is closer to 18%.

What it learned

Depth matters, and the last layer is a useful summary

Remove any middle layer, lose ≈2%

Taking out any one of the middle convolution layers raised top-1 error by about 2%. That was early evidence that depth itself carries the performance. The paper shows that it matters but does not explain why.

Similar images, similar vectors

Each image produces 4,096 numbers in the last hidden layer. Images that are close in this space show the same kind of object, even when their pixels look very different (paper Figure 4). The bars above are an illustration. The same idea powers embedding search today.

Claims vs evidence

What the paper says, what it shows, and what held up

ClaimEvidence in the paperSince then
ReLU trains much faster6× faster to 25% error, small net on CIFAR-10Held up
Dropout prevents overfittingWithout it, "substantial overfitting"Held up
Depth is importantRemoving a middle layer costs ≈2% top-1Held up VGG, ResNet
Augmentation reduces overfittingPCA color shift −1% top-1; crops needed to avoid overfittingHeld up
Response normalization helps−1.4% top-1, single comparisonFaded BatchNorm
Overlapping pooling helps−0.4% top-1, no error barsFaded
Two-GPU split helps−1.7% vs a smaller one-GPU netFaded hardware workaround