Most work on FER-2013 chases the leaderboard. This went the other way: how small can the network get before the accuracy stops being worth it?
The architecture
Three downsampling stages, each two 3×3 convolutions with batch norm and ReLU followed by a 2×2 max pool, taking the image from 64 down to 8. Then a fully-connected head — 512 to 128 to 7 — with dropout at 0.5 and 0.3. Trained with Adam at a learning rate of 0.001 and cross-entropy loss, batch size 64, for 25 epochs on a GPU.
The dataset is 35,887 images across seven emotions, upsampled from 48×48 grayscale to 64×64 RGB by duplicating the channel.
What the numbers actually say
The model reaches 57.3% test accuracy. A fine-tuned ResNet-18 on the same data reaches about 68.5%. That is eleven points given up — for roughly seventy percent fewer parameters.
Whether that trade is good depends entirely on where the model runs. On a server it is a bad trade. On a device with a fixed memory budget it may be the only trade available.
Where it fails, and why
The confusion matrix is more interesting than the accuracy. Disgust is never predicted at all — it is five percent of the data against Happy's twenty-eight, and the network learns that guessing it is never worth the risk. Neutral and Sad are confused constantly, which is a harder problem, because the boundary between them is not always clear to a human annotator either.
