Archived
  • PyTorch
  • CNN
  • FER-2013
  • Computer vision

Facial emotion CNN

Eleven points of accuracy traded away for seventy percent fewer parameters

A lightweight convolutional network for seven-class facial emotion recognition on FER-2013, built to be small rather than accurate, and measured honestly against a ResNet-18 baseline.

Most work on FER-2013 chases the leaderboard. This went the other way: how small can the network get before the accuracy stops being worth it?

The architecture

Three downsampling stages, each two 3×3 convolutions with batch norm and ReLU followed by a 2×2 max pool, taking the image from 64 down to 8. Then a fully-connected head — 512 to 128 to 7 — with dropout at 0.5 and 0.3. Trained with Adam at a learning rate of 0.001 and cross-entropy loss, batch size 64, for 25 epochs on a GPU.

The dataset is 35,887 images across seven emotions, upsampled from 48×48 grayscale to 64×64 RGB by duplicating the channel.

What the numbers actually say

The model reaches 57.3% test accuracy. A fine-tuned ResNet-18 on the same data reaches about 68.5%. That is eleven points given up — for roughly seventy percent fewer parameters.

Whether that trade is good depends entirely on where the model runs. On a server it is a bad trade. On a device with a fixed memory budget it may be the only trade available.

Where it fails, and why

The confusion matrix is more interesting than the accuracy. Disgust is never predicted at all — it is five percent of the data against Happy's twenty-eight, and the network learns that guessing it is never worth the risk. Neutral and Sad are confused constantly, which is a harder problem, because the boundary between them is not always clear to a human annotator either.

Your browser will not display this PDF inline.

Open Lightweight CNN for Facial Emotion Recognition
Lightweight CNN for Facial Emotion Recognition — 4 pagesOpen the PDF