A nine-person team project. My part was the deep learning pipeline — the shared architecture across all four recurrent variants, the evaluation charts, and the results section of the paper.
The comparison
82,077 labelled emails, aggregated from seven public datasets. Four recurrent architectures — Simple RNN, LSTM, GRU, and a bidirectional RNN — all sharing one structure of embedding, recurrent layer, dense ReLU, dense sigmoid. Against them, five classical baselines on TF-IDF features: naive Bayes, logistic regression, SGD, random forest, and an MLP.
The classical model won
The MLP baseline came out ahead at 98.92% accuracy and 98.96% F1, marginally past LSTM, the strongest deep model, at 98.68% F1. On a well-separated text classification task with enough labelled data, TF-IDF and a dense network are extremely hard to beat, and it is worth reporting that rather than burying it.
GRU is the one worth arguing for anyway: it had the lowest false-negative rate of the nine at 0.45%, missing only 29 phishing emails. In this problem a false negative is the expensive error.
The failure that taught the most
The Simple RNN collapsed — 61.33% accuracy with a 78.05% false-positive rate. It had converged to something close to guessing.
The cause was vanishing gradients across 200-token sequences. The gated architectures on the identical data did not have the problem, because gating is precisely the mechanism that fixes it. Textbooks assert this. Here it is as a measured 37-point accuracy gap between two models that differ only in whether their recurrent cell has gates.
