Menu
Home Program Lecturers Important Dates Venue Sponsors Past Editions Speakers Alumni Versions GAI2026 Contact
SMS Spam Detection: Classical, Neural, and Transformer — Project #7 | Summer School on Generative AI
Project 7

SMS Spam Detection: Classical, Neural, and Transformer

CPU OK
Note. This page lays out one full version of the project — the goal, a sample input/output, suggested tools, and a step-by-step plan. Treat it as a reference, not a script. Your team can pick a different angle, swap libraries, narrow the scope, or take the project somewhere we did not anticipate. As long as the final deliverable makes sense for the goal, you are on track.
Task. Use the SMS Spam Collection (5,574 messages, free). Build three spam classifiers — a classical baseline, a small neural model, and a fine-tuned transformer — and put them head-to-head on accuracy and F1 on the spam class. The lesson is that the classical baseline is almost perfect here — sometimes BERT is overkill. You could use sklearn for the classical model (TF-IDF + Naive Bayes is the textbook recipe), PyTorch for the neural model (BiLSTM, GRU, or CNN), and HuggingFace transformers for the transformer. Going further (optional). Build a small UI that flags messages as a user types, or extend with an adversarial test where the spammer uses common evasion tricks (zero-width characters, letter substitutions) and see which model holds up.
Resources: CPU is enough for all three stages on this tiny dataset.

What you'll build

Build three spam classifiers for short text messages (SMS) and put them in a head-to-head comparison. The dataset is small enough to fit on a laptop with no GPU, which makes it perfect for showing that sometimes the textbook classical baseline is already so good that fancier models are overkill. The takeaway is a clear table comparing classical, neural, and transformer approaches on the same data.

What goes in, what comes out

Input

SMS Spam Collection: 5,574 short text messages labelled "ham" (legitimate) or "spam". Around 13% spam, so the dataset is imbalanced.

Output

Three classifiers plus a comparison table with accuracy, F1 on the spam class, training time, and inference latency.

A few rows from the dataset
label  message
-----  -------
 ham   Ok lar... Joking wif u oni...
 ham   Are you free for dinner tonight? Let me know.
spam   FREE entry in 2 a wkly comp to win FA Cup final tkts.
       Text FA to 87121 to receive entry. T&Cs apply.
spam   URGENT! Your mobile has been awarded a £2000 cash prize.
       Call 09058091870 from a landline. Box95QU.
 ham   Sorry I'll call you later, in a meeting.
Head-to-head comparison on the test set
Model                          Accuracy  Spam F1  Train time  Inference (ms/ex)
----------------------------   --------  -------  ----------  -----------------
TF-IDF + Multinomial NB          0.985     0.94      4 sec         0.02
GRU on word embeddings           0.987     0.95     90 sec         1.5
DistilBERT fine-tune             0.988     0.96    4 min          4.2

Lesson: classical baseline is already at 98.5% accuracy with F1=0.94
on spam. The fancier models add fractions of a point at much higher
training and inference cost. This is when BERT is overkill.

Datasets

SMS Spam Collection ↗

Tiny (5,574 messages), clean, and famous. The textbook example of a problem where a Naive Bayes baseline is nearly perfect.

How to get it: from datasets import load_dataset; ds = load_dataset("ucirvine/sms_spam").

License: Free for research use.

Tools you'll need

These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.

Python: Python 3.10 or newer. Compute: A laptop CPU is enough for all three stages on this tiny dataset.

Stage 1 — classical
  • scikit-learn — TF-IDF vectoriser and Multinomial Naive Bayes. The textbook recipe.
Stage 2 — neural
  • torch — PyTorch. A small GRU or CNN is enough — do not over-engineer.
Stage 3 — transformer
  • transformers — Load and fine-tune DistilBERT.
  • datasets — Streams the SMS data into the training loop.
  • accelerate — Hardware-agnostic training loop.
Evaluation
  • evaluate — Accuracy + F1 wrappers.

How to approach it

One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.

  1. Load. Pull the SMS dataset. Check the class balance (around 13% spam) and decide that F1 on the spam class is the metric to watch, not accuracy.
  2. Split. 80/20 train/test, stratified by label so the test set also has 13% spam.
  3. Stage 1 — Classical. Fit a TF-IDF vectoriser on train, train a Multinomial Naive Bayes classifier, evaluate on test.
  4. Stage 2 — Neural. Build a tiny PyTorch model: embedding layer + GRU (32 hidden) + linear head. Train for a few epochs.
  5. Stage 3 — Transformer. Fine-tune distilbert-base-uncased for 2 epochs.
  6. Measure. For each model: test accuracy, F1 on spam, training time, inference latency per message.
  7. Compare. Put it all in one table. Discuss when classical is "good enough".

What to deliver

  • Three notebooks (one per stage).
  • A comparison table with accuracy, spam F1, training time, inference latency, model size.
  • A one-paragraph conclusion: would the team deploy the classical, neural, or transformer model in a real spam filter? Why?

References