sklearn for the classical model (TF-IDF + Naive Bayes is the textbook recipe), PyTorch for the neural model (BiLSTM, GRU, or CNN), and HuggingFace transformers for the transformer. Going further (optional). Build a small UI that flags messages as a user types, or extend with an adversarial test where the spammer uses common evasion tricks (zero-width characters, letter substitutions) and see which model holds up.What you'll build
Build three spam classifiers for short text messages (SMS) and put them in a head-to-head comparison. The dataset is small enough to fit on a laptop with no GPU, which makes it perfect for showing that sometimes the textbook classical baseline is already so good that fancier models are overkill. The takeaway is a clear table comparing classical, neural, and transformer approaches on the same data.
What goes in, what comes out
Input
SMS Spam Collection: 5,574 short text messages labelled "ham" (legitimate) or "spam". Around 13% spam, so the dataset is imbalanced.
Output
Three classifiers plus a comparison table with accuracy, F1 on the spam class, training time, and inference latency.
label message
----- -------
ham Ok lar... Joking wif u oni...
ham Are you free for dinner tonight? Let me know.
spam FREE entry in 2 a wkly comp to win FA Cup final tkts.
Text FA to 87121 to receive entry. T&Cs apply.
spam URGENT! Your mobile has been awarded a £2000 cash prize.
Call 09058091870 from a landline. Box95QU.
ham Sorry I'll call you later, in a meeting.
Model Accuracy Spam F1 Train time Inference (ms/ex)
---------------------------- -------- ------- ---------- -----------------
TF-IDF + Multinomial NB 0.985 0.94 4 sec 0.02
GRU on word embeddings 0.987 0.95 90 sec 1.5
DistilBERT fine-tune 0.988 0.96 4 min 4.2
Lesson: classical baseline is already at 98.5% accuracy with F1=0.94
on spam. The fancier models add fractions of a point at much higher
training and inference cost. This is when BERT is overkill.
Datasets
SMS Spam Collection ↗
Tiny (5,574 messages), clean, and famous. The textbook example of a problem where a Naive Bayes baseline is nearly perfect.
How to get it: from datasets import load_dataset; ds = load_dataset("ucirvine/sms_spam").
Tools you'll need
These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.
Python: Python 3.10 or newer. Compute: A laptop CPU is enough for all three stages on this tiny dataset.
scikit-learn— TF-IDF vectoriser and Multinomial Naive Bayes. The textbook recipe.
torch— PyTorch. A small GRU or CNN is enough — do not over-engineer.
transformers— Load and fine-tune DistilBERT.datasets— Streams the SMS data into the training loop.accelerate— Hardware-agnostic training loop.
evaluate— Accuracy + F1 wrappers.
How to approach it
One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.
- Load. Pull the SMS dataset. Check the class balance (around 13% spam) and decide that F1 on the spam class is the metric to watch, not accuracy.
- Split. 80/20 train/test, stratified by label so the test set also has 13% spam.
- Stage 1 — Classical. Fit a TF-IDF vectoriser on train, train a Multinomial Naive Bayes classifier, evaluate on test.
- Stage 2 — Neural. Build a tiny PyTorch model: embedding layer + GRU (32 hidden) + linear head. Train for a few epochs.
- Stage 3 — Transformer. Fine-tune
distilbert-base-uncasedfor 2 epochs. - Measure. For each model: test accuracy, F1 on spam, training time, inference latency per message.
- Compare. Put it all in one table. Discuss when classical is "good enough".
What to deliver
- Three notebooks (one per stage).
- A comparison table with accuracy, spam F1, training time, inference latency, model size.
- A one-paragraph conclusion: would the team deploy the classical, neural, or transformer model in a real spam filter? Why?