sklearn for the classical baseline (TF-IDF + Logistic Regression, Naive Bayes, or SVM), PyTorch or Keras for the neural model (BiLSTM, CNN, anything), and HuggingFace transformers for the transformer (DistilBERT, BERT, RoBERTa). Resource-friendly mode. If the team is short on RAM or compute, subsample with a fixed seed and keep the classes balanced — 50k tweets on a laptop, 200k on a mid-range GPU, the full 1.6M on Colab T4 or better. The comparison stays meaningful at any size. Going further (optional). Build a small UI where someone types a tweet and sees all three predictions, train at three subsample sizes to plot accuracy vs data, or add an LLM zero-shot baseline so you can compare four stages of NLP, not three.What you'll build
Build three sentiment classifiers for short English tweets and compare them head-to-head. The point is to see what each generation of NLP buys you on the same task: a classical TF-IDF + logistic regression baseline, a small neural model (BiLSTM or CNN), and a fine-tuned transformer. Report accuracy, training time, inference time per example, and disk size. The dataset is large (1.6M tweets) but easily subsamples to whatever the team's hardware can handle.
What goes in, what comes out
Input
Sentiment140 tweets (around 1.6M total, balanced binary labels). Each tweet is short (often under 25 words). Labels are 0 (negative) or 4 (positive); remap 4 → 1 for convenience.
Output
Three classifiers plus a comparison table with accuracy, training time, inference latency, and model size.
label tweet
----- -----
0 @kennethcole missing you so much! I hope you come back soon :(
1 Just had the best pancakes ever at this little diner downtown. Going to be a great day!
0 My internet has been out for 3 hours. So frustrating, I have a deadline tonight.
1 Thanks @sarahsmith for the birthday wishes! You made my day :)
0 Long lines at the airport, missed my flight. Worst Monday in a while.
Model Accuracy Train time Inference (ms/ex) Model size
------------------------------- -------- ---------- ----------------- ----------
TF-IDF + Logistic Regression 0.78 1 min 0.02 4 MB
BiLSTM on word embeddings 0.81 8 min 1.2 12 MB
DistilBERT fine-tune 0.84 10 min 3.8 260 MB
Lessons:
- Sentiment140 caps at around 84-85% because the labels themselves
are noisy (auto-derived from emoticons).
- Classical baseline is competitive thanks to the short input.
- DistilBERT helps on negation and sarcasm-ish phrasing — not on
the noisy half of the labels.
Datasets
Sentiment140 ↗
1.6M English tweets labelled positive or negative. Labels were auto-derived from emoticons, so they are noisy but the dataset is large and balanced. Tweets are short, which makes all three stages train fast.
How to get it: from datasets import load_dataset; ds = load_dataset("stanfordnlp/sentiment140"). Remap labels 4 → 1.
Subset for low-resource teams
If the team is short on RAM or compute, do not train on the full 1.6M. A balanced subsample works just as well for the comparison. Suggested sizes by hardware: laptop CPU only → 50k tweets; mid-range laptop with GPU → 200k tweets; Colab T4 / Kaggle → 500k tweets or the full set.
How to get it: After loading, shuffle with a fixed seed and slice: ds["train"].shuffle(seed=42).select(range(50_000)). Make sure both classes stay balanced — sample 25k from each.
Tools you'll need
These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.
Python: Python 3.10 or newer. Compute: A laptop CPU is fine for stages 1 and 2 on a 50k subsample. Stage 3 (DistilBERT) is much faster on a Colab T4, but it works on CPU with smaller subsets too.
scikit-learn— TF-IDF vectoriser and logistic regression. Two lines of code.
torch— PyTorch for the BiLSTM or CNN. Train your own embedding layer or load pretrained GloVe / fastText.datasets— Stream tweets and labels.
transformers— Load and fine-tune DistilBERT.accelerate— Hardware-agnostic training loop.
evaluate— Accuracy and confusion matrix wrappers.
How to approach it
One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.
- Load. Pull Sentiment140. Remap labels 4 → 1. Decide on a subsample size based on your hardware (50k is enough to make the comparison meaningful).
- Light cleanup. Strip URLs (
http[s]?://...), replace@usernameswith@user, leave hashtags as-is (they often carry sentiment). - Stage 1 — Classical. Fit a TF-IDF vectoriser on train (word + character bigrams), train a logistic regression, evaluate on test.
- Stage 2 — Neural. Build a PyTorch model: an embedding layer + BiLSTM + linear head. Train for 3 epochs.
- Stage 3 — Transformer. Fine-tune
distilbert-base-uncasedfor 2 epochs with the HuggingFaceTrainer. Truncate to 64 tokens — tweets rarely need more. - Measure. For each model record: test accuracy, training wall-clock time, inference latency per example, on-disk size.
- Compare. Put it all in one table. Discuss which approach you would pick for which use case, and how the curve might change with the full 1.6M.
What to deliver
- Three notebooks (one per stage) that produce reproducible numbers.
- A single comparison table with accuracy, training time, inference latency, disk size.
- A one-page write-up: in which scenario would the team pick each approach? And how would the picture change at 10x more data?
- A documented note on the subsample size used — so the comparison is reproducible.