Menu
Home Program Lecturers Important Dates Venue Sponsors Past Editions Speakers Alumni Versions GAI2026 Contact
Sentiment Classification on Sentiment140: Classical, Neural, and Transformer — Project #6 | Summer School on Generative AI
Project 6

Sentiment Classification on Sentiment140: Classical, Neural, and Transformer

CPU / GPU
Note. This page lays out one full version of the project — the goal, a sample input/output, suggested tools, and a step-by-step plan. Treat it as a reference, not a script. Your team can pick a different angle, swap libraries, narrow the scope, or take the project somewhere we did not anticipate. As long as the final deliverable makes sense for the goal, you are on track.
Task. Use the Sentiment140 dataset (1.6M English tweets, binary positive / negative, balanced). Build three sentiment classifiers on the same split — a classical model, a small neural model, and a fine-tuned transformer — and compare accuracy, F1, training time, and inference latency. The point is to feel what each stage of NLP buys you on the same task. You could use sklearn for the classical baseline (TF-IDF + Logistic Regression, Naive Bayes, or SVM), PyTorch or Keras for the neural model (BiLSTM, CNN, anything), and HuggingFace transformers for the transformer (DistilBERT, BERT, RoBERTa). Resource-friendly mode. If the team is short on RAM or compute, subsample with a fixed seed and keep the classes balanced — 50k tweets on a laptop, 200k on a mid-range GPU, the full 1.6M on Colab T4 or better. The comparison stays meaningful at any size. Going further (optional). Build a small UI where someone types a tweet and sees all three predictions, train at three subsample sizes to plot accuracy vs data, or add an LLM zero-shot baseline so you can compare four stages of NLP, not three.
Resources: CPU works for the sklearn baseline and a small subsample; Colab T4 or 8GB GPU for the LSTM and the BERT fine-tune. The dataset is large (1.6M tweets) but easily subsamples — pick a size that fits your hardware.

What you'll build

Build three sentiment classifiers for short English tweets and compare them head-to-head. The point is to see what each generation of NLP buys you on the same task: a classical TF-IDF + logistic regression baseline, a small neural model (BiLSTM or CNN), and a fine-tuned transformer. Report accuracy, training time, inference time per example, and disk size. The dataset is large (1.6M tweets) but easily subsamples to whatever the team's hardware can handle.

What goes in, what comes out

Input

Sentiment140 tweets (around 1.6M total, balanced binary labels). Each tweet is short (often under 25 words). Labels are 0 (negative) or 4 (positive); remap 4 → 1 for convenience.

Output

Three classifiers plus a comparison table with accuracy, training time, inference latency, and model size.

A few rows from the train set
label  tweet
-----  -----
  0    @kennethcole missing you so much! I hope you come back soon :(
  1    Just had the best pancakes ever at this little diner downtown. Going to be a great day!
  0    My internet has been out for 3 hours. So frustrating, I have a deadline tonight.
  1    Thanks @sarahsmith for the birthday wishes! You made my day :)
  0    Long lines at the airport, missed my flight. Worst Monday in a while.
Head-to-head comparison (trained on a 200k subsample, evaluated on a 50k held-out)
Model                              Accuracy  Train time  Inference (ms/ex)  Model size
-------------------------------    --------  ----------  -----------------  ----------
TF-IDF + Logistic Regression         0.78        1 min         0.02            4 MB
BiLSTM on word embeddings            0.81        8 min         1.2            12 MB
DistilBERT fine-tune                 0.84       10 min         3.8           260 MB

Lessons:
  - Sentiment140 caps at around 84-85% because the labels themselves
    are noisy (auto-derived from emoticons).
  - Classical baseline is competitive thanks to the short input.
  - DistilBERT helps on negation and sarcasm-ish phrasing — not on
    the noisy half of the labels.

Datasets

Sentiment140 ↗

1.6M English tweets labelled positive or negative. Labels were auto-derived from emoticons, so they are noisy but the dataset is large and balanced. Tweets are short, which makes all three stages train fast.

How to get it: from datasets import load_dataset; ds = load_dataset("stanfordnlp/sentiment140"). Remap labels 4 → 1.

License: Free for research use.

Subset for low-resource teams

If the team is short on RAM or compute, do not train on the full 1.6M. A balanced subsample works just as well for the comparison. Suggested sizes by hardware: laptop CPU only → 50k tweets; mid-range laptop with GPU → 200k tweets; Colab T4 / Kaggle → 500k tweets or the full set.

How to get it: After loading, shuffle with a fixed seed and slice: ds["train"].shuffle(seed=42).select(range(50_000)). Make sure both classes stay balanced — sample 25k from each.

Tools you'll need

These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.

Python: Python 3.10 or newer. Compute: A laptop CPU is fine for stages 1 and 2 on a 50k subsample. Stage 3 (DistilBERT) is much faster on a Colab T4, but it works on CPU with smaller subsets too.

Stage 1 — classical
  • scikit-learn — TF-IDF vectoriser and logistic regression. Two lines of code.
Stage 2 — neural
  • torch — PyTorch for the BiLSTM or CNN. Train your own embedding layer or load pretrained GloVe / fastText.
  • datasets — Stream tweets and labels.
Stage 3 — transformer
  • transformers — Load and fine-tune DistilBERT.
  • accelerate — Hardware-agnostic training loop.
Evaluation
  • evaluate — Accuracy and confusion matrix wrappers.

How to approach it

One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.

  1. Load. Pull Sentiment140. Remap labels 4 → 1. Decide on a subsample size based on your hardware (50k is enough to make the comparison meaningful).
  2. Light cleanup. Strip URLs (http[s]?://...), replace @usernames with @user, leave hashtags as-is (they often carry sentiment).
  3. Stage 1 — Classical. Fit a TF-IDF vectoriser on train (word + character bigrams), train a logistic regression, evaluate on test.
  4. Stage 2 — Neural. Build a PyTorch model: an embedding layer + BiLSTM + linear head. Train for 3 epochs.
  5. Stage 3 — Transformer. Fine-tune distilbert-base-uncased for 2 epochs with the HuggingFace Trainer. Truncate to 64 tokens — tweets rarely need more.
  6. Measure. For each model record: test accuracy, training wall-clock time, inference latency per example, on-disk size.
  7. Compare. Put it all in one table. Discuss which approach you would pick for which use case, and how the curve might change with the full 1.6M.

What to deliver

  • Three notebooks (one per stage) that produce reproducible numbers.
  • A single comparison table with accuracy, training time, inference latency, disk size.
  • A one-page write-up: in which scenario would the team pick each approach? And how would the picture change at 10x more data?
  • A documented note on the subsample size used — so the comparison is reproducible.

References