Menu
Home Program Lecturers Important Dates Venue Sponsors Past Editions Speakers Alumni Versions GAI2026 Contact
Sentiment Classification on SST-2 — Project #4 | Summer School on Generative AI
Project 4

Sentiment Classification on SST-2

CPU / GPU
Note. This page lays out one full version of the project — the goal, a sample input/output, suggested tools, and a step-by-step plan. Treat it as a reference, not a script. Your team can pick a different angle, swap libraries, narrow the scope, or take the project somewhere we did not anticipate. As long as the final deliverable makes sense for the goal, you are on track.
Task. Use SST-2 (binary movie review sentiment, ~67k train, 872 dev). Build a fine-tuned sentiment classifier and compare it against an LLM with zero-shot and few-shot prompts on the same dev set. Report accuracy, latency per example, and a confusion table — especially for the negation cases SST-2 is known for. You could use HuggingFace transformers for fine-tuning (distilbert-base-uncased is a good starting point), any LLM for prompting, and HuggingFace datasets for the data. Going further (optional). Build a small UI where someone types a review and sees both predictions side by side, or extend with a third model — sentence embeddings + logistic regression — to see where the classical baseline lands.
Resources: Colab T4 for the fine-tuning baseline. CPU works for the LLM prompting variants.

What you'll build

Build a sentiment classifier for short English movie review sentences. The system reads one sentence and predicts whether it expresses positive or negative sentiment. The team will compare a fine-tuned encoder against a hosted language model used zero-shot, on the same dev set, and report accuracy plus a few illustrative failure cases.

What goes in, what comes out

Input

SST-2 movie review sentences (around 67,000 in train, 872 in dev). Each labelled 0 (negative) or 1 (positive).

Output

A classifier that, given a sentence, returns a probability for each class. Plus an evaluation report with accuracy, a confusion matrix, and a few example failure cases.

A few rows from the dev set
label  sentence
-----  --------
  1    a powerful, deeply moving film with two stunning performances.
  0    a chaotic, often unfunny mess that wastes a great cast.
  1    quietly profound and beautifully shot.
  0    feels like a series of disconnected sketches strung together.
  0    the dialogue is wooden and the plot makes no sense.
Dev-set results from both approaches
Approach              Accuracy  Latency/example
-------------------   --------  ---------------
DistilBERT fine-tune    0.91      ~5 ms (T4 GPU)
LLM zero-shot           0.88      ~600 ms (hosted)
LLM 4-shot              0.91      ~700 ms (hosted)

Confusion matrix (DistilBERT, dev):
                  predicted
              neg       pos
  actual neg  398        30
         pos   46       398
  (16 negative misses, 30 positive misses)

Datasets

SST-2 (Stanford Sentiment Treebank, binary) ↗

Sentences from movie reviews labelled positive or negative. Part of the GLUE benchmark.

How to get it: from datasets import load_dataset; ds = load_dataset("stanfordnlp/sst2").

License: Free for research use.

Tools you'll need

These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.

Python: Python 3.10 or newer. Compute: Colab T4 or any 8 GB GPU for fine-tuning. A laptop CPU works for the zero-shot LLM baseline.

Core
  • datasets — Loads SST-2.
  • transformers — Loads and fine-tunes DistilBERT.
  • evaluate — Wrapper around accuracy and F1 metrics.
  • scikit-learn — For the confusion matrix and classification report.
Optional (LLM baselines)
  • openai / anthropic / groq — A hosted LLM to compare against the fine-tuned model.

How to approach it

One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.

  1. Load. Pull SST-2 from HuggingFace. Print 10 random examples per class so the team understands the data.
  2. Tokenise. Run the train and dev splits through the DistilBERT tokeniser with a max length around 64.
  3. Fine-tune. Train distilbert-base-uncased with the HuggingFace Trainer. 2–3 epochs is usually enough.
  4. Score the supervised model. Compute dev accuracy and a confusion matrix.
  5. Prompt the LLM. Run zero-shot and few-shot prompts on the same dev set. Be deterministic (temperature 0).
  6. Compare. Put the three numbers in a table. Look at where each approach fails.

What to deliver

  • A reproducible notebook that fine-tunes DistilBERT and runs the LLM baseline on the dev set.
  • A results table comparing the three approaches on accuracy, latency, and cost (if hosted).
  • A short error analysis: 10 examples each approach gets wrong, with the team's guess at why.

References