transformers for fine-tuning (distilbert-base-uncased is a good starting point), any LLM for prompting, and HuggingFace datasets for the data. Going further (optional). Build a small UI where someone types a review and sees both predictions side by side, or extend with a third model — sentence embeddings + logistic regression — to see where the classical baseline lands.What you'll build
Build a sentiment classifier for short English movie review sentences. The system reads one sentence and predicts whether it expresses positive or negative sentiment. The team will compare a fine-tuned encoder against a hosted language model used zero-shot, on the same dev set, and report accuracy plus a few illustrative failure cases.
What goes in, what comes out
Input
SST-2 movie review sentences (around 67,000 in train, 872 in dev). Each labelled 0 (negative) or 1 (positive).
Output
A classifier that, given a sentence, returns a probability for each class. Plus an evaluation report with accuracy, a confusion matrix, and a few example failure cases.
label sentence
----- --------
1 a powerful, deeply moving film with two stunning performances.
0 a chaotic, often unfunny mess that wastes a great cast.
1 quietly profound and beautifully shot.
0 feels like a series of disconnected sketches strung together.
0 the dialogue is wooden and the plot makes no sense.
Approach Accuracy Latency/example
------------------- -------- ---------------
DistilBERT fine-tune 0.91 ~5 ms (T4 GPU)
LLM zero-shot 0.88 ~600 ms (hosted)
LLM 4-shot 0.91 ~700 ms (hosted)
Confusion matrix (DistilBERT, dev):
predicted
neg pos
actual neg 398 30
pos 46 398
(16 negative misses, 30 positive misses)
Datasets
SST-2 (Stanford Sentiment Treebank, binary) ↗
Sentences from movie reviews labelled positive or negative. Part of the GLUE benchmark.
How to get it: from datasets import load_dataset; ds = load_dataset("stanfordnlp/sst2").
Tools you'll need
These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.
Python: Python 3.10 or newer. Compute: Colab T4 or any 8 GB GPU for fine-tuning. A laptop CPU works for the zero-shot LLM baseline.
datasets— Loads SST-2.transformers— Loads and fine-tunes DistilBERT.evaluate— Wrapper around accuracy and F1 metrics.scikit-learn— For the confusion matrix and classification report.
openai / anthropic / groq— A hosted LLM to compare against the fine-tuned model.
How to approach it
One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.
- Load. Pull SST-2 from HuggingFace. Print 10 random examples per class so the team understands the data.
- Tokenise. Run the train and dev splits through the DistilBERT tokeniser with a max length around 64.
- Fine-tune. Train
distilbert-base-uncasedwith the HuggingFaceTrainer. 2–3 epochs is usually enough. - Score the supervised model. Compute dev accuracy and a confusion matrix.
- Prompt the LLM. Run zero-shot and few-shot prompts on the same dev set. Be deterministic (temperature 0).
- Compare. Put the three numbers in a table. Look at where each approach fails.
What to deliver
- A reproducible notebook that fine-tunes DistilBERT and runs the LLM baseline on the dev set.
- A results table comparing the three approaches on accuracy, latency, and cost (if hosted).
- A short error analysis: 10 examples each approach gets wrong, with the team's guess at why.