sklearn for the classical baseline (logistic regression on lexical overlap, length difference, negation cues), PyTorch for the neural model (Siamese BiLSTM is a classic choice), and HuggingFace transformers with roberta-base or DeBERTa. Going further (optional). Build a small UI where someone types two sentences and sees the predicted relation, or extend with an LLM zero-shot baseline and compare it to the supervised numbers.What you'll build
Build three Natural Language Inference (NLI) classifiers and compare them on the same dev set. Given two sentences — a premise and a hypothesis — the system decides whether the hypothesis is entailed, contradicted, or neutral with respect to the premise. NLI is the canonical sentence-pair task: it is where bag-of-words is hopeless, small neural models start to work, and transformers genuinely earn their cost.
What goes in, what comes out
Input
SNLI (Stanford Natural Language Inference) sentence pairs. Around 550k train, 10k dev, 10k test. Each pair labelled entailment, contradiction, or neutral.
Output
Three classifiers plus a results table with overall accuracy and per-class F1, plus a confusion matrix from the transformer.
label premise hypothesis
------------- ----------------------------------------------- -----------------------------
entailment A soccer game with multiple males playing. Some men are playing a sport.
contradiction A soccer game with multiple males playing. The men are sitting on a bench.
neutral A soccer game with multiple males playing. The men are playing in a league match.
entailment A woman in a blue jacket walking a dog. A person is outside with an animal.
contradiction A woman in a blue jacket walking a dog. A woman is asleep at home.
Approach Accuracy F1 (ent) F1 (con) F1 (neu)
------------------------- -------- -------- -------- --------
TF-IDF + Logistic Reg. 0.55 0.58 0.56 0.51
BiLSTM + word embeddings 0.74 0.76 0.75 0.71
DistilBERT fine-tune 0.87 0.89 0.88 0.84
Confusion matrix (DistilBERT, dev):
predicted
ent con neu
actual ent 2980 85 275
con 93 2950 298
neu 330 312 2677
(neutral is the hardest class — it gets confused with both ends)
Datasets
SNLI (Stanford Natural Language Inference) ↗
570k human-written sentence pairs covering everyday scenes, labelled by crowdworkers. The canonical NLI benchmark.
How to get it: from datasets import load_dataset; ds = load_dataset("stanfordnlp/snli"). Filter out rows where label == -1 (no consensus).
Tools you'll need
These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.
Python: Python 3.10 or newer. Compute: Colab T4 or any 8 GB GPU for the transformer. Laptop CPU is enough for stages 1 and 2 on a subsample.
scikit-learn— TF-IDF on concatenated premise + hypothesis, plus logistic regression. The "bag of words has no chance" baseline.
torch— PyTorch. A small BiLSTM that encodes each sentence and combines them
transformers— Fine-tune DistilBERT on the sentence pair.datasets— Streams SNLI in.accelerate— Hardware-agnostic training loop.
evaluate— Accuracy and per-class F1 wrappers.scikit-learn— Confusion matrix and classification report.
How to approach it
One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.
- Load. Pull SNLI. Drop the
label == -1rows (no annotator consensus). - Subsample (optional). 550k is more than you need. Train on a 100k subsample if you are time-bounded.
- Stage 1 — Classical. TF-IDF on premise + " " + hypothesis, train a 3-class logistic regression. Expect ~55% accuracy — barely above random (33%). That is the point.
- Stage 2 — Neural. Build a Siamese BiLSTM in PyTorch. Encode premise and hypothesis separately, project to 3 classes.
- Stage 3 — Transformer. Tokenise as a pair ([CLS] p [SEP] h [SEP]). Fine-tune
distilbert-base-uncasedfor 2 epochs. - Score. Overall accuracy plus per-class F1 for all three. Confusion matrix for the transformer.
What to deliver
- Three notebooks (one per stage) producing the comparison numbers.
- A single results table with overall accuracy and per-class F1.
- A confusion matrix for the transformer plus a short error analysis on the neutral class (it is reliably the hardest).