Menu
Home Program Lecturers Important Dates Venue Sponsors Past Editions Speakers Alumni Versions GAI2026 Contact
Natural Language Inference on SNLI: Hand-Crafted, Siamese LSTM, and RoBERTa — Project #8 | Summer School on Generative AI
Project 8

Natural Language Inference on SNLI: Hand-Crafted, Siamese LSTM, and RoBERTa

GPU
Note. This page lays out one full version of the project — the goal, a sample input/output, suggested tools, and a step-by-step plan. Treat it as a reference, not a script. Your team can pick a different angle, swap libraries, narrow the scope, or take the project somewhere we did not anticipate. As long as the final deliverable makes sense for the goal, you are on track.
Task. Use SNLI (premise–hypothesis pairs, 3 classes: entailment, contradiction, neutral). Build three classifiers — a classical model with hand-crafted features, a small neural sentence-pair model, and a fine-tuned transformer — and report accuracy on the dev set. As a bonus, train each model on hypotheses only (no premise) to reveal the famous SNLI dataset artifact. You could use sklearn for the classical baseline (logistic regression on lexical overlap, length difference, negation cues), PyTorch for the neural model (Siamese BiLSTM is a classic choice), and HuggingFace transformers with roberta-base or DeBERTa. Going further (optional). Build a small UI where someone types two sentences and sees the predicted relation, or extend with an LLM zero-shot baseline and compare it to the supervised numbers.
Resources: CPU works for the hand-crafted baseline; Colab T4 or 8GB GPU for the LSTM and RoBERTa.

What you'll build

Build three Natural Language Inference (NLI) classifiers and compare them on the same dev set. Given two sentences — a premise and a hypothesis — the system decides whether the hypothesis is entailed, contradicted, or neutral with respect to the premise. NLI is the canonical sentence-pair task: it is where bag-of-words is hopeless, small neural models start to work, and transformers genuinely earn their cost.

What goes in, what comes out

Input

SNLI (Stanford Natural Language Inference) sentence pairs. Around 550k train, 10k dev, 10k test. Each pair labelled entailment, contradiction, or neutral.

Output

Three classifiers plus a results table with overall accuracy and per-class F1, plus a confusion matrix from the transformer.

A few rows from the train set
label          premise                                          hypothesis
-------------  -----------------------------------------------  -----------------------------
entailment     A soccer game with multiple males playing.       Some men are playing a sport.
contradiction  A soccer game with multiple males playing.       The men are sitting on a bench.
neutral        A soccer game with multiple males playing.       The men are playing in a league match.
entailment     A woman in a blue jacket walking a dog.          A person is outside with an animal.
contradiction  A woman in a blue jacket walking a dog.          A woman is asleep at home.
Dev-set results from the three approaches
Approach                    Accuracy  F1 (ent)  F1 (con)  F1 (neu)
-------------------------   --------  --------  --------  --------
TF-IDF + Logistic Reg.        0.55      0.58      0.56      0.51
BiLSTM + word embeddings      0.74      0.76      0.75      0.71
DistilBERT fine-tune          0.87      0.89      0.88      0.84

Confusion matrix (DistilBERT, dev):
                           predicted
                     ent     con     neu
  actual ent        2980     85      275
         con          93   2950      298
         neu         330    312     2677
  (neutral is the hardest class — it gets confused with both ends)

Datasets

SNLI (Stanford Natural Language Inference) ↗

570k human-written sentence pairs covering everyday scenes, labelled by crowdworkers. The canonical NLI benchmark.

How to get it: from datasets import load_dataset; ds = load_dataset("stanfordnlp/snli"). Filter out rows where label == -1 (no consensus).

License: Creative Commons Attribution-ShareAlike 4.0.

Tools you'll need

These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.

Python: Python 3.10 or newer. Compute: Colab T4 or any 8 GB GPU for the transformer. Laptop CPU is enough for stages 1 and 2 on a subsample.

Stage 1 — classical
  • scikit-learn — TF-IDF on concatenated premise + hypothesis, plus logistic regression. The "bag of words has no chance" baseline.
Stage 2 — neural
  • torch — PyTorch. A small BiLSTM that encodes each sentence and combines them
Stage 3 — transformer
  • transformers — Fine-tune DistilBERT on the sentence pair.
  • datasets — Streams SNLI in.
  • accelerate — Hardware-agnostic training loop.
Evaluation
  • evaluate — Accuracy and per-class F1 wrappers.
  • scikit-learn — Confusion matrix and classification report.

How to approach it

One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.

  1. Load. Pull SNLI. Drop the label == -1 rows (no annotator consensus).
  2. Subsample (optional). 550k is more than you need. Train on a 100k subsample if you are time-bounded.
  3. Stage 1 — Classical. TF-IDF on premise + " " + hypothesis, train a 3-class logistic regression. Expect ~55% accuracy — barely above random (33%). That is the point.
  4. Stage 2 — Neural. Build a Siamese BiLSTM in PyTorch. Encode premise and hypothesis separately, project to 3 classes.
  5. Stage 3 — Transformer. Tokenise as a pair ([CLS] p [SEP] h [SEP]). Fine-tune distilbert-base-uncased for 2 epochs.
  6. Score. Overall accuracy plus per-class F1 for all three. Confusion matrix for the transformer.

What to deliver

  • Three notebooks (one per stage) producing the comparison numbers.
  • A single results table with overall accuracy and per-class F1.
  • A confusion matrix for the transformer plus a short error analysis on the neutral class (it is reliably the hardest).

References