Menu
Home Program Lecturers Important Dates Venue Sponsors Past Editions Speakers Alumni Versions GAI2026 Contact
Paraphrase Detection on MRPC — Project #5 | Summer School on Generative AI
Project 5

Paraphrase Detection on MRPC

CPU / GPU
Note. This page lays out one full version of the project — the goal, a sample input/output, suggested tools, and a step-by-step plan. Treat it as a reference, not a script. Your team can pick a different angle, swap libraries, narrow the scope, or take the project somewhere we did not anticipate. As long as the final deliverable makes sense for the goal, you are on track.
Task. Use MRPC from GLUE (binary paraphrase classification, ~3.7k train, 408 dev). Build paraphrase classifiers and report accuracy and F1 on the dev set. MRPC is small and class-imbalanced — watch the F1 on the minority class, not just accuracy. You could use cosine similarity over sentence-transformers embeddings with a logistic-regression threshold, a fine-tuned encoder (bert-base-uncased, RoBERTa), an LLM zero-shot baseline, or any combination of these. Going further (optional). Build a small UI where someone types two sentences and sees the paraphrase probability, or extend with a hard-negative mining script that finds the cases your best model gets wrong.
Resources: CPU works for sentence-embedding baselines; Colab T4 if you also fine-tune.

What you'll build

Build a binary classifier that decides whether two sentences are paraphrases of each other. The team will build three approaches and compare them on the same dev set: cosine similarity between sentence embeddings (with a tuned threshold), a fine-tuned encoder on the sentence pair, and a zero-shot language model. The headline number is F1 on the paraphrase class, because MRPC is class-imbalanced.

What goes in, what comes out

Input

MRPC sentence pairs from the GLUE benchmark (around 3,700 in train, 408 in dev). Each pair labelled 0 (not paraphrase) or 1 (paraphrase). The training set is ~67% positive — imbalanced.

Output

Three classifiers (embedding + threshold, BERT fine-tuned, LLM zero-shot) plus a results table with accuracy and F1 on the paraphrase class.

A few rows from the train set
label  sentence_1                                              sentence_2
-----  ------------------------------------------------------  --------------------------------------------
  1    Amrozi accused his brother of deliberately misleading   Referring to him as only "the witness",
       the court.                                              Amrozi accused his brother of deliberately
                                                               distorting his evidence.
  0    Yucaipa owned Dominick's before selling the chain       Yucaipa bought Dominick's in 1995 for $693
       to Safeway in 1998 for $2.5 billion.                    million and sold it to Safeway for $1.8 billion
                                                               in 1998.
  1    Around 0334 GMT, Nikkei futures suggested it would       The September Nikkei futures were quoted at
       open up around 9,250 yen.                              9,250 yen, suggesting an open up around 50.
Dev-set results from the three approaches
Approach                   Accuracy  F1 (paraphrase)  F1 (not-paraphrase)
------------------------   --------  ---------------  -------------------
Embedding + threshold        0.74         0.83              0.45
BERT fine-tune               0.85         0.89              0.74
LLM zero-shot                0.79         0.85              0.61

Note: training set is ~67% positive. Predicting "always paraphrase"
already gives 67% accuracy with F1=0.80 on the majority class.

Datasets

MRPC (Microsoft Research Paraphrase Corpus) ↗

Sentence pairs from news articles, hand-labelled for paraphrase. Small (3,700 + 408 + 1,725) but enough for fine-tuning.

How to get it: from datasets import load_dataset; ds = load_dataset("nyu-mll/glue", "mrpc").

License: Free for research use; part of the GLUE benchmark.

Tools you'll need

These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.

Python: Python 3.10 or newer. Compute: A laptop CPU is enough for the embedding-similarity baseline. Colab T4 for the BERT fine-tune.

Core
  • datasets — Loads MRPC.
  • transformers — Loads BERT for the fine-tune.
  • sentence-transformers — Pre-built sentence embedders. all-MiniLM-L6-v2 is a good small choice.
  • scikit-learn — Logistic regression on top of cosine similarity, plus the metrics report.
  • evaluate — Easy accuracy and F1 wrappers.
Optional (LLM baseline)
  • openai / anthropic / groq — A hosted LLM for the zero-shot baseline. Prompt it to reply YES or NO.

How to approach it

One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.

  1. Load. Pull MRPC. Check class balance (around 67% positive in train) and decide the metric: F1 on the paraphrase class.
  2. Baseline 1 — Embedding + threshold. Embed each sentence with all-MiniLM-L6-v2. Compute cosine similarity per pair. Fit a logistic regression (1 feature) on train to learn the threshold.
  3. Baseline 2 — BERT fine-tune. Tokenise the sentence pair as one input ("[CLS] sent1 [SEP] sent2 [SEP]"). Fine-tune bert-base-uncased with the HuggingFace Trainer.
  4. Baseline 3 — LLM zero-shot. Prompt a hosted LLM with both sentences and ask "Are these paraphrases? YES or NO." Be deterministic.
  5. Score. Compute accuracy, F1 on positive, F1 on negative for all three. Put them in a single table.
  6. Analyse. Look at the cases where the three approaches disagree. Pick 10 examples and discuss.

What to deliver

  • A reproducible notebook for all three approaches.
  • A results table comparing accuracy and per-class F1.
  • A short error analysis: 10 disagreement cases with the team's read on which approach was right.

References