sentence-transformers embeddings with a logistic-regression threshold, a fine-tuned encoder (bert-base-uncased, RoBERTa), an LLM zero-shot baseline, or any combination of these. Going further (optional). Build a small UI where someone types two sentences and sees the paraphrase probability, or extend with a hard-negative mining script that finds the cases your best model gets wrong.What you'll build
Build a binary classifier that decides whether two sentences are paraphrases of each other. The team will build three approaches and compare them on the same dev set: cosine similarity between sentence embeddings (with a tuned threshold), a fine-tuned encoder on the sentence pair, and a zero-shot language model. The headline number is F1 on the paraphrase class, because MRPC is class-imbalanced.
What goes in, what comes out
Input
MRPC sentence pairs from the GLUE benchmark (around 3,700 in train, 408 in dev). Each pair labelled 0 (not paraphrase) or 1 (paraphrase). The training set is ~67% positive — imbalanced.
Output
Three classifiers (embedding + threshold, BERT fine-tuned, LLM zero-shot) plus a results table with accuracy and F1 on the paraphrase class.
label sentence_1 sentence_2
----- ------------------------------------------------------ --------------------------------------------
1 Amrozi accused his brother of deliberately misleading Referring to him as only "the witness",
the court. Amrozi accused his brother of deliberately
distorting his evidence.
0 Yucaipa owned Dominick's before selling the chain Yucaipa bought Dominick's in 1995 for $693
to Safeway in 1998 for $2.5 billion. million and sold it to Safeway for $1.8 billion
in 1998.
1 Around 0334 GMT, Nikkei futures suggested it would The September Nikkei futures were quoted at
open up around 9,250 yen. 9,250 yen, suggesting an open up around 50.
Approach Accuracy F1 (paraphrase) F1 (not-paraphrase)
------------------------ -------- --------------- -------------------
Embedding + threshold 0.74 0.83 0.45
BERT fine-tune 0.85 0.89 0.74
LLM zero-shot 0.79 0.85 0.61
Note: training set is ~67% positive. Predicting "always paraphrase"
already gives 67% accuracy with F1=0.80 on the majority class.
Datasets
MRPC (Microsoft Research Paraphrase Corpus) ↗
Sentence pairs from news articles, hand-labelled for paraphrase. Small (3,700 + 408 + 1,725) but enough for fine-tuning.
How to get it: from datasets import load_dataset; ds = load_dataset("nyu-mll/glue", "mrpc").
Tools you'll need
These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.
Python: Python 3.10 or newer. Compute: A laptop CPU is enough for the embedding-similarity baseline. Colab T4 for the BERT fine-tune.
datasets— Loads MRPC.transformers— Loads BERT for the fine-tune.sentence-transformers— Pre-built sentence embedders.all-MiniLM-L6-v2is a good small choice.scikit-learn— Logistic regression on top of cosine similarity, plus the metrics report.evaluate— Easy accuracy and F1 wrappers.
openai / anthropic / groq— A hosted LLM for the zero-shot baseline. Prompt it to reply YES or NO.
How to approach it
One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.
- Load. Pull MRPC. Check class balance (around 67% positive in train) and decide the metric: F1 on the paraphrase class.
- Baseline 1 — Embedding + threshold. Embed each sentence with
all-MiniLM-L6-v2. Compute cosine similarity per pair. Fit a logistic regression (1 feature) on train to learn the threshold. - Baseline 2 — BERT fine-tune. Tokenise the sentence pair as one input ("[CLS] sent1 [SEP] sent2 [SEP]"). Fine-tune
bert-base-uncasedwith the HuggingFaceTrainer. - Baseline 3 — LLM zero-shot. Prompt a hosted LLM with both sentences and ask "Are these paraphrases? YES or NO." Be deterministic.
- Score. Compute accuracy, F1 on positive, F1 on negative for all three. Put them in a single table.
- Analyse. Look at the cases where the three approaches disagree. Pick 10 examples and discuss.
What to deliver
- A reproducible notebook for all three approaches.
- A results table comparing accuracy and per-class F1.
- A short error analysis: 10 disagreement cases with the team's read on which approach was right.