transformers with a model like sshleifer/distilbart-cnn-12-6, facebook/bart-large-cnn, or google/pegasus-xsum, the evaluate library for ROUGE, an NLI model for the faithfulness pass, and any LLM for the zero-shot baseline. Going further (optional). Build a Streamlit UI where the user pastes an article and sees both summaries side by side with unsupported spans highlighted, or extend with controllable length ("one sentence" / "a paragraph") via prompts or special tokens.What you'll build
Build a text-summarisation system for English news articles and compare two fundamentally different approaches on the same evaluation set: an extractive summariser (picks the most important original sentences) and an abstractive summariser (a fine-tuned encoder-decoder model that writes a new short paragraph). Optionally add a zero-shot Large Language Model (LLM) baseline. Report ROUGE scores and a small human-rated quality study.
What goes in, what comes out
Input
CNN/DailyMail news articles (around 287k train, 13k validation, 11k test). Each article has a multi-sentence reference summary written by the original journalists.
Output
A summariser that, given an article, produces a 2–4 sentence summary. Plus a results table with ROUGE-1 / ROUGE-2 / ROUGE-L scores and the team's human ratings.
article (truncated):
London (CNN) — The Brexit transition period will expire at the end
of December, the UK government confirmed on Friday, with negotiators
in Brussels still working through three sticking points: fishing
rights, state-aid rules, and the dispute-resolution mechanism. The
prime minister told reporters that progress had been "limited" but
added that he was "hopeful" a deal could be reached in the coming
days. [continues for ~700 words]
reference_summary:
The UK government confirmed the Brexit transition will end in December
with three issues — fishing, state aid, and dispute resolution —
still unresolved. The prime minister said progress was limited but
he was hopeful of a deal.
Approach ROUGE-1 ROUGE-2 ROUGE-L Human rating (avg of 50)
------------------- ------- ------- ------- ------------------------
Lead-3 (extractive) 0.40 0.17 0.36 3.4 / 5
BART-base fine-tune 0.43 0.20 0.40 4.0 / 5
LLM zero-shot 0.41 0.18 0.38 4.2 / 5
Note: ROUGE rewards word overlap with the reference; the LLM
summary often paraphrases the reference style and scores lower
on ROUGE but higher with human raters.
Datasets
CNN/DailyMail ↗
English news articles with reference summaries written by journalists. The standard benchmark for English summarisation.
How to get it: from datasets import load_dataset; ds = load_dataset("abisee/cnn_dailymail", "3.0.0").
Tools you'll need
These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.
Python: Python 3.10 or newer. Compute: Colab T4 or any 12 GB GPU for fine-tuning BART-base. Laptop CPU is enough for the extractive baseline and the LLM zero-shot baseline.
transformers— Loads BART or T5 for sequence-to-sequence fine-tuning.datasets— Loads CNN/DailyMail.accelerate— Handles device placement.sentencepiece— Required tokeniser dependency for T5 and many other sequence-to-sequence models.
nltk— Sentence splitter for the Lead-3 baseline (first 3 sentences as summary).
rouge-score— Computes ROUGE-1, ROUGE-2, ROUGE-L. The standard summarisation metric.evaluate— Wraps rouge-score with the HuggingFace API.
openai / anthropic / groq— A hosted LLM. Prompt: "Summarise the following article in 2–3 sentences."
How to approach it
One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.
- Load. Pull CNN/DailyMail. Use the 3.0.0 split. Subsample to ~20k articles for training if you are time-bounded — fine-tuning on the full 287k is overnight on a T4.
- Baseline 1 — Lead-3. Take the first 3 sentences as the "summary". Score with ROUGE. This is the bar that abstractive systems must beat.
- Baseline 2 — BART fine-tune. Fine-tune
facebook/bart-baseon (article, summary) pairs. 2 epochs. Use a max input length of 1024 and max output length of 128. - Baseline 3 — LLM zero-shot (optional). Prompt a hosted LLM to summarise each test article in 2–3 sentences. Deterministic (temperature 0).
- Score with ROUGE. All three on the same 1,000-article test subset.
- Human evaluation. Sample 50 articles. For each, show the team the article + the three summaries (blind to which is which). Rate each summary 1–5 on fluency and faithfulness.
- Compare. ROUGE and human ratings rarely agree. Discuss the gap.
What to deliver
- A reproducible notebook for the extractive and abstractive systems (and the LLM baseline if added).
- A ROUGE table on a fixed test subset.
- A blind human-evaluation report on 50 articles with fluency and faithfulness ratings.
- A short discussion: ROUGE vs human. Where do they disagree and why?