Menu
Home Program Lecturers Important Dates Venue Sponsors Past Editions Speakers Alumni Versions GAI2026 Contact
Automatic Text Summarization — Project #17 | Summer School on Generative AI
Project 17

Automatic Text Summarization

CPU / GPU
Note. This page lays out one full version of the project — the goal, a sample input/output, suggested tools, and a step-by-step plan. Treat it as a reference, not a script. Your team can pick a different angle, swap libraries, narrow the scope, or take the project somewhere we did not anticipate. As long as the final deliverable makes sense for the goal, you are on track.
Task. Build a summarization system that takes a longer text (a news article, a research abstract, a blog post) and produces a short, faithful summary. Compare two approaches on the same ~200 test documents: a small fine-tuned summariser and a zero-shot LLM. Report ROUGE-1/2/L, output length, and a small faithfulness check using an NLI model to flag claims not supported by the source. You could use HuggingFace transformers with a model like sshleifer/distilbart-cnn-12-6, facebook/bart-large-cnn, or google/pegasus-xsum, the evaluate library for ROUGE, an NLI model for the faithfulness pass, and any LLM for the zero-shot baseline. Going further (optional). Build a Streamlit UI where the user pastes an article and sees both summaries side by side with unsupported spans highlighted, or extend with controllable length ("one sentence" / "a paragraph") via prompts or special tokens.
Resources: CPU works for hosted LLMs and small models; Colab T4 / 8GB GPU if you fine-tune.

What you'll build

Build a text-summarisation system for English news articles and compare two fundamentally different approaches on the same evaluation set: an extractive summariser (picks the most important original sentences) and an abstractive summariser (a fine-tuned encoder-decoder model that writes a new short paragraph). Optionally add a zero-shot Large Language Model (LLM) baseline. Report ROUGE scores and a small human-rated quality study.

What goes in, what comes out

Input

CNN/DailyMail news articles (around 287k train, 13k validation, 11k test). Each article has a multi-sentence reference summary written by the original journalists.

Output

A summariser that, given an article, produces a 2–4 sentence summary. Plus a results table with ROUGE-1 / ROUGE-2 / ROUGE-L scores and the team's human ratings.

One row from the train set (article truncated)
article (truncated):
  London (CNN) — The Brexit transition period will expire at the end
  of December, the UK government confirmed on Friday, with negotiators
  in Brussels still working through three sticking points: fishing
  rights, state-aid rules, and the dispute-resolution mechanism. The
  prime minister told reporters that progress had been "limited" but
  added that he was "hopeful" a deal could be reached in the coming
  days. [continues for ~700 words]

reference_summary:
  The UK government confirmed the Brexit transition will end in December
  with three issues — fishing, state aid, and dispute resolution —
  still unresolved. The prime minister said progress was limited but
  he was hopeful of a deal.
Test-set results from the two approaches
Approach              ROUGE-1   ROUGE-2   ROUGE-L   Human rating (avg of 50)
-------------------   -------   -------   -------   ------------------------
Lead-3 (extractive)    0.40      0.17      0.36       3.4 / 5
BART-base fine-tune    0.43      0.20      0.40       4.0 / 5
LLM zero-shot          0.41      0.18      0.38       4.2 / 5

Note: ROUGE rewards word overlap with the reference; the LLM
summary often paraphrases the reference style and scores lower
on ROUGE but higher with human raters.

Datasets

CNN/DailyMail ↗

English news articles with reference summaries written by journalists. The standard benchmark for English summarisation.

How to get it: from datasets import load_dataset; ds = load_dataset("abisee/cnn_dailymail", "3.0.0").

License: Free for research use.

Tools you'll need

These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.

Python: Python 3.10 or newer. Compute: Colab T4 or any 12 GB GPU for fine-tuning BART-base. Laptop CPU is enough for the extractive baseline and the LLM zero-shot baseline.

Core
  • transformers — Loads BART or T5 for sequence-to-sequence fine-tuning.
  • datasets — Loads CNN/DailyMail.
  • accelerate — Handles device placement.
  • sentencepiece — Required tokeniser dependency for T5 and many other sequence-to-sequence models.
Extractive baseline
  • nltk — Sentence splitter for the Lead-3 baseline (first 3 sentences as summary).
Evaluation
  • rouge-score — Computes ROUGE-1, ROUGE-2, ROUGE-L. The standard summarisation metric.
  • evaluate — Wraps rouge-score with the HuggingFace API.
Optional (LLM baseline)
  • openai / anthropic / groq — A hosted LLM. Prompt: "Summarise the following article in 2–3 sentences."

How to approach it

One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.

  1. Load. Pull CNN/DailyMail. Use the 3.0.0 split. Subsample to ~20k articles for training if you are time-bounded — fine-tuning on the full 287k is overnight on a T4.
  2. Baseline 1 — Lead-3. Take the first 3 sentences as the "summary". Score with ROUGE. This is the bar that abstractive systems must beat.
  3. Baseline 2 — BART fine-tune. Fine-tune facebook/bart-base on (article, summary) pairs. 2 epochs. Use a max input length of 1024 and max output length of 128.
  4. Baseline 3 — LLM zero-shot (optional). Prompt a hosted LLM to summarise each test article in 2–3 sentences. Deterministic (temperature 0).
  5. Score with ROUGE. All three on the same 1,000-article test subset.
  6. Human evaluation. Sample 50 articles. For each, show the team the article + the three summaries (blind to which is which). Rate each summary 1–5 on fluency and faithfulness.
  7. Compare. ROUGE and human ratings rarely agree. Discuss the gap.

What to deliver

  • A reproducible notebook for the extractive and abstractive systems (and the LLM baseline if added).
  • A ROUGE table on a fixed test subset.
  • A blind human-evaluation report on 50 articles with fluency and faithfulness ratings.
  • A short discussion: ROUGE vs human. Where do they disagree and why?

References