Menu
Home Program Lecturers Important Dates Venue Sponsors Past Editions Speakers Alumni Versions GAI2026 Contact
Headline Rewriter with a Small Language Model — Project #21 | Summer School on Generative AI
Project 21

Headline Rewriter with a Small Language Model

CPU / GPU
Note. This page lays out one full version of the project — the goal, a sample input/output, suggested tools, and a step-by-step plan. Treat it as a reference, not a script. Your team can pick a different angle, swap libraries, narrow the scope, or take the project somewhere we did not anticipate. As long as the final deliverable makes sense for the goal, you are on track.
Task. Fine-tune a small language model (SmolLM-360M / SmolLM2-1.7B / Qwen2.5-0.5B) to rewrite news article leads into headlines in three styles: factual, attention-grabbing, and neutral. The team curates a small paired dataset (article → headline) from open news sources, fine-tunes with full SFT or LoRA, and builds a side-by-side demo where one input produces all three variants. Constrained tasks like this are exactly where a small LLM shines — the search space is narrow, the iteration loop is fast, and the demo is fun. You could use HuggingFace transformers + trl.SFTTrainer, any small base (SmolLM2-360M is a good default), and a hand-curated paired set of ~1000–3000 (article, headline) rows scraped from open news archives or CNN/DailyMail. Evaluation: ROUGE on the factual style, blind human ratings on the three styles. Going further (optional). Add a fourth style for a specific outlet's voice (Reuters, BBC), expose the demo as a small web tool, or evaluate against a hosted LLM zero-shot baseline.
Resources: Colab T4 or 8GB GPU for fine-tuning; CPU works for inference with a sub-1B model.

What you'll build

Fine-tune a small open-source language model (SmolLM2-360M, SmolLM2-1.7B, or Qwen2.5-0.5B) to rewrite the lead paragraph of a news article into a headline. Train it to do this in three distinct styles — factual, attention-grabbing, and neutral — selected by a control token in the prompt. The team curates a small paired dataset of (article lead, reference headline) rows, fine-tunes with Supervised Fine-Tuning (SFT), and builds a side-by-side demo where one input produces all three variants. Constrained rewriting is exactly where a small Large Language Model (LLM) shines: narrow output space, fast iteration loop, fun demo.

What goes in, what comes out

Input

A news-article lead paragraph (the first paragraph or first 200 words) plus a control token specifying the desired style.

Output

A short headline (5–15 words) in the requested style. The demo produces all three variants from one input.

A few rows from the team-curated dataset
[
  {
    "style": "factual",
    "lead": "The European Central Bank raised interest rates by 25 basis points on Thursday, citing persistent inflation pressures across the eurozone, while signalling that further increases would depend on incoming data.",
    "headline": "ECB raises rates by 25 basis points, eyes data for next move"
  },
  {
    "style": "attention",
    "lead": "The European Central Bank raised interest rates by 25 basis points on Thursday...",
    "headline": "ECB stuns markets with another rate hike — more pain ahead?"
  },
  {
    "style": "neutral",
    "lead": "The European Central Bank raised interest rates by 25 basis points on Thursday...",
    "headline": "European Central Bank announces interest-rate increase"
  }
]
Demo output — same lead, three styles, plus evaluation
Input lead (truncated):
  "Researchers at MIT have built a robotic gripper that can pick up
   delicate objects ranging from a single grape to a 3-kilogram box,
   adapting its grip strength in real time using a network of tactile
   sensors. The system, trained entirely in simulation, generalised
   to real-world objects on the first try."

Generated headlines:
  [factual]   "MIT robotic gripper adapts grip strength using tactile sensors"
  [attention] "This MIT robot can pick up anything — from a grape to a box"
  [neutral]   "New tactile-sensor robotic gripper from MIT"

Evaluation on a 200-row held-out test set:
  Approach                       ROUGE-L (factual)   Blind human preference
  ----------------------------   -----------------   ----------------------
  Zero-shot prompted LLM             0.31                  35%
  SmolLM2-360M SFT (this project)    0.38                  41%
  Reference (journalist-written)     1.00                  24%
  (human raters preferred the tuned model 41% of the time)

Datasets

CNN/DailyMail (factual seed) ↗

Use the first paragraph as the lead and the article title as the reference headline. Provides a strong factual-style baseline for training.

How to get it: from datasets import load_dataset; ds = load_dataset("abisee/cnn_dailymail", "3.0.0"). Subsample 2,000–5,000 rows for the factual style.

License: Free for research use.

Team-curated attention-grabbing and neutral sets

For the other two styles you will not find paired data off the shelf. Pick 300–500 leads from the factual set and either (a) rewrite the headline by hand, or (b) generate a draft with a hosted LLM and have the team curate.

How to get it: Spend half a day on this. Two team members write, two review. Save as JSON with style tags.

A small blind eval set

Hold out 200 article leads. These never enter training. For evaluation, the team rates blind side-by-side: tuned model vs zero-shot baseline vs reference.

How to get it: Split off before any training begins. Fix the seed in the README.

Tools you'll need

These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.

Python: Python 3.10 or newer. Compute: Colab T4 / Kaggle P100 (16 GB GPU) is comfortable for SmolLM2-360M or SmolLM2-1.7B SFT. Inference with the tuned 360M model runs on a laptop CPU in real time.

Base model
  • transformers — Loads HuggingFaceTB/SmolLM2-360M (or SmolLM2-1.7B / Qwen2.5-0.5B) and the tokeniser with its chat template.
Training
  • trl — Provides SFTTrainer — wraps the full SFT loop including chat-template formatting and packing.
  • accelerate — Mixed-precision + device placement.
  • torch — Training backend.
Data + UI
  • datasets — Loads CNN/DailyMail and streams the curated paired set.
  • streamlit — Builds the three-styles-side-by-side demo.
  • sentencepiece — Tokeniser dependency.
Evaluation
  • rouge-score — For the ROUGE comparison on the factual style.
  • evaluate — HuggingFace metric wrapper.

How to approach it

One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.

  1. Pick the base model. SmolLM2-360M is a good default. SmolLM2-1.7B gives more headroom if your GPU has the room.
  2. Build the paired dataset. Pull 2,000–5,000 (lead, headline) pairs from CNN/DailyMail for the factual style. Hand-curate 300–500 for the attention and neutral styles.
  3. Format with the chat template. Each training row becomes a user message "Rewrite this lead as a [style] headline:\n[lead]" and an assistant message with the target headline. Use tokenizer.apply_chat_template.
  4. Run SFT. Use trl.SFTTrainer. 2–3 epochs, learning rate 2e-5, batch size 8, bf16. About 30–60 minutes on a T4.
  5. Build the demo. Streamlit app: paste a lead, hit Generate, see all three styles side by side. Cache the model to avoid reloading per request.
  6. Evaluate. On the held-out 200 leads, generate all three styles. Score the factual style with ROUGE-L. Run a blind side-by-side: tuned vs hosted-LLM zero-shot vs reference. Rate 1–5 on style match and quality.
  7. Report. ROUGE numbers + human-preference rates + a small gallery of example outputs.

What to deliver

  • The fine-tuned model checkpoint (HuggingFace repo or saved artifacts).
  • A Streamlit demo that produces the three styles from one input.
  • An eval report on the 200 held-out leads (ROUGE + blind side-by-side ratings).
  • A short README documenting base model, dataset composition, hyperparameters, and findings.

References