transformers + trl.SFTTrainer, any small base (SmolLM2-360M is a good default), and a hand-curated paired set of ~1000–3000 (article, headline) rows scraped from open news archives or CNN/DailyMail. Evaluation: ROUGE on the factual style, blind human ratings on the three styles. Going further (optional). Add a fourth style for a specific outlet's voice (Reuters, BBC), expose the demo as a small web tool, or evaluate against a hosted LLM zero-shot baseline.What you'll build
Fine-tune a small open-source language model (SmolLM2-360M, SmolLM2-1.7B, or Qwen2.5-0.5B) to rewrite the lead paragraph of a news article into a headline. Train it to do this in three distinct styles — factual, attention-grabbing, and neutral — selected by a control token in the prompt. The team curates a small paired dataset of (article lead, reference headline) rows, fine-tunes with Supervised Fine-Tuning (SFT), and builds a side-by-side demo where one input produces all three variants. Constrained rewriting is exactly where a small Large Language Model (LLM) shines: narrow output space, fast iteration loop, fun demo.
What goes in, what comes out
Input
A news-article lead paragraph (the first paragraph or first 200 words) plus a control token specifying the desired style.
Output
A short headline (5–15 words) in the requested style. The demo produces all three variants from one input.
[
{
"style": "factual",
"lead": "The European Central Bank raised interest rates by 25 basis points on Thursday, citing persistent inflation pressures across the eurozone, while signalling that further increases would depend on incoming data.",
"headline": "ECB raises rates by 25 basis points, eyes data for next move"
},
{
"style": "attention",
"lead": "The European Central Bank raised interest rates by 25 basis points on Thursday...",
"headline": "ECB stuns markets with another rate hike — more pain ahead?"
},
{
"style": "neutral",
"lead": "The European Central Bank raised interest rates by 25 basis points on Thursday...",
"headline": "European Central Bank announces interest-rate increase"
}
]
Input lead (truncated):
"Researchers at MIT have built a robotic gripper that can pick up
delicate objects ranging from a single grape to a 3-kilogram box,
adapting its grip strength in real time using a network of tactile
sensors. The system, trained entirely in simulation, generalised
to real-world objects on the first try."
Generated headlines:
[factual] "MIT robotic gripper adapts grip strength using tactile sensors"
[attention] "This MIT robot can pick up anything — from a grape to a box"
[neutral] "New tactile-sensor robotic gripper from MIT"
Evaluation on a 200-row held-out test set:
Approach ROUGE-L (factual) Blind human preference
---------------------------- ----------------- ----------------------
Zero-shot prompted LLM 0.31 35%
SmolLM2-360M SFT (this project) 0.38 41%
Reference (journalist-written) 1.00 24%
(human raters preferred the tuned model 41% of the time)
Datasets
CNN/DailyMail (factual seed) ↗
Use the first paragraph as the lead and the article title as the reference headline. Provides a strong factual-style baseline for training.
How to get it: from datasets import load_dataset; ds = load_dataset("abisee/cnn_dailymail", "3.0.0"). Subsample 2,000–5,000 rows for the factual style.
Team-curated attention-grabbing and neutral sets
For the other two styles you will not find paired data off the shelf. Pick 300–500 leads from the factual set and either (a) rewrite the headline by hand, or (b) generate a draft with a hosted LLM and have the team curate.
How to get it: Spend half a day on this. Two team members write, two review. Save as JSON with style tags.
A small blind eval set
Hold out 200 article leads. These never enter training. For evaluation, the team rates blind side-by-side: tuned model vs zero-shot baseline vs reference.
How to get it: Split off before any training begins. Fix the seed in the README.
Tools you'll need
These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.
Python: Python 3.10 or newer. Compute: Colab T4 / Kaggle P100 (16 GB GPU) is comfortable for SmolLM2-360M or SmolLM2-1.7B SFT. Inference with the tuned 360M model runs on a laptop CPU in real time.
transformers— LoadsHuggingFaceTB/SmolLM2-360M(orSmolLM2-1.7B/Qwen2.5-0.5B) and the tokeniser with its chat template.
trl— ProvidesSFTTrainer— wraps the full SFT loop including chat-template formatting and packing.accelerate— Mixed-precision + device placement.torch— Training backend.
datasets— Loads CNN/DailyMail and streams the curated paired set.streamlit— Builds the three-styles-side-by-side demo.sentencepiece— Tokeniser dependency.
rouge-score— For the ROUGE comparison on the factual style.evaluate— HuggingFace metric wrapper.
How to approach it
One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.
- Pick the base model. SmolLM2-360M is a good default. SmolLM2-1.7B gives more headroom if your GPU has the room.
- Build the paired dataset. Pull 2,000–5,000 (lead, headline) pairs from CNN/DailyMail for the factual style. Hand-curate 300–500 for the attention and neutral styles.
- Format with the chat template. Each training row becomes a user message "Rewrite this lead as a [style] headline:\n[lead]" and an assistant message with the target headline. Use
tokenizer.apply_chat_template. - Run SFT. Use
trl.SFTTrainer. 2–3 epochs, learning rate 2e-5, batch size 8, bf16. About 30–60 minutes on a T4. - Build the demo. Streamlit app: paste a lead, hit Generate, see all three styles side by side. Cache the model to avoid reloading per request.
- Evaluate. On the held-out 200 leads, generate all three styles. Score the factual style with ROUGE-L. Run a blind side-by-side: tuned vs hosted-LLM zero-shot vs reference. Rate 1–5 on style match and quality.
- Report. ROUGE numbers + human-preference rates + a small gallery of example outputs.
What to deliver
- The fine-tuned model checkpoint (HuggingFace repo or saved artifacts).
- A Streamlit demo that produces the three styles from one input.
- An eval report on the 200 held-out leads (ROUGE + blind side-by-side ratings).
- A short README documenting base model, dataset composition, hyperparameters, and findings.