Menu
Home Program Lecturers Important Dates Venue Sponsors Past Editions Speakers Alumni Versions GAI2026 Contact
Machine Translation with Transformers — Project #18 | Summer School on Generative AI
Project 18

Machine Translation with Transformers

GPU
Note. This page lays out one full version of the project — the goal, a sample input/output, suggested tools, and a step-by-step plan. Treat it as a reference, not a script. Your team can pick a different angle, swap libraries, narrow the scope, or take the project somewhere we did not anticipate. As long as the final deliverable makes sense for the goal, you are on track.
Task. Build a machine-translation system between two languages of your choice. Pick a parallel corpus, fine-tune a pretrained translation model, and compare it against a strong general-purpose LLM used zero-shot for the same translations. Report BLEU and chrF on a held-out test set and look at the qualitative differences: who is better at idioms, named entities, rare vocabulary? You could use HuggingFace transformers with a base model like Helsinki-NLP/opus-mt-*, facebook/mbart-large-50, or facebook/nllb-200-distilled-600M, the sacrebleu library for metrics, and any LLM for the zero-shot baseline. Going further (optional). Build a Streamlit UI for live side-by-side translation, or extend to a language pair where parallel data is scarce and study how performance degrades with less training data.
Resources: Colab T4 or 8GB GPU for fine-tuning; CPU works for hosted-model baselines.

What you'll build

Build an English-to-German Machine Translation (MT) system and compare three approaches on the same test set: a small encoder-decoder transformer trained from scratch on the WMT data, a fine-tuned pretrained Helsinki-NLP MarianMT model, and a zero-shot Large Language Model (LLM). Report BLEU and chrF scores plus latency, and run a small human side-by-side evaluation.

What goes in, what comes out

Input

English sentences, output is German translations. Training data is parallel sentence pairs from WMT14 or similar.

Output

Three MT systems plus a results table with BLEU, chrF, latency, and the team's human side-by-side preferences.

A few rows from WMT14 en-de
en: The European Central Bank announced new measures to support the eurozone economy.
de: Die Europäische Zentralbank kündigte neue Maßnahmen zur Unterstützung der Eurozonen-Wirtschaft an.

en: Researchers say the new treatment reduced symptoms in 70% of patients.
de: Forscher sagen, die neue Behandlung habe die Symptome bei 70% der Patienten reduziert.

en: The new high-speed rail line will connect the two cities in under three hours.
de: Die neue Hochgeschwindigkeitsbahnlinie wird die beiden Städte in weniger als drei Stunden verbinden.
Test-set results from the three approaches
Approach                          BLEU    chrF    Latency (ms/sent)   Human win-rate vs reference
-------------------------------   -----   -----   -----------------   ---------------------------
Transformer from scratch (1ep)     19.4   46.2          90                  18%
MarianMT fine-tune (en-de)         33.7   60.1          70                  47%
LLM zero-shot                      31.2   58.4         420                  51%

WMT14 newstest (1000 sentences, beam 5):
  - Fine-tuned MarianMT and LLM are roughly tied on auto metrics.
  - On human side-by-side, the LLM wins on natural phrasing,
    Marian wins on technical / domain-specific terms.

Datasets

WMT14 English-German ↗

Parallel English-German data from news articles, books, and crawled web text. The standard MT benchmark.

How to get it: from datasets import load_dataset; ds = load_dataset("wmt/wmt14", "de-en"). The full dataset is huge — subsample 200k for from-scratch training.

License: Free for research use.

Tools you'll need

These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.

Python: Python 3.10 or newer. Compute: A 16 GB GPU is comfortable. Training a transformer from scratch is the heavy step (~8h on T4 for 200k pairs).

Core
  • transformers — Loads MarianMT, T5, mBART. Also provides EncoderDecoderModel for the from-scratch baseline.
  • datasets — Loads WMT14 with parallel pairs already aligned.
  • accelerate — Handles device placement and mixed-precision training.
  • sentencepiece — Tokeniser dependency for most MT models.
Training
  • torch — PyTorch for custom training loops and the from-scratch baseline.
Evaluation
  • sacrebleu — The standard BLEU and chrF implementation for MT — reproducible across papers.
  • evaluate — HuggingFace wrapper around sacrebleu.
Optional (LLM baseline)
  • openai / anthropic / groq — A hosted LLM. Prompt: "Translate the following English sentence to German. Reply with only the German translation."

How to approach it

One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.

  1. Load. Pull WMT14 en-de. Subsample 200k pairs for training. Use the official newstest14 set (3,003 sentences) as the test set.
  2. Baseline 1 — from scratch. Train a 6-layer encoder-decoder transformer from scratch on the 200k pairs for 1 epoch. Expect modest BLEU. The point is to feel why pretraining matters.
  3. Baseline 2 — MarianMT fine-tune. Load Helsinki-NLP/opus-mt-en-de. Fine-tune for 1 epoch on the 200k pairs. This will be much better.
  4. Baseline 3 — LLM zero-shot. Prompt a hosted LLM to translate each of the 3,003 test sentences. Be deterministic.
  5. Generate. Beam search with num_beams=5, max_length=128, for the two transformer models.
  6. Score automatic metrics. sacrebleu BLEU and chrF on all three. Record per-sentence latency too.
  7. Human side-by-side. Sample 100 sentences. For each, show all three translations + the reference (blind). The team picks the best.

What to deliver

  • Three notebooks (from-scratch, MarianMT fine-tune, LLM baseline).
  • A results table with BLEU, chrF, latency.
  • A blind human side-by-side report on 100 sentences with win-rates.
  • A short discussion: where does each approach fail? When is the LLM's "natural" style actually wrong?

References