Menu
Home Program Lecturers Important Dates Venue Sponsors Past Editions Speakers Alumni Versions GAI2026 Contact
Translation and Tools for a Low-Resource Language — Project #20 | Summer School on Generative AI
Project 20

Translation and Tools for a Low-Resource Language

GPU
Note. This page lays out one full version of the project — the goal, a sample input/output, suggested tools, and a step-by-step plan. Treat it as a reference, not a script. Your team can pick a different angle, swap libraries, narrow the scope, or take the project somewhere we did not anticipate. As long as the final deliverable makes sense for the goal, you are on track.
Task. Pick a language the team speaks where NLP tools are weak — Darija, Sicilian, Berber, Tigrinya, Quechua, a regional dialect. Build a small useful tool: a translation system to and from English, a basic spell checker, a sentence-level classifier, or whatever the language most needs. Collect or curate a small parallel or labelled corpus (~500–2000 examples), fine-tune a multilingual model on it, and evaluate against a strong off-the-shelf baseline. The goal is to show measurable improvement over what existed before for that language. You could use HuggingFace transformers with a multilingual base (facebook/nllb-200-distilled-600M, facebook/mbart-large-50, google/madlad400-3b-mt), peft for LoRA if memory is tight, sacrebleu or chrF for metrics, and Common Voice / FLORES-200 if the language is partially covered. Going further (optional). Build a small Streamlit translator the wider community can use, or share the curated corpus on HuggingFace so others can build on top of it.
Resources: Colab T4 or 8GB GPU for fine-tuning.

What you'll build

Build a small Machine Translation (MT) system for a low-resource language. The team picks one language they have access to (an Italian regional language like Sardinian or Neapolitan, a North-African dialect like Tunisian Arabic, a small Berber language like Tamazight, or any minority language they know personally). The team curates a tiny parallel corpus, fine-tunes a massively multilingual MT model (NLLB-200), and compares against the same model used zero-shot. Reports BLEU/chrF plus an honest writeup of what is hard about low-resource MT.

What goes in, what comes out

Input

English (or another high-resource language) source sentences. Fine-tuning data: parallel sentence pairs (a few thousand).

Output

A translation system into the low-resource target language. A results table comparing zero-shot NLLB-200, fine-tuned NLLB-200, and an LLM zero-shot baseline.

A few rows from a hand-curated English-Sardinian corpus
en: Good morning. How are you today?
sc: Bonu mangianu. Comente ses oe?

en: The library opens at nine in the morning and closes at six in the evening.
sc: Sa biblioteca apertat a sas noe de manzanu e tancat a sas ses de sero.

en: My grandmother used to tell me stories about the old village.
sc: Mia ajaja mi contaiat istorias de sa idda antiga.
Results on a 200-sentence held-out test set
Setup:
  Target language: Sardinian (sc) — chosen by the team
  Parallel corpus: 1,800 train pairs + 200 test pairs
  Sources: open community translations + 5 short stories + Wikipedia
           parallel extraction
  Base model: facebook/nllb-200-distilled-600M

Approach                       BLEU    chrF   Human acceptability (1-5)
----------------------------   -----   -----  -------------------------
NLLB zero-shot                  6.8    27.4         2.4
NLLB fine-tuned (1.8k pairs)   18.1    44.7         3.5
LLM zero-shot                   9.4    32.1         2.9

Notes:
  - Fine-tune lifts BLEU by ~11 points on top of zero-shot.
  - Orthographic variation across dialects is the largest residual
    source of error.
  - LLM produces fluent-sounding output that is often wrong
    grammatically. Humans rate it higher than BLEU suggests.

Datasets

Build a small parallel corpus

Pick 2–3 open sources: Wikipedia articles in both languages (extract aligned sentences), open community translation projects, public-domain literature, religious texts (if openly licensed), or song lyrics with translations. Aim for 1,000–3,000 sentence pairs.

How to get it: For Wikipedia: use the Wikipedia API to find articles in both languages, extract paragraph-aligned sentences, hand-verify the alignment.

Optional: existing low-resource benchmark ↗

FLORES-200 has dev / devtest splits for 200+ languages including many low-resource ones. Use it as the test set if your target language is covered.

How to get it: from datasets import load_dataset; ds = load_dataset("facebook/flores") and filter to your language pair.

License: Creative Commons Attribution-ShareAlike 4.0.

Tools you'll need

These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.

Python: Python 3.10 or newer. Compute: A 16 GB GPU is comfortable for fine-tuning NLLB-200 distilled 600M. CPU works for the zero-shot baseline only.

MT model
  • transformers — Loads NLLB-200 and other multilingual MT models.
  • sentencepiece — NLLB tokeniser dependency.
  • accelerate — Device placement and mixed precision.
Corpus building
  • datasets — Streams parallel pairs into training.
  • requests / wikipedia-api — For Wikipedia sentence extraction.
Evaluation
  • sacrebleu — BLEU and chrF — chrF is more reliable than BLEU on morphologically rich low-resource languages.
  • evaluate — HuggingFace wrapper.
Optional (LLM baseline)
  • openai / anthropic / groq — For the LLM zero-shot baseline.

How to approach it

One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.

  1. Pick the target language. Pick one the team has at least one member who can read it (verification is impossible otherwise).
  2. Build the corpus. Spend a full day on this. Extract from Wikipedia, transcribe from open community sources, hand-translate 100 sentences from English. Aim for ~2,000 parallel pairs.
  3. Clean and split. Remove duplicates, fix obvious alignment errors, normalise orthography if needed. Split into 80% train / 10% dev / 10% test.
  4. Baseline 1 — NLLB zero-shot. Load facebook/nllb-200-distilled-600M. Translate the test set using the language code for your target (if NLLB knows it). Score.
  5. Baseline 2 — NLLB fine-tune. Fine-tune the same model on the train set for 3–5 epochs.
  6. Baseline 3 — LLM zero-shot. Prompt a hosted LLM ("Translate this English sentence to [target language]. Reply with only the translation.") on the test set.
  7. Score. sacrebleu BLEU and chrF on all three. Plus a human rating (1–5) from the team member who reads the language.
  8. Discuss honestly. Where does each system fail? Dialect variation, named entities, rare grammar? Write up findings.

What to deliver

  • The cleaned parallel corpus (train + dev + test JSON files), with sources documented.
  • A reproducible notebook for the fine-tune and the three baselines.
  • A results table with BLEU, chrF, and human ratings.
  • A short writeup: what was hard, what surprised the team, what is still wrong.

References