transformers with a multilingual base (facebook/nllb-200-distilled-600M, facebook/mbart-large-50, google/madlad400-3b-mt), peft for LoRA if memory is tight, sacrebleu or chrF for metrics, and Common Voice / FLORES-200 if the language is partially covered. Going further (optional). Build a small Streamlit translator the wider community can use, or share the curated corpus on HuggingFace so others can build on top of it.What you'll build
Build a small Machine Translation (MT) system for a low-resource language. The team picks one language they have access to (an Italian regional language like Sardinian or Neapolitan, a North-African dialect like Tunisian Arabic, a small Berber language like Tamazight, or any minority language they know personally). The team curates a tiny parallel corpus, fine-tunes a massively multilingual MT model (NLLB-200), and compares against the same model used zero-shot. Reports BLEU/chrF plus an honest writeup of what is hard about low-resource MT.
What goes in, what comes out
Input
English (or another high-resource language) source sentences. Fine-tuning data: parallel sentence pairs (a few thousand).
Output
A translation system into the low-resource target language. A results table comparing zero-shot NLLB-200, fine-tuned NLLB-200, and an LLM zero-shot baseline.
en: Good morning. How are you today?
sc: Bonu mangianu. Comente ses oe?
en: The library opens at nine in the morning and closes at six in the evening.
sc: Sa biblioteca apertat a sas noe de manzanu e tancat a sas ses de sero.
en: My grandmother used to tell me stories about the old village.
sc: Mia ajaja mi contaiat istorias de sa idda antiga.
Setup:
Target language: Sardinian (sc) — chosen by the team
Parallel corpus: 1,800 train pairs + 200 test pairs
Sources: open community translations + 5 short stories + Wikipedia
parallel extraction
Base model: facebook/nllb-200-distilled-600M
Approach BLEU chrF Human acceptability (1-5)
---------------------------- ----- ----- -------------------------
NLLB zero-shot 6.8 27.4 2.4
NLLB fine-tuned (1.8k pairs) 18.1 44.7 3.5
LLM zero-shot 9.4 32.1 2.9
Notes:
- Fine-tune lifts BLEU by ~11 points on top of zero-shot.
- Orthographic variation across dialects is the largest residual
source of error.
- LLM produces fluent-sounding output that is often wrong
grammatically. Humans rate it higher than BLEU suggests.
Datasets
Build a small parallel corpus
Pick 2–3 open sources: Wikipedia articles in both languages (extract aligned sentences), open community translation projects, public-domain literature, religious texts (if openly licensed), or song lyrics with translations. Aim for 1,000–3,000 sentence pairs.
How to get it: For Wikipedia: use the Wikipedia API to find articles in both languages, extract paragraph-aligned sentences, hand-verify the alignment.
Optional: existing low-resource benchmark ↗
FLORES-200 has dev / devtest splits for 200+ languages including many low-resource ones. Use it as the test set if your target language is covered.
How to get it: from datasets import load_dataset; ds = load_dataset("facebook/flores") and filter to your language pair.
Tools you'll need
These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.
Python: Python 3.10 or newer. Compute: A 16 GB GPU is comfortable for fine-tuning NLLB-200 distilled 600M. CPU works for the zero-shot baseline only.
transformers— Loads NLLB-200 and other multilingual MT models.sentencepiece— NLLB tokeniser dependency.accelerate— Device placement and mixed precision.
datasets— Streams parallel pairs into training.requests / wikipedia-api— For Wikipedia sentence extraction.
sacrebleu— BLEU and chrF — chrF is more reliable than BLEU on morphologically rich low-resource languages.evaluate— HuggingFace wrapper.
openai / anthropic / groq— For the LLM zero-shot baseline.
How to approach it
One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.
- Pick the target language. Pick one the team has at least one member who can read it (verification is impossible otherwise).
- Build the corpus. Spend a full day on this. Extract from Wikipedia, transcribe from open community sources, hand-translate 100 sentences from English. Aim for ~2,000 parallel pairs.
- Clean and split. Remove duplicates, fix obvious alignment errors, normalise orthography if needed. Split into 80% train / 10% dev / 10% test.
- Baseline 1 — NLLB zero-shot. Load
facebook/nllb-200-distilled-600M. Translate the test set using the language code for your target (if NLLB knows it). Score. - Baseline 2 — NLLB fine-tune. Fine-tune the same model on the train set for 3–5 epochs.
- Baseline 3 — LLM zero-shot. Prompt a hosted LLM ("Translate this English sentence to [target language]. Reply with only the translation.") on the test set.
- Score. sacrebleu BLEU and chrF on all three. Plus a human rating (1–5) from the team member who reads the language.
- Discuss honestly. Where does each system fail? Dialect variation, named entities, rare grammar? Write up findings.
What to deliver
- The cleaned parallel corpus (train + dev + test JSON files), with sources documented.
- A reproducible notebook for the fine-tune and the three baselines.
- A results table with BLEU, chrF, and human ratings.
- A short writeup: what was hard, what surprised the team, what is still wrong.