transformers + trl.SFTTrainer, a small open base (SmolLM2-360M, Qwen2.5-0.5B), the errant library for M2 scoring, and language-tool-python for the rule-based baseline. Going further (optional). Build a small browser-extension or Streamlit demo, evaluate per error type (missing article, wrong tense, agreement) instead of just overall score, or specialise on errors a specific L1 (Italian, French, Arabic) makes when writing English.What you'll build
Fine-tune a small open-source language model (SmolLM2-360M or Qwen2.5-0.5B) to correct grammatical errors in short English sentences written by learners. Train on a learner-English corpus (JFLEG or BEA-2019), evaluate properly with M2 score (the official Grammatical Error Correction metric) and GLEU, and compare against an open rule-based tool (LanguageTool) and a hosted Large Language Model (LLM) used zero-shot. The clear before/after demo plus the existence of well-defined metrics makes this an ideal first NLP project.
What goes in, what comes out
Input
A learner-written English sentence (possibly with grammatical errors).
Output
The corrected version of the sentence. Plus an M2 / GLEU score on the held-out test set, and a confusion breakdown by error type.
source: I am very happy because I do not has anything to do.
refs: [
"I am very happy because I do not have anything to do.",
"I am very happy because I have nothing to do.",
"I am very happy because I do not have anything to do.",
"I'm very happy because I have nothing to do."
]
source: She go to school every day with her friend, but yesterday she don't go.
refs: [
"She goes to school every day with her friend, but yesterday she did not go.",
"She goes to school every day with her friend, but yesterday she didn't go.",
"She goes to school with her friend every day, but yesterday she did not go.",
"She goes to school every day with her friend, but yesterday she didn't."
]
Approach GLEU M2 F0.5 Latency (ms/sent)
---------------------------- ----- -------- -----------------
No-change (copy input) 40.5 0.00 ~0
LanguageTool (rule-based) 49.2 34.1 ~30
SmolLM2-360M SFT 55.7 49.4 ~80
Qwen2.5-0.5B SFT 57.1 51.0 ~95
LLM zero-shot 58.4 53.2 ~600
Per error type (SmolLM2-360M SFT, F0.5):
Missing article ("a/the") 0.62 ← strong
Wrong verb tense 0.51
Subject-verb agreement 0.55
Wrong preposition 0.32 ← still weak
Spelling errors 0.41
Word order 0.28 ← hardest
Datasets
JFLEG ↗
A small (1,511 sentences) but high-quality grammatical error correction dataset with 4 references per source. Standard test set for GEC.
How to get it: from datasets import load_dataset; ds = load_dataset("jfleg").
BEA-2019 shared task ↗
Larger and more recent GEC benchmark. Includes the W&I + LOCNESS corpus with native and learner English at three CEFR levels. Use as additional training data.
How to get it: Download from the shared-task page after accepting terms.
Lang-8 (optional, for pretraining) ↗
Much larger noisy corpus of language-learner posts with corrections from native speakers. Use to bootstrap before fine-tuning on JFLEG/BEA.
How to get it: Request access via the NAIST corpora page.
Tools you'll need
These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.
Python: Python 3.10 or newer. Compute: Colab T4 / Kaggle P100 (16 GB GPU) is comfortable for SmolLM2-360M / Qwen2.5-0.5B SFT. Inference runs on a laptop CPU in real time.
transformers— LoadsHuggingFaceTB/SmolLM2-360MorQwen/Qwen2.5-0.5Bwith its tokeniser.trl— ProvidesSFTTrainerfor the supervised fine-tuning loop.accelerate— Mixed-precision and device placement.
language-tool-python— Wraps the LanguageTool rule-based grammar checker. The traditional baseline.openai / anthropic / groq— For the zero-shot hosted-LLM baseline.
errant— The official ERRANT scorer for GEC. Produces M2 scores and a per-error-type breakdown.sacrebleu— GLEU is implemented inside sacrebleu (or via the original gleu.py script).evaluate— Wrapper for ROUGE-style metrics.
How to approach it
One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.
- Pick the base model. SmolLM2-360M as the default. Qwen2.5-0.5B is a strong alternative.
- Load training data. BEA-2019 + W&I + LOCNESS for training (around 40k pairs); JFLEG for dev / test. Optionally pretrain on Lang-8 first if compute allows.
- Format with the chat template. Each row becomes user "Correct the grammar in this sentence:\n[source]" → assistant "[target]". Use
tokenizer.apply_chat_template. - Run SFT.
trl.SFTTrainer, 2–3 epochs, learning rate 2e-5, bf16 if supported. - Baseline 1 — Copy. Score the no-change baseline (output = input). This is the absolute floor.
- Baseline 2 — LanguageTool. Run LanguageTool on the test set, take its top suggestion, score.
- Baseline 3 — Zero-shot LLM. Prompt a hosted LLM with the same instruction. Be deterministic.
- Score with ERRANT. Run
errant_parallelto align corrections, then compute M2 F0.5 against gold references. Also compute GLEU. - Break down per error type. ERRANT tags edits as M (missing), R (replace), U (unnecessary) per type. Compute per-type F0.5 to see what the tuned model is good and bad at.
What to deliver
- The fine-tuned model checkpoint (HuggingFace repo or local artifacts).
- A reproducible training + inference notebook.
- A results table with GLEU, M2 F0.5, and latency for all four approaches.
- A per-error-type breakdown for the tuned model, with at least 5 example corrections (good and bad) per type.
- A short README documenting base model, training data, hyperparameters, and the team's observations.