Menu
Home Program Lecturers Important Dates Venue Sponsors Past Editions Speakers Alumni Versions GAI2026 Contact
Grammar Corrector with a Small Language Model — Project #22 | Summer School on Generative AI
Project 22

Grammar Corrector with a Small Language Model

CPU / GPU
Note. This page lays out one full version of the project — the goal, a sample input/output, suggested tools, and a step-by-step plan. Treat it as a reference, not a script. Your team can pick a different angle, swap libraries, narrow the scope, or take the project somewhere we did not anticipate. As long as the final deliverable makes sense for the goal, you are on track.
Task. Fine-tune a small language model to correct grammatical errors in short English sentences written by learners. Take an existing learner-English corpus (Lang-8, JFLEG, BEA-2019), tune SmolLM2-360M / Qwen2.5-0.5B on the (ungrammatical, corrected) pairs, and evaluate properly with M2 (the official grammatical-error-correction metric) or GLEU. Compare against an open rule-based tool like LanguageTool and a zero-shot hosted LLM baseline. The clear before/after and the existence of a real metric make this an ideal first NLP project. You could use HuggingFace transformers + trl.SFTTrainer, a small open base (SmolLM2-360M, Qwen2.5-0.5B), the errant library for M2 scoring, and language-tool-python for the rule-based baseline. Going further (optional). Build a small browser-extension or Streamlit demo, evaluate per error type (missing article, wrong tense, agreement) instead of just overall score, or specialise on errors a specific L1 (Italian, French, Arabic) makes when writing English.
Resources: Colab T4 or 8GB GPU for fine-tuning; CPU is fine for inference with a sub-1B model.

What you'll build

Fine-tune a small open-source language model (SmolLM2-360M or Qwen2.5-0.5B) to correct grammatical errors in short English sentences written by learners. Train on a learner-English corpus (JFLEG or BEA-2019), evaluate properly with M2 score (the official Grammatical Error Correction metric) and GLEU, and compare against an open rule-based tool (LanguageTool) and a hosted Large Language Model (LLM) used zero-shot. The clear before/after demo plus the existence of well-defined metrics makes this an ideal first NLP project.

What goes in, what comes out

Input

A learner-written English sentence (possibly with grammatical errors).

Output

The corrected version of the sentence. Plus an M2 / GLEU score on the held-out test set, and a confusion breakdown by error type.

A few rows from JFLEG dev
source:  I am very happy because I do not has anything to do.
refs:    [
  "I am very happy because I do not have anything to do.",
  "I am very happy because I have nothing to do.",
  "I am very happy because I do not have anything to do.",
  "I'm very happy because I have nothing to do."
]

source:  She go to school every day with her friend, but yesterday she don't go.
refs:    [
  "She goes to school every day with her friend, but yesterday she did not go.",
  "She goes to school every day with her friend, but yesterday she didn't go.",
  "She goes to school with her friend every day, but yesterday she did not go.",
  "She goes to school every day with her friend, but yesterday she didn't."
]
Comparison on the JFLEG test set
Approach                        GLEU    M2 F0.5    Latency (ms/sent)
----------------------------    -----   --------   -----------------
No-change (copy input)          40.5     0.00            ~0
LanguageTool (rule-based)       49.2     34.1            ~30
SmolLM2-360M SFT                55.7     49.4            ~80
Qwen2.5-0.5B SFT                57.1     51.0            ~95
LLM zero-shot                   58.4     53.2           ~600

Per error type (SmolLM2-360M SFT, F0.5):
  Missing article ("a/the")         0.62  ← strong
  Wrong verb tense                  0.51
  Subject-verb agreement            0.55
  Wrong preposition                 0.32  ← still weak
  Spelling errors                   0.41
  Word order                        0.28  ← hardest

Datasets

JFLEG ↗

A small (1,511 sentences) but high-quality grammatical error correction dataset with 4 references per source. Standard test set for GEC.

How to get it: from datasets import load_dataset; ds = load_dataset("jfleg").

License: Free for research use.

BEA-2019 shared task ↗

Larger and more recent GEC benchmark. Includes the W&I + LOCNESS corpus with native and learner English at three CEFR levels. Use as additional training data.

How to get it: Download from the shared-task page after accepting terms.

License: Free for research use; requires accepting terms.

Lang-8 (optional, for pretraining) ↗

Much larger noisy corpus of language-learner posts with corrections from native speakers. Use to bootstrap before fine-tuning on JFLEG/BEA.

How to get it: Request access via the NAIST corpora page.

License: Free for research use.

Tools you'll need

These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.

Python: Python 3.10 or newer. Compute: Colab T4 / Kaggle P100 (16 GB GPU) is comfortable for SmolLM2-360M / Qwen2.5-0.5B SFT. Inference runs on a laptop CPU in real time.

Base model + training
  • transformers — Loads HuggingFaceTB/SmolLM2-360M or Qwen/Qwen2.5-0.5B with its tokeniser.
  • trl — Provides SFTTrainer for the supervised fine-tuning loop.
  • accelerate — Mixed-precision and device placement.
Baselines
  • language-tool-python — Wraps the LanguageTool rule-based grammar checker. The traditional baseline.
  • openai / anthropic / groq — For the zero-shot hosted-LLM baseline.
Evaluation
  • errant — The official ERRANT scorer for GEC. Produces M2 scores and a per-error-type breakdown.
  • sacrebleu — GLEU is implemented inside sacrebleu (or via the original gleu.py script).
  • evaluate — Wrapper for ROUGE-style metrics.

How to approach it

One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.

  1. Pick the base model. SmolLM2-360M as the default. Qwen2.5-0.5B is a strong alternative.
  2. Load training data. BEA-2019 + W&I + LOCNESS for training (around 40k pairs); JFLEG for dev / test. Optionally pretrain on Lang-8 first if compute allows.
  3. Format with the chat template. Each row becomes user "Correct the grammar in this sentence:\n[source]" → assistant "[target]". Use tokenizer.apply_chat_template.
  4. Run SFT. trl.SFTTrainer, 2–3 epochs, learning rate 2e-5, bf16 if supported.
  5. Baseline 1 — Copy. Score the no-change baseline (output = input). This is the absolute floor.
  6. Baseline 2 — LanguageTool. Run LanguageTool on the test set, take its top suggestion, score.
  7. Baseline 3 — Zero-shot LLM. Prompt a hosted LLM with the same instruction. Be deterministic.
  8. Score with ERRANT. Run errant_parallel to align corrections, then compute M2 F0.5 against gold references. Also compute GLEU.
  9. Break down per error type. ERRANT tags edits as M (missing), R (replace), U (unnecessary) per type. Compute per-type F0.5 to see what the tuned model is good and bad at.

What to deliver

  • The fine-tuned model checkpoint (HuggingFace repo or local artifacts).
  • A reproducible training + inference notebook.
  • A results table with GLEU, M2 F0.5, and latency for all four approaches.
  • A per-error-type breakdown for the tuned model, with at least 5 example corrections (good and bad) per type.
  • A short README documenting base model, training data, hyperparameters, and the team's observations.

References