Menu
Home Program Lecturers Important Dates Venue Sponsors Past Editions Speakers Alumni Versions GAI2026 Contact
Question Classification on TREC: SVM, MLP, and BERT — Project #9 | Summer School on Generative AI
Project 9

Question Classification on TREC: SVM, MLP, and BERT

CPU / GPU
Note. This page lays out one full version of the project — the goal, a sample input/output, suggested tools, and a step-by-step plan. Treat it as a reference, not a script. Your team can pick a different angle, swap libraries, narrow the scope, or take the project somewhere we did not anticipate. As long as the final deliverable makes sense for the goal, you are on track.
Task. Use the TREC question dataset (~5,500 questions, 6 coarse categories: ABBR, ENTY, DESC, HUM, LOC, NUM). Build three classifiers — a classical baseline, a small neural model, and a fine-tuned transformer — and report accuracy plus the per-class confusion matrix. The dataset is small enough that all three stages fit in a single session. You could use sklearn for the classical baseline (TF-IDF + Linear SVM or Logistic Regression), PyTorch for the neural model (an MLP on bag-of-word-embeddings, or a small BiLSTM), and HuggingFace transformers with distilbert-base-uncased. Going further (optional). Build a small UI where someone types a question and sees the predicted category from all three, or extend to the fine-grained TREC labels (50 classes) to see how each stage handles the harder version.
Resources: CPU works for the sklearn baseline and the MLP; Colab T4 for the BERT fine-tune.

What you'll build

Build three classifiers that take a question and predict what kind of answer it expects (a person, a location, a number, a description, an entity, or an abbreviation). The dataset is small (5,500 train / 500 test), so this is the perfect tiny task to show that on short, well-defined text, a TF-IDF + Logistic Regression baseline is already very strong, and that the deltas from neural and transformer models are real but small.

What goes in, what comes out

Input

TREC question dataset. 5,452 train + 500 test, English questions labelled with 6 coarse types: DESC (description), ENTY (entity), ABBR (abbreviation), HUM (person), LOC (location), NUM (number).

Output

Three classifiers plus a results table with overall accuracy and per-class precision / recall.

A few rows from the train set
label  question
-----  --------
 DESC   What is the origin of the word 'rugby'?
 ENTY   What primary colors do you mix to make orange?
 ABBR   What does NASA stand for?
  HUM   Who painted the ceiling of the Sistine Chapel?
  LOC   Where is the Great Barrier Reef?
  NUM   How many time zones are there in Russia?
Test-set results from the three approaches
Model                          Accuracy   Macro F1
----------------------------   --------   --------
TF-IDF + Logistic Reg.           0.88       0.79
GRU on word embeddings           0.91       0.84
DistilBERT fine-tune             0.96       0.92

Per-class F1 (DistilBERT):
  ABBR   0.78  (small support — only 9 in test)
  DESC   0.96
  ENTY   0.91
  HUM    0.98
  LOC    0.96
  NUM    0.96

Datasets

TREC question classification ↗

Hand-labelled English questions in 6 coarse categories (and 50 fine-grained ones). Tiny but classic.

How to get it: from datasets import load_dataset; ds = load_dataset("CogComp/trec"). Use coarse_label for the 6-way task.

License: Free for research use.

Tools you'll need

These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.

Python: Python 3.10 or newer. Compute: A laptop CPU is enough for all three stages on this small dataset.

Stage 1 — classical
  • scikit-learn — TF-IDF + Logistic Regression. Add a tiny bit of character-n-gram features and you are nearly at 90%.
Stage 2 — neural
  • torch — PyTorch. A GRU with 64 hidden units is plenty. Do not use a giant model on 5k training examples.
Stage 3 — transformer
  • transformers — Fine-tune DistilBERT for 3–4 epochs.
  • accelerate — Easy training loop.
Evaluation
  • evaluate — Accuracy and per-class F1.
  • scikit-learn — Classification report.

How to approach it

One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.

  1. Load. Pull TREC. Use the 6-way coarse labels. Print 20 examples per class.
  2. Stage 1 — Classical. TF-IDF (word + character n-grams) + logistic regression. Tune C with cross-validation.
  3. Stage 2 — Neural. A small GRU (or 1D CNN) in PyTorch, trained for ~10 epochs with early stopping.
  4. Stage 3 — Transformer. Fine-tune distilbert-base-uncased for 3–4 epochs. The training set is small enough to overfit, so watch the dev curve.
  5. Score. Test-set accuracy, macro F1, and per-class F1 for all three.
  6. Discuss. Look at the ABBR class (very few examples). Discuss what to do about minority classes when the dataset is tiny.

What to deliver

  • Three notebooks producing the comparison numbers.
  • A results table with overall accuracy, macro F1, and per-class F1.
  • A short paragraph on the ABBR class: why is it hard and what could the team do to make it better with only 86 training examples?

References