sklearn for the classical baseline (TF-IDF + Linear SVM or Logistic Regression), PyTorch for the neural model (an MLP on bag-of-word-embeddings, or a small BiLSTM), and HuggingFace transformers with distilbert-base-uncased. Going further (optional). Build a small UI where someone types a question and sees the predicted category from all three, or extend to the fine-grained TREC labels (50 classes) to see how each stage handles the harder version.What you'll build
Build three classifiers that take a question and predict what kind of answer it expects (a person, a location, a number, a description, an entity, or an abbreviation). The dataset is small (5,500 train / 500 test), so this is the perfect tiny task to show that on short, well-defined text, a TF-IDF + Logistic Regression baseline is already very strong, and that the deltas from neural and transformer models are real but small.
What goes in, what comes out
Input
TREC question dataset. 5,452 train + 500 test, English questions labelled with 6 coarse types: DESC (description), ENTY (entity), ABBR (abbreviation), HUM (person), LOC (location), NUM (number).
Output
Three classifiers plus a results table with overall accuracy and per-class precision / recall.
label question
----- --------
DESC What is the origin of the word 'rugby'?
ENTY What primary colors do you mix to make orange?
ABBR What does NASA stand for?
HUM Who painted the ceiling of the Sistine Chapel?
LOC Where is the Great Barrier Reef?
NUM How many time zones are there in Russia?
Model Accuracy Macro F1
---------------------------- -------- --------
TF-IDF + Logistic Reg. 0.88 0.79
GRU on word embeddings 0.91 0.84
DistilBERT fine-tune 0.96 0.92
Per-class F1 (DistilBERT):
ABBR 0.78 (small support — only 9 in test)
DESC 0.96
ENTY 0.91
HUM 0.98
LOC 0.96
NUM 0.96
Datasets
TREC question classification ↗
Hand-labelled English questions in 6 coarse categories (and 50 fine-grained ones). Tiny but classic.
How to get it: from datasets import load_dataset; ds = load_dataset("CogComp/trec"). Use coarse_label for the 6-way task.
Tools you'll need
These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.
Python: Python 3.10 or newer. Compute: A laptop CPU is enough for all three stages on this small dataset.
scikit-learn— TF-IDF + Logistic Regression. Add a tiny bit of character-n-gram features and you are nearly at 90%.
torch— PyTorch. A GRU with 64 hidden units is plenty. Do not use a giant model on 5k training examples.
transformers— Fine-tune DistilBERT for 3–4 epochs.accelerate— Easy training loop.
evaluate— Accuracy and per-class F1.scikit-learn— Classification report.
How to approach it
One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.
- Load. Pull TREC. Use the 6-way coarse labels. Print 20 examples per class.
- Stage 1 — Classical. TF-IDF (word + character n-grams) + logistic regression. Tune C with cross-validation.
- Stage 2 — Neural. A small GRU (or 1D CNN) in PyTorch, trained for ~10 epochs with early stopping.
- Stage 3 — Transformer. Fine-tune
distilbert-base-uncasedfor 3–4 epochs. The training set is small enough to overfit, so watch the dev curve. - Score. Test-set accuracy, macro F1, and per-class F1 for all three.
- Discuss. Look at the ABBR class (very few examples). Discuss what to do about minority classes when the dataset is tiny.
What to deliver
- Three notebooks producing the comparison numbers.
- A results table with overall accuracy, macro F1, and per-class F1.
- A short paragraph on the ABBR class: why is it hard and what could the team do to make it better with only 86 training examples?