transformers for fine-tuning (bert-base-cased, RoBERTa, or DeBERTa), seqeval for the metric, and any LLM (with a structured-output prompt) for the baseline. Going further (optional). Build a small UI where someone can paste a sentence and see highlighted entities, or wrap the model behind an agent that flags entities while the user is typing a longer text.What you'll build
Build a named-entity recogniser for English news text. The system reads a sentence and tags every word that is a person, organisation, location, or miscellaneous entity. Output is a labelled CoNLL-format file plus per-entity-type F1 scores so the team can see where the model is strong and where it is weak.
What goes in, what comes out
Input
English news sentences from the CoNLL-2003 dataset (around 14,000 sentences in train, 3,250 in dev, 3,453 in test).
Output
A model that, given a sentence, returns a list of entity spans with their types. Plus a final report with overall F1 and per-type F1 (PER, ORG, LOC, MISC).
Word POS Tag
--------------------------
Germany NNP B-LOC
's POS O
representative VBZ O
to TO O
the DT O
European JJ B-ORG
Union NNP I-ORG
veterinary JJ O
committee NN O
Werner NNP B-PER
Zwingmann NNP I-PER
said VBD O
. . O
entity precision recall F1 support
------ --------- ------ ---- -------
PER 0.96 0.96 0.96 1617
LOC 0.93 0.92 0.93 1668
ORG 0.88 0.87 0.88 1661
MISC 0.79 0.80 0.79 702
Overall (micro) 0.92 5648
Datasets
CoNLL-2003 (English) ↗
Four entity types (PER, ORG, LOC, MISC) over Reuters newswire text. The standard benchmark for English NER.
How to get it: One line of code: from datasets import load_dataset; ds = load_dataset("eriktks/conll2003").
Tools you'll need
These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.
Python: Python 3.10 or newer. Compute: Colab T4 or any 8 GB GPU is comfortable for fine-tuning BERT-base. A laptop CPU is enough for the zero-shot LLM baseline only.
datasets— Loads CoNLL-2003 with the gold BIO tags already parsed.transformers— HuggingFace library for loading and fine-tuning the encoder.accelerate— Handles device placement so you do not write boilerplate.
seqeval— The standard NER evaluator. Computes span-level precision/recall/F1, not token-level (which would be misleading).evaluate— HuggingFace wrapper around seqeval; integrates with the Trainer API.
openai / anthropic / groq— A hosted LLM to compare against. Just prompt it to return JSON spans.
How to approach it
One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.
- Load. Pull CoNLL-2003 from HuggingFace. Inspect the BIO tag scheme so the team understands what B- and I- mean.
- Tokenise. Run the dataset through the encoder's tokeniser. Be careful: a single word can split into multiple subword tokens, and you must align the BIO labels with the subwords (label the first subword, mark the rest with -100 so they are ignored in loss).
- Fine-tune. Train
bert-base-cased(orroberta-base) for token classification using the HuggingFaceTrainer. 3 epochs is usually enough. - Predict. Run the trained model on the dev set, decode back to BIO tags, and merge B- / I- sequences into spans.
- Score. Use
seqevalfor span-level F1 overall and per entity type. - Compare (optional). Prompt a hosted LLM with the same sentences and ask for JSON spans. Score with the same evaluator.
What to deliver
- A reproducible notebook that loads CoNLL-2003, fine-tunes the encoder, and produces a per-entity-type F1 table.
- A short error analysis: look at 20 wrong predictions and group them by failure mode (PER↔ORG confusion is the classic one).
- A short README explaining how to run the training and what the team found.