HuggingFaceTB/SmolLM-135M as the base, the trl.SFTTrainer for the training loop, and Dolly-15k or Alpaca as the dataset. Chat-template formatting, dataset filtering, and the eval prompts are your design choices. Going further (optional). Compare full fine-tune against LoRA on the same model, train on a narrow domain (cooking, code comments, customer support) and report the change in tone, or quantise the final model to 4-bit and measure the quality cost.What you'll build
Take the base SmolLM 135M model — a 135-million-parameter pretrained language model that does not follow instructions out of the box — and fine-tune it into a tiny chat assistant via Supervised Fine-Tuning (SFT) on an Alpaca-style instruction dataset. At this scale, a full fine-tune is genuinely feasible on a single small GPU (no LoRA needed), so the team gets to feel the complete instruction-tuning loop end-to-end: chat templates, the SFT training loop, generation, and a blind side-by-side evaluation against the un-tuned base.
What goes in, what comes out
Input
A pretrained base language model checkpoint (HuggingFaceTB/SmolLM-135M) plus an instruction dataset of (instruction, optional input, response) triples.
Output
A fine-tuned model checkpoint that follows instructions, plus a results table comparing base and tuned on a 50-prompt eval set with blind human ratings.
[
{
"instruction": "Explain what a recursive function is in one paragraph.",
"input": "",
"output": "A recursive function is a function that solves a problem by calling itself on a smaller version of the same problem, with a base case that stops the recursion. Each call works on a smaller input until it reaches the base case, at which point the partial results are combined back into the final answer."
},
{
"instruction": "Translate to French.",
"input": "The library closes at six in the evening.",
"output": "La bibliothèque ferme à six heures du soir."
},
{
"instruction": "List three reasons to use version control on a software project.",
"input": "",
"output": "1. It tracks every change so you can revert mistakes. 2. It lets multiple people work on the same codebase without overwriting each other. 3. It documents the history of why each change was made."
}
]
Prompt: "Explain what a list comprehension is in one sentence."
Base SmolLM-135M (no instruction tuning):
"What is a list comprehension? A list comprehension is a way to
create a list from another list. List comprehensions are useful
because they are more concise than for loops..."
(continues text without following the instruction format)
Fine-tuned SmolLM-135M:
"A list comprehension is a concise Python syntax for building a
list by writing the expression and the loop on a single line,
such as `[x*x for x in range(10)]`."
(follows the instruction, stops at the right place)
Blind ratings on 50 held-out prompts (1=bad ... 5=excellent):
Approach Avg. quality Follows instruction
----------------------------- ------------ -------------------
SmolLM-135M base 2.1 28%
SmolLM-135M instruction-tuned 3.6 92%
Datasets
Databricks Dolly 15k ↗
15,000 hand-written instruction/response pairs across eight task categories (open Q&A, closed Q&A, brainstorming, summarisation, etc). Openly licensed and good quality.
How to get it: from datasets import load_dataset; ds = load_dataset("databricks/databricks-dolly-15k").
Alternative: Alpaca-cleaned ↗
~52k instruction/response pairs (cleaner version of the original Alpaca set). Larger but noisier than Dolly. Pick one — do not mix on the first run.
How to get it: from datasets import load_dataset; ds = load_dataset("yahma/alpaca-cleaned").
Tools you'll need
These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.
Python: Python 3.10 or newer. Compute: Colab T4 / Kaggle P100 (16 GB GPU) is comfortable for a full fine-tune of a 135M model. CPU is realistic for inference only.
transformers— LoadsHuggingFaceTB/SmolLM-135Mand its tokeniser. The tokeniser carries the chat template.
trl— ProvidesSFTTrainer— handles instruction formatting, packing, and the training loop without boilerplate.accelerate— Mixed-precision and device placement.torch— Underlying training backend.
datasets— Loads Dolly / Alpaca and applies the chat template viamap.sentencepiece— Tokeniser dependency for SmolLM.
evaluate— For any automatic metrics on top of human ratings.
How to approach it
One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.
- Load the base model.
HuggingFaceTB/SmolLM-135Mplus its tokeniser. Generate a few responses to confirm it does not follow instructions — it just continues text. This is the "before" snapshot. - Prepare the dataset. Load Dolly-15k. Map each row to a chat-template-formatted string using
tokenizer.apply_chat_templatewith[{"role": "user", "content": instruction}, {"role": "assistant", "content": output}]. - Configure SFT. Use
trl.SFTTrainer. Suggested defaults: 2-3 epochs, learning rate 2e-5, batch size 8 (with gradient accumulation if needed), bf16 if your GPU supports it, max sequence length 1024. - Train. Should take ~30-60 minutes on a T4 for 15k examples × 2 epochs. Watch the loss curve.
- Save the checkpoint. The whole fine-tuned model is ~500 MB on disk (full weights, not adapters).
- Build a 50-prompt eval set. The team writes prompts that probe instruction following: definitions, list generation, format following, refusals, translations.
- Generate from both models. Same prompts, same decoding settings (greedy or low-temperature sampling), one notebook per model.
- Blind side-by-side rating. Shuffle the (base, tuned) pairs so raters do not know which is which. Each team member rates every response 1-5 on quality and on "did it follow the instruction".
What to deliver
- The fine-tuned SmolLM-135M checkpoint (upload to a HuggingFace repo or save in the project artifacts).
- A reproducible training notebook + an inference notebook that loads either model and chats interactively.
- A 50-prompt eval set with the team's blind ratings for base and tuned models, plus the average scores.
- A short README documenting the dataset, hyperparameters, training time, and what changed between base and tuned.