Menu
Home Program Lecturers Important Dates Venue Sponsors Past Editions Speakers Alumni Versions GAI2026 Contact
Instruction Fine-Tuning SmolLM 135M into a Tiny Assistant — Project #2 | Summer School on Generative AI
Project 2

Instruction Fine-Tuning SmolLM 135M into a Tiny Assistant

GPU
Note. This page lays out one full version of the project — the goal, a sample input/output, suggested tools, and a step-by-step plan. Treat it as a reference, not a script. Your team can pick a different angle, swap libraries, narrow the scope, or take the project somewhere we did not anticipate. As long as the final deliverable makes sense for the goal, you are on track.
Task. Take a base SmolLM 135M model (no instruction tuning) and turn it into a tiny chat assistant by full supervised fine-tuning on an Alpaca-style instruction dataset. At 135M parameters, a full fine-tune is genuinely feasible on a single small GPU — no LoRA needed. The team will compare the base model and the fine-tuned model on the same 50-prompt evaluation set and see, with their own eyes, what "instruction following" actually looks like at this scale. You could use HuggingFaceTB/SmolLM-135M as the base, the trl.SFTTrainer for the training loop, and Dolly-15k or Alpaca as the dataset. Chat-template formatting, dataset filtering, and the eval prompts are your design choices. Going further (optional). Compare full fine-tune against LoRA on the same model, train on a narrow domain (cooking, code comments, customer support) and report the change in tone, or quantise the final model to 4-bit and measure the quality cost.
Resources: Colab T4 / Kaggle P100 / any 8GB GPU is comfortable; CPU works for inference only and will be slow.

What you'll build

Take the base SmolLM 135M model — a 135-million-parameter pretrained language model that does not follow instructions out of the box — and fine-tune it into a tiny chat assistant via Supervised Fine-Tuning (SFT) on an Alpaca-style instruction dataset. At this scale, a full fine-tune is genuinely feasible on a single small GPU (no LoRA needed), so the team gets to feel the complete instruction-tuning loop end-to-end: chat templates, the SFT training loop, generation, and a blind side-by-side evaluation against the un-tuned base.

What goes in, what comes out

Input

A pretrained base language model checkpoint (HuggingFaceTB/SmolLM-135M) plus an instruction dataset of (instruction, optional input, response) triples.

Output

A fine-tuned model checkpoint that follows instructions, plus a results table comparing base and tuned on a 50-prompt eval set with blind human ratings.

A few rows from Dolly-15k
[
  {
    "instruction": "Explain what a recursive function is in one paragraph.",
    "input": "",
    "output": "A recursive function is a function that solves a problem by calling itself on a smaller version of the same problem, with a base case that stops the recursion. Each call works on a smaller input until it reaches the base case, at which point the partial results are combined back into the final answer."
  },
  {
    "instruction": "Translate to French.",
    "input": "The library closes at six in the evening.",
    "output": "La bibliothèque ferme à six heures du soir."
  },
  {
    "instruction": "List three reasons to use version control on a software project.",
    "input": "",
    "output": "1. It tracks every change so you can revert mistakes. 2. It lets multiple people work on the same codebase without overwriting each other. 3. It documents the history of why each change was made."
  }
]
Same prompt, base model vs fine-tuned model
Prompt: "Explain what a list comprehension is in one sentence."

Base SmolLM-135M (no instruction tuning):
  "What is a list comprehension? A list comprehension is a way to
   create a list from another list. List comprehensions are useful
   because they are more concise than for loops..."
  (continues text without following the instruction format)

Fine-tuned SmolLM-135M:
  "A list comprehension is a concise Python syntax for building a
   list by writing the expression and the loop on a single line,
   such as `[x*x for x in range(10)]`."
  (follows the instruction, stops at the right place)

Blind ratings on 50 held-out prompts (1=bad ... 5=excellent):
  Approach                        Avg. quality   Follows instruction
  -----------------------------   ------------   -------------------
  SmolLM-135M base                    2.1              28%
  SmolLM-135M instruction-tuned       3.6              92%

Datasets

Databricks Dolly 15k ↗

15,000 hand-written instruction/response pairs across eight task categories (open Q&A, closed Q&A, brainstorming, summarisation, etc). Openly licensed and good quality.

How to get it: from datasets import load_dataset; ds = load_dataset("databricks/databricks-dolly-15k").

License: Creative Commons Attribution-ShareAlike 3.0 — fully open.

Alternative: Alpaca-cleaned ↗

~52k instruction/response pairs (cleaner version of the original Alpaca set). Larger but noisier than Dolly. Pick one — do not mix on the first run.

How to get it: from datasets import load_dataset; ds = load_dataset("yahma/alpaca-cleaned").

License: Creative Commons Attribution-NonCommercial 4.0.

Tools you'll need

These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.

Python: Python 3.10 or newer. Compute: Colab T4 / Kaggle P100 (16 GB GPU) is comfortable for a full fine-tune of a 135M model. CPU is realistic for inference only.

Base model
  • transformers — Loads HuggingFaceTB/SmolLM-135M and its tokeniser. The tokeniser carries the chat template.
Training
  • trl — Provides SFTTrainer — handles instruction formatting, packing, and the training loop without boilerplate.
  • accelerate — Mixed-precision and device placement.
  • torch — Underlying training backend.
Data
  • datasets — Loads Dolly / Alpaca and applies the chat template via map.
  • sentencepiece — Tokeniser dependency for SmolLM.
Evaluation
  • evaluate — For any automatic metrics on top of human ratings.

How to approach it

One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.

  1. Load the base model. HuggingFaceTB/SmolLM-135M plus its tokeniser. Generate a few responses to confirm it does not follow instructions — it just continues text. This is the "before" snapshot.
  2. Prepare the dataset. Load Dolly-15k. Map each row to a chat-template-formatted string using tokenizer.apply_chat_template with [{"role": "user", "content": instruction}, {"role": "assistant", "content": output}].
  3. Configure SFT. Use trl.SFTTrainer. Suggested defaults: 2-3 epochs, learning rate 2e-5, batch size 8 (with gradient accumulation if needed), bf16 if your GPU supports it, max sequence length 1024.
  4. Train. Should take ~30-60 minutes on a T4 for 15k examples × 2 epochs. Watch the loss curve.
  5. Save the checkpoint. The whole fine-tuned model is ~500 MB on disk (full weights, not adapters).
  6. Build a 50-prompt eval set. The team writes prompts that probe instruction following: definitions, list generation, format following, refusals, translations.
  7. Generate from both models. Same prompts, same decoding settings (greedy or low-temperature sampling), one notebook per model.
  8. Blind side-by-side rating. Shuffle the (base, tuned) pairs so raters do not know which is which. Each team member rates every response 1-5 on quality and on "did it follow the instruction".

What to deliver

  • The fine-tuned SmolLM-135M checkpoint (upload to a HuggingFace repo or save in the project artifacts).
  • A reproducible training notebook + an inference notebook that loads either model and chats interactively.
  • A 50-prompt eval set with the team's blind ratings for base and tuned models, plus the average scores.
  • A short README documenting the dataset, hyperparameters, training time, and what changed between base and tuned.

References