Menu
Home Program Lecturers Important Dates Venue Sponsors Past Editions Speakers Alumni Versions GAI2026 Contact
Quantising and Distilling a Small Language Model for Laptop Deployment — Project #11 | Summer School on Generative AI
Project 11

Quantising and Distilling a Small Language Model for Laptop Deployment

CPU / GPU
Note. This page lays out one full version of the project — the goal, a sample input/output, suggested tools, and a step-by-step plan. Treat it as a reference, not a script. Your team can pick a different angle, swap libraries, narrow the scope, or take the project somewhere we did not anticipate. As long as the final deliverable makes sense for the goal, you are on track.
Task. Take a small language model (1–3B parameters) and produce a version that runs comfortably on a laptop without a GPU. Compare two paths: post-training quantisation (8-bit, 4-bit, GGUF) and knowledge distillation into an even smaller student. For each variant, measure quality on a held-out task, tokens-per-second on CPU, and on-disk size. Identify the sweet spot between speed and quality. You could use llama.cpp or bitsandbytes for quantisation; transformers + peft + trl for the distillation; any small base model (SmolLM2, Qwen2.5-0.5B / 1.5B, Phi-3-mini, TinyLlama-1.1B); llama-bench or your own timer for tokens-per-second. Going further (optional). Wrap the quantised model in a small offline chat UI (Streamlit or a terminal app) that runs entirely on the laptop, or compare against a 4-bit version served via Ollama for a realistic deployment.
Resources: Colab T4 or 8GB GPU for the distillation run. The final quantised model must run on a laptop CPU.

What you'll build

Take a small Large Language Model (LLM) and shrink it three ways: 8-bit quantisation, 4-bit quantisation, and a tiny student model trained by knowledge distillation from the original. Compare all four checkpoints (full precision base + three compressed variants) on the same eval set and report quality, model size, inference latency, and memory use. The point is to feel the trade-off curve between size, speed, and quality.

What goes in, what comes out

Input

A small base LLM (around 1B parameters) and a short, fixed eval prompt set (50–100 prompts) plus a held-out classification or generation task to score quality on.

Output

Four checkpoints (full precision, 8-bit, 4-bit, distilled student) plus a comparison table with size, latency, memory, and quality.

A few rows from the eval set
[
  {"id": 1, "prompt": "Classify this customer message: 'My package arrived broken!'\nLabel (complaint/question/praise):", "gold": "complaint"},
  {"id": 2, "prompt": "Classify: 'Do you ship to Germany?'\nLabel:",                                       "gold": "question"},
  {"id": 3, "prompt": "Classify: 'Loved the new packaging, looks great!'\nLabel:",                       "gold": "praise"}
]
Trade-off curve from the four checkpoints
Base model: Qwen2.5-1.5B (full precision FP16)

Checkpoint               Size    Peak VRAM   Latency (ms/token)   Quality
----------------------   -----   ---------   ------------------   -------
Full precision (FP16)    3.0 GB    4.2 GB           14              0.88
8-bit quantised          1.6 GB    2.4 GB           17              0.87
4-bit quantised          0.9 GB    1.6 GB           19              0.84
Distilled student        0.3 GB    0.6 GB            5              0.79
  (~150M params, trained on teacher's soft labels for 1 epoch)

Takeaways:
  - 8-bit is free: half the memory, same quality.
  - 4-bit costs ~4 quality points for another halving.
  - Distilled student is 10x smaller and 3x faster but loses 9 points.

Datasets

For the downstream task: any small benchmark ↗

Any classification task with a clear label is fine — SST-5, AG News, banking77. Pick something where you can score quality with one number.

How to get it: from datasets import load_dataset; ds = load_dataset("SetFit/sst5") (or similar).

License: Free for research use.

For distillation: ~50k unlabelled prompts

Distillation needs lots of (input, teacher logits) pairs. Generate them by running the teacher on prompts from any text dataset (Wikipedia, C4, OpenWebText) — the labels do not matter, only the teacher's logits do.

How to get it: Use a subset of a large unlabelled text dataset (e.g., 50k passages from C4) and run the teacher to produce soft labels.

Tools you'll need

These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.

Python: Python 3.10 or newer. Compute: A single 16 GB GPU is comfortable. Distillation is the expensive step — budget a few hours for the student training.

Base + quantisation
  • transformers — Loads the base model.
  • bitsandbytes — Drop-in 8-bit and 4-bit loading via BitsAndBytesConfig.
  • accelerate — Handles device placement.
Distillation
  • torch — You will write the distillation loss (KL between teacher and student logits) yourself.
  • datasets — Streams the unlabelled corpus through the teacher.
Measurement
  • evaluate — Quality metrics.
  • torch.cuda — Built-in. torch.cuda.max_memory_allocated() for VRAM measurement; torch.cuda.Event for latency.

How to approach it

One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.

  1. Pick a base model. A small instruct LLM around 1B–1.5B parameters: Qwen/Qwen2.5-1.5B-Instruct or meta-llama/Llama-3.2-1B-Instruct.
  2. Pick a downstream task. Choose something measurable in one number (a classification task with accuracy, or generation with an LLM-judge score).
  3. Establish the baseline. Run the base model in FP16. Record quality, model size on disk, peak VRAM, latency per generated token (median over 30 runs).
  4. 8-bit. Reload the base with load_in_8bit=True. Re-run the eval. Record the same four numbers.
  5. 4-bit. Same but with load_in_4bit=True. Re-run. Record.
  6. Distillation. Pick a small student architecture (e.g., a 6-layer transformer, ~150M params). Run the teacher on 50k unlabelled passages and save the teacher's top-k log-probabilities. Train the student to match the teacher's distribution (KL divergence) plus a small cross-entropy on ground-truth tokens.
  7. Evaluate the student. Run the same eval. Record.
  8. Compare. Put it all in one table. Pick a "best for which use case" recommendation.

What to deliver

  • A notebook that produces the four checkpoints and the comparison numbers.
  • A single table comparing size, peak VRAM, latency, and quality.
  • A short discussion: in which scenario would the team deploy each checkpoint?

References