llama.cpp or bitsandbytes for quantisation; transformers + peft + trl for the distillation; any small base model (SmolLM2, Qwen2.5-0.5B / 1.5B, Phi-3-mini, TinyLlama-1.1B); llama-bench or your own timer for tokens-per-second. Going further (optional). Wrap the quantised model in a small offline chat UI (Streamlit or a terminal app) that runs entirely on the laptop, or compare against a 4-bit version served via Ollama for a realistic deployment.What you'll build
Take a small Large Language Model (LLM) and shrink it three ways: 8-bit quantisation, 4-bit quantisation, and a tiny student model trained by knowledge distillation from the original. Compare all four checkpoints (full precision base + three compressed variants) on the same eval set and report quality, model size, inference latency, and memory use. The point is to feel the trade-off curve between size, speed, and quality.
What goes in, what comes out
Input
A small base LLM (around 1B parameters) and a short, fixed eval prompt set (50–100 prompts) plus a held-out classification or generation task to score quality on.
Output
Four checkpoints (full precision, 8-bit, 4-bit, distilled student) plus a comparison table with size, latency, memory, and quality.
[
{"id": 1, "prompt": "Classify this customer message: 'My package arrived broken!'\nLabel (complaint/question/praise):", "gold": "complaint"},
{"id": 2, "prompt": "Classify: 'Do you ship to Germany?'\nLabel:", "gold": "question"},
{"id": 3, "prompt": "Classify: 'Loved the new packaging, looks great!'\nLabel:", "gold": "praise"}
]
Base model: Qwen2.5-1.5B (full precision FP16)
Checkpoint Size Peak VRAM Latency (ms/token) Quality
---------------------- ----- --------- ------------------ -------
Full precision (FP16) 3.0 GB 4.2 GB 14 0.88
8-bit quantised 1.6 GB 2.4 GB 17 0.87
4-bit quantised 0.9 GB 1.6 GB 19 0.84
Distilled student 0.3 GB 0.6 GB 5 0.79
(~150M params, trained on teacher's soft labels for 1 epoch)
Takeaways:
- 8-bit is free: half the memory, same quality.
- 4-bit costs ~4 quality points for another halving.
- Distilled student is 10x smaller and 3x faster but loses 9 points.
Datasets
For the downstream task: any small benchmark ↗
Any classification task with a clear label is fine — SST-5, AG News, banking77. Pick something where you can score quality with one number.
How to get it: from datasets import load_dataset; ds = load_dataset("SetFit/sst5") (or similar).
For distillation: ~50k unlabelled prompts
Distillation needs lots of (input, teacher logits) pairs. Generate them by running the teacher on prompts from any text dataset (Wikipedia, C4, OpenWebText) — the labels do not matter, only the teacher's logits do.
How to get it: Use a subset of a large unlabelled text dataset (e.g., 50k passages from C4) and run the teacher to produce soft labels.
Tools you'll need
These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.
Python: Python 3.10 or newer. Compute: A single 16 GB GPU is comfortable. Distillation is the expensive step — budget a few hours for the student training.
transformers— Loads the base model.bitsandbytes— Drop-in 8-bit and 4-bit loading viaBitsAndBytesConfig.accelerate— Handles device placement.
torch— You will write the distillation loss (KL between teacher and student logits) yourself.datasets— Streams the unlabelled corpus through the teacher.
evaluate— Quality metrics.torch.cuda— Built-in.torch.cuda.max_memory_allocated()for VRAM measurement;torch.cuda.Eventfor latency.
How to approach it
One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.
- Pick a base model. A small instruct LLM around 1B–1.5B parameters:
Qwen/Qwen2.5-1.5B-Instructormeta-llama/Llama-3.2-1B-Instruct. - Pick a downstream task. Choose something measurable in one number (a classification task with accuracy, or generation with an LLM-judge score).
- Establish the baseline. Run the base model in FP16. Record quality, model size on disk, peak VRAM, latency per generated token (median over 30 runs).
- 8-bit. Reload the base with
load_in_8bit=True. Re-run the eval. Record the same four numbers. - 4-bit. Same but with
load_in_4bit=True. Re-run. Record. - Distillation. Pick a small student architecture (e.g., a 6-layer transformer, ~150M params). Run the teacher on 50k unlabelled passages and save the teacher's top-k log-probabilities. Train the student to match the teacher's distribution (KL divergence) plus a small cross-entropy on ground-truth tokens.
- Evaluate the student. Run the same eval. Record.
- Compare. Put it all in one table. Pick a "best for which use case" recommendation.
What to deliver
- A notebook that produces the four checkpoints and the comparison numbers.
- A single table comparing size, peak VRAM, latency, and quality.
- A short discussion: in which scenario would the team deploy each checkpoint?