SmolLM2-360M / 1.7B, Qwen2.5-0.5B / 1.5B, TinyLlama-1.1B, or Pythia-1B as the base; peft + transformers for LoRA; bitsandbytes for 8-bit Adam if you want to fit the run on smaller hardware. Going further (optional). Build a small UI to show side-by-side answers from base / few-shot / LoRA, or try two adapter sizes and report the quality-vs-parameter trade-off.What you'll build
Fine-tune a small open-source language model (around 1B–3B parameters) on a narrow domain task using LoRA (Low-Rank Adaptation), a Parameter-Efficient Fine-Tuning (PEFT) technique. The team will compare the LoRA-adapted model against the base model on the same eval set, measure the size of the adapter weights, and explain why LoRA lets you customise a model on a single small GPU.
What goes in, what comes out
Input
A small instruction dataset for the chosen domain (1k–10k examples). For example: Alpaca-style instruction/response pairs in a specific domain (cooking, customer support, legal Q&A).
Output
A LoRA adapter (typically 10–50 MB on disk) plus a results table comparing base and LoRA models on the same eval prompts.
[
{
"instruction": "Suggest a vegetarian substitute for pancetta in pasta carbonara.",
"input": "",
"output": "Smoked tofu or king-oyster mushrooms work well. Dice into 1cm cubes, pan-fry in olive oil with a pinch of smoked paprika until crisp on the edges, then fold into the egg-cheese sauce off-heat as you would with pancetta."
},
{
"instruction": "What can I do with leftover roast chicken?",
"input": "",
"output": "Shred and toss into a chicken-noodle soup, a Caesar salad, or a quick chicken tortilla wrap. The bones make a strong stock — simmer 2 hours with onion, carrot, and bay leaf."
}
]
Setup:
Base model: meta-llama/Llama-3.2-3B-Instruct
LoRA rank: 16, alpha=32, dropout=0.05
Target layers: q_proj, v_proj
Training: 3 epochs on 2,000 examples, 1x T4 (16GB)
Adapter size: 19 MB on disk
GPU peak: 11.4 GB
Training time: 38 minutes
Eval on 50 held-out instructions (1=bad ... 5=excellent):
Approach Avg. quality In-domain style Stays on-topic
----------- ------------ --------------- --------------
Base model 3.4 2.1 3.0
LoRA fine-tuned 4.1 4.3 4.5
Datasets
Alpaca / Dolly / a curated domain set ↗
Start with a general instruction dataset (Dolly-15k is open-licensed) and filter to your chosen domain, or build a small custom set of 500–2,000 instruction/response pairs by hand or with a synthetic LLM pipeline.
How to get it: from datasets import load_dataset; ds = load_dataset("databricks/databricks-dolly-15k"). Then filter by category or topic.
Tools you'll need
These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.
Python: Python 3.10 or newer. Compute: A single 16 GB GPU (Colab T4 / Kaggle P100) is enough for a 1B–3B model with 4-bit base + LoRA. CPU is not realistic for the training step.
transformers— Loads the base language model.bitsandbytes— 4-bit / 8-bit loading. Without this you cannot fit a 3B model on a 16 GB GPU during training.accelerate— Handles device placement automatically.
peft— HuggingFace library for LoRA, QLoRA, prefix tuning, and other PEFT methods.trl— ProvidesSFTTrainer— a turnkey supervised-fine-tuning trainer that handles instruction formatting and LoRA together.
datasets— Streams the instruction dataset.sentencepiece— Tokeniser dependency for Llama and many other models.
evaluate— For automatic metrics. Most of the eval here will be qualitative side-by-side.
How to approach it
One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.
- Pick a base model. A small Instruct model in the 1B–3B range:
meta-llama/Llama-3.2-3B-Instruct,Qwen/Qwen2.5-1.5B-Instruct, ormicrosoft/Phi-3-mini-4k-instruct. - Pick a domain and curate data. 1,000–2,000 high-quality instruction/response pairs in your domain. Quality matters more than quantity at this size.
- Load in 4-bit. Use
BitsAndBytesConfigwithload_in_4bit=True. This drops the base model to ~2 GB on disk and frees up VRAM for training. - Attach LoRA. Use
peft.LoraConfig(r=16, lora_alpha=32, target_modules=["q_proj", "v_proj"]). Only the adapter weights train. - Train. Use
trl.SFTTrainer. 2–3 epochs is usually enough. Watch the loss curve. - Save the adapter. The final artifact is just the adapter (10–50 MB), not the base model.
- Evaluate. Pick 50 held-out instructions. Run both the base model and the LoRA-tuned model. Have the team rate both, blind, on quality / in-domain style / staying on topic.
- Report. Adapter size on disk, GPU peak, training time, and the comparison numbers.
What to deliver
- The LoRA adapter checkpoint (uploaded to a HuggingFace repo or stored in the project artifacts).
- A reproducible notebook for training and inference.
- A side-by-side eval report on 50 held-out instructions with team ratings.
- A short README documenting base model, dataset, hyperparameters, GPU used, and adapter size.