Menu
Home Program Lecturers Important Dates Venue Sponsors Past Editions Speakers Alumni Versions GAI2026 Contact
LoRA Fine-Tuning a Small Language Model on a Niche Dataset — Project #10 | Summer School on Generative AI
Project 10

LoRA Fine-Tuning a Small Language Model on a Niche Dataset

CPU / GPU
Note. This page lays out one full version of the project — the goal, a sample input/output, suggested tools, and a step-by-step plan. Treat it as a reference, not a script. Your team can pick a different angle, swap libraries, narrow the scope, or take the project somewhere we did not anticipate. As long as the final deliverable makes sense for the goal, you are on track.
Task. Take a small open language model and adapt it to a narrow task using LoRA. Pick a niche dataset (a custom support-ticket set, a cooking corpus, code-switching examples, anything where the base model is mediocre out of the box). Compare the LoRA-tuned model against the same base model with zero-shot and few-shot prompting, on accuracy and inference latency. The goal is to feel how much a small fine-tune buys you. You could use SmolLM2-360M / 1.7B, Qwen2.5-0.5B / 1.5B, TinyLlama-1.1B, or Pythia-1B as the base; peft + transformers for LoRA; bitsandbytes for 8-bit Adam if you want to fit the run on smaller hardware. Going further (optional). Build a small UI to show side-by-side answers from base / few-shot / LoRA, or try two adapter sizes and report the quality-vs-parameter trade-off.
Resources: Colab T4 or 8GB GPU is comfortable; a laptop CPU also works for sub-1.5B models if you have a few hours.

What you'll build

Fine-tune a small open-source language model (around 1B–3B parameters) on a narrow domain task using LoRA (Low-Rank Adaptation), a Parameter-Efficient Fine-Tuning (PEFT) technique. The team will compare the LoRA-adapted model against the base model on the same eval set, measure the size of the adapter weights, and explain why LoRA lets you customise a model on a single small GPU.

What goes in, what comes out

Input

A small instruction dataset for the chosen domain (1k–10k examples). For example: Alpaca-style instruction/response pairs in a specific domain (cooking, customer support, legal Q&A).

Output

A LoRA adapter (typically 10–50 MB on disk) plus a results table comparing base and LoRA models on the same eval prompts.

A few rows from a domain dataset
[
  {
    "instruction": "Suggest a vegetarian substitute for pancetta in pasta carbonara.",
    "input": "",
    "output": "Smoked tofu or king-oyster mushrooms work well. Dice into 1cm cubes, pan-fry in olive oil with a pinch of smoked paprika until crisp on the edges, then fold into the egg-cheese sauce off-heat as you would with pancetta."
  },
  {
    "instruction": "What can I do with leftover roast chicken?",
    "input": "",
    "output": "Shred and toss into a chicken-noodle soup, a Caesar salad, or a quick chicken tortilla wrap. The bones make a strong stock — simmer 2 hours with onion, carrot, and bay leaf."
  }
]
Comparison on the eval set after LoRA fine-tuning
Setup:
  Base model:    meta-llama/Llama-3.2-3B-Instruct
  LoRA rank:     16, alpha=32, dropout=0.05
  Target layers: q_proj, v_proj
  Training:      3 epochs on 2,000 examples, 1x T4 (16GB)
  Adapter size:  19 MB on disk
  GPU peak:      11.4 GB
  Training time: 38 minutes

Eval on 50 held-out instructions (1=bad ... 5=excellent):
  Approach        Avg. quality   In-domain style   Stays on-topic
  -----------     ------------   ---------------   --------------
  Base model         3.4              2.1               3.0
  LoRA fine-tuned    4.1              4.3               4.5

Datasets

Alpaca / Dolly / a curated domain set ↗

Start with a general instruction dataset (Dolly-15k is open-licensed) and filter to your chosen domain, or build a small custom set of 500–2,000 instruction/response pairs by hand or with a synthetic LLM pipeline.

How to get it: from datasets import load_dataset; ds = load_dataset("databricks/databricks-dolly-15k"). Then filter by category or topic.

License: Mostly permissive; check the specific dataset card.

Tools you'll need

These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.

Python: Python 3.10 or newer. Compute: A single 16 GB GPU (Colab T4 / Kaggle P100) is enough for a 1B–3B model with 4-bit base + LoRA. CPU is not realistic for the training step.

Base model and loading
  • transformers — Loads the base language model.
  • bitsandbytes — 4-bit / 8-bit loading. Without this you cannot fit a 3B model on a 16 GB GPU during training.
  • accelerate — Handles device placement automatically.
LoRA and training
  • peft — HuggingFace library for LoRA, QLoRA, prefix tuning, and other PEFT methods.
  • trl — Provides SFTTrainer — a turnkey supervised-fine-tuning trainer that handles instruction formatting and LoRA together.
Data
  • datasets — Streams the instruction dataset.
  • sentencepiece — Tokeniser dependency for Llama and many other models.
Evaluation
  • evaluate — For automatic metrics. Most of the eval here will be qualitative side-by-side.

How to approach it

One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.

  1. Pick a base model. A small Instruct model in the 1B–3B range: meta-llama/Llama-3.2-3B-Instruct, Qwen/Qwen2.5-1.5B-Instruct, or microsoft/Phi-3-mini-4k-instruct.
  2. Pick a domain and curate data. 1,000–2,000 high-quality instruction/response pairs in your domain. Quality matters more than quantity at this size.
  3. Load in 4-bit. Use BitsAndBytesConfig with load_in_4bit=True. This drops the base model to ~2 GB on disk and frees up VRAM for training.
  4. Attach LoRA. Use peft.LoraConfig(r=16, lora_alpha=32, target_modules=["q_proj", "v_proj"]). Only the adapter weights train.
  5. Train. Use trl.SFTTrainer. 2–3 epochs is usually enough. Watch the loss curve.
  6. Save the adapter. The final artifact is just the adapter (10–50 MB), not the base model.
  7. Evaluate. Pick 50 held-out instructions. Run both the base model and the LoRA-tuned model. Have the team rate both, blind, on quality / in-domain style / staying on topic.
  8. Report. Adapter size on disk, GPU peak, training time, and the comparison numbers.

What to deliver

  • The LoRA adapter checkpoint (uploaded to a HuggingFace repo or stored in the project artifacts).
  • A reproducible notebook for training and inference.
  • A side-by-side eval report on 50 held-out instructions with team ratings.
  • A short README documenting base model, dataset, hyperparameters, GPU used, and adapter size.

References