BLIP-2, LLaVA, Qwen2-VL, Florence-2) via HuggingFace transformers, the evaluate library for BLEU/METEOR, and HuggingFace datasets for the test set. Going further (optional). Build a Streamlit UI where the user drops an image and sees the caption appear, or extend with style-conditioning so the model can produce "formal" / "funny" / "poetic" versions of the same caption.What you'll build
Build an image-captioning system. Given a photo, the model produces a short English sentence describing it. The team will build it in two ways and compare: (1) use an existing pretrained vision-language model (BLIP-2 or similar) zero-shot, and (2) fine-tune a small encoder-decoder model (Vision Transformer encoder + small GPT-2 decoder) on a captioning dataset. Score both with BLEU and CIDEr, and run a small human quality study.
What goes in, what comes out
Input
An image (a JPEG or PNG, any size). For the fine-tune: image + reference caption pairs from a captioning dataset.
Output
A short English sentence (5–15 words) describing the image.
{
"image_id": 391895,
"file_name": "COCO_val2014_000000391895.jpg",
"captions": [
"A man with a red helmet on a small moped on a dirt road.",
"Man riding a motor bike on a dirt road on the countryside.",
"A man riding on the back of a motorcycle.",
"A dirt path with a young person on a motor bike rests to the foreground of a verdant area with a bridge.",
"A man in a red shirt and a red hat is on a motorcycle on a hill side."
]
}
Approach BLEU-4 CIDEr Human rating (100 imgs)
------------------------- ------ ----- -----------------------
BLIP-2 zero-shot 0.31 1.05 4.0 / 5
ViT + GPT-2 fine-tune 0.28 0.98 3.6 / 5
Example caption (image: man on motorbike on a dirt road):
BLIP-2: "A man in a red shirt is riding a motorcycle on a dirt road."
Fine-tuned: "A man on a motorcycle riding down a dirt road."
Reference: "A man with a red helmet on a small moped on a dirt road."
Datasets
COCO Captions ↗
~123k images with 5 reference captions each. The standard captioning benchmark. Use the Karpathy splits (113k train / 5k val / 5k test).
How to get it: from datasets import load_dataset; ds = load_dataset("yerevann/coco-karpathy"). Or download images and JSON from cocodataset.org directly.
Tools you'll need
These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.
Python: Python 3.10 or newer. Compute: A 16 GB GPU is enough for inference with BLIP-2 (in 8-bit) and for fine-tuning a ViT+GPT-2 model. A CPU works for inference only and will be slow.
transformers— Provides BLIP, BLIP-2, ViT, and GPT-2 — everything the team needs.torchvision— Image transforms for fine-tuning.pillow— Image I/O.
torch— PyTorch for the training loop.accelerate— Handles mixed precision and device placement.bitsandbytes— For loading BLIP-2 in 8-bit if VRAM is tight.
evaluate— BLEU.pycocoevalcap— The official COCO Captions evaluator: BLEU, METEOR, CIDEr, SPICE.
How to approach it
One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.
- Load COCO Captions. Use the Karpathy splits. Download images locally (the JSON has filenames).
- Stage 1 — BLIP-2 zero-shot. Load
Salesforce/blip2-opt-2.7b(in 8-bit if needed). For each image in the val set, generate one caption. Save outputs. - Stage 2 — Encoder-decoder fine-tune. Combine a Vision Transformer encoder (
google/vit-base-patch16-224) with a small text decoder (gpt2). UseVisionEncoderDecoderModelfrom HuggingFace. Fine-tune on the train set for 2 epochs. - Generate on val. For both systems. Use beam search with num_beams=4.
- Score automatic metrics. Run pycocoevalcap on the generated captions vs the 5 references per image. Report BLEU-4 and CIDEr.
- Human evaluation. Sample 100 images. Have the team rate each system's caption (blind) on a 1–5 scale for accuracy and fluency.
- Compare. Automatic + human. Discuss the cases where the metrics disagree.
What to deliver
- Two captioning notebooks: BLIP-2 zero-shot and ViT+GPT-2 fine-tuned.
- A results table with BLEU-4, CIDEr, and the team's human ratings on 100 images.
- A short error analysis: 10 failure cases (the team's favourites) with explanation.