Menu
Home Program Lecturers Important Dates Venue Sponsors Past Editions Speakers Alumni Versions GAI2026 Contact
Image Caption Generator — Project #16 | Summer School on Generative AI
Project 16

Image Caption Generator

GPU
Note. This page lays out one full version of the project — the goal, a sample input/output, suggested tools, and a step-by-step plan. Treat it as a reference, not a script. Your team can pick a different angle, swap libraries, narrow the scope, or take the project somewhere we did not anticipate. As long as the final deliverable makes sense for the goal, you are on track.
Task. Build an image caption generator: given an image, output a one- or two-sentence caption that describes what's in it. Evaluate on ~100 images from a public test set and report BLEU and METEOR against the reference captions, plus a small human spot-check for naturalness. Look at the failure cases — does the model miss objects, hallucinate, get colors wrong, repeat itself? You could use any vision-language model (BLIP-2, LLaVA, Qwen2-VL, Florence-2) via HuggingFace transformers, the evaluate library for BLEU/METEOR, and HuggingFace datasets for the test set. Going further (optional). Build a Streamlit UI where the user drops an image and sees the caption appear, or extend with style-conditioning so the model can produce "formal" / "funny" / "poetic" versions of the same caption.
Resources: Colab T4 or 8GB GPU recommended for vision-language model inference.

What you'll build

Build an image-captioning system. Given a photo, the model produces a short English sentence describing it. The team will build it in two ways and compare: (1) use an existing pretrained vision-language model (BLIP-2 or similar) zero-shot, and (2) fine-tune a small encoder-decoder model (Vision Transformer encoder + small GPT-2 decoder) on a captioning dataset. Score both with BLEU and CIDEr, and run a small human quality study.

What goes in, what comes out

Input

An image (a JPEG or PNG, any size). For the fine-tune: image + reference caption pairs from a captioning dataset.

Output

A short English sentence (5–15 words) describing the image.

One image with its references
{
  "image_id": 391895,
  "file_name": "COCO_val2014_000000391895.jpg",
  "captions": [
    "A man with a red helmet on a small moped on a dirt road.",
    "Man riding a motor bike on a dirt road on the countryside.",
    "A man riding on the back of a motorcycle.",
    "A dirt path with a young person on a motor bike rests to the foreground of a verdant area with a bridge.",
    "A man in a red shirt and a red hat is on a motorcycle on a hill side."
  ]
}
Results from the two approaches on the COCO val set
Approach                   BLEU-4   CIDEr   Human rating (100 imgs)
-------------------------  ------   -----   -----------------------
BLIP-2 zero-shot            0.31    1.05         4.0 / 5
ViT + GPT-2 fine-tune       0.28    0.98         3.6 / 5

Example caption (image: man on motorbike on a dirt road):
  BLIP-2:     "A man in a red shirt is riding a motorcycle on a dirt road."
  Fine-tuned: "A man on a motorcycle riding down a dirt road."
  Reference:  "A man with a red helmet on a small moped on a dirt road."

Datasets

COCO Captions ↗

~123k images with 5 reference captions each. The standard captioning benchmark. Use the Karpathy splits (113k train / 5k val / 5k test).

How to get it: from datasets import load_dataset; ds = load_dataset("yerevann/coco-karpathy"). Or download images and JSON from cocodataset.org directly.

License: Creative Commons Attribution 4.0.

Tools you'll need

These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.

Python: Python 3.10 or newer. Compute: A 16 GB GPU is enough for inference with BLIP-2 (in 8-bit) and for fine-tuning a ViT+GPT-2 model. A CPU works for inference only and will be slow.

Vision-language models
  • transformers — Provides BLIP, BLIP-2, ViT, and GPT-2 — everything the team needs.
  • torchvision — Image transforms for fine-tuning.
  • pillow — Image I/O.
Fine-tuning
  • torch — PyTorch for the training loop.
  • accelerate — Handles mixed precision and device placement.
  • bitsandbytes — For loading BLIP-2 in 8-bit if VRAM is tight.
Evaluation
  • evaluate — BLEU.
  • pycocoevalcap — The official COCO Captions evaluator: BLEU, METEOR, CIDEr, SPICE.

How to approach it

One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.

  1. Load COCO Captions. Use the Karpathy splits. Download images locally (the JSON has filenames).
  2. Stage 1 — BLIP-2 zero-shot. Load Salesforce/blip2-opt-2.7b (in 8-bit if needed). For each image in the val set, generate one caption. Save outputs.
  3. Stage 2 — Encoder-decoder fine-tune. Combine a Vision Transformer encoder (google/vit-base-patch16-224) with a small text decoder (gpt2). Use VisionEncoderDecoderModel from HuggingFace. Fine-tune on the train set for 2 epochs.
  4. Generate on val. For both systems. Use beam search with num_beams=4.
  5. Score automatic metrics. Run pycocoevalcap on the generated captions vs the 5 references per image. Report BLEU-4 and CIDEr.
  6. Human evaluation. Sample 100 images. Have the team rate each system's caption (blind) on a 1–5 scale for accuracy and fluency.
  7. Compare. Automatic + human. Discuss the cases where the metrics disagree.

What to deliver

  • Two captioning notebooks: BLIP-2 zero-shot and ViT+GPT-2 fine-tuned.
  • A results table with BLEU-4, CIDEr, and the team's human ratings on 100 images.
  • A short error analysis: 10 failure cases (the team's favourites) with explanation.

References