Menu
Home Program Lecturers Important Dates Venue Sponsors Past Editions Speakers Alumni Versions GAI2026 Contact
Project Ideas - International Summer School on Generative AI 2027

Project Ideas

22 hands-on project ideas — teams of 4 students · extensible to short conference papers.

Filter by Hardware
CPU OK GPU CPU or GPU Clear all
Showing all 22 projects

Have your own idea?

You don't have to pick from this list. If you'd like to propose your own project, please discuss it with the organizers first so we can confirm scope, hardware, and mentorship.

1Topic Modeling on arXiv Abstracts

CPU OK
Resources: CPU-only (or laptop with any GPU).
Task. Collect 6-12 months of arXiv abstracts in NLP and machine learning, group them into topics, label each topic, and show which topics are growing or shrinking month by month. You could use the arxiv Python package to fetch papers, any sentence embedding model (e.g. all-MiniLM-L6-v2), and a clustering method of your choice (BERTopic, HDBSCAN, k-means). You're free to pick the labeling strategy, and the visualisation library. Going further (optional). Build a small Streamlit / Gradio dashboard for the demo, or turn it into an agent that watches arXiv weekly and writes a short LLM-generated summary of the week's rising topics.

2Instruction Fine-Tuning SmolLM 135M into a Tiny Assistant

GPU
Resources: Colab T4 / Kaggle P100 / any 8GB GPU is comfortable; CPU works for inference only and will be slow.
Task. Take a base SmolLM 135M model (no instruction tuning) and turn it into a tiny chat assistant by full supervised fine-tuning on an Alpaca-style instruction dataset. At 135M parameters, a full fine-tune is genuinely feasible on a single small GPU — no LoRA needed. The team will compare the base model and the fine-tuned model on the same 50-prompt evaluation set and see, with their own eyes, what "instruction following" actually looks like at this scale. You could use HuggingFaceTB/SmolLM-135M as the base, the trl.SFTTrainer for the training loop, and Dolly-15k or Alpaca as the dataset. Chat-template formatting, dataset filtering, and the eval prompts are your design choices. Going further (optional). Compare full fine-tune against LoRA on the same model, train on a narrow domain (cooking, code comments, customer support) and report the change in tone, or quantise the final model to 4-bit and measure the quality cost.
Datasets: Dolly-15k or Alpaca-cleaned (around 15k–50k instruction/response pairs); a hand-written set of 50 evaluation prompts the team picks to probe the difference between base and tuned model.

3Named Entity Recognition on CoNLL-2003

CPU / GPU
Resources: Colab T4 or 8GB GPU for BERT fine-tuning. CPU works for the LLM zero-shot baseline.
Task. Use the CoNLL-2003 English NER dataset (4 entity types: PER, ORG, LOC, MISC). Train an NER model and a zero-shot LLM baseline on the same data, then report F1 per entity type and the confusion patterns (PER vs ORG is the classic one). You could use HuggingFace transformers for fine-tuning (bert-base-cased, RoBERTa, or DeBERTa), seqeval for the metric, and any LLM (with a structured-output prompt) for the baseline. Going further (optional). Build a small UI where someone can paste a sentence and see highlighted entities, or wrap the model behind an agent that flags entities while the user is typing a longer text.
Datasets: CoNLL-2003 (English NER, public, standard benchmark) via HuggingFace datasets.

4Sentiment Classification on SST-2

CPU / GPU
Resources: Colab T4 for the fine-tuning baseline. CPU works for the LLM prompting variants.
Task. Use SST-2 (binary movie review sentiment, ~67k train, 872 dev). Build a fine-tuned sentiment classifier and compare it against an LLM with zero-shot and few-shot prompts on the same dev set. Report accuracy, latency per example, and a confusion table — especially for the negation cases SST-2 is known for. You could use HuggingFace transformers for fine-tuning (distilbert-base-uncased is a good starting point), any LLM for prompting, and HuggingFace datasets for the data. Going further (optional). Build a small UI where someone types a review and sees both predictions side by side, or extend with a third model — sentence embeddings + logistic regression — to see where the classical baseline lands.
Datasets: SST-2 from the GLUE benchmark via HuggingFace datasets (glue, sst2).

5Paraphrase Detection on MRPC

CPU / GPU
Resources: CPU works for sentence-embedding baselines; Colab T4 if you also fine-tune.
Task. Use MRPC from GLUE (binary paraphrase classification, ~3.7k train, 408 dev). Build paraphrase classifiers and report accuracy and F1 on the dev set. MRPC is small and class-imbalanced — watch the F1 on the minority class, not just accuracy. You could use cosine similarity over sentence-transformers embeddings with a logistic-regression threshold, a fine-tuned encoder (bert-base-uncased, RoBERTa), an LLM zero-shot baseline, or any combination of these. Going further (optional). Build a small UI where someone types two sentences and sees the paraphrase probability, or extend with a hard-negative mining script that finds the cases your best model gets wrong.
Datasets: MRPC from the GLUE benchmark via HuggingFace datasets (glue, mrpc).

6Sentiment Classification on Sentiment140: Classical, Neural, and Transformer

CPU / GPU
Resources: CPU works for the sklearn baseline and a small subsample; Colab T4 or 8GB GPU for the LSTM and the BERT fine-tune. The dataset is large (1.6M tweets) but easily subsamples — pick a size that fits your hardware.
Task. Use the Sentiment140 dataset (1.6M English tweets, binary positive / negative, balanced). Build three sentiment classifiers on the same split — a classical model, a small neural model, and a fine-tuned transformer — and compare accuracy, F1, training time, and inference latency. The point is to feel what each stage of NLP buys you on the same task. You could use sklearn for the classical baseline (TF-IDF + Logistic Regression, Naive Bayes, or SVM), PyTorch or Keras for the neural model (BiLSTM, CNN, anything), and HuggingFace transformers for the transformer (DistilBERT, BERT, RoBERTa). Resource-friendly mode. If the team is short on RAM or compute, subsample with a fixed seed and keep the classes balanced — 50k tweets on a laptop, 200k on a mid-range GPU, the full 1.6M on Colab T4 or better. The comparison stays meaningful at any size. Going further (optional). Build a small UI where someone types a tweet and sees all three predictions, train at three subsample sizes to plot accuracy vs data, or add an LLM zero-shot baseline so you can compare four stages of NLP, not three.
Datasets: Sentiment140 via HuggingFace datasets (stanfordnlp/sentiment140). Remap labels 4 → 1 after loading.

7SMS Spam Detection: Classical, Neural, and Transformer

CPU OK
Resources: CPU is enough for all three stages on this tiny dataset.
Task. Use the SMS Spam Collection (5,574 messages, free). Build three spam classifiers — a classical baseline, a small neural model, and a fine-tuned transformer — and put them head-to-head on accuracy and F1 on the spam class. The lesson is that the classical baseline is almost perfect here — sometimes BERT is overkill. You could use sklearn for the classical model (TF-IDF + Naive Bayes is the textbook recipe), PyTorch for the neural model (BiLSTM, GRU, or CNN), and HuggingFace transformers for the transformer. Going further (optional). Build a small UI that flags messages as a user types, or extend with an adversarial test where the spammer uses common evasion tricks (zero-width characters, letter substitutions) and see which model holds up.
Datasets: SMS Spam Collection via UCI or HuggingFace datasets (ucirvine/sms_spam).

8Natural Language Inference on SNLI: Hand-Crafted, Siamese LSTM, and RoBERTa

GPU
Resources: CPU works for the hand-crafted baseline; Colab T4 or 8GB GPU for the LSTM and RoBERTa.
Task. Use SNLI (premise–hypothesis pairs, 3 classes: entailment, contradiction, neutral). Build three classifiers — a classical model with hand-crafted features, a small neural sentence-pair model, and a fine-tuned transformer — and report accuracy on the dev set. As a bonus, train each model on hypotheses only (no premise) to reveal the famous SNLI dataset artifact. You could use sklearn for the classical baseline (logistic regression on lexical overlap, length difference, negation cues), PyTorch for the neural model (Siamese BiLSTM is a classic choice), and HuggingFace transformers with roberta-base or DeBERTa. Going further (optional). Build a small UI where someone types two sentences and sees the predicted relation, or extend with an LLM zero-shot baseline and compare it to the supervised numbers.
Datasets: SNLI via HuggingFace datasets (stanfordnlp/snli).

9Question Classification on TREC: SVM, MLP, and BERT

CPU / GPU
Resources: CPU works for the sklearn baseline and the MLP; Colab T4 for the BERT fine-tune.
Task. Use the TREC question dataset (~5,500 questions, 6 coarse categories: ABBR, ENTY, DESC, HUM, LOC, NUM). Build three classifiers — a classical baseline, a small neural model, and a fine-tuned transformer — and report accuracy plus the per-class confusion matrix. The dataset is small enough that all three stages fit in a single session. You could use sklearn for the classical baseline (TF-IDF + Linear SVM or Logistic Regression), PyTorch for the neural model (an MLP on bag-of-word-embeddings, or a small BiLSTM), and HuggingFace transformers with distilbert-base-uncased. Going further (optional). Build a small UI where someone types a question and sees the predicted category from all three, or extend to the fine-grained TREC labels (50 classes) to see how each stage handles the harder version.
Datasets: TREC via HuggingFace datasets (CogComp/trec).

10LoRA Fine-Tuning a Small Language Model on a Niche Dataset

CPU / GPU
Resources: Colab T4 or 8GB GPU is comfortable; a laptop CPU also works for sub-1.5B models if you have a few hours.
Task. Take a small open language model and adapt it to a narrow task using LoRA. Pick a niche dataset (a custom support-ticket set, a cooking corpus, code-switching examples, anything where the base model is mediocre out of the box). Compare the LoRA-tuned model against the same base model with zero-shot and few-shot prompting, on accuracy and inference latency. The goal is to feel how much a small fine-tune buys you. You could use SmolLM2-360M / 1.7B, Qwen2.5-0.5B / 1.5B, TinyLlama-1.1B, or Pythia-1B as the base; peft + transformers for LoRA; bitsandbytes for 8-bit Adam if you want to fit the run on smaller hardware. Going further (optional). Build a small UI to show side-by-side answers from base / few-shot / LoRA, or try two adapter sizes and report the quality-vs-parameter trade-off.
Datasets: A custom small dataset of your choice (~500–2000 examples), or a HuggingFace dataset for a niche domain (e.g. Dolly-15k subset filtered to one task type, Alpaca domain slice).

11Quantising and Distilling a Small Language Model for Laptop Deployment

CPU / GPU
Resources: Colab T4 or 8GB GPU for the distillation run. The final quantised model must run on a laptop CPU.
Task. Take a small language model (1–3B parameters) and produce a version that runs comfortably on a laptop without a GPU. Compare two paths: post-training quantisation (8-bit, 4-bit, GGUF) and knowledge distillation into an even smaller student. For each variant, measure quality on a held-out task, tokens-per-second on CPU, and on-disk size. Identify the sweet spot between speed and quality. You could use llama.cpp or bitsandbytes for quantisation; transformers + peft + trl for the distillation; any small base model (SmolLM2, Qwen2.5-0.5B / 1.5B, Phi-3-mini, TinyLlama-1.1B); llama-bench or your own timer for tokens-per-second. Going further (optional). Wrap the quantised model in a small offline chat UI (Streamlit or a terminal app) that runs entirely on the laptop, or compare against a 4-bit version served via Ollama for a realistic deployment.
Datasets: A small held-out evaluation set on whichever task you train for — instruction-following (Alpaca-eval mini), summarization (XSum-mini), or arithmetic (GSM8K-mini).

12Chatbot using a Large Language Model

CPU / GPU
Resources: CPU works with a hosted LLM; Colab T4 / 8GB GPU if you want to run a small open model locally.
Task. Build a chatbot powered by a language model. The bot should hold a coherent multi-turn conversation, remember previous turns, follow a clear persona or set of instructions, and refuse gracefully when asked something outside its scope. Define what your chatbot is for (general assistant, study buddy, recipe helper, language tutor — your choice) and write 20 test conversations to evaluate it on. You could use any hosted LLM (OpenAI, Anthropic, Groq) or open model (Llama-3.1-8B, Qwen2.5-7B, Mistral-7B) via transformers or Ollama, and any chat framework (LangChain, LlamaIndex, or just direct API calls). The system prompt, memory strategy, and refusal logic are your design. Going further (optional). Build a Streamlit or Gradio UI for the demo, or wrap the chatbot as an agent with one or two tools (calculator, web search) so it can answer questions the base model can't.
Datasets: No fixed dataset — write 20 hand-crafted test conversations covering normal use, edge cases, and refusal scenarios.

13Customer Support Agent

CPU / GPU
Resources: CPU works with a hosted LLM; Colab T4 / 8GB GPU if you run a local open model.
Task. Build a customer-support agent for an imaginary product (a software tool, an online shop, an airline — pick something concrete). The agent should answer common questions from a knowledge base, handle multi-turn clarifications, look up order or account information using a simple tool, and escalate to a human when uncertain. Write 30 test conversations covering FAQ-style questions, account lookups, and tricky cases that should escalate. You could use any LLM, any vector store for the knowledge base (FAISS, Chroma), a small SQLite "orders" table for the lookup tool, and any agent framework (LangChain, smolagents, or pure Python). Going further (optional). Build a Streamlit chat UI with a side panel showing the agent's reasoning and tool calls, or add a sentiment classifier that detects frustrated customers and escalates earlier.
Datasets: A small hand-written knowledge base (20–50 FAQ entries) and a mock \"orders\" SQLite table; 30 evaluation conversations.

14Medical Assistant

CPU / GPU
Resources: CPU works with a hosted LLM; Colab T4 / 8GB GPU for a local open model.
Task. Build a medical assistant that helps a user understand a symptom, a medication, or a health condition using only authoritative sources (WHO, NIH MedlinePlus, NICE). The assistant must cite every claim, refuse when no source supports the answer, and carry a clear "informational only, not medical advice" notice in the UI. Evaluate on 30 realistic consumer-health questions covering common, ambiguous, and out-of-scope cases. You could use any retrieval framework, any embedding model, any LLM, and any readability metric (Flesch-Kincaid is one option). Going further (optional). Build a Streamlit UI with a prominent disclaimer banner and clickable citations, or extend with a multilingual mode so users can ask in their own language and still see the cited source in English.
Datasets: Public health corpora: MedlinePlus (NIH, public domain), WHO factsheets (CC-BY-NC-SA), NICE guidelines (UK open licence).

15PDF Question Answering System

CPU / GPU
Resources: CPU works with a hosted LLM; Colab T4 / 8GB GPU for a local open model.
Task. Build a system that lets the user upload one or more PDFs and ask questions about them. The system extracts text, chunks it, indexes it, retrieves the most relevant chunks for each question, and answers grounded in the document with a citation back to the page or paragraph. Test on a mix of PDFs: a research paper, a textbook chapter, a long contract — anything where finding the right paragraph matters. You could use pypdf, marker, or docling for PDF parsing, any embedding model and vector store, and any LLM. Going further (optional). Build a Streamlit UI where the user clicks the citation and jumps to the supporting page in the rendered PDF, or extend to multi-document mode where the user uploads several PDFs and the system tells them which document an answer came from.
Datasets: Any PDFs your team finds interesting — research papers from arXiv, government documents, manuals, textbook chapters.

16Image Caption Generator

GPU
Resources: Colab T4 or 8GB GPU recommended for vision-language model inference.
Task. Build an image caption generator: given an image, output a one- or two-sentence caption that describes what's in it. Evaluate on ~100 images from a public test set and report BLEU and METEOR against the reference captions, plus a small human spot-check for naturalness. Look at the failure cases — does the model miss objects, hallucinate, get colors wrong, repeat itself? You could use any vision-language model (BLIP-2, LLaVA, Qwen2-VL, Florence-2) via HuggingFace transformers, the evaluate library for BLEU/METEOR, and HuggingFace datasets for the test set. Going further (optional). Build a Streamlit UI where the user drops an image and sees the caption appear, or extend with style-conditioning so the model can produce "formal" / "funny" / "poetic" versions of the same caption.
Datasets: MS COCO Captions (a small dev slice) or Flickr8k via HuggingFace datasets.

17Automatic Text Summarization

CPU / GPU
Resources: CPU works for hosted LLMs and small models; Colab T4 / 8GB GPU if you fine-tune.
Task. Build a summarization system that takes a longer text (a news article, a research abstract, a blog post) and produces a short, faithful summary. Compare two approaches on the same ~200 test documents: a small fine-tuned summariser and a zero-shot LLM. Report ROUGE-1/2/L, output length, and a small faithfulness check using an NLI model to flag claims not supported by the source. You could use HuggingFace transformers with a model like sshleifer/distilbart-cnn-12-6, facebook/bart-large-cnn, or google/pegasus-xsum, the evaluate library for ROUGE, an NLI model for the faithfulness pass, and any LLM for the zero-shot baseline. Going further (optional). Build a Streamlit UI where the user pastes an article and sees both summaries side by side with unsupported spans highlighted, or extend with controllable length ("one sentence" / "a paragraph") via prompts or special tokens.
Datasets: CNN/DailyMail or XSum via HuggingFace datasets.

18Machine Translation with Transformers

GPU
Resources: Colab T4 or 8GB GPU for fine-tuning; CPU works for hosted-model baselines.
Task. Build a machine-translation system between two languages of your choice. Pick a parallel corpus, fine-tune a pretrained translation model, and compare it against a strong general-purpose LLM used zero-shot for the same translations. Report BLEU and chrF on a held-out test set and look at the qualitative differences: who is better at idioms, named entities, rare vocabulary? You could use HuggingFace transformers with a base model like Helsinki-NLP/opus-mt-*, facebook/mbart-large-50, or facebook/nllb-200-distilled-600M, the sacrebleu library for metrics, and any LLM for the zero-shot baseline. Going further (optional). Build a Streamlit UI for live side-by-side translation, or extend to a language pair where parallel data is scarce and study how performance degrades with less training data.
Datasets: WMT tasks, OPUS-100, FLORES-200, or Tatoeba via HuggingFace datasets.

19Multimodal and Domain-Specific NLP

GPU
Resources: Colab T4 or 8GB GPU recommended for the vision-language stage.
Task. Build a small system that combines two modalities (image + text) in a specific domain. Pick something useful to a real audience: receipts where the user uploads a photo and gets back structured info (vendor, total, line items), scientific charts where the user uploads a figure and the system describes the trend, or product photos where the system writes a short marketing blurb. Test on ~50 examples and evaluate the structured fields or text output against a small hand-annotated ground truth. You could use a vision-language model (Qwen2-VL, LLaVA, Florence-2) via HuggingFace transformers, an OCR library if useful (tesseract, easyocr, docling), and any LLM for the post-processing step. Going further (optional). Build a Streamlit UI for the chosen use case, or extend to a third modality (audio for receipts read aloud, for example).
Datasets: Pick a small, public, domain-specific image set: SROIE for receipts, ChartQA for scientific figures, or Fashion-MNIST captions for product photos.

20Translation and Tools for a Low-Resource Language

GPU
Resources: Colab T4 or 8GB GPU for fine-tuning.
Task. Pick a language the team speaks where NLP tools are weak — Darija, Sicilian, Berber, Tigrinya, Quechua, a regional dialect. Build a small useful tool: a translation system to and from English, a basic spell checker, a sentence-level classifier, or whatever the language most needs. Collect or curate a small parallel or labelled corpus (~500–2000 examples), fine-tune a multilingual model on it, and evaluate against a strong off-the-shelf baseline. The goal is to show measurable improvement over what existed before for that language. You could use HuggingFace transformers with a multilingual base (facebook/nllb-200-distilled-600M, facebook/mbart-large-50, google/madlad400-3b-mt), peft for LoRA if memory is tight, sacrebleu or chrF for metrics, and Common Voice / FLORES-200 if the language is partially covered. Going further (optional). Build a small Streamlit translator the wider community can use, or share the curated corpus on HuggingFace so others can build on top of it.
Datasets: A team-curated mini corpus, plus any of FLORES-200, OPUS, Common Voice if the language is partially covered.

21Headline Rewriter with a Small Language Model

CPU / GPU
Resources: Colab T4 or 8GB GPU for fine-tuning; CPU works for inference with a sub-1B model.
Task. Fine-tune a small language model (SmolLM-360M / SmolLM2-1.7B / Qwen2.5-0.5B) to rewrite news article leads into headlines in three styles: factual, attention-grabbing, and neutral. The team curates a small paired dataset (article → headline) from open news sources, fine-tunes with full SFT or LoRA, and builds a side-by-side demo where one input produces all three variants. Constrained tasks like this are exactly where a small LLM shines — the search space is narrow, the iteration loop is fast, and the demo is fun. You could use HuggingFace transformers + trl.SFTTrainer, any small base (SmolLM2-360M is a good default), and a hand-curated paired set of ~1000–3000 (article, headline) rows scraped from open news archives or CNN/DailyMail. Evaluation: ROUGE on the factual style, blind human ratings on the three styles. Going further (optional). Add a fourth style for a specific outlet's voice (Reuters, BBC), expose the demo as a small web tool, or evaluate against a hosted LLM zero-shot baseline.
Datasets: A hand-curated paired set scraped from open news (CNN/DailyMail, BBC headlines via RSS, the team's own scraping); ~1000–3000 (article-lead, headline) rows.

22Grammar Corrector with a Small Language Model

CPU / GPU
Resources: Colab T4 or 8GB GPU for fine-tuning; CPU is fine for inference with a sub-1B model.
Task. Fine-tune a small language model to correct grammatical errors in short English sentences written by learners. Take an existing learner-English corpus (Lang-8, JFLEG, BEA-2019), tune SmolLM2-360M / Qwen2.5-0.5B on the (ungrammatical, corrected) pairs, and evaluate properly with M2 (the official grammatical-error-correction metric) or GLEU. Compare against an open rule-based tool like LanguageTool and a zero-shot hosted LLM baseline. The clear before/after and the existence of a real metric make this an ideal first NLP project. You could use HuggingFace transformers + trl.SFTTrainer, a small open base (SmolLM2-360M, Qwen2.5-0.5B), the errant library for M2 scoring, and language-tool-python for the rule-based baseline. Going further (optional). Build a small browser-extension or Streamlit demo, evaluate per error type (missing article, wrong tense, agreement) instead of just overall score, or specialise on errors a specific L1 (Italian, French, Arabic) makes when writing English.
Datasets: JFLEG (jfleg) and BEA-2019 shared task data via HuggingFace datasets; Lang-8 as a larger pretraining-style corpus.