Menu
Home Program Lecturers Important Dates Venue Sponsors Past Editions Speakers Alumni Versions GAI2026 Contact
Chatbot using a Large Language Model — Project #12 | Summer School on Generative AI
Project 12

Chatbot using a Large Language Model

CPU / GPU
Note. This page lays out one full version of the project — the goal, a sample input/output, suggested tools, and a step-by-step plan. Treat it as a reference, not a script. Your team can pick a different angle, swap libraries, narrow the scope, or take the project somewhere we did not anticipate. As long as the final deliverable makes sense for the goal, you are on track.
Task. Build a chatbot powered by a language model. The bot should hold a coherent multi-turn conversation, remember previous turns, follow a clear persona or set of instructions, and refuse gracefully when asked something outside its scope. Define what your chatbot is for (general assistant, study buddy, recipe helper, language tutor — your choice) and write 20 test conversations to evaluate it on. You could use any hosted LLM (OpenAI, Anthropic, Groq) or open model (Llama-3.1-8B, Qwen2.5-7B, Mistral-7B) via transformers or Ollama, and any chat framework (LangChain, LlamaIndex, or just direct API calls). The system prompt, memory strategy, and refusal logic are your design. Going further (optional). Build a Streamlit or Gradio UI for the demo, or wrap the chatbot as an agent with one or two tools (calculator, web search) so it can answer questions the base model can't.
Resources: CPU works with a hosted LLM; Colab T4 / 8GB GPU if you want to run a small open model locally.

What you'll build

Build a chatbot with a friendly web interface that talks to a hosted Large Language Model (LLM). The bot must hold multi-turn conversations (remember what was said earlier), follow a system persona ("You are a helpful study buddy for first-year university students"), and refuse a defined set of off-topic or unsafe requests. The point of the project is not the LLM call — that is one function — but everything around it: conversation history, persona prompting, safe refusal, latency, and a clean front-end.

What goes in, what comes out

Input

A user message typed in a web textbox. Optionally a fresh conversation or a continuation of an existing thread.

Output

An assistant reply, streamed token-by-token if possible, plus the running conversation history rendered above the input.

What you send to the LLM on turn 3
[
  {"role": "system", "content": "You are a helpful study buddy for first-year university students. You explain concepts simply, give examples, and never write a student's homework for them — instead you guide them through it."},
  {"role": "user",   "content": "Can you explain what a derivative is?"},
  {"role": "assistant", "content": "Sure! A derivative measures how quickly something changes. Imagine you're driving... [shortened]"},
  {"role": "user",   "content": "Cool. Can you do my calculus homework? I'll send a photo."}
]
Persona-consistent refusal followed by helpful redirect
Assistant: I can't write your homework for you — that's not what I'm
here for. But I'm happy to walk you through the problem step by step
so you can solve it yourself. Want to share the first question and
we can work on it together?

Datasets

No training dataset required

The chatbot uses a pretrained hosted LLM. The team's "dataset" is a list of ~30 hand-written test prompts: on-topic, off-topic, and unsafe. These drive evaluation, not training.

How to get it: Write 30 test prompts in a JSON file. 10 should be the kind of question the bot should answer well, 10 should be off-topic (the bot should redirect politely), 10 should be unsafe or homework-cheating attempts (the bot should refuse).

Tools you'll need

These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.

Python: Python 3.10 or newer. Compute: No GPU needed — the model runs on the hosting provider. A laptop is enough.

LLM client
  • openai — Client for OpenAI chat-completion models. Use this if the team has an OpenAI key.
  • anthropic — Client for Claude. Same chat shape.
  • groq — Free tier for hosted open models (Llama 3, Mixtral). A good zero-cost option.
Front-end
  • streamlit — The fastest way to build a chat UI in Python. st.chat_input and st.chat_message exist out of the box.
  • gradio — Alternative chat UI. Pick whichever the team finds nicer.
Token accounting
  • tiktoken — Count tokens so you can decide when to summarise the history (the context budget is finite).
Secrets
  • python-dotenv — Load the API key from a local .env file — never commit it.

How to approach it

One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.

  1. Write the system prompt. Define the persona, the scope, and the refusal rules. Put it in a config file, not in the code, so it is easy to iterate.
  2. Build the front-end. One page with a chat-style scroll area and an input box. Use Streamlit's st.chat_message components.
  3. Wire the LLM call. Build the message list from history + new user message + system prompt, call the API, render the response. Stream tokens if the API supports it.
  4. Add memory management. Keep the full conversation history. If the total token count crosses (say) 6,000, summarise the oldest turns into one assistant message and keep only the last 10 turns verbatim.
  5. Write the test prompt set. 30 prompts across on-topic / off-topic / unsafe. Save in JSON.
  6. Evaluate. Run every test prompt and label the response: did it stay in persona? Did it refuse what it should refuse? Did it help when it should help?
  7. Iterate on the prompt. Where the bot misbehaves, tighten the system prompt and re-run the test set. Aim for at least 26 / 30 acceptable behaviours.

What to deliver

  • A runnable chatbot (Streamlit or Gradio) that holds multi-turn conversations.
  • A test set of 30 prompts (JSON) with the team's manual ratings of the bot's replies.
  • A short README explaining the persona, the refusal rules, and the prompt iterations the team went through.

References