transformers or Ollama, and any chat framework (LangChain, LlamaIndex, or just direct API calls). The system prompt, memory strategy, and refusal logic are your design. Going further (optional). Build a Streamlit or Gradio UI for the demo, or wrap the chatbot as an agent with one or two tools (calculator, web search) so it can answer questions the base model can't.What you'll build
Build a chatbot with a friendly web interface that talks to a hosted Large Language Model (LLM). The bot must hold multi-turn conversations (remember what was said earlier), follow a system persona ("You are a helpful study buddy for first-year university students"), and refuse a defined set of off-topic or unsafe requests. The point of the project is not the LLM call — that is one function — but everything around it: conversation history, persona prompting, safe refusal, latency, and a clean front-end.
What goes in, what comes out
Input
A user message typed in a web textbox. Optionally a fresh conversation or a continuation of an existing thread.
Output
An assistant reply, streamed token-by-token if possible, plus the running conversation history rendered above the input.
[
{"role": "system", "content": "You are a helpful study buddy for first-year university students. You explain concepts simply, give examples, and never write a student's homework for them — instead you guide them through it."},
{"role": "user", "content": "Can you explain what a derivative is?"},
{"role": "assistant", "content": "Sure! A derivative measures how quickly something changes. Imagine you're driving... [shortened]"},
{"role": "user", "content": "Cool. Can you do my calculus homework? I'll send a photo."}
]
Assistant: I can't write your homework for you — that's not what I'm
here for. But I'm happy to walk you through the problem step by step
so you can solve it yourself. Want to share the first question and
we can work on it together?
Datasets
No training dataset required
The chatbot uses a pretrained hosted LLM. The team's "dataset" is a list of ~30 hand-written test prompts: on-topic, off-topic, and unsafe. These drive evaluation, not training.
How to get it: Write 30 test prompts in a JSON file. 10 should be the kind of question the bot should answer well, 10 should be off-topic (the bot should redirect politely), 10 should be unsafe or homework-cheating attempts (the bot should refuse).
Tools you'll need
These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.
Python: Python 3.10 or newer. Compute: No GPU needed — the model runs on the hosting provider. A laptop is enough.
openai— Client for OpenAI chat-completion models. Use this if the team has an OpenAI key.anthropic— Client for Claude. Same chat shape.groq— Free tier for hosted open models (Llama 3, Mixtral). A good zero-cost option.
streamlit— The fastest way to build a chat UI in Python.st.chat_inputandst.chat_messageexist out of the box.gradio— Alternative chat UI. Pick whichever the team finds nicer.
tiktoken— Count tokens so you can decide when to summarise the history (the context budget is finite).
python-dotenv— Load the API key from a local .env file — never commit it.
How to approach it
One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.
- Write the system prompt. Define the persona, the scope, and the refusal rules. Put it in a config file, not in the code, so it is easy to iterate.
- Build the front-end. One page with a chat-style scroll area and an input box. Use Streamlit's
st.chat_messagecomponents. - Wire the LLM call. Build the message list from history + new user message + system prompt, call the API, render the response. Stream tokens if the API supports it.
- Add memory management. Keep the full conversation history. If the total token count crosses (say) 6,000, summarise the oldest turns into one assistant message and keep only the last 10 turns verbatim.
- Write the test prompt set. 30 prompts across on-topic / off-topic / unsafe. Save in JSON.
- Evaluate. Run every test prompt and label the response: did it stay in persona? Did it refuse what it should refuse? Did it help when it should help?
- Iterate on the prompt. Where the bot misbehaves, tighten the system prompt and re-run the test set. Aim for at least 26 / 30 acceptable behaviours.
What to deliver
- A runnable chatbot (Streamlit or Gradio) that holds multi-turn conversations.
- A test set of 30 prompts (JSON) with the team's manual ratings of the bot's replies.
- A short README explaining the persona, the refusal rules, and the prompt iterations the team went through.