Menu
Home Program Lecturers Important Dates Venue Sponsors Past Editions Speakers Alumni Versions GAI2026 Contact
PDF Question Answering System — Project #15 | Summer School on Generative AI
Project 15

PDF Question Answering System

CPU / GPU
Note. This page lays out one full version of the project — the goal, a sample input/output, suggested tools, and a step-by-step plan. Treat it as a reference, not a script. Your team can pick a different angle, swap libraries, narrow the scope, or take the project somewhere we did not anticipate. As long as the final deliverable makes sense for the goal, you are on track.
Task. Build a system that lets the user upload one or more PDFs and ask questions about them. The system extracts text, chunks it, indexes it, retrieves the most relevant chunks for each question, and answers grounded in the document with a citation back to the page or paragraph. Test on a mix of PDFs: a research paper, a textbook chapter, a long contract — anything where finding the right paragraph matters. You could use pypdf, marker, or docling for PDF parsing, any embedding model and vector store, and any LLM. Going further (optional). Build a Streamlit UI where the user clicks the citation and jumps to the supporting page in the rendered PDF, or extend to multi-document mode where the user uploads several PDFs and the system tells them which document an answer came from.
Resources: CPU works with a hosted LLM; Colab T4 / 8GB GPU for a local open model.

What you'll build

Build a Retrieval-Augmented Generation (RAG) system that lets a user upload one or more PDF documents and ask questions about their content. The system must chunk the PDF, embed the chunks, retrieve the most relevant ones for a question, and then ask a Large Language Model (LLM) to answer using only those chunks. Every answer must show its sources (page number + short snippet) so the user can verify it.

What goes in, what comes out

Input

A PDF (typically 10–200 pages) and a natural-language question.

Output

A short text answer plus a list of source citations (page number + ~50-word snippet of the chunk the answer used).

A question and the retrieved chunks the LLM is grounded on
{
  "question": "What was the company's total revenue in 2024 and how did it change versus 2023?",
  "retrieved_chunks": [
    {"page": 4,  "score": 0.91, "text": "Total revenue for fiscal year 2024 was $2.34B, up 18% versus $1.98B in fiscal 2023."},
    {"page": 12, "score": 0.74, "text": "The growth was primarily driven by the enterprise SaaS segment, which contributed $890M."},
    {"page": 31, "score": 0.71, "text": "In contrast, hardware revenue declined 4% year over year due to inventory normalisation."}
  ]
}
Grounded answer with citations rendered in the UI
Answer:
  Total revenue in fiscal 2024 was $2.34B, an 18% increase versus
  $1.98B in fiscal 2023.

Sources:
  [page 4]  "Total revenue for fiscal year 2024 was $2.34B, up 18%
             versus $1.98B in fiscal 2023."
  [page 12] "The growth was primarily driven by the enterprise SaaS
             segment..."

Datasets

Bring your own PDFs

Pick 3–5 PDFs in different domains: a research paper, a company annual report, a textbook chapter, a long-form article. Different shapes test different things (tables, references, equations, footnotes).

How to get it: Download from arXiv, sec.gov, your university library, or a textbook PDF you have rights to use.

A small evaluation set of question / answer pairs

Write ~20 questions per PDF with the gold answer and the gold page number. This is what you score against.

How to get it: Have the team read the PDFs and write the eval set by hand. This usually takes 90 minutes per PDF and is the most valuable hour of the project.

Tools you'll need

These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.

Python: Python 3.10 or newer. Compute: A laptop is enough. The embedding model is small; the LLM runs on a hosting provider.

PDF parsing
  • pypdf — Extracts text from PDFs page by page. Simple and works for most documents.
  • pdfplumber — Better for PDFs with tables or multi-column layouts. Slower but more accurate.
Embeddings
  • sentence-transformers — Local sentence embeddings. all-MiniLM-L6-v2 is small and fast; BAAI/bge-small-en-v1.5 is a bit better.
Vector store
  • chromadb — Tiny embedded vector database. Runs in-process, no server.
  • faiss-cpu — Alternative; faster but no metadata filtering out of the box.
LLM
  • openai / anthropic / groq — A hosted LLM to generate the grounded answer.
UI
  • streamlit — PDF upload widget, chat-style question/answer, citation rendering — all easy.
Utilities
  • tiktoken — Count tokens before sending to the LLM so you stay under the context window.

How to approach it

One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.

  1. Parse. Extract text from the PDF page by page. Keep the page number with every chunk — you need it for citations.
  2. Chunk. Split each page into chunks of ~500 tokens with ~50 tokens of overlap. Try not to break in the middle of a sentence.
  3. Embed. Run every chunk through a sentence-embedding model. Store the embedding + chunk text + page number in the vector store.
  4. Retrieve. When the user asks a question, embed it the same way and find the top-k (k=5 is a good default) most similar chunks.
  5. Generate. Build the prompt: system instruction ("Answer using only the provided context. If the context does not contain the answer, say so."), the retrieved chunks, and the question. Call the LLM.
  6. Render. Show the answer plus the cited chunks (page number + snippet) so the user can verify.
  7. Evaluate. For each question in the eval set, check (a) did retrieval surface the correct page, (b) is the generated answer correct, (c) did the system cite the right page.

What to deliver

  • A Streamlit app where the user uploads a PDF and asks questions.
  • An evaluation report on the team's eval set: retrieval@5 (did the right page show up in the top 5?), answer accuracy, citation accuracy.
  • A short README documenting the chunking strategy, the embedding model, the system prompt, and the team's observations.

References