Menu
Home Program Lecturers Important Dates Venue Sponsors Past Editions Speakers Alumni Versions GAI2026 Contact
Topic Modeling on arXiv Abstracts — Project #1 | Summer School on Generative AI
Project 1

Topic Modeling on arXiv Abstracts

CPU OK
Note. This page lays out one full version of the project — the goal, a sample input/output, suggested tools, and a step-by-step plan. Treat it as a reference, not a script. Your team can pick a different angle, swap libraries, narrow the scope, or take the project somewhere we did not anticipate. As long as the final deliverable makes sense for the goal, you are on track.
Task. Collect 6-12 months of arXiv abstracts in NLP and machine learning, group them into topics, label each topic, and show which topics are growing or shrinking month by month. You could use the arxiv Python package to fetch papers, any sentence embedding model (e.g. all-MiniLM-L6-v2), and a clustering method of your choice (BERTopic, HDBSCAN, k-means). You're free to pick the labeling strategy, and the visualisation library. Going further (optional). Build a small Streamlit / Gradio dashboard for the demo, or turn it into an agent that watches arXiv weekly and writes a short LLM-generated summary of the week's rising topics.
Resources: CPU-only (or laptop with any GPU).

What you'll build

Build a small system that ingests recent research abstracts from arXiv, groups them into coherent topics, gives each topic a human-readable name, and shows which topics are growing month by month. The output is a short ranked list of fast-rising topics plus a simple dashboard a researcher could browse to see what the field is excited about right now.

What goes in, what comes out

Input

Around 2,000–8,000 arXiv abstracts from the chosen categories, over a 2–10 month time window.

Output

A ranked list of around 20 named topics, monthly trajectories per topic, and the top 10 fastest-rising topics — viewable in a Streamlit or Gradio dashboard.

One arXiv record after parsing (saved to data/abstracts.jsonl)
{
  "paper_id": "2509.12345",
  "title": "Self-Refining Agents for Long-Horizon Tool Use",
  "abstract": "We introduce a self-refining agent framework that combines tool use with periodic reflection. The agent maintains a scratchpad summarising what worked and what failed, and periodically rewrites its plan...",
  "categories": ["cs.CL", "cs.LG"],
  "submitted": "2025-09-14",
  "authors": ["A. Doe", "B. Smith"]
}
Top fastest-rising topics over the chosen window
rank  topic_label                                  papers (last 3mo)  growth
----  -------------------------------------------  -----------------  ------
   1  Self-refining and reflective LLM agents                  412   +218 %
   2  Long-context retrieval with reranking                    287   +154 %
   3  Speculative decoding for inference speedup               203   +120 %
   4  Mixture-of-experts routing and load balance              188    +96 %
   5  Synthetic data for instruction tuning                    174    +71 %

Datasets

arXiv abstracts (chosen categories, last 2–10 months) ↗

Each record contains a title, an abstract (~150 words), one or more category tags, and a submission date.

How to get it: Use the arxiv Python package. Query one category at a time, respect the 3-second-per-request rate limit, save to data/abstracts.jsonl.

License: Metadata is public domain (CC0); abstracts are free for non-commercial research.

Kaggle arXiv metadata dump (optional, larger) ↗

A snapshot of 2M+ arXiv records as a single JSON file. Use it if you want a bigger dataset without the API.

How to get it: Download the JSON (~3 GB), filter to the categories and date range you care about, save to CSV or parquet.

License: CC0

Tools you'll need

These are suggestions, not requirements. If your team is more comfortable with a different library, model, or framework that achieves the same goal, use it — and briefly explain the choice in your README.

Python: Python 3.10 or newer. Compute: A laptop CPU is enough. No GPU required. Around 2 GB of free disk for cached abstracts and embeddings.

Embeddings: BAAI/bge-small-en-v1.5. LLM (optional): any hosted option (e.g. GPT-4o-mini or GPT-4.1-mini (OpenAI)).

Core
  • pandas — Storing and slicing the abstracts (one row per paper).
  • numpy — Embedding matrix handling, monthly counts.
  • tqdm — Progress bars while you fetch and embed thousands of abstracts.
Data ingestion
  • arxiv — Friendly client for the arXiv API. Handles pagination and rate limits.
  • requests — For pulling the optional Kaggle metadata dump or any extra URLs.
Embeddings and clustering
  • sentence-transformers — Embeds each abstract. all-MiniLM-L6-v2 is the standard small choice.
  • bertopic — Wraps embeddings + UMAP + HDBSCAN + keyword labels in one object. Fastest baseline. (UMAP and HDBSCAN are defined below.)
  • umap-learn — UMAP stands for Uniform Manifold Approximation and Projection. It compresses the high-dimensional (384-dim) embeddings down to a few dimensions so the clusterer can find structure.
  • hdbscan — HDBSCAN stands for Hierarchical Density-Based Spatial Clustering of Applications with Noise. It groups points by density, finds clusters of variable size, and labels outliers as noise.
  • scikit-learn — Used here for TF-IDF (Term Frequency–Inverse Document Frequency) baselines and small helpers.
Language model access (for the optional topic-labelling step)
  • openai / anthropic / groq — Hosted Large Language Models (LLMs) for naming clusters. The model only sees the keywords plus 3 sample titles per cluster.
  • transformers — Optional, if you want to run a small local language model for labelling.
Dashboard and viz
  • plotly — Time-series and stacked-area charts.
  • streamlit — Quickest path to a clickable demo (~50 lines of Python).
  • gradio — Alternative dashboard, slightly nicer for ML demos.

How to approach it

One reasonable path through the project. Specific tools (UMAP, HDBSCAN, BERTopic, etc.) are examples — feel free to swap them for alternatives you know better.

  1. Fetch. Pull abstracts for the chosen categories and time window. Cache to data/abstracts.jsonl.
  2. Clean. Drop abstracts shorter than 20 words, strip LaTeX, normalise whitespace.
  3. Embed. Encode each abstract with bge-small-en-v1.5. Save the matrix to data/embeddings.npy.
  4. Reduce. Use UMAP (Uniform Manifold Approximation and Projection) to compress the embeddings from 384 dimensions down to 5.
  5. Cluster. Run HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) on the reduced vectors. Expect 20–40 clusters plus a noise bucket.
  6. Label. Extract the top 8 keywords per cluster using c-TF-IDF (class-based Term Frequency–Inverse Document Frequency: TF-IDF computed per cluster instead of per document). Optionally ask a language model to name each cluster from the keywords plus 3 sample titles.
  7. Trend. Bin per month, count per topic per month, compute a growth-rate score.
  8. Present. Build a Streamlit dashboard: top-rising table, monthly trajectories, sample papers per topic.

What to deliver

  • A reproducible script or notebook that takes a date range and produces the topic map.
  • A Streamlit (or Gradio) dashboard (~100 lines) showing rising topics, monthly trajectories, and example papers per topic.
  • A short README explaining how to run the pipeline and what the team found.

References