AI Study Notes Generator
Upload a PDF chapter and get summaries, flashcards, a glossary, and topic clusters back. A Streamlit NLP app that pairs an LLM with classical NLP — gpt-4o-mini for generation, KeyBERT + spaCy for terms, UMAP + HDBSCAN for clustering — and grades its own output with ROUGE.
01 Overview
Reading a dense AI/ML textbook chapter and turning it into usable study material is slow. This app does it in one pass: it chunks the PDF, generates a summary at the depth you choose, extracts a glossary, builds flashcards, and clusters the chapter into topics.
The design choice worth calling out is using the right tool for each job rather than throwing the LLM at everything: gpt-4o-mini writes prose, but KeyBERT and spaCy handle keyphrase extraction, and UMAP + HDBSCAN do the clustering. And because generation quality is hard to trust, the app measures itself with ROUGE and glossary accuracy against ground-truth datasets.
02 What it produces
03 Architecture
A PDF goes in; four kinds of study material and an evaluation report come out. The processing flow at a glance, then each step below.
Ingest & chunk
The PDF is read and split into ~800-character chunks, then embedded with a sentence-transformer (all-MiniLM-L6-v2) and indexed in FAISS for semantic lookup.
Summarize
gpt-4o-mini (via LangChain, low temperature) generates a summary at the requested depth — short, medium, or detailed — grounded in the chunked text.
Extract glossary & cards
KeyBERT surfaces chapter-specific keyphrases by semantic similarity, spaCy filters candidates, and the LLM turns them into definitions and flashcards.
Cluster topics
Chunk embeddings are reduced with UMAP and clustered with HDBSCAN to reveal the chapter's topic structure, with a 2-D projection for visualization.
Evaluate
An evaluation tab scores generated summaries with ROUGE and checks glossary accuracy against ground-truth datasets — so quality is measured, not assumed.
04 Design choices
Right tool per task
The LLM writes prose, but keyphrase extraction goes to KeyBERT and clustering to UMAP/HDBSCAN — cheaper, faster, and more controllable than prompting for everything.
Measured, not assumed
Generative output is easy to fake confidence about — so summaries are ROUGE-scored and the glossary is checked against ground truth, giving an honest quality signal.
Local embeddings
A 22 MB local sentence-transformer handles embeddings and KeyBERT — no extra API cost or latency for the retrieval and extraction paths.
One upload, four artifacts
Summary, flashcards, glossary, and topic clusters all come from a single PDF pass, organized into a tabbed Streamlit UI.
05 Stack
Run it on your own PDFs
The repo has the full Streamlit app, the NLP pipeline, and the evaluation harness.