A Retrieval-Augmented Generation (RAG) pipeline that answers questions based on your PDF notes using ChromaDB and Groq LLM.
Upload your PDF notes, ask any question, and get accurate answers with citations showing exactly which page the answer came from.
- Document Loading — PyMuPDF
- Text Splitting — LangChain RecursiveCharacterTextSplitter
- Embeddings — SentenceTransformers (all-MiniLM-L6-v2)
- Vector Store — ChromaDB
- LLM — Groq (LLaMA 3.3 70b)
- Clone the repo
git clone https://github.com/AkshitaSharma211/RAG.git
cd RAG- Install dependencies
pip install -r requirements.txt- Add your Groq API key
echo "GROQ_API_KEY=your_key_here" > .env- Run the notebook
Open
pdf_loader.ipynband run all cells top to bottom.
- PDFs are loaded and split into chunks
- Each chunk is converted to embeddings (384-dim vectors)
- Embeddings stored in ChromaDB
- User question is embedded and matched against stored chunks
- Top matching chunks sent to Groq LLM for answer generation
result = adv_rag.query("What is video compression?", top_k=3, min_score=0.1)
print(result['answer'])