LLM APPLICATION
DocQuery RAG
A local document question-answering prototype that retrieves relevant text from ChromaDB and uses Ollama to generate concise, context-grounded answers.
- Python
- Ollama
- ChromaDB
- RAG
At a glance
- Included source collection
- 21 text articles
- Local model
- Gemma 2B
- Retrieved context
- Top 2 chunks per question
- Storage
- Persistent ChromaDB
The problem
A language model needs access to your documents to answer questions about their contents. Sending an entire collection in every prompt is impractical.
What I built
Built a Python RAG pipeline that loads local news articles, splits them into overlapping chunks, generates embeddings with Ollama, and stores them in persistent ChromaDB. Questions retrieve relevant chunks before answer generation.
How it works
- Load UTF-8 .txt files from the news_articles directory.
- Split documents into 1,000-character chunks with a 20-character overlap and generate Gemma 2B embeddings.
- Upsert text and embeddings into a persistent Chroma collection; retrieve the top two chunks for a question.
- Pass the retrieved context to Gemma 2B with instructions to answer concisely and acknowledge missing information.
Outcome
An end-to-end text ingestion, vector retrieval, and local answer-generation pipeline.
Limitations
- The current app is a Python script with an example question, rather than a browser chat interface or document-upload service.
- Ingestion supports .txt files. PDF parsing and explicit source citations in generated answers are not implemented.
- A local Ollama installation is required. Retrieval and generation quality have not been benchmarked in the repository.