Skip to content
aranyaadheu
← All projects

LLM APPLICATION

DocQuery RAG

A local document question-answering prototype that retrieves relevant text from ChromaDB and uses Ollama to generate concise, context-grounded answers.

View GitHub repository →

At a glance

Included source collection
21 text articles
Local model
Gemma 2B
Retrieved context
Top 2 chunks per question
Storage
Persistent ChromaDB

The problem

A language model needs access to your documents to answer questions about their contents. Sending an entire collection in every prompt is impractical.

What I built

Built a Python RAG pipeline that loads local news articles, splits them into overlapping chunks, generates embeddings with Ollama, and stores them in persistent ChromaDB. Questions retrieve relevant chunks before answer generation.

How it works

  1. Load UTF-8 .txt files from the news_articles directory.
  2. Split documents into 1,000-character chunks with a 20-character overlap and generate Gemma 2B embeddings.
  3. Upsert text and embeddings into a persistent Chroma collection; retrieve the top two chunks for a question.
  4. Pass the retrieved context to Gemma 2B with instructions to answer concisely and acknowledge missing information.

Outcome

An end-to-end text ingestion, vector retrieval, and local answer-generation pipeline.

Limitations

Explore the work