All work

Generative AI / Retrieval + generation

Enterprise RAG Knowledge Assistant

A document-grounded knowledge assistant that ingests common business files, retrieves relevant context, and generates grounded answers.

  • Streamlit
  • LangGraph
  • FAISS
  • BAAI/bge-small-en-v1.5
  • Gemini 2.5 Flash
Enterprise RAG Knowledge Assistant application overview
Overview
Enterprise RAG Knowledge Assistant knowledge base screen
Knowledge base
Enterprise RAG Knowledge Assistant chat screen
Chat
Retrieval workflow verified flow
Documents
Chunking
Embeddings
FAISS
Top-4 retrieval
Context
Gemini
Answer

Problem

Users need a practical way to ask questions over documents without sending the entire document collection into every model request.

Approach

The application loads supported documents, splits them into overlapping chunks, creates normalized BGE-small-en-v1.5 embeddings on CPU, stores them in FAISS, retrieves the top four matches, and supplies the built context to Gemini 2.5 Flash.

Architecture

01User
02Streamlit
03Document loader
04RecursiveCharacterTextSplitter
05BGE-small-en-v1.5 embeddings
06FAISS
07Top-4 similarity retrieval
08Context building
09Gemini 2.5 Flash
10Grounded answer

Capabilities

  • PDF, DOCX, TXT, Markdown, and CSV ingestion
  • Chunk size 1000
  • Chunk overlap 200
  • CPU embeddings
  • Normalized embeddings
  • Top-k retrieval of 4
  • Thread-based state

Problem

Knowledge spread across PDFs, DOCX files, text, Markdown, and CSV needs a focused retrieval workflow before an LLM can answer questions over it.

Solution

The assistant turns uploaded documents into searchable chunks, retrieves the most similar context, and generates an answer using that context.

Document ingestion

The loader supports PDF, DOCX, TXT, Markdown, and CSV documents.

Chunking and embeddings

Documents use a chunk size of 1000 with 200-token overlap. BAAI/bge-small-en-v1.5 creates normalized CPU embeddings.

Retrieval and generation

FAISS retrieves the top four similar chunks. The context builder passes those results to Gemini 2.5 Flash at temperature 0.3.

LangGraph orchestration

The workflow is organized as retrieve, build_context, and generate_answer, with thread-based state.

Challenges and learnings

The implementation makes the ingestion, retrieval, and generation stages explicit so the grounded-answer path remains easy to inspect.