Skip to content

Latest commit

Β 

History

57 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

askdocs-rag-agent

Tests Advanced RAG

Ask questions to your documents and get grounded, cited answers β€” a production-ready Document Q&A service built with FastAPI, PostgreSQL+pgvector, and advanced RAG techniques.

Stack: Python 3.12 Β· FastAPI Β· PostgreSQL + pgvector Β· Nuxt 4 Β· Tailwind CSS Β· Docker Β· Ollama (offline LLM)

Latest Update: Production-ready with 216/218 tests passing, HNSW indexing, duplicate detection, and API key authentication

Description

A production-ready RAG (Retrieval-Augmented Generation) system that enables natural language Q&A over document collections with guaranteed citation accuracy. Built with enterprise-grade architecture featuring custom query routing, HNSW-indexed vector similarity search using pgvector, and multi-LLM support (Gemini, Ollama, Azure OpenAI).

Key Differentiators:

  • Grounded-or-refuse architecture - Never hallucinates; returns "not_found" when answers aren't in documents
  • Citation tracking - Every answer includes exact document and page references
  • Duplicate detection - SHA-256 content hashing prevents duplicate uploads
  • HNSW indexing - 10-100x faster vector search with PostgreSQL+pgvector
  • API key authentication - All endpoints secured with X-API-Key header validation
  • Production-ready - 216/218 tests passing (99.1%), cloud deployment docs (GCP/Azure)
  • Flexible LLM backend - Swappable providers via adapter pattern (Gemini, Ollama offline, Azure OpenAI)

Target Use Cases: HR knowledge bases, customer support documentation, legal/compliance document search, IT helpdesk automation, sales enablement.

Market Position: Developer-first, self-hosted alternative to enterprise search products (Glean, Writer) β€” full data control and per-query costs instead of per-seat licensing. See docs/business for cost model.


πŸ“Έ Demo Screenshots

Chat Interface - Welcome Screen

Chat Welcome

Chat interface with filters (department, grade, type) and "New Chat" button. Clean welcome message guiding users to ask questions.


Document Upload with Metadata

Documents Upload

Upload PDF documents with optional metadata (department, grade level, document type, tags). Shows uploaded documents list with chunk counts.


Chat Q&A with Citations

Chat Q&A

Ask questions and get answers with exact source citations (document name, page number, similarity scores). Delete button to remove individual messages or entire chat.


Document Management

Documents List

View all uploaded documents with chunk counts and upload dates. Delete documents as needed.


Data Extraction - Schema Builder

Extraction Schema

Define custom extraction schemas with field names and types (Text, Number, List). Quick templates for common use cases (Job Posting, Invoice, Resume).


Data Extraction - Results

Extraction Results

Extracted structured data from documents with confidence scores. Export results as JSON or CSV.


πŸ“‹ Quick Demo: Getting Started | Sample Questions | Ollama Local LLM

Testing: Works with Ollama (free, 100% offline) or Gemini (cloud, best quality).


What It Does

Upload PDF documents β†’ Ask questions in natural language β†’ Get answers grounded in those documents with citations, or an honest "not found."

No hallucinations. Every answer either cites the exact source (document + page) or explicitly says the information doesn't exist in your documents.

Example:

Q: "What is the refund policy?"
A: "Refunds are processed within 14 days of purchase."
   Sources: [terms.pdf, page 7]

Q: "What's the weather today?"
A: "not_found - This question cannot be answered from the uploaded documents."

Why Build This?

The Problem:

  • Organizations have thousands of policy documents, manuals, handbooks
  • Employees and customers waste hours searching for answers
  • Generic AI chat tools hallucinate facts about your specific policies

The Solution:

  • Grounded answers only - responses use retrieved document chunks, with confidence thresholds
  • Persistent knowledge base - upload once, query forever (unlike ChatGPT's per-conversation uploads)
  • API-first - integrate into Slack, web apps, customer support tools
  • Production-style - typed code, tests, CI/CD, cloud deployment paths

What makes this different from ChatGPT/Claude? See Why Not Just Use ChatGPT? for detailed comparison.


Key Features

Currently Working:

  • Grounded Q&A - Answers only from retrieved chunks, with [doc, page] citations
  • Honest refusal - Returns "not_found" if confidence is too low (no guessing)
  • Query routing - Classifies queries: answer / clarify / refuse based on confidence
  • Duplicate detection - SHA-256 hashing prevents uploading same document twice
  • Two-stage retrieval - Vector search (30 candidates) β†’ Cross-encoder reranking (top 5)
  • HNSW indexing - Fast vector similarity search (10-100x speedup)
  • API key authentication - Secure endpoints with X-API-Key header
  • Swappable LLM - Gemini, Ollama (offline), Azure OpenAI via adapter pattern
  • pgvector - Vector embeddings in PostgreSQL (no separate vector DB)
  • Web UI - Nuxt 4 interface for chat, document management, data extraction

Will be implemented later:

  • Slack Bot integration
  • Structured data extraction backend endpoint (UI exists)
  • Multi-turn chat sessions API (database models exist)
  • MCP integration for Claude Desktop

Quick Start (Local)

Prerequisites:

  • Docker & Docker Compose
  • LLM provider (pick one):
    • Gemini API (free tier) - best quality
    • Ollama (local) - 100% offline, zero cost

Run Backend API:

git clone https://github.com/dinkar1708/askdocs-rag-agent.git
cd askdocs-rag-agent

# Configure LLM provider
cp .env.example .env
# Edit .env: set LLM_PROVIDER=gemini and add your GEMINI_API_KEY
# OR set LLM_PROVIDER=ollama for fully offline mode

# Start backend services
docker compose up --build

# API available at http://localhost:8000
# Swagger UI at http://localhost:8000/docs

Run Web UI (optional):

cd web-ui
npm install
npm run dev

# Web UI available at http://localhost:3000

Note: Slack bot integration is documented but not yet implemented. See Slack Integration Guide for the planned implementation.

Test the service:

  1. Upload a document - POST /documents with a PDF file
  2. Ask a question - POST /ask with {"question": "what is X?"}
  3. Verify grounding - Check the sources array in the response
  4. Try the Web UI - Open http://localhost:3000

Try the demo with sample data:

# Upload sample company policy document
curl -X POST http://localhost:8000/documents/ \
  -H "X-API-Key: test-api-key-not-for-production" \
  -F "file=@app/samples/company_policy.pdf"

# Ask a test question
curl -X POST http://localhost:8000/ask/ \
  -H "Content-Type: application/json" \
  -H "X-API-Key: test-api-key-not-for-production" \
  -d '{"question": "How many vacation days do employees get?"}'

# Expected: "15 days of paid vacation per year" with citations

Note: All API endpoints require X-API-Key header for authentication.

πŸ“‹ Quick Demo: See Getting Started Guide for copy-paste questions with expected answers.

Verify it works:

# Run tests
docker compose exec api pytest

# Check test results
# Expected: 216/218 tests passing (99.1%)

Note: Evaluation harness (retrieval quality metrics) will be implemented later.

See Local Development Guide for detailed setup.


Testing

Total: 240 tests - 216/218 backend passing (99.1%), 24 E2E tests

  • 216/218 Unit/Integration Tests (Backend)

    • API endpoints (health, documents, Q&A)
    • RAG retrieval & reranking
    • LLM adapters (Gemini, Ollama, Azure OpenAI, Mock)
    • Document ingestion & chunking
    • Database operations
    • Query routing
  • 24 End-to-End Tests (Frontend + Backend)

    • Document upload & management
    • Question answering with citations
    • Multi-turn conversations (frontend only, API pending)
    • Source citation verification
    • Data extraction UI (backend endpoint pending)

Run Tests:

# Backend tests (pytest)
docker compose exec api pytest                    # All tests
docker compose exec api pytest app/tests/api/     # API tests only
docker compose exec api pytest -v                 # Verbose output

# E2E tests (Playwright)
cd web-ui
./run-e2e-tests.sh           # Run with mock LLM (fast, for CI/CD)
./run-e2e-with-ollama.sh     # Run with Ollama (real LLM, slower)
npm run test:e2e:ui          # Interactive mode (debug tests)

Test Isolation:

  • E2E tests run on isolated ports (8001/3001) - won't affect development server (8000/3000)
  • Separate test database on port 5433
  • Mock LLM provider for fast, deterministic tests

See Testing Guide for details.


How It Works

Ingestion Pipeline:

PDF β†’ Extract text (with page numbers) β†’ Chunk (512 tokens, 128 overlap)
    β†’ Embed (sentence-transformers) β†’ Store (PostgreSQL + pgvector)

Query Pipeline:

Question β†’ Embed β†’ Vector search (top-k chunks) β†’ LangGraph router
    β”œβ”€ High confidence β†’ Generate answer + citations
    β”œβ”€ Ambiguous β†’ Ask for clarification
    └─ Low confidence β†’ Return "not_found"

Key Design Decisions:

  • Grounded-or-refuse - Trust is the product; never improvise answers
  • LLM adapter pattern - Cloud/model choice is config, not code change
  • pgvector in PostgreSQL - Single database for relational + vector data
  • MCP-first - API endpoints also exposed as AI assistant tools

See Architecture Guide for deep dive.


Project Structure

askdocs-rag-agent/
β”œβ”€β”€ app/                   # Backend (Python/FastAPI)
β”‚   β”œβ”€β”€ api/               # FastAPI routes
β”‚   β”‚   β”œβ”€β”€ documents.py   # Document upload endpoints
β”‚   β”‚   β”œβ”€β”€ questions.py   # Q&A endpoints
β”‚   β”‚   β”œβ”€β”€ slack.py       # Slack webhook endpoints
β”‚   β”‚   └── ...
β”‚   β”œβ”€β”€ services/          # Business logic
β”‚   β”‚   β”œβ”€β”€ slack_bot.py   # Slack bot service
β”‚   β”‚   β”œβ”€β”€ retriever.py   # RAG retrieval
β”‚   β”‚   └── ...
β”‚   β”œβ”€β”€ ingest/            # PDF extraction, chunking, embedding
β”‚   β”œβ”€β”€ graph/             # LangGraph query router
β”‚   β”œβ”€β”€ llm/               # Provider adapters (Gemini/Ollama/Azure)
β”‚   β”œβ”€β”€ mcp/               # MCP server
β”‚   β”œβ”€β”€ db/                # SQLAlchemy models, pgvector setup
β”‚   β”œβ”€β”€ core/              # Config, logging
β”‚   └── tests/             # pytest suites with auto-generated API docs
β”œβ”€β”€ web-ui/                # Frontend (Nuxt 4/Vue/Tailwind)
β”‚   β”œβ”€β”€ app/               # Vue components and pages
β”‚   β”œβ”€β”€ composables/       # Vue composables and API services
β”‚   β”œβ”€β”€ public/            # Static assets
β”‚   └── nuxt.config.ts     # Nuxt configuration
β”œβ”€β”€ docs/
β”‚   β”œβ”€β”€ testing/
β”‚   β”‚   └── api-results/   # Auto-generated API request/response examples
β”‚   β”œβ”€β”€ core/              # Architecture, deployment guides
β”‚   └── interfaces/        # API, Web UI, Slack bot docs
β”œβ”€β”€ samples/               # Sample PDFs for testing
└── docker-compose.yml

Deployment

GCP (Primary):

  • Cloud Run (stateless API, scales to zero)
  • Cloud SQL (PostgreSQL + pgvector)
  • Gemini API / Vertex AI

See Deployment Guide for step-by-step.

Azure (Supported):

  • Azure Container Apps
  • Azure Database for PostgreSQL Flexible Server
  • Azure OpenAI

Brief setup: docs/core/deployment/AZURE.md


Documentation

All documentation is in the /docs folder.

Document Description
Documentation Index Start here - Complete navigation guide
Architecture System design, data flow, key decisions
Development Developer quick reference
Local Setup Detailed setup, testing, debugging
API Guide API integration guide with examples
Web UI Browser interface for end users
Slack Bot Slack integration setup & usage
Configuration Environment variables, tuning
Deployment GCP (detailed), Azure (brief)
Features User-focused feature docs
Security Security guidelines & checklist
Business Sales materials, pricing, ROI
Why This? vs ChatGPT/Claude

Roadmap

  • Core RAG API (ingest, ask, grounded answers)
  • LangGraph router (answer/clarify/refuse)
  • Multi-turn chat with memory
  • MCP server tools
  • Evaluation harness
  • Web UI (Nuxt 4 + Tailwind CSS)
  • User authentication & authorization
  • GCP Cloud Run deployment + CI/CD
  • Azure Container Apps deployment
  • Multi-tenant support
  • Japanese document support

Contributing

Contributions welcome! See CONTRIBUTING.md


License

MIT License - see LICENSE


Author

Dinakar Maurya β€” Solution Architect / AI Engineer, Tokyo


Questions? Open an issue on GitHub.

About

A production-ready RAG (Retrieval-Augmented Generation) system that enables natural language Q&A over document collections with guaranteed citation accuracy. Built with enterprise-grade architecture featuring LangGraph-powered query routing, vector similarity search using pgvector, and multi-LLM support (Gemini, Ollama, Azure OpenAI).

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages