Ask questions to your documents and get grounded, cited answers β a production-ready Document Q&A service built with FastAPI, PostgreSQL+pgvector, and advanced RAG techniques.
Stack: Python 3.12 Β· FastAPI Β· PostgreSQL + pgvector Β· Nuxt 4 Β· Tailwind CSS Β· Docker Β· Ollama (offline LLM)
Latest Update: Production-ready with 216/218 tests passing, HNSW indexing, duplicate detection, and API key authentication
A production-ready RAG (Retrieval-Augmented Generation) system that enables natural language Q&A over document collections with guaranteed citation accuracy. Built with enterprise-grade architecture featuring custom query routing, HNSW-indexed vector similarity search using pgvector, and multi-LLM support (Gemini, Ollama, Azure OpenAI).
Key Differentiators:
- Grounded-or-refuse architecture - Never hallucinates; returns "not_found" when answers aren't in documents
- Citation tracking - Every answer includes exact document and page references
- Duplicate detection - SHA-256 content hashing prevents duplicate uploads
- HNSW indexing - 10-100x faster vector search with PostgreSQL+pgvector
- API key authentication - All endpoints secured with X-API-Key header validation
- Production-ready - 216/218 tests passing (99.1%), cloud deployment docs (GCP/Azure)
- Flexible LLM backend - Swappable providers via adapter pattern (Gemini, Ollama offline, Azure OpenAI)
Target Use Cases: HR knowledge bases, customer support documentation, legal/compliance document search, IT helpdesk automation, sales enablement.
Market Position: Developer-first, self-hosted alternative to enterprise search products (Glean, Writer) β full data control and per-query costs instead of per-seat licensing. See docs/business for cost model.
Chat interface with filters (department, grade, type) and "New Chat" button. Clean welcome message guiding users to ask questions.
Upload PDF documents with optional metadata (department, grade level, document type, tags). Shows uploaded documents list with chunk counts.
Ask questions and get answers with exact source citations (document name, page number, similarity scores). Delete button to remove individual messages or entire chat.
View all uploaded documents with chunk counts and upload dates. Delete documents as needed.
Define custom extraction schemas with field names and types (Text, Number, List). Quick templates for common use cases (Job Posting, Invoice, Resume).
Extracted structured data from documents with confidence scores. Export results as JSON or CSV.
π Quick Demo: Getting Started | Sample Questions | Ollama Local LLM
Testing: Works with Ollama (free, 100% offline) or Gemini (cloud, best quality).
Upload PDF documents β Ask questions in natural language β Get answers grounded in those documents with citations, or an honest "not found."
No hallucinations. Every answer either cites the exact source (document + page) or explicitly says the information doesn't exist in your documents.
Example:
Q: "What is the refund policy?"
A: "Refunds are processed within 14 days of purchase."
Sources: [terms.pdf, page 7]
Q: "What's the weather today?"
A: "not_found - This question cannot be answered from the uploaded documents."
The Problem:
- Organizations have thousands of policy documents, manuals, handbooks
- Employees and customers waste hours searching for answers
- Generic AI chat tools hallucinate facts about your specific policies
The Solution:
- Grounded answers only - responses use retrieved document chunks, with confidence thresholds
- Persistent knowledge base - upload once, query forever (unlike ChatGPT's per-conversation uploads)
- API-first - integrate into Slack, web apps, customer support tools
- Production-style - typed code, tests, CI/CD, cloud deployment paths
What makes this different from ChatGPT/Claude? See Why Not Just Use ChatGPT? for detailed comparison.
Currently Working:
- Grounded Q&A - Answers only from retrieved chunks, with
[doc, page]citations - Honest refusal - Returns "not_found" if confidence is too low (no guessing)
- Query routing - Classifies queries: answer / clarify / refuse based on confidence
- Duplicate detection - SHA-256 hashing prevents uploading same document twice
- Two-stage retrieval - Vector search (30 candidates) β Cross-encoder reranking (top 5)
- HNSW indexing - Fast vector similarity search (10-100x speedup)
- API key authentication - Secure endpoints with X-API-Key header
- Swappable LLM - Gemini, Ollama (offline), Azure OpenAI via adapter pattern
- pgvector - Vector embeddings in PostgreSQL (no separate vector DB)
- Web UI - Nuxt 4 interface for chat, document management, data extraction
Will be implemented later:
- Slack Bot integration
- Structured data extraction backend endpoint (UI exists)
- Multi-turn chat sessions API (database models exist)
- MCP integration for Claude Desktop
Prerequisites:
- Docker & Docker Compose
- LLM provider (pick one):
- Gemini API (free tier) - best quality
- Ollama (local) - 100% offline, zero cost
Run Backend API:
git clone https://github.com/dinkar1708/askdocs-rag-agent.git
cd askdocs-rag-agent
# Configure LLM provider
cp .env.example .env
# Edit .env: set LLM_PROVIDER=gemini and add your GEMINI_API_KEY
# OR set LLM_PROVIDER=ollama for fully offline mode
# Start backend services
docker compose up --build
# API available at http://localhost:8000
# Swagger UI at http://localhost:8000/docsRun Web UI (optional):
cd web-ui
npm install
npm run dev
# Web UI available at http://localhost:3000Note: Slack bot integration is documented but not yet implemented. See Slack Integration Guide for the planned implementation.
Test the service:
- Upload a document -
POST /documentswith a PDF file - Ask a question -
POST /askwith{"question": "what is X?"} - Verify grounding - Check the
sourcesarray in the response - Try the Web UI - Open http://localhost:3000
Try the demo with sample data:
# Upload sample company policy document
curl -X POST http://localhost:8000/documents/ \
-H "X-API-Key: test-api-key-not-for-production" \
-F "file=@app/samples/company_policy.pdf"
# Ask a test question
curl -X POST http://localhost:8000/ask/ \
-H "Content-Type: application/json" \
-H "X-API-Key: test-api-key-not-for-production" \
-d '{"question": "How many vacation days do employees get?"}'
# Expected: "15 days of paid vacation per year" with citationsNote: All API endpoints require X-API-Key header for authentication.
π Quick Demo: See Getting Started Guide for copy-paste questions with expected answers.
Verify it works:
# Run tests
docker compose exec api pytest
# Check test results
# Expected: 216/218 tests passing (99.1%)Note: Evaluation harness (retrieval quality metrics) will be implemented later.
See Local Development Guide for detailed setup.
Total: 240 tests - 216/218 backend passing (99.1%), 24 E2E tests
-
216/218 Unit/Integration Tests (Backend)
- API endpoints (health, documents, Q&A)
- RAG retrieval & reranking
- LLM adapters (Gemini, Ollama, Azure OpenAI, Mock)
- Document ingestion & chunking
- Database operations
- Query routing
-
24 End-to-End Tests (Frontend + Backend)
- Document upload & management
- Question answering with citations
- Multi-turn conversations (frontend only, API pending)
- Source citation verification
- Data extraction UI (backend endpoint pending)
Run Tests:
# Backend tests (pytest)
docker compose exec api pytest # All tests
docker compose exec api pytest app/tests/api/ # API tests only
docker compose exec api pytest -v # Verbose output
# E2E tests (Playwright)
cd web-ui
./run-e2e-tests.sh # Run with mock LLM (fast, for CI/CD)
./run-e2e-with-ollama.sh # Run with Ollama (real LLM, slower)
npm run test:e2e:ui # Interactive mode (debug tests)Test Isolation:
- E2E tests run on isolated ports (8001/3001) - won't affect development server (8000/3000)
- Separate test database on port 5433
- Mock LLM provider for fast, deterministic tests
See Testing Guide for details.
Ingestion Pipeline:
PDF β Extract text (with page numbers) β Chunk (512 tokens, 128 overlap)
β Embed (sentence-transformers) β Store (PostgreSQL + pgvector)
Query Pipeline:
Question β Embed β Vector search (top-k chunks) β LangGraph router
ββ High confidence β Generate answer + citations
ββ Ambiguous β Ask for clarification
ββ Low confidence β Return "not_found"
Key Design Decisions:
- Grounded-or-refuse - Trust is the product; never improvise answers
- LLM adapter pattern - Cloud/model choice is config, not code change
- pgvector in PostgreSQL - Single database for relational + vector data
- MCP-first - API endpoints also exposed as AI assistant tools
See Architecture Guide for deep dive.
askdocs-rag-agent/
βββ app/ # Backend (Python/FastAPI)
β βββ api/ # FastAPI routes
β β βββ documents.py # Document upload endpoints
β β βββ questions.py # Q&A endpoints
β β βββ slack.py # Slack webhook endpoints
β β βββ ...
β βββ services/ # Business logic
β β βββ slack_bot.py # Slack bot service
β β βββ retriever.py # RAG retrieval
β β βββ ...
β βββ ingest/ # PDF extraction, chunking, embedding
β βββ graph/ # LangGraph query router
β βββ llm/ # Provider adapters (Gemini/Ollama/Azure)
β βββ mcp/ # MCP server
β βββ db/ # SQLAlchemy models, pgvector setup
β βββ core/ # Config, logging
β βββ tests/ # pytest suites with auto-generated API docs
βββ web-ui/ # Frontend (Nuxt 4/Vue/Tailwind)
β βββ app/ # Vue components and pages
β βββ composables/ # Vue composables and API services
β βββ public/ # Static assets
β βββ nuxt.config.ts # Nuxt configuration
βββ docs/
β βββ testing/
β β βββ api-results/ # Auto-generated API request/response examples
β βββ core/ # Architecture, deployment guides
β βββ interfaces/ # API, Web UI, Slack bot docs
βββ samples/ # Sample PDFs for testing
βββ docker-compose.yml
GCP (Primary):
- Cloud Run (stateless API, scales to zero)
- Cloud SQL (PostgreSQL + pgvector)
- Gemini API / Vertex AI
See Deployment Guide for step-by-step.
Azure (Supported):
- Azure Container Apps
- Azure Database for PostgreSQL Flexible Server
- Azure OpenAI
Brief setup: docs/core/deployment/AZURE.md
All documentation is in the /docs folder.
| Document | Description |
|---|---|
| Documentation Index | Start here - Complete navigation guide |
| Architecture | System design, data flow, key decisions |
| Development | Developer quick reference |
| Local Setup | Detailed setup, testing, debugging |
| API Guide | API integration guide with examples |
| Web UI | Browser interface for end users |
| Slack Bot | Slack integration setup & usage |
| Configuration | Environment variables, tuning |
| Deployment | GCP (detailed), Azure (brief) |
| Features | User-focused feature docs |
| Security | Security guidelines & checklist |
| Business | Sales materials, pricing, ROI |
| Why This? | vs ChatGPT/Claude |
- Core RAG API (ingest, ask, grounded answers)
- LangGraph router (answer/clarify/refuse)
- Multi-turn chat with memory
- MCP server tools
- Evaluation harness
- Web UI (Nuxt 4 + Tailwind CSS)
- User authentication & authorization
- GCP Cloud Run deployment + CI/CD
- Azure Container Apps deployment
- Multi-tenant support
- Japanese document support
Contributions welcome! See CONTRIBUTING.md
MIT License - see LICENSE
Dinakar Maurya β Solution Architect / AI Engineer, Tokyo
- GitHub: @dinkar1708
- Medium: @dinkar1708
- LinkedIn: in/dinkar1708
Questions? Open an issue on GitHub.





