Live Demo: https://smartdocq.vercel.app
SmartDocQ is a full-stack document intelligence platform for uploading, indexing, searching, editing, and querying documents. It combines semantic vector retrieval, BM25 lexical search, and Reciprocal Rank Fusion (RRF), with a task-aware LLM routing layer for generation and provider fallback.
- AI-Powered Chat: Hybrid RAG-based question answering using Vector Search + BM25 + RRF Fusion.
- Interactive Spreadsheet Editing: Edit CSV and Excel (XLSX) documents directly in the browser. SmartDocQ incrementally synchronizes affected table and paragraph chunks without rebuilding the entire document index.
- Quiz Generation: Automatic creation of multiple-choice, true/false, and short-answer questions from document content.
- Flashcard Creation: Smart extraction of key concepts and definitions for effective learning and revision.
- Text Summarization: Concise summaries of document content for quick comprehension.
- Contextual document embeddings: Prepend document title, hierarchical headings, and page ranges before embedding, improving retrieval quality while preserving the original chunk text for generation.
- Multi-Format Support: Upload and process PDF, DOCX, TXT, CSV, and XLSX files.
- Token-Aware Chunking: Heading-aware section tracing and token bounds packing with tiktoken.
-
Hybrid Retrieval: Dense vector search + version-isolated BM25 lexical search combined with Reciprocal Rank Fusion (RRF,
$k=60$ ). - Spreadsheet-Aware Indexing: Cell-level incremental sync updates affected semantic vectors and lexical indexes without full re-indexing.
- Atomic Shadow Indexing: Versioned builds with compare-and-swap (CAS) activation and automated rollback.
- Index-State Bloom Filter: Process-local Counting Bloom filter for non-authoritative fast negative rejection.
- HTTP-only Cookie Authentication: JWT session tokens are stored in HTTP-only cookies and validated against server-side session state, with role-based access control.
- CSRF Protection: State-changing requests use CSRF protection with Origin/Referer validation.
- Server-Side Input Validation: API inputs are validated before business logic or database access.
- Authentication & Rate Limiting: Authentication and sensitive endpoints use dedicated rate limits and authorization controls.
- Internal AI Service Authentication: Browser clients cannot directly access Flask; protected AI routes require a shared
SERVICE_TOKEN. - Sensitive-Data Detection & Consent: Detects sensitive information such as Aadhaar, PAN, credit cards, emails, and phone numbers and requires consent before processing.
- Prompt Injection & Context Sanitization: User questions are checked against injection/jailbreak heuristics, while retrieved document content is treated as untrusted input and sanitized before LLM processing.
- Document Hashing & Deduplication: SHA-256 document fingerprints prevent unnecessary duplicate processing.
Detailed security architecture, implementation, threat considerations, and testing are documented in
SECURITY.md.
- User Management: Comprehensive admin dashboard for user oversight and role assignment.
- Document Analytics: Track document uploads, processing status, and usage statistics.
- Report Management: Handle user feedback and support inquiries efficiently.
- System Monitoring: Structured logging of HTTP requests and security events (Pino), Prometheus metrics counters, health checks, and background watchdog maintenance jobs.
graph TD
%% Styles
classDef client fill:#1E293B,stroke:#38BDF8,stroke-width:2px,color:#F8FAFC;
classDef business fill:#064E3B,stroke:#34D399,stroke-width:2px,color:#F0FDF4;
classDef aiservice fill:#1E1B4B,stroke:#818CF8,stroke-width:2px,color:#EEF2FF;
classDef storage fill:#451A03,stroke:#F59E0B,stroke-width:2px,color:#FFFBEB;
classDef external fill:#14532D,stroke:#4ADE80,stroke-width:2px,color:#F0FDF4;
%% Client Layer
subgraph Client_Layer ["Presentation Layer"]
ReactSPA["React SPA<br/>(i18next / GSAP / Lottie)"]:::client
end
%% Business Logic Layer
subgraph Middleware_Layer ["Business Logic Layer (Express Server)"]
ExpressRouter["Express API Router<br/>(Session & Auth Validation)<br/>(CSRF Protection)<br/>(Rate Limiting)<br/>(Structured Logging)<br/>(Service Token Proxy)"]:::business
AuthGuard["Auth & Session Middleware<br/>(JWT httpOnly Cookie + UserSession Validation)"]:::business
ZodValidator["Input Validation<br/>(Zod Schemas)"]:::business
MongooseDB["Mongoose ODM<br/>(User, Document, Chat, DocChunk models)"]:::business
end
%% AI Processing Layer
subgraph AI_Layer ["AI Processing Layer (Flask Service)"]
FlaskApp["Flask API Router<br/>(/api/index-from-atlas, /api/document/ask)"]:::aiservice
Parser["Document Parser & Table Extractor<br/>(PyMuPDF4LLM → Markdown → Block Parser → Chunker)"]:::aiservice
IndexSynchronizer["Index Synchronizer<br/>(Incremental Cell Edit Sync)"]:::aiservice
ShadowVersionManager["Shadow Version Manager<br/>(CAS Activation & Rollbacks)"]:::aiservice
RetrievalPipeline["Hybrid Retrieval Engine<br/>(Vector Search + BM25 + RRF)"]:::aiservice
Sanitizer["Prompt Injection Sanitizer<br/>(sanitize_context)"]:::aiservice
LLMRouter["LLM Router<br/>(Task-aware Routing)<br/>(Provider Fallback)<br/>(Model Fallback)<br/>(Timeout Budgeting)"]:::aiservice
end
%% Storage Layer
subgraph Storage_Layer ["Data & Storage Layer"]
MongoDB["MongoDB Atlas (Cloud)<br/>(Accounts, Metadata, Binary Files, Chunks)"]:::storage
ChromaDB["ChromaDB (Local Disk)<br/>(Vector embeddings & metadata)"]:::storage
BM25Cache["BM25 Index (In-Memory)<br/>(Tokenized Lexical Cache)"]:::storage
end
%% External
subgraph External_APIs ["External API Layer"]
LLM_APIs["Multi-Provider LLM APIs<br/>(Gemini / Groq / Cerebras)"]:::external
end
%% Flow/Connections
ReactSPA <-->|"HTTPS API Calls<br/>(JWT Cookie + X-CSRF-Token)"| AuthGuard
AuthGuard --> ZodValidator
ZodValidator --> ExpressRouter
ExpressRouter <-->|"CRUD Operations"| MongooseDB
MongooseDB <-->|"TCP / Driver"| MongoDB
ExpressRouter -->|"Server-to-Server POST<br/>(Service Token Auth)"| FlaskApp
FlaskApp --> Parser
Parser --> ShadowVersionManager
FlaskApp --> IndexSynchronizer
IndexSynchronizer --> ShadowVersionManager
ShadowVersionManager --> RetrievalPipeline
RetrievalPipeline --> Sanitizer
FlaskApp <-->|"Download Document Binary"| ExpressRouter
FlaskApp <-->|"Vector query / write"| ChromaDB
RetrievalPipeline <-->|"Lexical query"| BM25Cache
FlaskApp --> LLMRouter
LLMRouter <-->|"HTTPS / REST"| LLM_APIs
- React 19: SPA component architecture with React Router 7, i18next, GSAP, Lottie, and Focus Trap React.
- Node.js & Express 5: API server, Mongoose 8 ODM, Pino logging, Helmet security, compression, and express-rate-limit.
- Authentication: JWT stored in an
httpOnlycookie verified against server-side session state (UserSession).
- Flask 3: Microservice for document processing, RAG pipelines, and LLM routing.
- LLM Router: Task-aware provider routing with provider and model fallback.
-
Vector & Lexical Retrieval: ChromaDB (
gemini-embedding-2), in-memory BM25 lexical search, and Reciprocal Rank Fusion (RRF,$k=60$ ).
- PyMuPDF4LLM & PyMuPDF (fitz): Layout-aware structural Markdown extraction.
- PyPDF2: Backup plain-text PDF parser.
- python-docx & openpyxl: Microsoft Word and Excel spreadsheet parsing.
- tiktoken: BPE token packing and section chunking bounds.
- MongoDB Atlas: User accounts, document metadata, file chunks, and server session state.
- ChromaDB: Local persistent vector database for document embeddings.
SmartDocQ processes PDF documents through a multi-stage indexing pipeline:
flowchart TD
A([PDF Upload])
--> B["Three-tier Extraction Chain\n(PyMuPDF4LLM → PyMuPDF → PyPDF2 fallback)"]
--> C["Markdown Normalization\n(Bulleted fixes, noise removals, line merging)"]
--> D["Extensible Block Parsing\n(Paragraph, List, Table, Code, Blockquote blocks)"]
--> E["Heading Extraction\n(H1–H5 nested path extraction)"]
--> F["Section-aware Chunking\n(Isolated table/code chunks, snapped text bounds)"]
--> G["Token-aware Packing\n(tiktoken bounds packing with overlap bounds mapping)"]
--> H["Contextual Headers\n(Prepending Document, Section, Subsection, Page Range)"]
--> I["Gemini Embeddings\n(models/gemini-embedding-2)"]
--> J[("ChromaDB Storage\nClean text documents + detailed metadata")]
SmartDocQ uses versioned shadow indexes to safely evolve embeddings and indexing logic without interrupting retrieval.
Each generation tracks the embedding model, pipeline version, chunking version, source file hash, and indexing timestamp. New generations are built independently and activated using compare-and-swap (CAS). Failed builds leave the active generation unchanged, while stale indexing jobs are recovered in the background.
Retrieval is always performed against the currently active index generation, ensuring background reindexing never interrupts user queries.
SmartDocQ uses a Hybrid RAG pipeline that combines:
- Semantic vector retrieval (ChromaDB + gemini-embedding-2)
- Version-isolated BM25 lexical retrieval with in-memory caching
- Reciprocal Rank Fusion (RRF)
- Table-aware ranking
- Contextual document embeddings (Document, Section, Subsection, Page Ranges)
This approach improves both semantic understanding and exact-match retrieval for identifiers, spreadsheet data, and structured documents.
SmartDocQ includes separate empirical benchmark suites for retrieval quality, index decision path performance, and PDF extraction:
- Hybrid Dense + BM25 + RRF (Baseline): nDCG@5 = 0.8294, MRR@5 = 0.8092, Recall@5 = 0.9146, Retrieval Latency = 149.09 ms.
- RRF + BGE Reranker (
BAAI/bge-reranker-base): nDCG@5 = 0.7337 (-11.53%), MRR@5 = 0.7107 (-12.18%), Recall@5 = 0.8331 (-8.91%), Retrieval Latency = 15,373.12 ms (+15.24s cross-encoder latency). - Candidate Pool Recall: 0.9867 (relevant documents were already captured in top RRF candidates prior to reranking).
- Decision: Cross-encoder reranking decreased retrieval metrics while adding substantial latency; baseline Hybrid RRF is enabled by default.
- Observed Sample False-Positive Rate: 0.00% (0/1,000 false positives; exact 95% upper confidence bound
< 0.30%). - Rejection Latency: 0.0052 ms (Bloom ON) vs 0.9613 ms (Bloom OFF) median lookup latency for absent documents.
- P95 Tail Latency: Up to 10.42x speedup on negative lookups.
Detailed benchmark methodology, configurations, and reports:
- Information Retrieval & Reranker Report
- Index Bloom Filter Benchmark Report
- PDF Extractor Benchmark Overview
To set up SmartDocQ locally, you'll need:
- Node.js: Version 20.x or higher
- Python: Version 3.9 or higher
- MongoDB: Local installation or MongoDB Atlas account
- AI Provider API Keys: Gemini, Groq, and Cerebras API credentials for LLM generation
- Git: Version control
git clone https://github.com/SmartDocQ/SmartDocQ.git
cd SmartDocQcd servers
npm install
# Create .env file with the following variables:
# PORT=5000
# MONGO_URI=your_mongodb_connection_string
# JWT_SECRET=your_jwt_secret_key
# FRONTEND_ORIGINS=http://localhost:3000
# DNS_SERVERS=1.1.1.1,8.8.8.8
# SERVICE_TOKEN=shared_strong_secret
# FLASK_ASK_URL=http://localhost:5001/api/document/ask
# FLASK_INDEX_URL=http://localhost:5001/api/index-from-atlas
# FLASK_CONVERT_URL=http://localhost:5001/api/convert/word-to-pdf
# MAX_UPLOAD_SIZE_MB=15
# MAIL_USER=your_gmail_address (required in development only)
# MAIL_PASS=your_gmail_app_password (required in development only)
# BREVO_API_KEY=your_brevo_api_key (required in production only)
# BREVO_SENDER_EMAIL=your_brevo_sender_email (required in production only)
# BREVO_SENDER_NAME=your_brevo_sender_name (optional in production only)
npm startReact communicates only with Node.js. The Flask service is intended for internal server-to-server communication and should not be called directly by browser clients.
cd ../backend
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
pip install -r requirements.txt
# Create .env file with:
# PORT=5001
# FRONTEND_ORIGINS=http://localhost:3000
# NODE_BASE_URL=http://localhost:5000
# SERVICE_TOKEN=shared_strong_secret (must be identical to the SERVICE_TOKEN in servers/.env)
# GEMINI_API_KEY=your_google_ai_api_key
# GROQ_API_KEY=your_groq_api_key
# CEREBRAS_API_KEY=your_cerebras_api_key
# INDEX_BATCH_SIZE=64
# MAX_UPLOAD_SIZE_MB=15
# Optional chunking configurations:
# CHUNK_TARGET_TOKENS=512
# CHUNK_SOFT_LIMIT=600
# CHUNK_HARD_LIMIT=800
# CHUNK_OVERLAP_TOKENS=80
# IGNORE_REFERENCE_SECTIONS=True
python main.pycd ../my-app
npm install
# Create .env file with:
# REACT_APP_API_URL=http://localhost:5000
# REACT_APP_GOOGLE_CLIENT_ID=your_google_oauth_client_id
npm startRun the Python test suite (Flask AI service):
cd backend
# Set SERVICE_TOKEN (PowerShell: $env:SERVICE_TOKEN="dev-token")
python -m pytest tests/ -vSmartDocQ includes built-in latency instrumentation for every retrieval request.
Captured metrics include:
- embedding latency
- Chroma retrieval latency
- BM25 latency
- RRF fusion latency
- LLM latency
- LLM provider/model selection
- fallback activation and fallback reason
- total request latency
Additional backend metrics include:
- HTTP request latency histograms
- Node.js process metrics
- Session auto-upgrade counter
- CSRF refresh counter
Internal timings are logged server-side while remaining hidden from client responses.
Thanks to all the contributors who have helped build SmartDocQ:
Dr-Venom29 |
ANIRUDH-7600 |
Sameeksha270905 |
ananya-1507 |
srithi-05 |