Comprehensive system architecture and design for OpenWebAnswer.
- Overview
- System Components
- Data Flow
- Technology Stack
- Database Schema
- Caching Strategy
- Vector Store Architecture
- API Design
- Background Jobs
- Error Handling & Resilience
OpenWebAnswer is a production-grade Question-Answering system that combines:
- Real-time web search with semantic relevance
- Intelligent caching to reduce redundant searches
- Quality scoring to filter low-quality sources
- Followup question generation to enhance user engagement
- Comprehensive analytics for system monitoring
| Feature | Purpose | Implementation |
|---|---|---|
| Semantic Search | Find relevant documents via embeddings | FAISS + All-MiniLM-L6-v2 (384-dim) |
| Smart Caching | Reduce API calls and improve response time | Multi-level with SQLite + FAISS |
| Quality Filtering | Exclude low-quality sources | Domain authority + content analysis |
| Followup Generation | Suggest related questions | Template + embedding similarity |
| Analytics Tracking | Monitor performance and quality | QueryAnalytics table + insights |
| Background Jobs | System maintenance | APScheduler (7 scheduled jobs) |
File: backend/api/routes/
api/
├── query.py # Main QA endpoint
├── health.py # Health status check
├── analytics.py # Analytics & monitoring endpoints
└── schemas.py # Request/response models
Endpoints:
POST /api/query→ QA with cachingGET /api/health→ System healthGET /api/analytics/*→ 9 monitoring endpoints
File: backend/pipelines/realtime_qa.py
Orchestrates the full QA flow:
Request → Cache Check → Web Search → Quality Scoring → Answer Generation
↓ ↓ ↓ ↓
(Cache Hit?) (Fetch Docs) (Filter Low-Q) (Generate + Followups)
↓ ↓ ↓ ↓
Return Cached Cache+Index Score Track AnalyticsKey Methods:
run()- Main QA pipeline with caching_search()- Web search integration_fetch_documents()- Document retrieval with quality scoring_generate_answer()- LLM-based answer generation
File: backend/database/
database/
├── schema.sql # Database tables & indexes
├── models.py # SQLAlchemy ORM models
├── init_db.py # Initialization & migrations
└── repositories/
├── document_repo.py # Document CRUD (11 methods)
├── chunk_repo.py # Chunk CRUD (13 methods)
├── query_repo.py # Query CRUD (11 methods)
├── analytics_repo.py # Analytics CRUD (14 methods)
└── followup_repo.py # Followup CRUD (10 methods)
Tables (6 total):
- documents - Cached web documents (metadata only, content in chunks)
- document_chunks - Document text chunks for searching (single source of truth for chunk_text)
- queries - Cached query results with source tracking
- query_sources - Maps queries to specific source chunks used in answers
- query_analytics - Performance metrics & timing breakdowns (per query_id)
- followup_questions - Generated followup suggestions with question types
File: indexer/
indexer/
├── vector_store_manager.py # FAISS management
├── vector_store.py # Persistence & integrity
└── vector_stores/
└── faiss_store.py # FAISS implementation
Features:
- FAISS index for efficient similarity search
- JSON metadata sidecar for ID mapping
- Integrity verification & rebuilding
- Statistics and diagnostics
File: cache/
cache/
├── cache_manager.py # Multi-level caching
├── deduplication.py # Query & document deduplication
├── invalidation.py # TTL-based cache expiration
├── query_analyzer.py # Cache hit prediction
└── models.py # Cache data structures
Cache Levels:
- Query Cache - Full QA responses (7-day TTL)
- Document Cache - Fetched documents (30-day TTL)
- Vector Cache - Embeddings in FAISS (synced with docs)
File: search_fetch/source_scorer.py
Evaluates source quality across 3 dimensions:
Quality Score = Domain Authority (40%) + Content Quality (35%) + Freshness (25%)
Range: 0.0 - 1.0
Threshold: >= 0.3 (30%)
Components:
- Domain Authority: Trusted (+0.4), Unknown (+0.1), Blacklist (-0.5)
- Content Quality: Length, structure, keyword presence (0.0-0.3)
- Freshness: Recent (+0.3), Old (-0.1)
File: llm_interface/
llm_interface/
├── llm_engine.py # Multi-provider LLM support
├── followup_generator.py # Follow-up question generation
├── prompts.py # System prompts & templates
└── models.py # Response models
LLM Providers: OpenAI, Anthropic, Ollama, HuggingFace
Followup Generation:
- 4 question types: expansion, clarification, related, deeper
- 16 templates total (4 × 4 categories)
- Relevance scoring via embedding similarity
File: backend/analytics/tracker.py & backend/api/routes/analytics.py
Tracks metrics and generates insights:
Tracked Metrics:
- Query frequency & patterns
- Response time (P50, P95, P99)
- Cache hit rate (%)
- Document quality scores
- Error rates
Health Score (0-100):
Health = (CacheHitRate × 30%) + (ResponseTime × 30%) +
(DocQuality × 20%) + (SuccessRate × 20%)
File: backend/jobs/
jobs/
├── cleanup_scheduler.py # Job implementations
└── scheduler.py # APScheduler registration
7 Scheduled Jobs:
- Cache cleanup (every 6h)
- Query cleanup (daily 2am)
- Document cleanup (weekly Sun 2am)
- Analytics cleanup (weekly Sun 3am)
- Vector index rebuild (weekly Sun 4am)
- Database vacuum (weekly Sun 5am)
- Health check (hourly)
┌─────────────────────────────────────────────────────────────────┐
│ User Submits Question │
│ POST /api/query │
└─────────────────┬───────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────┐
│ Step 0: Check Query Cache │
│ - Hash question with MD5 │
│ - Look up in SQLite queries table │
│ - Check expiration (7 days) │
└──────┬──────────────────────────────────┬──────────────────────┘
│ │
CACHE HIT CACHE MISS
(40-50%) (50-60%)
│ │
▼ ▼
┌──────────────────────┐ ┌─────────────────────────────────────┐
│ Return Cached │ │ Step 1: Web Search │
│ - Response │ │ - Query search engine │
│ - Sources │ │ - Get top N results │
│ - Followups │ │ - Return URLs + snippets │
│ - Analytics (hit) │ │ │
└────────┬─────────────┘ │ │
│ └──────────┬──────────────────────────┘
│ │
│ ▼
│ ┌────────────────────────────────────────┐
│ │ Step 2: Fetch & Score Documents │
│ │ - Download full content │
│ │ - Calculate quality score: │
│ │ * Domain authority │
│ │ * Content quality │
│ │ * Freshness │
│ │ - Filter low quality (< 0.3) │
│ │ - Rerank by relevance │
│ └────────────┬─────────────────────────┘
│ │
│ ▼
│ ┌────────────────────────────────────────┐
│ │ Step 2.5: Cache Documents │
│ │ - Embed with All-MiniLM-L6-v2 │
│ │ - Save to SQLite documents table │
│ │ - Index in FAISS vector store │
│ │ - Save chunks for future search │
│ └────────────┬─────────────────────────┘
│ │
│ ▼
│ ┌────────────────────────────────────────┐
│ │ Step 3: Generate Answer │
│ │ - Build context from documents │
│ │ - Call LLM (OpenAI/Anthropic/etc.) │
│ │ - Generate structured answer │
│ └────────────┬─────────────────────────┘
│ │
│ ▼
│ ┌────────────────────────────────────────┐
│ │ Step 3.5: Generate Followups │
│ │ - Template-based generation │
│ │ - Score relevance via embeddings │
│ │ - Filter duplicates │
│ │ - Return top 3 questions │
│ └────────────┬─────────────────────────┘
│ │
│ ▼
│ ┌────────────────────────────────────────┐
│ │ Step 5: Cache Response │
│ │ - Save to queries table │
│ │ - Set TTL: 7 days │
│ │ - Record source_type (cache|web) │
│ │ - Track used chunks in query_sources │
│ │ - Store model_used for analysis │
│ └────────────┬─────────────────────────┘
│ │
│ ▼
│ ┌────────────────────────────────────────┐
│ │ Step 6: Track Analytics │
│ │ - Record response time │
│ │ - Log cache miss reason │
│ │ - Average quality score │
│ │ - Success/failure status │
│ └────────────┬─────────────────────────┘
│ │
└─────────────┬───────────┘
│
▼
┌─────────────────────────────────────────────────────────────────┐
│ Return QueryResponse │
│ { │
│ "id": "query_abc123", │
│ "question": "What is Python?", │
│ "answer": "Python is a...", │
│ "sources": [{url, title, quality_score, ...}, ...], │
│ "followups": ["What are Python applications?", ...], │
│ "response_time_ms": 2500, │
│ "cache_hit": false, │
│ "average_quality_score": 0.85 │
│ } │
└─────────────────────────────────────────────────────────────────┘
| Layer | Technology | Version | Purpose |
|---|---|---|---|
| Framework | FastAPI | 0.109+ | REST API & async support |
| Database | SQLite 3 | 3.50+ | Persistent cache storage |
| ORM | SQLAlchemy | 2.0+ | Database abstraction |
| Vector DB | FAISS | 1.7+ | Similarity search |
| Embeddings | Sentence Transformers | 2.2+ | Text embeddings (384-dim) |
| LLM | Multiple | - | GPT-4, Claude, Ollama |
| Scheduler | APScheduler | 3.10+ | Background jobs |
| Server | Uvicorn | 0.24+ | ASGI server |
| Component | Technology | Purpose |
|---|---|---|
| Framework | React 18 | UI components |
| Build Tool | Vite | Fast bundling |
| API Client | Axios | HTTP requests |
| Styling | Tailwind CSS | Styling |
| State | React Hooks | State management |
| Component | Technology | Purpose |
|---|---|---|
| Containerization | Docker | Application packaging |
| Orchestration | Docker Compose | Multi-container setup |
| Reverse Proxy | Nginx | Load balancing |
| Server | Uvicorn | Python ASGI |
documents (web document metadata)
├── id: INTEGER PRIMARY KEY
├── url: TEXT UNIQUE
├── title: TEXT
├── domain: TEXT (indexed)
├── quality_score: FLOAT (0.0-1.0)
├── is_archived: BOOLEAN
├── chunk_count: INTEGER
└── indexes: [domain], [quality_score], [created_at]
document_chunks (embedding context - SINGLE SOURCE OF TRUTH)
├── id: INTEGER PRIMARY KEY
├── doc_id: INTEGER FK
├── chunk_text: TEXT (ONLY source for chunk text)
├── embedding_id: TEXT UNIQUE (FAISS reference)
├── token_count: INTEGER
├── position: INTEGER (order in document)
└── indexes: [doc_id], [embedding_id], [doc_id, position]
queries (cached query results with source tracking)
├── id: INTEGER PRIMARY KEY
├── query_text: TEXT
├── query_hash: TEXT UNIQUE (deduplication)
├── answer: TEXT
├── response_time_ms: INTEGER
├── source_type: TEXT (cache|vector|web)
├── model_used: TEXT (huggingface|openai|etc)
├── expires_at: TIMESTAMP (7-day TTL)
└── indexes: [query_hash], [source_type], [expires_at]
query_sources (NEW - Phase 7: Track answer sources)
├── id: INTEGER PRIMARY KEY
├── query_id: INTEGER FK
├── doc_id: INTEGER FK
├── chunk_id: INTEGER FK
├── position_in_answer: INTEGER (order in answer)
└── indexes: [query_id], [doc_id], [chunk_id]
query_analytics (performance metrics per query_id)
├── id: INTEGER PRIMARY KEY
├── query_id: INTEGER FK (indexed)
├── vector_search_time_ms: INTEGER
├── web_search_time_ms: INTEGER
├── embedding_time_ms: INTEGER
├── llm_generation_time_ms: INTEGER
├── chunk_retrieval_time_ms: INTEGER
├── total_time_ms: INTEGER
├── cache_hit_count: INTEGER
├── quality_score: FLOAT
└── indexes: [query_id], [created_at DESC]
followup_questions (engagement with question types)
├── id: INTEGER PRIMARY KEY
├── query_id: INTEGER FK
├── followup_text: TEXT
├── question_type: TEXT (expansion|clarification|related|deeper)
├── relevance_score: FLOAT (0.0-1.0)
└── indexes: [query_id], [relevance_score]-
Document Content Storage: Moved from documents table to document_chunks
- documents table: Metadata only (URL, title, domain, quality_score, is_archived)
- document_chunks: Single source of truth for chunk_text
- Rationale: Enables efficient chunk-level searching and deduplication
-
Source Tracking: New query_sources table
- Maps each query to specific chunks used in generating the answer
- Enables complete source attribution and verification
- Tracks position_in_answer for reconstruction of source order
-
Query Analytics Redesign: Changed from query_hash to query_id based
- Now tracks detailed timing breakdowns per query execution
- Separate fields: vector_search_time_ms, web_search_time_ms, embedding_time_ms, llm_generation_time_ms
- Enables performance bottleneck identification
-
Deduplication: Hash-based (query_hash, content_hash where applicable)
- Prevents duplicate web searches for identical queries
- Enables cache reuse across similar queries
-
TTL-Based Expiration: Automatic cleanup via scheduled jobs
- Queries: 7-day TTL
- Documents: 30-day TTL
- Analytics: 90-day TTL
-
Question Type Classification
- FollowupQuestion.question_type: expansion, clarification, related, deeper
- Enables categorization and analysis of followup engagement
-
Cascading Deletes: Document deletion cascades to chunks and query_sources
- Maintains referential integrity
- Automatic cleanup when documents expire
┌─────────────────────────────────────────────────────────────┐
│ Query Cache (Level 1) │
│ - Full QA responses │
│ - SQLite queries table │
│ - 7-day TTL │
│ - Hit rate: 40-50% (identical queries) │
└─────────────────────────────────────────────────────────────┘
↓ (miss)
┌─────────────────────────────────────────────────────────────┐
│ Document Cache (Level 2) │
│ - Fetched web documents │
│ - SQLite documents table │
│ - 30-day TTL │
│ - Hit rate: 30-40% (related queries) │
│ - Includes: URL, title, content, quality score │
└─────────────────────────────────────────────────────────────┘
↓ (miss)
┌─────────────────────────────────────────────────────────────┐
│ Vector Cache (Level 3) │
│ - FAISS index of document embeddings │
│ - JSON metadata sidecar │
│ - Enables semantic search without web fetch │
│ - Hit rate: 50%+ (similar queries) │
└─────────────────────────────────────────────────────────────┘
↓ (miss)
┌─────────────────────────────────────────────────────────────┐
│ Web Search (Miss) │
│ - Call external search engine │
│ - Fetch full documents │
│ - All three cache levels populated │
└─────────────────────────────────────────────────────────────┘
Strategy: TTL-based (lazy deletion)
# Query Cache: Expires after 7 days
expires_at = datetime.utcnow() + timedelta(days=7)
# Document Cache: Expires after 30 days
expires_at = datetime.utcnow() + timedelta(days=30)
# Cleanup Job: Runs daily, removes expired entries
# Reclaims ~50-70% of old data per weekPerformance Impact:
- Cache hits: 1-50ms (no web search)
- Cache misses: 2-10 seconds (web search + LLM)
- Average response time: 2-3 seconds (with 45% cache hit rate)
VectorStore (FAISS)
├── Index Type: HNSW32
│ └── Efficient for ~1M documents
│ - Query time: 10-100ms
│ - Memory: ~4GB per 1M vectors
│
├── Dimension: 384 (All-MiniLM-L6-v2)
│ └── Trade-off: quality vs. memory
│ - 384-dim: Balanced (recommended)
│ - 768-dim: Better quality (more memory)
│
└── Metadata Sidecar (JSON)
└── Maps embedding_id → document_id, chunk_index
- Used for document retrieval after searchUser Query
↓
Embed with All-MiniLM-L6-v2 (384-dim)
↓
FAISS search (k=50, nprobe=32)
↓
Score results via relevance ranking
↓
Filter by quality score (threshold: 0.3)
↓
Return top-k (default: 5) documents
Periodic Rebuild: Weekly (Sunday 4am)
- Compacts index structure
- Improves search performance
- Reclaims memory from deleted vectors
- Takes 5-30 minutes for 1M vectors
Request:
POST /api/query
{
"question": "What is machine learning?",
"mode": "hybrid",
"top_k": 5,
"stream": false
}Response:
{
"id": "query_abc123",
"question": "What is machine learning?",
"answer": "Machine learning is...",
"sources": [
{
"url": "https://example.com",
"title": "ML Guide",
"snippet": "...",
"relevance_score": 0.95,
"quality_score": 0.85,
"quality_reason": "High authority domain"
}
],
"followups": [
"What are supervised learning algorithms?",
"How does neural networks work?"
],
"mode": "realtime",
"response_time_ms": 2534.5,
"sources_count": 5,
"average_quality_score": 0.82,
"cache_hit": false
}9 Endpoints:
GET /api/analytics/insights- Top queries & patternsGET /api/analytics/top-queries- Most frequent questionsGET /api/analytics/slowest-queries- Performance bottlenecksGET /api/analytics/performance- System metricsGET /api/analytics/documents- Document statisticsGET /api/analytics/health- System health scoreGET /api/analytics/cache-hit-rate- Cache efficiencyPOST /api/analytics/rate-query- User ratings (1-5)GET /api/analytics/report- Comprehensive report
Every Hour: Health Check
Every 6 Hours: Cache Cleanup
Daily 2 AM: Query Cleanup (>7 days)
Weekly Sun 2 AM: Document Cleanup (>30 days)
Weekly Sun 3 AM: Analytics Cleanup (>90 days)
Weekly Sun 4 AM: Vector Index Rebuild
Weekly Sun 5 AM: Database Vacuum (VACUUM)
| Job | Frequency | Duration | Impact |
|---|---|---|---|
| Health Check | 1h | 5-10s | Read-only, no impact |
| Cache Cleanup | 6h | 30-60s | Removes cache entries |
| Query Cleanup | 1/day | 1-2m | Removes old responses |
| Doc Cleanup | 1/week | 5-10m | Removes documents + vectors |
| Analytics Cleanup | 1/week | 2-5m | Removes old metrics |
| Vector Rebuild | 1/week | 10-30m | Improves search perf |
| DB Vacuum | 1/week | 5-15m | Reclaims disk space |
# Web Search Failures
- Retry up to 3 times
- Exponential backoff (1s, 2s, 4s)
- Fallback to cache if all retries fail
# LLM API Failures
- Retry with exponential backoff
- Timeout: 30 seconds
- Fallback to generic response
# Database Errors
- Retry transactions up to 5 times
- Log and continue on persistent errors
- Health check reports database statusExternal Services:
├── Search Engine
│ └── Circuit opens after 5 consecutive failures
│ Fallback: Return cached documents only
│
├── LLM API
│ └── Circuit opens after 3 consecutive timeouts
│ Fallback: Return summarized documents
│
└── Vector Store
└── Circuit opens on index corruption
Fallback: Use SQLite search onlyFull Functionality: Cache + Web + LLM
↓ (LLM fails)
Partial Functionality: Cache + Web, no LLM
↓ (Web fails)
Limited Functionality: Cache only
↓ (Cache fails)
Minimal Functionality: Error message
| Scenario | Target | Actual |
|---|---|---|
| Cache hit | 50ms | 30-80ms |
| Cache miss (web) | 5000ms | 2-10s |
| Analytics lookup | 100ms | 50-200ms |
| Vector search | 50ms | 20-100ms |
| Average (45% hit rate) | 2500ms | 2-3s |
Data Volume:
├── 1M documents → 4GB FAISS index, 500MB SQLite
├── 10M documents → 40GB FAISS index, 5GB SQLite
└── 100M documents → 400GB FAISS (distributed needed)
Requests/Second:
├── 1-10 RPS → Single instance (8GB RAM)
├── 10-100 RPS → 2-3 instances + load balancer
└── 100+ RPS → Kubernetes cluster
- Question length: 1-2000 characters
- Rate limiting: 100 requests/hour per IP
- SQL injection prevention via SQLAlchemy ORM
- HTML escaping for web content
- CORS restricted to configured frontend
- HTTPS in production
- No API key required (public API)
- Sensitive data excluded from logs
- No user identification required
- Queries stored for analytics only
- Cache cleanup removes old data
- Compliance: GDPR-friendly (auto-deletion)
┌──────────────────────────────────┐
│ Windows/Linux/macOS │
│ ┌────────────────────────────┐ │
│ │ Python Backend │ │
│ │ - FastAPI (Uvicorn) │ │
│ │ - SQLite │ │
│ │ - FAISS │ │
│ └────────────────────────────┘ │
│ ┌────────────────────────────┐ │
│ │ Node.js Frontend │ │
│ │ - React + Vite │ │
│ │ - Tailwind CSS │ │
│ └────────────────────────────┘ │
└──────────────────────────────────┘
┌──────────────────────────────────────────┐
│ Docker Compose / Kubernetes │
│ ┌──────────────────────────────────┐ │
│ │ Nginx (Load Balancer) │ │
│ └────────────┬─────────────────────┘ │
│ │ │
│ ┌───────┴────────┐ │
│ │ │ │
│ ┌────▼─────┐ ┌────▼─────┐ │
│ │Backend 1 │ │Backend 2 │ │
│ │(Uvicorn) │ │(Uvicorn) │ │
│ └──────────┘ └──────────┘ │
│ │ │ │
│ └────────┬───────┘ │
│ ┌─▼──┐ │
│ │SQLite + Shared Volumes │
│ │FAISS + Metadata │
│ └────┘ │
│ │
│ ┌──────────────────────────────────┐ │
│ │ Frontend (Node) │ │
│ └──────────────────────────────────┘ │
└──────────────────────────────────────────┘
- Distributed Vector Store: PostgreSQL with pgvector
- Multi-Region Caching: Redis cluster
- Advanced Ranking: Learning-to-rank (LambdaMART)
- Custom Fine-tuning: Domain-specific embeddings
- Real-time Indexing: Streaming document updates
- Advanced Analytics: ML-based anomaly detection
- API Versioning: Backward compatibility
- GraphQL Support: Alternative to REST
Last Updated: December 13, 2025 Version: 1.0.0