Skip to content

Latest commit

 

History

History
826 lines (680 loc) · 31.8 KB

File metadata and controls

826 lines (680 loc) · 31.8 KB

Architecture Documentation

Comprehensive system architecture and design for OpenWebAnswer.

Table of Contents

  1. Overview
  2. System Components
  3. Data Flow
  4. Technology Stack
  5. Database Schema
  6. Caching Strategy
  7. Vector Store Architecture
  8. API Design
  9. Background Jobs
  10. Error Handling & Resilience

Overview

OpenWebAnswer is a production-grade Question-Answering system that combines:

  • Real-time web search with semantic relevance
  • Intelligent caching to reduce redundant searches
  • Quality scoring to filter low-quality sources
  • Followup question generation to enhance user engagement
  • Comprehensive analytics for system monitoring

Key Features

Feature Purpose Implementation
Semantic Search Find relevant documents via embeddings FAISS + All-MiniLM-L6-v2 (384-dim)
Smart Caching Reduce API calls and improve response time Multi-level with SQLite + FAISS
Quality Filtering Exclude low-quality sources Domain authority + content analysis
Followup Generation Suggest related questions Template + embedding similarity
Analytics Tracking Monitor performance and quality QueryAnalytics table + insights
Background Jobs System maintenance APScheduler (7 scheduled jobs)

System Components

1. API Layer

File: backend/api/routes/

api/
├── query.py          # Main QA endpoint
├── health.py         # Health status check
├── analytics.py      # Analytics & monitoring endpoints
└── schemas.py        # Request/response models

Endpoints:

  • POST /api/query → QA with caching
  • GET /api/health → System health
  • GET /api/analytics/* → 9 monitoring endpoints

2. Pipeline Layer

File: backend/pipelines/realtime_qa.py

Orchestrates the full QA flow:

Request → Cache Check → Web Search → Quality Scoring → Answer Generation
         ↓              ↓           ↓                 ↓
      (Cache Hit?)  (Fetch Docs) (Filter Low-Q)  (Generate + Followups)
         ↓              ↓           ↓                 ↓
      Return Cached   Cache+Index Score             Track Analytics

Key Methods:

  • run() - Main QA pipeline with caching
  • _search() - Web search integration
  • _fetch_documents() - Document retrieval with quality scoring
  • _generate_answer() - LLM-based answer generation

3. Database Layer

File: backend/database/

database/
├── schema.sql              # Database tables & indexes
├── models.py               # SQLAlchemy ORM models
├── init_db.py              # Initialization & migrations
└── repositories/
    ├── document_repo.py    # Document CRUD (11 methods)
    ├── chunk_repo.py       # Chunk CRUD (13 methods)
    ├── query_repo.py       # Query CRUD (11 methods)
    ├── analytics_repo.py   # Analytics CRUD (14 methods)
    └── followup_repo.py    # Followup CRUD (10 methods)

Tables (6 total):

  1. documents - Cached web documents (metadata only, content in chunks)
  2. document_chunks - Document text chunks for searching (single source of truth for chunk_text)
  3. queries - Cached query results with source tracking
  4. query_sources - Maps queries to specific source chunks used in answers
  5. query_analytics - Performance metrics & timing breakdowns (per query_id)
  6. followup_questions - Generated followup suggestions with question types

4. Vector Store Layer

File: indexer/

indexer/
├── vector_store_manager.py  # FAISS management
├── vector_store.py          # Persistence & integrity
└── vector_stores/
    └── faiss_store.py       # FAISS implementation

Features:

  • FAISS index for efficient similarity search
  • JSON metadata sidecar for ID mapping
  • Integrity verification & rebuilding
  • Statistics and diagnostics

5. Cache Layer

File: cache/

cache/
├── cache_manager.py         # Multi-level caching
├── deduplication.py         # Query & document deduplication
├── invalidation.py          # TTL-based cache expiration
├── query_analyzer.py        # Cache hit prediction
└── models.py                # Cache data structures

Cache Levels:

  1. Query Cache - Full QA responses (7-day TTL)
  2. Document Cache - Fetched documents (30-day TTL)
  3. Vector Cache - Embeddings in FAISS (synced with docs)

6. Scoring & Quality Layer

File: search_fetch/source_scorer.py

Evaluates source quality across 3 dimensions:

Quality Score = Domain Authority (40%) + Content Quality (35%) + Freshness (25%)
  Range: 0.0 - 1.0
  Threshold: >= 0.3 (30%)

Components:

  • Domain Authority: Trusted (+0.4), Unknown (+0.1), Blacklist (-0.5)
  • Content Quality: Length, structure, keyword presence (0.0-0.3)
  • Freshness: Recent (+0.3), Old (-0.1)

7. LLM & Generation Layer

File: llm_interface/

llm_interface/
├── llm_engine.py            # Multi-provider LLM support
├── followup_generator.py    # Follow-up question generation
├── prompts.py               # System prompts & templates
└── models.py                # Response models

LLM Providers: OpenAI, Anthropic, Ollama, HuggingFace

Followup Generation:

  • 4 question types: expansion, clarification, related, deeper
  • 16 templates total (4 × 4 categories)
  • Relevance scoring via embedding similarity

8. Analytics Layer

File: backend/analytics/tracker.py & backend/api/routes/analytics.py

Tracks metrics and generates insights:

Tracked Metrics:

  • Query frequency & patterns
  • Response time (P50, P95, P99)
  • Cache hit rate (%)
  • Document quality scores
  • Error rates

Health Score (0-100):

Health = (CacheHitRate × 30%) + (ResponseTime × 30%) + 
         (DocQuality × 20%) + (SuccessRate × 20%)

9. Background Jobs

File: backend/jobs/

jobs/
├── cleanup_scheduler.py     # Job implementations
└── scheduler.py             # APScheduler registration

7 Scheduled Jobs:

  1. Cache cleanup (every 6h)
  2. Query cleanup (daily 2am)
  3. Document cleanup (weekly Sun 2am)
  4. Analytics cleanup (weekly Sun 3am)
  5. Vector index rebuild (weekly Sun 4am)
  6. Database vacuum (weekly Sun 5am)
  7. Health check (hourly)

Data Flow

Complete QA Request Flow

┌─────────────────────────────────────────────────────────────────┐
│                    User Submits Question                         │
│                   POST /api/query                                │
└─────────────────┬───────────────────────────────────────────────┘
                  │
                  ▼
┌─────────────────────────────────────────────────────────────────┐
│              Step 0: Check Query Cache                           │
│  - Hash question with MD5                                        │
│  - Look up in SQLite queries table                               │
│  - Check expiration (7 days)                                     │
└──────┬──────────────────────────────────┬──────────────────────┘
       │                                  │
   CACHE HIT                         CACHE MISS
   (40-50%)                          (50-60%)
       │                                  │
       ▼                                  ▼
┌──────────────────────┐   ┌─────────────────────────────────────┐
│  Return Cached       │   │  Step 1: Web Search                 │
│  - Response          │   │  - Query search engine              │
│  - Sources           │   │  - Get top N results                │
│  - Followups         │   │  - Return URLs + snippets           │
│  - Analytics (hit)   │   │                                     │
└────────┬─────────────┘   │                                     │
         │                  └──────────┬──────────────────────────┘
         │                             │
         │                             ▼
         │            ┌────────────────────────────────────────┐
         │            │  Step 2: Fetch & Score Documents       │
         │            │  - Download full content               │
         │            │  - Calculate quality score:            │
         │            │    * Domain authority                  │
         │            │    * Content quality                   │
         │            │    * Freshness                         │
         │            │  - Filter low quality (< 0.3)          │
         │            │  - Rerank by relevance                 │
         │            └────────────┬─────────────────────────┘
         │                         │
         │                         ▼
         │            ┌────────────────────────────────────────┐
         │            │  Step 2.5: Cache Documents             │
         │            │  - Embed with All-MiniLM-L6-v2         │
         │            │  - Save to SQLite documents table      │
         │            │  - Index in FAISS vector store         │
         │            │  - Save chunks for future search       │
         │            └────────────┬─────────────────────────┘
         │                         │
         │                         ▼
         │            ┌────────────────────────────────────────┐
         │            │  Step 3: Generate Answer               │
         │            │  - Build context from documents        │
         │            │  - Call LLM (OpenAI/Anthropic/etc.)    │
         │            │  - Generate structured answer          │
         │            └────────────┬─────────────────────────┘
         │                         │
         │                         ▼
         │            ┌────────────────────────────────────────┐
         │            │  Step 3.5: Generate Followups          │
         │            │  - Template-based generation           │
         │            │  - Score relevance via embeddings      │
         │            │  - Filter duplicates                   │
         │            │  - Return top 3 questions              │
         │            └────────────┬─────────────────────────┘
         │                         │
         │                         ▼
         │            ┌────────────────────────────────────────┐
         │            │  Step 5: Cache Response              │
         │            │  - Save to queries table               │
         │            │  - Set TTL: 7 days                     │
         │            │  - Record source_type (cache|web)      │
         │            │  - Track used chunks in query_sources  │
         │            │  - Store model_used for analysis       │
         │            └────────────┬─────────────────────────┘
         │                         │
         │                         ▼
         │            ┌────────────────────────────────────────┐
         │            │  Step 6: Track Analytics               │
         │            │  - Record response time                │
         │            │  - Log cache miss reason               │
         │            │  - Average quality score               │
         │            │  - Success/failure status              │
         │            └────────────┬─────────────────────────┘
         │                         │
         └─────────────┬───────────┘
                       │
                       ▼
┌─────────────────────────────────────────────────────────────────┐
│              Return QueryResponse                               │
│  {                                                              │
│    "id": "query_abc123",                                        │
│    "question": "What is Python?",                              │
│    "answer": "Python is a...",                                 │
│    "sources": [{url, title, quality_score, ...}, ...],        │
│    "followups": ["What are Python applications?", ...],        │
│    "response_time_ms": 2500,                                   │
│    "cache_hit": false,                                         │
│    "average_quality_score": 0.85                               │
│  }                                                              │
└─────────────────────────────────────────────────────────────────┘

Technology Stack

Backend

Layer Technology Version Purpose
Framework FastAPI 0.109+ REST API & async support
Database SQLite 3 3.50+ Persistent cache storage
ORM SQLAlchemy 2.0+ Database abstraction
Vector DB FAISS 1.7+ Similarity search
Embeddings Sentence Transformers 2.2+ Text embeddings (384-dim)
LLM Multiple - GPT-4, Claude, Ollama
Scheduler APScheduler 3.10+ Background jobs
Server Uvicorn 0.24+ ASGI server

Frontend

Component Technology Purpose
Framework React 18 UI components
Build Tool Vite Fast bundling
API Client Axios HTTP requests
Styling Tailwind CSS Styling
State React Hooks State management

Deployment

Component Technology Purpose
Containerization Docker Application packaging
Orchestration Docker Compose Multi-container setup
Reverse Proxy Nginx Load balancing
Server Uvicorn Python ASGI

Database Schema

Tables Overview

documents (web document metadata)
├── id: INTEGER PRIMARY KEY
├── url: TEXT UNIQUE
├── title: TEXT
├── domain: TEXT (indexed)
├── quality_score: FLOAT (0.0-1.0)
├── is_archived: BOOLEAN 
├── chunk_count: INTEGER 
└── indexes: [domain], [quality_score], [created_at]

document_chunks (embedding context - SINGLE SOURCE OF TRUTH)
├── id: INTEGER PRIMARY KEY
├── doc_id: INTEGER FK
├── chunk_text: TEXT (ONLY source for chunk text)
├── embedding_id: TEXT UNIQUE (FAISS reference)
├── token_count: INTEGER
├── position: INTEGER (order in document)
└── indexes: [doc_id], [embedding_id], [doc_id, position]

queries (cached query results with source tracking)
├── id: INTEGER PRIMARY KEY
├── query_text: TEXT
├── query_hash: TEXT UNIQUE (deduplication)
├── answer: TEXT
├── response_time_ms: INTEGER
├── source_type: TEXT (cache|vector|web) 
├── model_used: TEXT (huggingface|openai|etc) 
├── expires_at: TIMESTAMP (7-day TTL)
└── indexes: [query_hash], [source_type], [expires_at]

query_sources (NEW - Phase 7: Track answer sources)
├── id: INTEGER PRIMARY KEY
├── query_id: INTEGER FK
├── doc_id: INTEGER FK
├── chunk_id: INTEGER FK
├── position_in_answer: INTEGER (order in answer)
└── indexes: [query_id], [doc_id], [chunk_id]

query_analytics (performance metrics per query_id)
├── id: INTEGER PRIMARY KEY
├── query_id: INTEGER FK (indexed)
├── vector_search_time_ms: INTEGER
├── web_search_time_ms: INTEGER
├── embedding_time_ms: INTEGER
├── llm_generation_time_ms: INTEGER
├── chunk_retrieval_time_ms: INTEGER
├── total_time_ms: INTEGER
├── cache_hit_count: INTEGER
├── quality_score: FLOAT
└── indexes: [query_id], [created_at DESC]

followup_questions (engagement with question types)
├── id: INTEGER PRIMARY KEY
├── query_id: INTEGER FK
├── followup_text: TEXT
├── question_type: TEXT (expansion|clarification|related|deeper) 
├── relevance_score: FLOAT (0.0-1.0)
└── indexes: [query_id], [relevance_score]

Key Design Decisions

  1. Document Content Storage: Moved from documents table to document_chunks

    • documents table: Metadata only (URL, title, domain, quality_score, is_archived)
    • document_chunks: Single source of truth for chunk_text
    • Rationale: Enables efficient chunk-level searching and deduplication
  2. Source Tracking: New query_sources table

    • Maps each query to specific chunks used in generating the answer
    • Enables complete source attribution and verification
    • Tracks position_in_answer for reconstruction of source order
  3. Query Analytics Redesign: Changed from query_hash to query_id based

    • Now tracks detailed timing breakdowns per query execution
    • Separate fields: vector_search_time_ms, web_search_time_ms, embedding_time_ms, llm_generation_time_ms
    • Enables performance bottleneck identification
  4. Deduplication: Hash-based (query_hash, content_hash where applicable)

    • Prevents duplicate web searches for identical queries
    • Enables cache reuse across similar queries
  5. TTL-Based Expiration: Automatic cleanup via scheduled jobs

    • Queries: 7-day TTL
    • Documents: 30-day TTL
    • Analytics: 90-day TTL
  6. Question Type Classification

    • FollowupQuestion.question_type: expansion, clarification, related, deeper
    • Enables categorization and analysis of followup engagement
  7. Cascading Deletes: Document deletion cascades to chunks and query_sources

    • Maintains referential integrity
    • Automatic cleanup when documents expire

Caching Strategy

Three-Level Cache Architecture

┌─────────────────────────────────────────────────────────────┐
│                   Query Cache (Level 1)                     │
│  - Full QA responses                                         │
│  - SQLite queries table                                      │
│  - 7-day TTL                                                 │
│  - Hit rate: 40-50% (identical queries)                      │
└─────────────────────────────────────────────────────────────┘
                        ↓ (miss)
┌─────────────────────────────────────────────────────────────┐
│                 Document Cache (Level 2)                    │
│  - Fetched web documents                                     │
│  - SQLite documents table                                    │
│  - 30-day TTL                                                │
│  - Hit rate: 30-40% (related queries)                        │
│  - Includes: URL, title, content, quality score              │
└─────────────────────────────────────────────────────────────┘
                        ↓ (miss)
┌─────────────────────────────────────────────────────────────┐
│                  Vector Cache (Level 3)                      │
│  - FAISS index of document embeddings                        │
│  - JSON metadata sidecar                                     │
│  - Enables semantic search without web fetch                 │
│  - Hit rate: 50%+ (similar queries)                          │
└─────────────────────────────────────────────────────────────┘
                        ↓ (miss)
┌─────────────────────────────────────────────────────────────┐
│                    Web Search (Miss)                         │
│  - Call external search engine                               │
│  - Fetch full documents                                      │
│  - All three cache levels populated                          │
└─────────────────────────────────────────────────────────────┘

Cache Invalidation

Strategy: TTL-based (lazy deletion)

# Query Cache: Expires after 7 days
expires_at = datetime.utcnow() + timedelta(days=7)

# Document Cache: Expires after 30 days
expires_at = datetime.utcnow() + timedelta(days=30)

# Cleanup Job: Runs daily, removes expired entries
# Reclaims ~50-70% of old data per week

Performance Impact:

  • Cache hits: 1-50ms (no web search)
  • Cache misses: 2-10 seconds (web search + LLM)
  • Average response time: 2-3 seconds (with 45% cache hit rate)

Vector Store Architecture

FAISS Configuration

VectorStore (FAISS)
├── Index Type: HNSW32
│   └── Efficient for ~1M documents
│       - Query time: 10-100ms
│       - Memory: ~4GB per 1M vectors
│   
├── Dimension: 384 (All-MiniLM-L6-v2)
│   └── Trade-off: quality vs. memory
│       - 384-dim: Balanced (recommended)
│       - 768-dim: Better quality (more memory)
│   
└── Metadata Sidecar (JSON)
    └── Maps embedding_id → document_id, chunk_index
        - Used for document retrieval after search

Search Process

User Query
    ↓
Embed with All-MiniLM-L6-v2 (384-dim)
    ↓
FAISS search (k=50, nprobe=32)
    ↓
Score results via relevance ranking
    ↓
Filter by quality score (threshold: 0.3)
    ↓
Return top-k (default: 5) documents

Rebuild Strategy

Periodic Rebuild: Weekly (Sunday 4am)

  • Compacts index structure
  • Improves search performance
  • Reclaims memory from deleted vectors
  • Takes 5-30 minutes for 1M vectors

API Design

Query Endpoint

Request:

POST /api/query
{
  "question": "What is machine learning?",
  "mode": "hybrid",
  "top_k": 5,
  "stream": false
}

Response:

{
  "id": "query_abc123",
  "question": "What is machine learning?",
  "answer": "Machine learning is...",
  "sources": [
    {
      "url": "https://example.com",
      "title": "ML Guide",
      "snippet": "...",
      "relevance_score": 0.95,
      "quality_score": 0.85,
      "quality_reason": "High authority domain"
    }
  ],
  "followups": [
    "What are supervised learning algorithms?",
    "How does neural networks work?"
  ],
  "mode": "realtime",
  "response_time_ms": 2534.5,
  "sources_count": 5,
  "average_quality_score": 0.82,
  "cache_hit": false
}

Analytics Endpoints

9 Endpoints:

  1. GET /api/analytics/insights - Top queries & patterns
  2. GET /api/analytics/top-queries - Most frequent questions
  3. GET /api/analytics/slowest-queries - Performance bottlenecks
  4. GET /api/analytics/performance - System metrics
  5. GET /api/analytics/documents - Document statistics
  6. GET /api/analytics/health - System health score
  7. GET /api/analytics/cache-hit-rate - Cache efficiency
  8. POST /api/analytics/rate-query - User ratings (1-5)
  9. GET /api/analytics/report - Comprehensive report

Background Jobs

Job Schedule

Every Hour:        Health Check
Every 6 Hours:     Cache Cleanup
Daily 2 AM:        Query Cleanup (>7 days)
Weekly Sun 2 AM:   Document Cleanup (>30 days)
Weekly Sun 3 AM:   Analytics Cleanup (>90 days)
Weekly Sun 4 AM:   Vector Index Rebuild
Weekly Sun 5 AM:   Database Vacuum (VACUUM)

Job Details

Job Frequency Duration Impact
Health Check 1h 5-10s Read-only, no impact
Cache Cleanup 6h 30-60s Removes cache entries
Query Cleanup 1/day 1-2m Removes old responses
Doc Cleanup 1/week 5-10m Removes documents + vectors
Analytics Cleanup 1/week 2-5m Removes old metrics
Vector Rebuild 1/week 10-30m Improves search perf
DB Vacuum 1/week 5-15m Reclaims disk space

Error Handling & Resilience

Retry Strategy

# Web Search Failures
- Retry up to 3 times
- Exponential backoff (1s, 2s, 4s)
- Fallback to cache if all retries fail

# LLM API Failures
- Retry with exponential backoff
- Timeout: 30 seconds
- Fallback to generic response

# Database Errors
- Retry transactions up to 5 times
- Log and continue on persistent errors
- Health check reports database status

Circuit Breaker Pattern

External Services:
├── Search Engine
│   └── Circuit opens after 5 consecutive failures
│       Fallback: Return cached documents only
│
├── LLM API
│   └── Circuit opens after 3 consecutive timeouts
│       Fallback: Return summarized documents
│
└── Vector Store
    └── Circuit opens on index corruption
        Fallback: Use SQLite search only

Graceful Degradation

Full Functionality:    Cache + Web + LLM
  ↓ (LLM fails)
Partial Functionality: Cache + Web, no LLM
  ↓ (Web fails)
Limited Functionality: Cache only
  ↓ (Cache fails)
Minimal Functionality: Error message

Performance Metrics

Target Response Times

Scenario Target Actual
Cache hit 50ms 30-80ms
Cache miss (web) 5000ms 2-10s
Analytics lookup 100ms 50-200ms
Vector search 50ms 20-100ms
Average (45% hit rate) 2500ms 2-3s

Scalability

Data Volume:
├── 1M documents      → 4GB FAISS index, 500MB SQLite
├── 10M documents     → 40GB FAISS index, 5GB SQLite
└── 100M documents    → 400GB FAISS (distributed needed)

Requests/Second:
├── 1-10 RPS          → Single instance (8GB RAM)
├── 10-100 RPS        → 2-3 instances + load balancer
└── 100+ RPS          → Kubernetes cluster

Security Considerations

Input Validation

  • Question length: 1-2000 characters
  • Rate limiting: 100 requests/hour per IP
  • SQL injection prevention via SQLAlchemy ORM
  • HTML escaping for web content

API Security

  • CORS restricted to configured frontend
  • HTTPS in production
  • No API key required (public API)
  • Sensitive data excluded from logs

Data Privacy

  • No user identification required
  • Queries stored for analytics only
  • Cache cleanup removes old data
  • Compliance: GDPR-friendly (auto-deletion)

Deployment Architecture

Development (Single Machine)

┌──────────────────────────────────┐
│    Windows/Linux/macOS           │
│  ┌────────────────────────────┐  │
│  │    Python Backend          │  │
│  │  - FastAPI (Uvicorn)       │  │
│  │  - SQLite                  │  │
│  │  - FAISS                   │  │
│  └────────────────────────────┘  │
│  ┌────────────────────────────┐  │
│  │    Node.js Frontend        │  │
│  │  - React + Vite            │  │
│  │  - Tailwind CSS            │  │
│  └────────────────────────────┘  │
└──────────────────────────────────┘

Production (Docker)

┌──────────────────────────────────────────┐
│        Docker Compose / Kubernetes       │
│  ┌──────────────────────────────────┐   │
│  │  Nginx (Load Balancer)           │   │
│  └────────────┬─────────────────────┘   │
│               │                          │
│       ┌───────┴────────┐                 │
│       │                │                 │
│  ┌────▼─────┐    ┌────▼─────┐           │
│  │Backend 1 │    │Backend 2 │           │
│  │(Uvicorn) │    │(Uvicorn) │           │
│  └──────────┘    └──────────┘           │
│       │                │                 │
│       └────────┬───────┘                 │
│              ┌─▼──┐                      │
│              │SQLite + Shared Volumes   │
│              │FAISS + Metadata          │
│              └────┘                      │
│                                         │
│  ┌──────────────────────────────────┐  │
│  │  Frontend (Node)                 │  │
│  └──────────────────────────────────┘  │
└──────────────────────────────────────────┘

Future Enhancements

  1. Distributed Vector Store: PostgreSQL with pgvector
  2. Multi-Region Caching: Redis cluster
  3. Advanced Ranking: Learning-to-rank (LambdaMART)
  4. Custom Fine-tuning: Domain-specific embeddings
  5. Real-time Indexing: Streaming document updates
  6. Advanced Analytics: ML-based anomaly detection
  7. API Versioning: Backward compatibility
  8. GraphQL Support: Alternative to REST

Last Updated: December 13, 2025 Version: 1.0.0