Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

CodeLens

Hybrid Graph-RAG for codebase intelligence — point it at a public GitHub repository, ask questions in plain English, and get answers grounded in retrieved code, a structural knowledge graph, and clickable source evidence.

https://github.com/expressjs/express
        ↓
"What calls Router?"  →  answer + sources + graph + impact analysis

Problem

Plain vector RAG underperforms on codebases. Code retrieval is not just about textual similarity — it is about structure:

  • "What happens when a user logs in?" needs the call chain (login → authenticate → createSession), not just files that mention "login".
  • "What breaks if I change validateToken()?" needs reverse dependencies — callers, importers, subclasses — which embeddings cannot represent.
  • Top-k chunks are fetched by surface similarity, so structurally central code is often missed entirely.
Vanilla RAG                    CodeLens

query                          query
  ↓                              |
embedding                        +--> semantic retrieval (what looks similar)
  ↓                              |
top-k chunks                     +--> graph retrieval (how it is connected)
  ↓                                        |
LLM                              hybrid ranking
                                   ↓
                                 context builder
                                   ↓
                                   LLM

Embeddings answer what code looks similar. The graph answers how code is connected.

Solution

CodeLens combines both evidence streams:

  1. Semantic retrieval — a local sentence-transformers model embeds AST-derived code chunks (functions, methods, classes — not arbitrary token windows) into a FAISS index.
  2. Graph retrieval — Tree-sitter parsing produces a NetworkX knowledge graph (File, Function, Class nodes; CONTAINS, IMPORTS, CALLS, EXTENDS, REFERENCES edges). Identifiers in the question are matched against graph nodes and the neighborhood is traversed deterministically.
  3. Hybrid ranking — final_score = 0.6 × semantic + 0.4 × graph, deduplicated per chunk, top-k kept for the context builder. The weighting is configurable and explainable — no black-box reranker.

Architecture

flowchart TB
    subgraph Ingestion
        A[GitHub repository] -->|git clone --depth 1| B[File filter]
        B -->|TypeScript / JavaScript / Python| C[Tree-sitter AST parser]
        C --> D[Code knowledge graph - NetworkX]
        C --> E[Structural chunks]
        E --> F[Embedding provider]
        F --> G[FAISS vector store]
        D --> H[(data/repositories/...)]
        G --> H
    end
    subgraph Query
        Q[Question] --> S[Semantic retrieval]
        Q --> T[Graph retrieval]
        S --> U[Hybrid merge + dedup + weighted ranking]
        T --> U
        U --> V[Context builder]
        V --> W[LLM generation]
        W --> X[Answer + sources + subgraph]
    end
    H --> T
    H --> S
Loading

RAG Pipeline

Stage What happens
1. Ingestion Shallow git clone into a temp directory. Repository code is never executed (npm install & friends never run; git runs with a scrubbed environment).
2. File filtering Skip node_modules, dist, build, binaries, *.min.*, .d.ts, oversized files.
3. Parsing Tree-sitter extracts functions, classes, methods, imports, call sites, and inheritance (with a regex fallback parser if grammars are unavailable).
4. Graph construction Cross-file resolution: imports map specifiers to repo files; calls resolve through import bindings; classes link by EXTENDS/REFERENCES.
5. Chunking One chunk per function/method/class with file_path, symbol, and line range. Files without symbols fall back to 80-line windows.
6. Embeddings all-MiniLM-L6-v2 locally (default) or the OpenAI embeddings API, behind one provider interface.
7. Vector retrieval Exact cosine search (IndexFlatIP over normalized vectors), top-k = VECTOR_TOP_K.
8. Graph retrieval Identifier extraction → node matching → BFS to depth GRAPH_DEPTH with score decay, plus shortest structural paths between matched symbols.
9. Hybrid ranking Weighted score merge, dedup by chunk key, provenance tracked (semantic / graph / semantic+graph).
10. Context + generation Compact context (code + provenance + relationships, capped by MAX_CONTEXT_CHARS) → OpenAI answer with citations. Without an API key the system degrades to retrieval-only output instead of failing.

Example

Real output from indexing https://github.com/expressjs/express (141 files → 218 functions, 407 relationships, 377 embedded chunks):

Question: What calls Router?

  • Graph retrieval matched test/Router.js (0.80) and traversed its symbol neighborhood (6 nodes, 8 edges).
  • Hybrid ranking kept 8 of 12 unique candidates; sources cite route · lib/application.js:256, param · lib/application.js:322, …
  • Clicking the citation opens the evidence viewer on exactly those lines:
FILE: lib/application.js
SYMBOL: route
LINES: 256-258

app.route = function route(path) {
  return this.router.route(path);
};

Impact analysis on tryRender (derived from real graph edges, not the LLM):

Direct callers: render (lib/application.js:522)
Impact: MEDIUM
"tryRender (lib/application.js:625) is called directly by 1 symbol(s). Impact level: MEDIUM."

Retrieval — why both streams

Semantic search alone ranks fn2 in test/app.router.js (the word "router" appears everywhere in tests). Graph traversal alone only answers questions that name an exact symbol. Together: embeddings pull in thematically relevant code, the graph supplies the call/contains/imports structure around the matched symbols, and the context builder gives the LLM both the code and how it connects — enabling execution-flow answers ("A calls B calls C") that neither stream supports alone.

Tech Stack

  • Backend: Python, FastAPI, Tree-sitter (via tree-sitter-language-pack), NetworkX, FAISS, sentence-transformers, OpenAI / Groq (OpenAI-compatible)
  • Frontend: Next.js 14 (App Router), TypeScript, Tailwind CSS, React Flow (@xyflow/react), Prism

Running Locally

Prerequisites: Python 3.12+, Node 18.17+, git.

Backend (port 8000):

cd backend
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
uvicorn app.main:app --reload

Frontend (port 3000):

cd frontend
npm install
npm run dev

Open http://localhost:3000, paste a repository URL (e.g. https://github.com/expressjs/express), and click Analyze repository.

Optionally copy .env.example to backend/.env and set OPENAI_API_KEY=... (or GROQ_API_KEY=...) to enable LLM answers and impact explanations. Groq uses an OpenAI-compatible endpoint — a free key from console.groq.com works out of the box with GROQ_MODEL=openai/gpt-oss-20b. Without any key everything else works in retrieval-only mode (local embeddings, no network calls for retrieval).

You can also supply a key at runtime: click API key in the top bar, pick OpenAI or Groq, and paste it. The key is kept in your browser's localStorage only and sent as a request header to your local backend — it takes precedence over backend/.env and can be removed anytime. Keys pasted without a provider choice are auto-detected by prefix (gsk_… → Groq, sk-… → OpenAI).

Tests (backend):

cd backend && source .venv/bin/activate
pytest tests/ -q

API

Method Path Purpose
POST /repositories/analyze Start indexing a public GitHub repository
GET /repositories/{id}/status Indexing progress (stage + real stats)
POST /repositories/{id}/query Hybrid retrieval + answer + sources + subgraph
GET /repositories/{id}/graph Focused subgraph (`?focus=<node
GET /repositories/{id}/symbols Symbol search for impact/focus pickers
POST /repositories/{id}/impact "What breaks if I change this?"
GET /repositories/{id}/source Read-only source lines for the evidence viewer

Design Decisions

  • Tree-sitter — real ASTs for TS/JS/Python with one API, fault-tolerant parsing of imperfect code, and a clean seam for adding languages. A regex fallback keeps indexing alive if grammars are unavailable.
  • NetworkX — the graph is small (thousands of nodes), in-memory, and pickle-persisted; BFS/shortest-path/impact queries are one-liners. A graph database would add infrastructure without adding capability here.
  • FAISS (IndexFlatIP) — exact cosine similarity over normalized vectors is fast at this scale and removes approximate-index tuning as a moving part. Chunk metadata lives in a JSON sidecar next to the index.
  • Hybrid scoring over a learned reranker — 0.6/0.4 weighted fusion is deterministic, explainable in an interview, and configurable via SEMANTIC_WEIGHT/GRAPH_WEIGHT. A cross-encoder would add latency and opacity for marginal gain at this corpus size.
  • Structural chunking — chunks map 1:1 to symbols, so a citation is always a real function/method/class with exact line ranges, and graph nodes join retrieval results without a fuzzy join step.
  • Retrieval works without an API key — local embeddings by default; the LLM is optional plumbing, not a prerequisite.

Limitations

  • Supported languages: TypeScript, JavaScript, Python. Other files are ignored (the parser abstraction exists for more).
  • Approximate call graph: no type inference. Calls resolve through same-file symbols, import bindings, and class receivers (this/self, imported classes). Dynamic dispatch (obj.method() on an untyped local) does not resolve.
  • Public repositories only, size-capped (MAX_REPOSITORY_SIZE_MB).
  • Repository code is never executed — analysis is purely static, so runtime behavior is only inferred, never observed.
  • Entity matching is lexical: graph retrieval fires when the question names a symbol/file (or close variant); purely conceptual questions lean on the semantic stream.

Project Structure

backend/
    app/
        api/            # FastAPI routes (repositories, query, graph, impact)
        core/           # configuration
        ingestion/      # GitHub loader + file filtering
        parsing/        # Tree-sitter parser, regex fallback, chunker
        graph/          # graph construction + deterministic queries
        retrieval/      # embeddings, FAISS store, graph/hybrid retrievers, reranker
        generation/     # prompts, context builder, LLM provider
        models/         # API schemas
        services/       # indexing + query orchestration
    tests/              # focused unit tests over a fixture repository
frontend/
    app/                # Next.js App Router (workspace page, layout, tokens)
    components/         # repository / query / graph / impact / code / layout
    lib/                # api client, shared types, UI constants

CodeLens

About

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages