Hybrid Graph-RAG for codebase intelligence — point it at a public GitHub repository, ask questions in plain English, and get answers grounded in retrieved code, a structural knowledge graph, and clickable source evidence.
https://github.com/expressjs/express
↓
"What calls Router?" → answer + sources + graph + impact analysis
Plain vector RAG underperforms on codebases. Code retrieval is not just about textual similarity — it is about structure:
- "What happens when a user logs in?" needs the call chain (
login → authenticate → createSession), not just files that mention "login". - "What breaks if I change
validateToken()?" needs reverse dependencies — callers, importers, subclasses — which embeddings cannot represent. - Top-k chunks are fetched by surface similarity, so structurally central code is often missed entirely.
Vanilla RAG CodeLens
query query
↓ |
embedding +--> semantic retrieval (what looks similar)
↓ |
top-k chunks +--> graph retrieval (how it is connected)
↓ |
LLM hybrid ranking
↓
context builder
↓
LLM
Embeddings answer what code looks similar. The graph answers how code is connected.
CodeLens combines both evidence streams:
- Semantic retrieval — a local sentence-transformers model embeds AST-derived code chunks (functions, methods, classes — not arbitrary token windows) into a FAISS index.
- Graph retrieval — Tree-sitter parsing produces a NetworkX knowledge graph (
File,Function,Classnodes;CONTAINS,IMPORTS,CALLS,EXTENDS,REFERENCESedges). Identifiers in the question are matched against graph nodes and the neighborhood is traversed deterministically. - Hybrid ranking —
final_score = 0.6 × semantic + 0.4 × graph, deduplicated per chunk, top-k kept for the context builder. The weighting is configurable and explainable — no black-box reranker.
flowchart TB
subgraph Ingestion
A[GitHub repository] -->|git clone --depth 1| B[File filter]
B -->|TypeScript / JavaScript / Python| C[Tree-sitter AST parser]
C --> D[Code knowledge graph - NetworkX]
C --> E[Structural chunks]
E --> F[Embedding provider]
F --> G[FAISS vector store]
D --> H[(data/repositories/...)]
G --> H
end
subgraph Query
Q[Question] --> S[Semantic retrieval]
Q --> T[Graph retrieval]
S --> U[Hybrid merge + dedup + weighted ranking]
T --> U
U --> V[Context builder]
V --> W[LLM generation]
W --> X[Answer + sources + subgraph]
end
H --> T
H --> S
| Stage | What happens |
|---|---|
| 1. Ingestion | Shallow git clone into a temp directory. Repository code is never executed (npm install & friends never run; git runs with a scrubbed environment). |
| 2. File filtering | Skip node_modules, dist, build, binaries, *.min.*, .d.ts, oversized files. |
| 3. Parsing | Tree-sitter extracts functions, classes, methods, imports, call sites, and inheritance (with a regex fallback parser if grammars are unavailable). |
| 4. Graph construction | Cross-file resolution: imports map specifiers to repo files; calls resolve through import bindings; classes link by EXTENDS/REFERENCES. |
| 5. Chunking | One chunk per function/method/class with file_path, symbol, and line range. Files without symbols fall back to 80-line windows. |
| 6. Embeddings | all-MiniLM-L6-v2 locally (default) or the OpenAI embeddings API, behind one provider interface. |
| 7. Vector retrieval | Exact cosine search (IndexFlatIP over normalized vectors), top-k = VECTOR_TOP_K. |
| 8. Graph retrieval | Identifier extraction → node matching → BFS to depth GRAPH_DEPTH with score decay, plus shortest structural paths between matched symbols. |
| 9. Hybrid ranking | Weighted score merge, dedup by chunk key, provenance tracked (semantic / graph / semantic+graph). |
| 10. Context + generation | Compact context (code + provenance + relationships, capped by MAX_CONTEXT_CHARS) → OpenAI answer with citations. Without an API key the system degrades to retrieval-only output instead of failing. |
Real output from indexing https://github.com/expressjs/express (141 files → 218 functions, 407 relationships, 377 embedded chunks):
Question: What calls Router?
- Graph retrieval matched
test/Router.js(0.80) and traversed its symbol neighborhood (6 nodes, 8 edges). - Hybrid ranking kept 8 of 12 unique candidates; sources cite
route · lib/application.js:256,param · lib/application.js:322, … - Clicking the citation opens the evidence viewer on exactly those lines:
FILE: lib/application.js
SYMBOL: route
LINES: 256-258
app.route = function route(path) {
return this.router.route(path);
};
Impact analysis on tryRender (derived from real graph edges, not the LLM):
Direct callers: render (lib/application.js:522)
Impact: MEDIUM
"tryRender (lib/application.js:625) is called directly by 1 symbol(s). Impact level: MEDIUM."
Semantic search alone ranks fn2 in test/app.router.js (the word "router" appears everywhere in tests). Graph traversal alone only answers questions that name an exact symbol. Together: embeddings pull in thematically relevant code, the graph supplies the call/contains/imports structure around the matched symbols, and the context builder gives the LLM both the code and how it connects — enabling execution-flow answers ("A calls B calls C") that neither stream supports alone.
- Backend: Python, FastAPI, Tree-sitter (via
tree-sitter-language-pack), NetworkX, FAISS, sentence-transformers, OpenAI / Groq (OpenAI-compatible) - Frontend: Next.js 14 (App Router), TypeScript, Tailwind CSS, React Flow (
@xyflow/react), Prism
Prerequisites: Python 3.12+, Node 18.17+, git.
Backend (port 8000):
cd backend
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
uvicorn app.main:app --reloadFrontend (port 3000):
cd frontend
npm install
npm run devOpen http://localhost:3000, paste a repository URL (e.g. https://github.com/expressjs/express), and click Analyze repository.
Optionally copy .env.example to backend/.env and set OPENAI_API_KEY=... (or GROQ_API_KEY=...) to enable LLM answers and impact explanations. Groq uses an OpenAI-compatible endpoint — a free key from console.groq.com works out of the box with GROQ_MODEL=openai/gpt-oss-20b. Without any key everything else works in retrieval-only mode (local embeddings, no network calls for retrieval).
You can also supply a key at runtime: click API key in the top bar, pick OpenAI or Groq, and paste it. The key is kept in your browser's localStorage only and sent as a request header to your local backend — it takes precedence over backend/.env and can be removed anytime. Keys pasted without a provider choice are auto-detected by prefix (gsk_… → Groq, sk-… → OpenAI).
Tests (backend):
cd backend && source .venv/bin/activate
pytest tests/ -q| Method | Path | Purpose |
|---|---|---|
POST |
/repositories/analyze |
Start indexing a public GitHub repository |
GET |
/repositories/{id}/status |
Indexing progress (stage + real stats) |
POST |
/repositories/{id}/query |
Hybrid retrieval + answer + sources + subgraph |
GET |
/repositories/{id}/graph |
Focused subgraph (`?focus=<node |
GET |
/repositories/{id}/symbols |
Symbol search for impact/focus pickers |
POST |
/repositories/{id}/impact |
"What breaks if I change this?" |
GET |
/repositories/{id}/source |
Read-only source lines for the evidence viewer |
- Tree-sitter — real ASTs for TS/JS/Python with one API, fault-tolerant parsing of imperfect code, and a clean seam for adding languages. A regex fallback keeps indexing alive if grammars are unavailable.
- NetworkX — the graph is small (thousands of nodes), in-memory, and pickle-persisted; BFS/shortest-path/impact queries are one-liners. A graph database would add infrastructure without adding capability here.
- FAISS (IndexFlatIP) — exact cosine similarity over normalized vectors is fast at this scale and removes approximate-index tuning as a moving part. Chunk metadata lives in a JSON sidecar next to the index.
- Hybrid scoring over a learned reranker —
0.6/0.4weighted fusion is deterministic, explainable in an interview, and configurable viaSEMANTIC_WEIGHT/GRAPH_WEIGHT. A cross-encoder would add latency and opacity for marginal gain at this corpus size. - Structural chunking — chunks map 1:1 to symbols, so a citation is always a real function/method/class with exact line ranges, and graph nodes join retrieval results without a fuzzy join step.
- Retrieval works without an API key — local embeddings by default; the LLM is optional plumbing, not a prerequisite.
- Supported languages: TypeScript, JavaScript, Python. Other files are ignored (the parser abstraction exists for more).
- Approximate call graph: no type inference. Calls resolve through same-file symbols, import bindings, and class receivers (
this/self, imported classes). Dynamic dispatch (obj.method()on an untyped local) does not resolve. - Public repositories only, size-capped (
MAX_REPOSITORY_SIZE_MB). - Repository code is never executed — analysis is purely static, so runtime behavior is only inferred, never observed.
- Entity matching is lexical: graph retrieval fires when the question names a symbol/file (or close variant); purely conceptual questions lean on the semantic stream.
backend/
app/
api/ # FastAPI routes (repositories, query, graph, impact)
core/ # configuration
ingestion/ # GitHub loader + file filtering
parsing/ # Tree-sitter parser, regex fallback, chunker
graph/ # graph construction + deterministic queries
retrieval/ # embeddings, FAISS store, graph/hybrid retrievers, reranker
generation/ # prompts, context builder, LLM provider
models/ # API schemas
services/ # indexing + query orchestration
tests/ # focused unit tests over a fixture repository
frontend/
app/ # Next.js App Router (workspace page, layout, tokens)
components/ # repository / query / graph / impact / code / layout
lib/ # api client, shared types, UI constants