A code intelligence agent that indexes repositories into a PostgreSQL knowledge graph (symbols + edges + optional pgvector embeddings), then answers questions through a LangGraph pipeline: classify intent → retrieve evidence → decide whether to call an LLM → synthesize an answer with latency metrics.
Single entrypoint for “the agent”: main.run_query() in main.py. The FastAPI app (ui/api.py) and CLI both call it so behavior stays consistent.
For day-to-day commands, benchmarks, and API curl examples, see sections below. For a longer runbook (ingest internals, indexes, verification checklists), see ARCHITECTURE_AND_RUNBOOK.md.
- High-level design (HLD)
- Low-level design (LLD)
- Project structure
- How the code is organized (for learning)
- Sample queries
- Quick start
- Common questions (FAQ)
- Tests & build
| Phase | Role |
|---|---|
| Ingest | Walk files → parse (Tree-sitter) → extract nodes/edges → embed nodes → upsert into kg_nodes / kg_edges. |
| Query | Classify question → run retrieval (SQL graph and/or vectors) → after seeing hits, pick kg_only vs kg_plus_llm → optional LLM → synthesize answer + metrics. |
CLI or HTTP API
|
v
run_query() ---> LangGraph (graph.py): router -> retriever -> post-router -> LLM? -> synthesiser
|
+--> PostgreSQL (kg_nodes, kg_edges, pgvector)
+--> Optional LLM (Ollama / OpenAI-compatible)
+--> Telemetry (latencies, CPU time, evidence counts)
flowchart LR
A[User question] --> B[Pre-router: intent + query_type]
B --> C[KG retriever]
C --> D[Post-router: route_taken]
D -->|kg_only| E[Synthesiser]
D -->|kg_plus_llm| F[LLM]
F --> E
E --> G[Answer + metrics]
Why two routers? The pre-router only estimates what kind of question it is (structural vs semantic, callers, imports, etc.). The post-KG router chooses the final route using mode, use_llm, query_type, and kg_hits, so cpu_first can stay KG-only when the graph already has evidence.
| Mode | Retrieval bias | Typical use |
|---|---|---|
cpu_first |
Structural SQL on the graph; no vector in this mode | Fast, deterministic symbol/call questions |
gpu_first |
KG first, then always LLM | Narration / explanation with evidence |
vector_rag |
pgvector similarity + optional LIKE fallback | “Concept” questions |
hybrid |
Vector seeds + bounded graph expansion | Richer context than vector-only |
All projects share the same tables. Rows are scoped by repo_id + commit_hash. Use the same pair on ingest and on every query, or you will see empty or wrong results.
CLI / API default: main.py and /query use repo_id=default if you omit it. The runbook often ingests with --repo-id local. If those differ, you get 0 KG hits even though the database has data—always pass --repo-id (and commit_hash) to match ingest, or run python -m tools.verify_kg --repo-id <yours> to confirm counts.
| Node | File | Responsibility |
|---|---|---|
router_node |
nodes/router.py |
query_type, router_intent, router_confidence, t_after_router — no final route_taken. |
kg_retriever_node |
nodes/kg_retriever.py |
SQL + optional vector/hybrid expansion → kg_results, semantic_results, kg_hits, evidence_files, t_after_kg. |
post_kg_router_node |
nodes/post_kg_router.py |
Deterministic route_taken: kg_only or kg_plus_llm. |
llm_node |
nodes/llm_node.py |
Optional streaming; builds prompt from question + top KG rows. |
synthesiser_node |
nodes/synthesiser.py |
Final text + metrics (e2e / router / kg / llm / synth / cpu_process_time_ms). |
State contract: state.py (AgentState TypedDict) — ids, query, mode, retrieval results, timings, optional stream_llm / llm_stream_sink for streaming.
Wiring: graph.py — router → kg_retriever → post_kg_router → (llm | synthesiser) → synthesiser.
| Step | File | Notes |
|---|---|---|
| Scan + checksums | nodes/ingest.py |
Incremental skip/rewrite |
| Parse + extract | nodes/ast_builder.py, nodes/extractors/* |
Per-language Tree-sitter |
| Persist | nodes/kg_writer.py, db/queries.py |
Upsert nodes, resolve edges, ingest_log |
| Table | Purpose |
|---|---|
kg_nodes |
Symbols/files: fqn, file_path, lines, embedding (384-d), scoped by repo_id, commit_hash |
kg_edges |
CALLS, IMPORTS, CONTAINS, … — src_id / dst_id |
ingest_log |
Per-file checksum / status for incremental ingest |
Schema and ANN indexes: db/schema.py (IVFFlat by default; optional HNSW attempt when pgvector supports it).
| Endpoint | Purpose |
|---|---|
GET /health |
Liveness |
POST /query |
JSON in → full JSON out (same graph as CLI) |
POST /query/stream |
SSE: chunks + final payload (LLM streaming when routed to LLM) |
GET /metrics |
Simple counters from in-process events |
evaluation/benchmark_queries.yaml— explicit mode × query ×use_llmmatrix.evaluation/benchmark_runner.py— runsrun_queryper row; can writebenchmark_report.json.evaluation/metrics.py— aggregate latency + per-mode stats.
python -m tools.verify_kg— counts and optional vector probe (POSTGRES_DSN).
Source layout (ignore build/ for development; it is build output):
CodeGraphh/
├── main.py # CLI: ingest | query | chat; run_query()
├── graph.py # LangGraph: nodes + conditional edges
├── state.py # AgentState
├── config.py # pydantic-settings / .env
├── pyproject.toml
├── README.md # This file (overview + HLD/LLD + FAQ)
├── ARCHITECTURE_AND_RUNBOOK.md # Deep runbook & checklists
│
├── db/
│ ├── connection.py # psycopg2 from DSN
│ ├── schema.py # DDL + vector indexes
│ └── queries.py # Upsert / incremental helpers
│
├── nodes/
│ ├── ingest.py
│ ├── ast_builder.py
│ ├── kg_writer.py
│ ├── router.py # Pre-KG classification
│ ├── post_kg_router.py # Post-KG route table
│ ├── kg_retriever.py # SQL + vector + hybrid expansion
│ ├── llm_node.py
│ ├── synthesiser.py
│ ├── embeddings.py
│ └── extractors/ # Tree-sitter: python, java, ts, …
│
├── evaluation/
│ ├── benchmark_runner.py
│ ├── benchmark_queries.yaml
│ └── metrics.py
│
├── tools/
│ └── verify_kg.py
│
├── ui/
│ ├── api.py # FastAPI
│ └── streamlit_app.py
│
└── tests/
├── unit/
└── integration/
- Separation of “workflow” vs “business logic” —
graph.pyonly connects nodes; each node file stays testable (e.g.decide_route_after_kginpost_kg_router.py). - Graph-first, then LLM — Evidence is fetched before committing to an LLM call; routing after retrieval avoids promising
kg_onlywhen there are no hits (unless policy says otherwise). - Modes are explicit — Retrieval behavior differs by
modeinkg_retriever.pyso benchmarks can show real differences (cpu_firstvsvector_ragvshybrid). - One agent entrypoint — Avoid duplicating orchestration in
ui/api.py; callrun_queryso CLI and HTTP stay aligned. - Typed state —
AgentStatedocuments what flows through the graph and helps when extending nodes.
Suggested reading order for a walkthrough: state.py → graph.py → nodes/router.py → nodes/kg_retriever.py → nodes/post_kg_router.py → nodes/llm_node.py → nodes/synthesiser.py → main.py → ui/api.py.
Use after ingesting the same codebase with matching repo_id / commit_hash (examples below assume this repo was ingested as myrepo + HEAD).
| Question |
|---|
Where is router_node defined? |
Who calls run_query? |
Callees of router_node |
What imports kg_writer? |
Where is post_kg_router_node defined? |
| Question |
|---|
| Postgres database connection settings |
| LLM provider configuration |
| Embedding vector similarity retrieval |
| How does hybrid mode expand from vector seeds? |
| Question |
|---|
| How does the LangGraph pipeline wire router and retriever? |
| Question |
|---|
Who calls nonexistent_symbol_xyz123? |
CLI examples:
python main.py query --repo-id myrepo --commit-hash HEAD --mode cpu_first "who calls run_query"
python main.py query --repo-id myrepo --mode vector_rag "postgres database connection settings"
python main.py chat --repo-id myrepo --mode hybrid- Copy
.env.example→.env(setPOSTGRES_DSN, optional LLM vars). pip install -e ".[dev]"python db/schema.py- Ingest:
python main.py ingest --project-root "e:/path/to/repo" --repo-id myrepo --commit-hash HEAD - Query / chat / API as above—use the same
repo_id/commit_hashas step 4.
Useful flags: --mode, --use-llm, -v / --verbose, --json, --stream (LLM token streaming when the LLM path runs). In chat: help, menu, /mode, /llm, /stream.
Servers:
python -m uvicorn ui.api:app --host 127.0.0.1 --port 8004
streamlit run ui/streamlit_app.py --server.port 8502 --server.headless trueBenchmarks: python -m evaluation.benchmark_runner --repo-id myrepo --commit-hash HEAD
Verify KG: python -m tools.verify_kg --repo-id myrepo --commit-hash HEAD
What are repo_id and commit_hash?
Labels that scope all KG rows in Postgres. Think: which project and which snapshot. Ingest and query must use the same values.
Is the “graph” stored in Postgres?
Yes — nodes and edges live in kg_nodes / kg_edges. LangGraph is only the orchestration graph in Python, not stored in the DB.
If I index multiple projects, does one search hit all of them?
No. Each query is filtered by the repo_id and commit_hash you pass in.
Why is my structural query returning zero hits?
Most often: wrong repo_id (CLI defaults to default; many ingests use local or myrepo). Also check commit_hash, that files were ingested, and the symbol exists. Run python -m tools.verify_kg --repo-id YOUR_ID --commit-hash HEAD for scoped counts.
When does the LLM run?
Only when post_kg_router sets route_taken to kg_plus_llm (depends on mode, use_llm, structural vs semantic, and kg_hits). See nodes/post_kg_router.py and the routing table in ARCHITECTURE_AND_RUNBOOK.md.
What is streaming?
When the LLM path runs, llm_node can stream tokens (CLI --stream, chat /stream on, or POST /query/stream). Non-LLM paths have nothing to stream.
Where is the “source of truth” for behavior?
Code paths in nodes/*.py and graph.py; narrative + ops detail in ARCHITECTURE_AND_RUNBOOK.md.
pytest -q tests/unit tests/integration(Full pytest -q may fail on legacy codegraph.* tests if that layout is not installed.)
python -m build --wheel --sdistArtifacts under dist/.