Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

CodeGraph Agent (POC)

A code intelligence agent that indexes repositories into a PostgreSQL knowledge graph (symbols + edges + optional pgvector embeddings), then answers questions through a LangGraph pipeline: classify intent → retrieve evidence → decide whether to call an LLM → synthesize an answer with latency metrics.

Single entrypoint for “the agent”: main.run_query() in main.py. The FastAPI app (ui/api.py) and CLI both call it so behavior stays consistent.

For day-to-day commands, benchmarks, and API curl examples, see sections below. For a longer runbook (ingest internals, indexes, verification checklists), see ARCHITECTURE_AND_RUNBOOK.md.


Table of contents

  1. High-level design (HLD)
  2. Low-level design (LLD)
  3. Project structure
  4. How the code is organized (for learning)
  5. Sample queries
  6. Quick start
  7. Common questions (FAQ)
  8. Tests & build

High-level design (HLD)

What the system does

Phase Role
Ingest Walk files → parse (Tree-sitter) → extract nodes/edges → embed nodes → upsert into kg_nodes / kg_edges.
Query Classify question → run retrieval (SQL graph and/or vectors) → after seeing hits, pick kg_only vs kg_plus_llm → optional LLM → synthesize answer + metrics.

Logical components

  CLI or HTTP API
        |
        v
   run_query()  --->  LangGraph (graph.py): router -> retriever -> post-router -> LLM? -> synthesiser
        |
        +--> PostgreSQL (kg_nodes, kg_edges, pgvector)
        +--> Optional LLM (Ollama / OpenAI-compatible)
        +--> Telemetry (latencies, CPU time, evidence counts)

Query pipeline (conceptual)

flowchart LR
  A[User question] --> B[Pre-router: intent + query_type]
  B --> C[KG retriever]
  C --> D[Post-router: route_taken]
  D -->|kg_only| E[Synthesiser]
  D -->|kg_plus_llm| F[LLM]
  F --> E
  E --> G[Answer + metrics]
Loading

Why two routers? The pre-router only estimates what kind of question it is (structural vs semantic, callers, imports, etc.). The post-KG router chooses the final route using mode, use_llm, query_type, and kg_hits, so cpu_first can stay KG-only when the graph already has evidence.

Runtime modes (product knobs)

Mode Retrieval bias Typical use
cpu_first Structural SQL on the graph; no vector in this mode Fast, deterministic symbol/call questions
gpu_first KG first, then always LLM Narration / explanation with evidence
vector_rag pgvector similarity + optional LIKE fallback “Concept” questions
hybrid Vector seeds + bounded graph expansion Richer context than vector-only

Multi-project data in one database

All projects share the same tables. Rows are scoped by repo_id + commit_hash. Use the same pair on ingest and on every query, or you will see empty or wrong results.

CLI / API default: main.py and /query use repo_id=default if you omit it. The runbook often ingests with --repo-id local. If those differ, you get 0 KG hits even though the database has data—always pass --repo-id (and commit_hash) to match ingest, or run python -m tools.verify_kg --repo-id <yours> to confirm counts.


Low-level design (LLD)

LangGraph nodes and state

Node File Responsibility
router_node nodes/router.py query_type, router_intent, router_confidence, t_after_router — no final route_taken.
kg_retriever_node nodes/kg_retriever.py SQL + optional vector/hybrid expansion → kg_results, semantic_results, kg_hits, evidence_files, t_after_kg.
post_kg_router_node nodes/post_kg_router.py Deterministic route_taken: kg_only or kg_plus_llm.
llm_node nodes/llm_node.py Optional streaming; builds prompt from question + top KG rows.
synthesiser_node nodes/synthesiser.py Final text + metrics (e2e / router / kg / llm / synth / cpu_process_time_ms).

State contract: state.py (AgentState TypedDict) — ids, query, mode, retrieval results, timings, optional stream_llm / llm_stream_sink for streaming.

Wiring: graph.py — router → kg_retriever → post_kg_router → (llm | synthesiser) → synthesiser.

Ingest path

Step File Notes
Scan + checksums nodes/ingest.py Incremental skip/rewrite
Parse + extract nodes/ast_builder.py, nodes/extractors/* Per-language Tree-sitter
Persist nodes/kg_writer.py, db/queries.py Upsert nodes, resolve edges, ingest_log

Storage (PostgreSQL + pgvector)

Table Purpose
kg_nodes Symbols/files: fqn, file_path, lines, embedding (384-d), scoped by repo_id, commit_hash
kg_edges CALLS, IMPORTS, CONTAINS, … — src_id / dst_id
ingest_log Per-file checksum / status for incremental ingest

Schema and ANN indexes: db/schema.py (IVFFlat by default; optional HNSW attempt when pgvector supports it).

API surface

Endpoint Purpose
GET /health Liveness
POST /query JSON in → full JSON out (same graph as CLI)
POST /query/stream SSE: chunks + final payload (LLM streaming when routed to LLM)
GET /metrics Simple counters from in-process events

Evaluation

  • evaluation/benchmark_queries.yaml — explicit mode × query × use_llm matrix.
  • evaluation/benchmark_runner.py — runs run_query per row; can write benchmark_report.json.
  • evaluation/metrics.py — aggregate latency + per-mode stats.

Operator tools

  • python -m tools.verify_kg — counts and optional vector probe (POSTGRES_DSN).

Project structure

Source layout (ignore build/ for development; it is build output):

CodeGraphh/
├── main.py                 # CLI: ingest | query | chat; run_query()
├── graph.py                # LangGraph: nodes + conditional edges
├── state.py                # AgentState
├── config.py               # pydantic-settings / .env
├── pyproject.toml
├── README.md               # This file (overview + HLD/LLD + FAQ)
├── ARCHITECTURE_AND_RUNBOOK.md   # Deep runbook & checklists
│
├── db/
│   ├── connection.py       # psycopg2 from DSN
│   ├── schema.py           # DDL + vector indexes
│   └── queries.py          # Upsert / incremental helpers
│
├── nodes/
│   ├── ingest.py
│   ├── ast_builder.py
│   ├── kg_writer.py
│   ├── router.py           # Pre-KG classification
│   ├── post_kg_router.py   # Post-KG route table
│   ├── kg_retriever.py # SQL + vector + hybrid expansion
│   ├── llm_node.py
│   ├── synthesiser.py
│   ├── embeddings.py
│   └── extractors/         # Tree-sitter: python, java, ts, …
│
├── evaluation/
│   ├── benchmark_runner.py
│   ├── benchmark_queries.yaml
│   └── metrics.py
│
├── tools/
│   └── verify_kg.py
│
├── ui/
│   ├── api.py              # FastAPI
│   └── streamlit_app.py
│
└── tests/
    ├── unit/
    └── integration/

How the code is organized (for learning)

  1. Separation of “workflow” vs “business logic” — graph.py only connects nodes; each node file stays testable (e.g. decide_route_after_kg in post_kg_router.py).
  2. Graph-first, then LLM — Evidence is fetched before committing to an LLM call; routing after retrieval avoids promising kg_only when there are no hits (unless policy says otherwise).
  3. Modes are explicit — Retrieval behavior differs by mode in kg_retriever.py so benchmarks can show real differences (cpu_first vs vector_rag vs hybrid).
  4. One agent entrypoint — Avoid duplicating orchestration in ui/api.py; call run_query so CLI and HTTP stay aligned.
  5. Typed state — AgentState documents what flows through the graph and helps when extending nodes.

Suggested reading order for a walkthrough: state.py → graph.py → nodes/router.py → nodes/kg_retriever.py → nodes/post_kg_router.py → nodes/llm_node.py → nodes/synthesiser.py → main.py → ui/api.py.


Sample queries

Use after ingesting the same codebase with matching repo_id / commit_hash (examples below assume this repo was ingested as myrepo + HEAD).

Structural (good with cpu_first, use_llm=false)

Question
Where is router_node defined?
Who calls run_query?
Callees of router_node
What imports kg_writer?
Where is post_kg_router_node defined?

Semantic / vector (good with vector_rag or hybrid)

Question
Postgres database connection settings
LLM provider configuration
Embedding vector similarity retrieval
How does hybrid mode expand from vector seeds?

Explanation (good with gpu_first or hybrid + use_llm=true)

Question
How does the LangGraph pipeline wire router and retriever?

Negative test (expect no KG hits unless symbol exists)

Question
Who calls nonexistent_symbol_xyz123?

CLI examples:

python main.py query --repo-id myrepo --commit-hash HEAD --mode cpu_first "who calls run_query"
python main.py query --repo-id myrepo --mode vector_rag "postgres database connection settings"
python main.py chat --repo-id myrepo --mode hybrid

Quick start

  1. Copy .env.example → .env (set POSTGRES_DSN, optional LLM vars).
  2. pip install -e ".[dev]"
  3. python db/schema.py
  4. Ingest:
    python main.py ingest --project-root "e:/path/to/repo" --repo-id myrepo --commit-hash HEAD
  5. Query / chat / API as above—use the same repo_id / commit_hash as step 4.

Useful flags: --mode, --use-llm, -v / --verbose, --json, --stream (LLM token streaming when the LLM path runs). In chat: help, menu, /mode, /llm, /stream.

Servers:

python -m uvicorn ui.api:app --host 127.0.0.1 --port 8004
streamlit run ui/streamlit_app.py --server.port 8502 --server.headless true

Benchmarks: python -m evaluation.benchmark_runner --repo-id myrepo --commit-hash HEAD
Verify KG: python -m tools.verify_kg --repo-id myrepo --commit-hash HEAD


Common questions (FAQ)

What are repo_id and commit_hash?
Labels that scope all KG rows in Postgres. Think: which project and which snapshot. Ingest and query must use the same values.

Is the “graph” stored in Postgres?
Yes — nodes and edges live in kg_nodes / kg_edges. LangGraph is only the orchestration graph in Python, not stored in the DB.

If I index multiple projects, does one search hit all of them?
No. Each query is filtered by the repo_id and commit_hash you pass in.

Why is my structural query returning zero hits?
Most often: wrong repo_id (CLI defaults to default; many ingests use local or myrepo). Also check commit_hash, that files were ingested, and the symbol exists. Run python -m tools.verify_kg --repo-id YOUR_ID --commit-hash HEAD for scoped counts.

When does the LLM run?
Only when post_kg_router sets route_taken to kg_plus_llm (depends on mode, use_llm, structural vs semantic, and kg_hits). See nodes/post_kg_router.py and the routing table in ARCHITECTURE_AND_RUNBOOK.md.

What is streaming?
When the LLM path runs, llm_node can stream tokens (CLI --stream, chat /stream on, or POST /query/stream). Non-LLM paths have nothing to stream.

Where is the “source of truth” for behavior?
Code paths in nodes/*.py and graph.py; narrative + ops detail in ARCHITECTURE_AND_RUNBOOK.md.


Tests & build

pytest -q tests/unit tests/integration

(Full pytest -q may fail on legacy codegraph.* tests if that layout is not installed.)

python -m build --wheel --sdist

Artifacts under dist/.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages