Overview
codetopo is a live software-topology engine for AI coding agents built in C++26. It parses multi-language source trees using Tree-sitter (13 languages), builds and persists a relational symbol graph in SQLite, and continuously synchronizes that graph in real time as files change. It fuses static AST relationships, protocol-level HTTP interactions, and empirical runtime traces into an epistemically categorized topology model.
Exposing 44 MCP tools over stdio (JSON-RPC 2.0), codetopo empowers AI coding agents (such as GitHub Copilot, Cursor, Windsurf, Claude, and Antigravity) to perform deterministic structural reasoning — navigating call graphs, calculating blast radius, summarizing architectural boundaries, computing semantic graph diffs, and evaluating measurable graph quality — without hallucinating code structure or blowing up context windows with raw source dumps.
Benchmarked on enterprise codebases exceeding 450,000 files, codetopo indexes 100,000 files in ~10 minutes with zero crashes.
Key Features
- Multi-Language AST Parsing — Native Tree-sitter grammar support for 13 languages: C, C++, C#, Go, YAML, TypeScript, JavaScript, Python, Rust, Java, Bash, PowerShell, and Batch.
- Structural-First Persistence — Relational symbol graph stored in SQLite with secondary query indexes, FTS5 symbol search, and trigram content search.
- Continuously Synchronized — Native OS file system watching (
FSEventson macOS,inotifyon Linux,ReadDirectoryChangesWon Windows) with debounce, in-flight coalescing, and non-blocking committed WAL reads. - Targeted Symbol Attraction — Distinguishes file parsing from graph linking: unchanged files with dangling or ambiguous references are selectively rebound in O(|Dchanged| · d̄) time without repository-wide rescans.
- Epistemic Edge Provenance (Schema v15) — Explicit certainty and evidence provenance for every edge:
RESOLVED(deterministic static AST),PROBABLE(type/arity heuristics),POSSIBLE(syntactic candidates), andOBSERVED(runtime-verified). - Runtime Trace Fusion — Ingests empirical OpenTelemetry/OTLP traces and extracts protocol-level HTTP client calls to validate and boost static call graph edges.
- Measurable Graph Quality —
codetopo qualityreports exact call resolution percentages by language, edge kind distributions, confidence histograms, and reference disambiguation metrics. - PR Blast Radius & Semantic Diff — BFS dependency traversal (
impact_of), Git diff blast-radius analysis (detect_changes), and semantic graph diffing between commits or working tree (diff/graph_diff). - Architectural Intelligence — Structural PageRank centrality, directory-based clustering, cohesion, coupling, and extractability scoring (
get_architecture,dependency_cluster). - 44 MCP Tools — Built for agents following “structure before source” and “batch, don’t bounce” workflows with durable stable keys and candidate fallbacks.
- Crash-Resilient Supervisor — Multi-process architecture with automatic restart, file quarantine, progress tracking, and single-thread fallback (up to 10 retries).
- Zero-Config Editor Setup — One command writes project-scoped
.mcp.json(cross-agent standard),.vscode/mcp.json, Cursor/Windsurf/Copilot configs, and.github/skills/.
Core Architectural Concepts
Structural-First + Continuously Synchronized
codetopo maintains a live structural model of a software system while the codebase is actively modified by human engineers and AI coding agents:
- File Parsing: Detects file modifications using xxHash64 content hashes and mtime metadata. Tree-sitter parses only modified files into CSTs in parallel worker threads using dedicated per-thread arena allocators.
- Graph Maintenance (Symbol Attraction): When a file introduces, renames, or deletes a symbol definition, unchanged files in the repository may contain unresolved or ambiguous references that can now resolve (or that were broken). Rather than scanning the entire repository, codetopo’s Symbol Attraction mechanism queries the SQLite index for dangling or ambiguous references matching the changed symbol names, selectively re-resolving only the affected call sites in O(|Dchanged| · d̄) time (scaling with changed definitions and average reference degree, rather than repository size).
Epistemic Edge Provenance (Schema v15)
The code graph combines structural facts from multiple origins, associating each edge with an explicit confidence score (c ∈ [0.0, 1.0]) and evidence provenance:
RESOLVED(c = 1.0): Deterministic static AST resolution (unambiguous symbol in scope or unique global symbol).PROBABLE(c ∈ [0.80, 0.99]): High-confidence heuristic resolution (receiver type inference, matching parameter arity, or single-candidate namespace lookup).POSSIBLE(c ∈ [0.30, 0.79]): Syntactically plausible ambiguous candidate (common method names like.get()or.parse()).OBSERVED: Empirically confirmed by runtime call traces (OpenTelemetry / OTLP) or protocol-level HTTP client observations, boosting edge confidence and resolving dynamic dispatch ambiguity.
Quick Start
Build from Source
# Prerequisites: CMake >= 3.20, C++26 compiler (Clang 19+, GCC 14+, MSVC 19.40+), vcpkg
cmake --preset release
cmake --build build --config Release
One-Command Setup
The fastest way to index a project and configure editor MCP settings:
codetopo init --root /path/to/repo
This single command will:
- Scan and index the entire repository into
.codetopo/index.sqlite. - Generate project-scoped MCP configuration files (
.mcp.jsonand.vscode/mcp.json). - Install agent skills into
.github/skills/(auto-discovered by GitHub Copilot). - Write non-destructive agent guidance markers into
AGENTS.mdand.github/copilot-instructions.md.
CLI Commands
| Command | Description |
|---|---|
codetopo init | Index a repo and configure editor MCP settings |
codetopo index | Build or incrementally update the code graph |
codetopo mcp | Start the MCP server over stdio with optional --watch |
codetopo watch | Watch for file changes and continuously re-index |
codetopo quality | Report graph quality, language resolution rates, and confidence metrics |
codetopo diff | Compute semantic graph diff between working tree, commits, or index |
codetopo workspace | Add, remove, refresh, or list extra workspace roots |
codetopo query | Run an ad-hoc MCP tool query directly from the terminal |
codetopo parse-file | Parse a single file and output AST/symbol diagnostics |
codetopo skills | List or install agent skill files |
codetopo doctor | Verify database health and index integrity |
MCP Tools (44 Tools across 6 Categories)
codetopo exposes 44 specialized tools following the “structure before source” philosophy:
1. Discovery & Directory Navigation
dir_tree— Return full directory subtree up to depth N with file sizes and detected languages.dir_list— List files and subdirectories in a given directory (one level).file_search— Search files by GLOB path pattern (e.g.*numa*,src/**/*.h).file_deps— File-level include, import, and module dependencies.
2. Symbol Inspection & Structural Context
context_by_name— Resolve symbol by name and return full structural context or disambiguation candidates.context_for— Full structural context: definition, source snippet, callers, callees, container, siblings, and bases.symbol_search— Full-text symbol search (SQLite FTS5) with kind, file, and match filters.symbol_list— Filter symbols by kind, file, or name glob without FTS.symbols_in_path— List symbols under a directory subtree in lean format.symbol_get/symbol_get_batch— Get detailed symbol records by durablestable_keyornode_id.file_summary/file_overview— Detailed symbol listings and signatures defined in a file.entrypoints— Find natural entry points (main,DllMain, etc.).
3. Call Graph, Flow & Blast Radius
callers_approx/callees_approx— Groupable inbound and outbound call relationships with candidate fallback.references— Find all usages of a symbol across the workspace.impact_of— Transitive blast radius (BFS traversal) of changing a symbol.detect_changes— Git diff blast-radius analysis: maps git diff to changed symbols and walks reverse call edges.subgraph— Extract local dependency neighborhood graph around seed symbol IDs.shortest_path— Shortest dependency path connecting two symbols.find_implementations— Types implementing an interface or inheriting from a base class.
4. Code, State & Architecture
get_architecture— Architectural summary: clusters, hotspots, boundaries, cohesion, and coupling.method_fields— Classifythis.Xfield accesses and internal/external outgoing calls.dependency_cluster— Group methods by shared field access patterns with extractability scores.find_similar— Near-duplicate function detection using MinHash fingerprints of AST leaf trigrams.source_at— Targeted raw source line reader by line range.code_search— Substring search across all indexed source files using trigram FTS.list_http_calls— Extracted HTTP client call references with endpoints and verbs.
5. Runtime Traces, Evidence & Quality
ingest_traces— Ingest empirical runtime call traces (OTLP/OpenTelemetry JSON) to boost edge confidence.get_traces— Query ingested runtime traces by caller/callee and call count.get_edge_evidence— Query epistemic provenance and runtime observation evidence for call edges.graph_quality— Comprehensive fidelity report: resolution rates by language, edge kinds, confidence buckets.graph_diff— Semantic graph diff between working tree, commits, or index.
6. Server, Health & Workspace Management
server_info/server_health— Server capabilities, schema version (v15), uptime, and readiness probes.repo_stats— File counts and index metadata.reindex— Trigger a background re-index with non-blocking committed reads.workspace_add/workspace_refresh/workspace_remove/workspace_list— Multi-root workspace indexing and remapped merges.workspace_job_status/workspace_job_cancel— Monitor and manage asynchronous background workspace jobs.
Performance & Enterprise Scale
Tested on an enterprise C++/C# monorepo containing over 450,000 files (Windows 11, 16 threads, NVMe SSD):
| Dataset Scale | Cold Index Time | Throughput | Indexed Symbols | Indexed Edges |
|---|---|---|---|---|
| 100,000 files | 659 seconds (~10 min) | 240 files/sec | 3,100,000 | 2,500,000 |
| 162,000 files | 971 seconds (~16 min) | 260 files/sec | 4,200,000 | 3,500,000 |
Memory footprint remains bounded during continuous re-indexing thanks to dedicated arena allocation pools, while committed SQLite WAL snapshots guarantee conflict-free, concurrent query reads for AI clients during active background indexing.