Home › Leiden Community Detection

Leiden Community Detection Without Embeddings

Graphify groups related code, docs and diagrams into communities using the Leiden algorithm over graph topology alone. No vector embeddings, no vector database, no separate similarity index.

Why not embeddings?

Most code-RAG systems chunk a repo, embed each chunk, and cluster the resulting vectors. That pipeline has three failure modes Graphify is designed to avoid:

Graphify sidesteps all three by running Leiden directly on the graph the Tree-sitter AST pass and the semantic pass produce. Edge density is the clustering signal.

How semantic similarity still participates

The semantic pass emits semantically_similar_to edges between nodes that look conceptually related but have no structural connection — a function in code and a concept in a paper describing the same algorithm, for example. These edges are marked INFERRED with a confidence score, and they live in the same graph as the structural edges. Leiden sees them, edge density goes up where it should, and communities form around conceptual affinity as well as call structure. No vector index required.

What a community looks like in the output

After clustering, each community becomes a section in GRAPH_REPORT.md and, optionally, a standalone article when you pass --wiki. The report lists:

For a worked example, the httpx corpus yields 6 communities with god nodes Client, AsyncClient, Response and Request, and surfaces the surprise edge DigestAuth → Response.

Tech choice

Leiden is implemented via graspologic. The rest of the graph layer is NetworkX. The entire clustering stage is pure-Python and runs locally, consistent with Graphify's no-server, no-telemetry posture.

Related topics