Home › Tree-sitter AST Extraction

Tree-sitter AST Extraction Across 19 Languages

The first pass of the Graphify pipeline is a deterministic Tree-sitter walk over every code file it finds. No LLM, no embeddings, no network — just AST nodes converted directly into graph nodes and edges.

Why Tree-sitter

Tree-sitter is an incremental parser generator with battle-tested grammars for every mainstream language. Graphify uses it because:

Languages supported

Graphify ships grammars for 19 languages out of the box:

FamilyLanguages
ScriptingPython, JavaScript, TypeScript, Ruby, PHP, Lua, PowerShell
SystemsGo, Rust, C, C++, Zig, Swift, Objective-C
JVM / .NETJava, Kotlin, Scala, C#
BEAMElixir

Adding a language is a matter of dropping its Tree-sitter grammar into the extractor and writing a small node-to-concept mapper — see ARCHITECTURE.md in the repo for the exact steps.

What Graphify pulls out of the AST

Deterministic vs semantic

Everything the AST pass emits is marked EXTRACTED (confidence 1.0). It's not a guess — the token is in the file. The second pass, run by Claude subagents against docs/papers/images, emits INFERRED edges with confidence scores, and anything the model is unsure about is tagged AMBIGUOUS. Those tags survive all the way into graph.json so a reviewer can always separate ground truth from model judgment.

Related topics