Fast static Python call graphs in Rust.
Parses Python source files and produces a directed graph of defines/uses relationships between modules, classes, functions, and methods. No Python runtime required — uses ruff's parser for AST parsing.
- CLI for DOT, TGF, text, and JSON output
- Module-level and symbol-level graph modes
- Real-world corpus smoke testing and a published GitHub Pages report
GitHub Pages report: https://nwyin.github.io/pycg-rs/
cargo install --git https://github.com/nwyin/pycg-rs --bin pycgFor local development from a checkout:
cargo install --path . --forceThis installs the CLI as pycg, typically at ~/.cargo/bin/pycg.
If pycg is not on your PATH, add this to your shell config:
export PATH="$HOME/.cargo/bin:$PATH"# Analyze a package and print a plain-text dependency list
pycg analyze mypackage/ --format text
# Emit machine-readable JSON
pycg analyze mypackage/ --format json > graph.json
# Render an SVG with GraphViz
pycg analyze mypackage/ --defines --uses --grouped --annotated mypackage/ | dot -Tsvg -o callgraph.svg# Analyze Python files, output DOT format (default)
pycg analyze src/**/*.py
# Output as plain text dependency list
pycg analyze mypackage/ --format text
# Analyze a package and keep module names rooted at the repo src dir
pycg analyze mypackage/ --root .
# Show both defines and uses edges, colored by file
pycg analyze mypackage/ -d -u --colored --grouped
# Pipe DOT to graphviz for SVG
pycg analyze mypackage/ | dot -Tsvg -o callgraph.svgpycg [OPTIONS] <FILES>...
Arguments:
<FILES>... Python source files or directories to analyze
Options:
-d, --defines Draw defines edges
-u, --uses Draw uses edges
-m, --modules Show module-level import dependencies instead of
symbol-level call graph
-c, --colored Color nodes by file
-g, --grouped Group nodes by namespace
-a, --annotated Annotate nodes with file:line info
-r, --root <ROOT> Root directory for module name resolution
--format <FORMAT> Output format: dot, tgf, text, json [default: dot]
--rankdir <DIR> GraphViz rank direction [default: TB]
-v, --verbose Enable verbose logging (-vv for debug)
If neither --defines nor --uses is specified, uses edges are shown by default.
# Inspect call dependencies in a package
pycg analyze src/
# Show only defines edges
pycg analyze src/ --defines
# Render a grouped, annotated SVG
pycg analyze src/ --root . --defines --uses --grouped --annotated | dot -Tsvg -o callgraph.svg
# Module-level import dependency graph
pycg analyze src/ --modules | dot -Tsvg -o imports.svg
# Debug analyzer decisions
pycg analyze src/ -vv- dot — GraphViz DOT format, suitable for rendering with
dot,neato, etc. - tgf — Trivial Graph Format
- text — Plain text dependency list with
[D]/[U]tags - json — Machine-readable graph output with nodes, edges, stats, and diagnostics for unresolved, external, ambiguous, and approximated analysis results
Example JSON workflow:
pycg analyze mypackage/ --format json > graph.json
jq '.stats' graph.jsonThe machine-readable contract is documented in
docs/json-contract.md.
Planned query-oriented CLI result shapes are sketched in
docs/query-contracts.md.
Known limitations, confidence guidance, and how to interpret diagnostics are
documented in docs/limitations.md.
Initial query JSON Schemas live in docs/json-schema/.
The corresponding JSON Schema lives at
docs/json-schema/pycg-graph-v1.schema.json.
- Walks the given files/directories collecting
.pyfiles - Parses each file into an AST using
ruff_python_parser - Runs two-pass analysis:
- Pass 1: Collect all definitions (modules, classes, functions), track name bindings, imports, and attribute access
- Between passes: Resolve base classes, compute MRO (C3 linearization)
- Pass 2: Re-analyze with full inheritance info to resolve forward references
- Postprocessing: expand wildcard references, resolve imports, contract undefined nodes, cull inherited edges, collapse inner scopes
- Build visual graph and write to the selected output format
- Module, class, function/method definitions and nesting
- Function calls (including class instantiation →
__init__) - Attribute access through
selfand resolved objects - Imports (absolute and relative, with aliases)
- Decorator analysis (
@staticmethod,@classmethod,@property) - Inheritance with MRO-aware attribute lookup
- Assignments, augmented assignments, annotated assignments
- Return-value propagation across calls, including multi-return cases
- Tuple/list destructuring and shallow list/dict subscript flow for statically-known literals
- For-loop bindings, comprehensions, lambdas
- Context-manager, iterator,
str,repr,del obj.attr, anddel obj[key]protocol edges - Match statement patterns (Python 3.10+)
- Type alias statements (Python 3.12+)
cargo build --releasecargo test --all-targetsSemantic accuracy fixtures are tracked separately from the broader integration suite. To run the declarative fixture report against the CLI:
python3 scripts/accuracy_report.py --pycg ./target/release/pycgCorpus-scale smoke tests run the full analysis pipeline over vendored real-world packages (requests, flask, rich) and assert non-degenerate graph statistics. They skip automatically when the corpora are absent (e.g. a fresh clone), so cargo test --all-targets stays green without them.
To clone the corpora (and the pyan/PyCG reference repos) locally:
./scripts/bootstrap-corpora.shThis is idempotent and safe to re-run. After cloning, cargo test will include the corpus smoke tests.
Wall-clock comparison against code2flow (v2.5.1) on real-world codebases. Both tools produce DOT output; timings include process startup. Measured on Apple M4 Pro, 5 runs after 1 warmup, interleaved tool order per round.
| Corpus | Files | pycg | code2flow | Speedup |
|---|---|---|---|---|
| requests | 18 | 14ms | 74ms | 5.4x |
| flask | 24 | 20ms | 94ms | 4.8x |
| black | 25 | 33ms | 161ms | 4.9x |
| httpx | 23 | 23ms | 111ms | 4.8x |
| rich | 78 | 63ms | 434ms | 6.9x |
These tools are not equivalent — pycg performs deeper static analysis (MRO resolution, return-value propagation, protocol edges) while code2flow does lightweight control-flow extraction. The comparison is wall-clock only and says nothing about output quality.
To reproduce:
./benchmarks/setup.sh
python3 benchmarks/bench.py --pycg ./target/release/pycgThe benchmark harness bootstraps the same corpora used by the smoke tests. Results are intended as trend data, not as a semantic correctness claim.
- Dynamic Python patterns, framework-driven indirection, and external-library behavior are only partially modeled.
- Corpus runs are smoke tests, not full semantic validation against a gold standard.
- Benchmark numbers are wall-clock comparisons and should be treated as directional.
- See
docs/limitations.mdfor confidence guidance and how to interpret JSON/querydiagnostics.
MIT