When I started working with Java, even relatively small projects felt confusing. I could understand what individual classes and methods were doing, but once the code started spreading across multiple files, I found it surprisingly difficult to see how everything fit together.
I kept wondering, What actually connects all of this? Is there some way to trace those connections instead of trying to understand the whole program one file at a time?
That made me curious about what the program looks like when you stop looking at it as a collection of source files. Because somewhere between the code I write and the code the machine executes, all of these relationships have to become concrete. Classes have to exist somewhere, methods have to be represented somehow, and the instructions inside those methods have to follow some structure.
Eventually, I started looking at the code from a different level, not as source files, but as something the compiler had already broken down into a much more explicit structure.
So where does that information actually exist?
Part of the answer is in the compiled .class files. They contain the structure the JVM needs to load a class and execute its methods, along with the bytecode that makes up those methods.
Instead of trying to understand a large project only through its source files, I could start from what the compiler produced and work my way down.
Classes connected to other classes, methods connected to the methods they call, and instructions connected by the order in which they can execute. The compiler doesn't just take a collection of unrelated files and somehow make them run.
The first useful thing to inspect is the method itself. At the source level, a method looks like one continuous block of code. After compilation, that same method is represented as a sequence of bytecode instructions, each describing a much smaller operation:
- load a value (
aload_0) - store a value (
dstore_2) - call another method (
invokevirtual) - compare a value and jump somewhere else (
ifeq)
Consider a service that processes an order. At the source level, the flow looks straightforward: the service receives an order, validates it, calculates the total, and finally stores it.
class OrderService {
private final Repository repository = new Repository();
void process(Order order) {
if (validate(order)) {
double total = calculateTotal(order);
repository.save(order, total);
}
}
boolean validate(Order order) {
return order != null && order.isValid();
}
double calculateTotal(Order order) {
return order.price() * order.quantity();
}
}
process() is a fair method to trace: one condition, two calls that only happen if that condition holds, and nothing manufactured to showcase branching.
Compiled with javap -c -p, process() looks like this:
void process(Order);
Code:
0: aload_0
1: aload_1
2: invokevirtual #16 // Method validate:(LOrder;)Z
5: ifeq 23
8: aload_0
9: aload_1
10: invokevirtual #20 // Method calculateTotal:(LOrder;)D
13: dstore_2
14: aload_0
15: getfield #10 // Field repository:LRepository;
18: aload_1
19: dload_2
20: invokevirtual #24 // Method Repository.save:(LOrder;D)V
23: return
The bytecode makes something else visible; these instructions aren't just a list.
The source-level if has disappeared. What remains is a method invocation, a conditional branch, and the sequence of operations that branch either skips or executes.
Most of these instructions don't matter individually. aload_0, aload_1, dstore_2 are just moving values around, doing exactly what the source code requires.
But one instruction is doing something different: ifeq 23 can send execution somewhere other than the next instruction. This raises a more useful question: when a method branches, where can execution actually go?
Even a method this small benefits from grouping. Once a method has more branches, doing this by eye stops being practical. A more useful representation is to group consecutive instructions that execute together without branching in the middle.
These groups are called basic blocks. A block has a single entry point and, except for its final instruction, execution proceeds sequentially through the instructions inside it.
In process(), that grouping follows directly from where ifeq 23 can land:
B0 [0–5] aload_0, aload_1, invokevirtual validate, ifeq 23
B1 [8–20] aload_0, aload_1, invokevirtual calculateTotal, dstore_2, aload_0, getfield repository, aload_1, dload_2, invokevirtual save
B2 [23] return
The important part is determining which instructions can be the beginning of a new block. These are usually called leaders.
An instruction is a leader if it is:
- the first instruction of the method
- the target of a jump
- the instruction immediately following a conditional or unconditional jump
- the start of an exception handler
In process(), that gives exactly three leaders:
- offset
0, the start of the method - offset
8, immediately afterifeq - offset
23, the jump target ofifeq
Every instruction from one leader up to the instruction before the next leader belongs to the same block.
Once the blocks have been identified, the next problem is determining how execution moves between them:
-
B0ends with a conditional jump, so it has two successors:B1, the block right afterifeq, ifvalidate()returns true, andB2, the jump target, if it returns false -
B1ends with an instruction that does not branch, so control falls through toB2 -
B2ends withreturn, so it has no successor at all
Nothing in the bytecode explicitly says "this is an if statement." Instead, the compiler translated that decision into a method invocation and a conditional branch.
The higher-level structure can be reconstructed from those lower-level operations, and this is the whole method's control flow, recovered without reading a line of the source above.
Real Java bytecode doesn't only produce this kind of simple branch. Loops, switch statements, and exception handlers can introduce more complex control flow paths.
process() also calls three other methods: validate, calculateTotal, and Repository.save. Each one is an invocation instruction, and at the bytecode level that relationship is explicit. The instruction refers, through the constant pool, to the class and method being invoked.
Instead of searching through source files to find every call site, an analyzer can inspect these invocation instructions and record which method is being referenced.
Those references can be represented as edges between methods, giving us a call graph alongside the control flow graph. The CFG shows where execution can move within a method, while the call graph shows where one method may lead to another.
None of this requires reading raw bytes by hand. Tools such as javap expose bytecode for inspection, while libraries such as ASM let programs inspect class files and their instructions directly.
The same idea extends beyond method calls. In process(), for example, getfield references the repository field, giving an analyzer another way to recover relationships between parts of the program. But the important part isn't any individual instruction. It's what happens when these relationships are collected across an entire project.
A method can have its own control flow graph, while calls from that method connect it to other methods. Those methods may belong to different classes, access shared fields, create objects, or inherit behavior from another class.
These graphs are built from the compiled structure, not from running the program, so some things stay out of reach:
- Dynamic dispatch: the runtime target of a method call can depend on the object's runtime type
- Reflection and other runtime behavior: these can introduce relationships that are not obvious from the bytecode alone
So the graph shows what the code can do, not what it will do at runtime.
This is the part that actually answered the question I started with. With a call graph, I wouldn't have to open process(), then go find validate(), then go find wherever Repository is defined, just to see how one method fits into the rest of the service. The call graph already has that shape, with process() reaching out to validate(), calculateTotal(), and save(), without me tracing any of it by hand.
Here that's three edges you could read off the source in seconds, but across a real project, it's the view I couldn't get from the source alone.
The source is still where I'd go to actually change the code. But it doesn't always give me a complete view of how those pieces connect. That's what the graph was for.
The tool I built around this idea constructs method-level control flow graphs from Java source and exports them as Graphviz DOT. This article goes one level lower, to bytecode.
rounakkm
/
bytecode-cfg
A Java engine that parses compiled (.class) and (.jar) files to reconstruct program structure. It will build class dependency and method level control flow graphs, will detect modules, will compute complexity metrics, and will generate visual reports for analyzing java programs.
Bytecode CFG
About
Bytecode CFG is a static analysis and control flow graph generation tool for Java source code. Powered by JavaParser, the tool inspects Java source files (.java), constructs Abstract Syntax Trees (ASTs), executes configurable static analysis rules, generates structured reports (JSON and HTML), and reconstructs method-level Control Flow Graphs (CFGs) exported in Graphviz DOT format.
The primary goal is to make it easier to analyze unfamiliar Java codebases, inspect execution flow and cyclomatic complexity, enforce coding standards, and identify potential bugs such as null dereferences.
Processing Pipeline
The analysis pipeline operates in sequential stages:
Stage 1: Input Discovery
- The target path specified via
--inputis resolved. - If a single
.javafile is provided, it is analyzed directly. - If a directory is provided, it is scanned recursively to collect all
.javasource files.
Stage 2: AST Parsing
- Source files are parsed into Abstract Syntax Trees using…
Top comments (0)