Official implementation of the TMLR paper Understanding Transformer-Based Vision Models via Modular Feature Inversion.
This repository provides training and evaluation code for modular feature inversion in modern Transformer vision architectures. The method learns lightweight inverse modules that map internal representations back toward earlier representations or image space, enabling direct visual inspection of how information evolves across DETR, ViT, DeiT, and Swin Transformer models.
Feature inversion is a useful lens for understanding what a neural representation preserves, discards, or transforms. Instead of training a single monolithic inverse model, this project studies a modular inversion pipeline: each inverse module targets one transition in the forward model. This makes inversion more scalable, and exposes stage-wise behavior.
The codebase includes:
- modular inverse networks for DETR backbone, encoder, decoder, and prediction head;
- modular inverse networks for ViT patch/backbone and encoder representations;
- inverse pipelines for Swin Transformer stages;
- parallel inverse-training baselines;
- fine-tuning scripts for DETR and ViT with mixed detection/classification and reconstruction objectives;
You can download all checkpoints here; separate checkpoint links are also provided in the table below.
| Family | Module | Representation Inverted | Dataset | Checkpoint |
|---|---|---|---|---|
| DETR | Inverse backbone | backbone embedding -> image | COCO 2017 | DETR inverse backbone |
| DETR | Inverse encoder | encoder embedding -> backbone embedding | COCO 2017 | DETR inverse encoder |
| DETR | Inverse decoder | decoder embedding -> encoder embedding | COCO 2017 | DETR inverse decoder |
| DETR | Inverse prediction head | DETR predictions -> decoder embedding | COCO 2017 | DETR inverse prediction checkpoint |
| ViT | Inverse backbone | patch/backbone embedding -> image | ImageNet-1k | ViT inverse backbone |
| ViT | Inverse encoder | encoder embedding -> patch/backbone embedding | ImageNet-1k | ViT inverse encoder |
| Swin | All inverse stages | stage features -> earlier features/image | ImageNet-1k | Swin inverse stages |
inverse-tvm/
|-- config.py # Local paths and runtime configuration
|-- figures/ # Paper/repository figures
|-- modules/
| |-- detr/ # DETR implementation used by the experiments
| |-- inv_detr/ # DETR inverse modules with individual, as well as modular, and end-to-end parallel training.
| | |-- inv_bb/
| | |-- inv_enc/
| | |-- inv_dec/
| | |-- inv_pred/
| | |-- parallel_training/
| |-- inv_vit/ # ViT inverse modules with individual, as well as modular, and end-to-end parallel training.
| | |-- inv_bb/
| | |-- inv_enc/
| | |-- parallel_training/
| |-- inv_swin/ # Swin inverse modules with individual, as well as modular, and end-to-end parallel training.
| | |-- models.py
| | |-- train.py
| | |-- utils.py
| | |-- parallel_training/
| |-- finetuned_detr/ # DETR fine-tuning with reconstruction objectives
| | |-- train.py
| | |-- utils.py
| |-- finetuned_vit/ # ViT fine-tuning with reconstruction objectives
| | |-- train.py
| | |-- utils.py
|-- tools/ # Dataset, model, training, and logging utilities
|-- requirements.txt
|-- README.md
Before running experiments, update config.py with local dataset and output paths:
Expected COCO 2017 layout:
coco/
|-- annotations/
| |-- instances_train2017.json
| |-- instances_val2017.json
|-- train2017/
|-- val2017/
ImageNet experiments support extracted class-folder splits:
imagenet/
|-- train/
| |-- n01440764/
| |-- ...
|-- val/
|-- n01440764/
|-- ...
If extracted ImageNet split folders are not present, the loader falls back to torchvision.datasets.ImageNet.
Train a DETR inverse backbone module:
python modules/inv_detr/inv_bb/train.py --epochs 100 --batch_size 32Train a ViT inverse backbone module:
python modules/inv_vit/inv_bb/train.py --epochs 100 --batch_size 128Each run creates a local directory under RUNS_DIR/<experiment-id>/. Use the printed experiment id as --run_id to resume a run or as --inv_bb_id, --inv_enc_id, and related arguments for downstream experiments.
python modules/inv_detr/inv_bb/train.py --epochs 100 --batch_size 32
python modules/inv_detr/inv_enc/train.py --epochs 100 --batch_size 128
python modules/inv_detr/inv_dec/train.py --epochs 100 --batch_size 128
python modules/inv_detr/inv_pred/train.py --epochs 100 --batch_size 128Parallel DETR inverse-training baselines:
python modules/inv_detr/parallel_training/modular_in_parallel.py
python modules/inv_detr/parallel_training/e2e_in_parallel.pypython modules/inv_vit/inv_bb/train.py --epochs 100 --batch_size 128
python modules/inv_vit/inv_enc/train.py --epochs 100 --batch_size 128Parallel ViT inverse-training baselines:
python modules/inv_vit/parallel_training/modular_in_parallel.py
python modules/inv_vit/parallel_training/e2e_in_parallel.pypython modules/inv_swin/train.py --epochs 1000 --batch_size 128
python modules/inv_swin/parallel_training/modular_in_parallel.py
python modules/inv_swin/parallel_training/e2e_in_parallel.pyFine-tune downstream models with reconstruction objectives from trained inverse modules:
python modules/finetuned_detr/train.py --inv_bb_id <experiment-id> --inv_enc_id <experiment-id> --inv_dec_id <experiment-id> --epochs 100 --batch_size 16
python modules/finetuned_vit/train.py --inv_bb_id <experiment-id> --inv_enc_id <experiment-id> --epochs 100 --batch_size 512All experiments log locally through tools/logging_utils.py.
Each run is stored as:
runs/<experiment-id>/
|-- config.json
|-- config.yaml
|-- metrics.jsonl
|-- checkpoints/
`-- model_states/
Checkpoints include model states, optimizer states, best validation loss, and training/evaluation step counters.
If you use this repository or build on modular feature inversion, please cite the paper:
@article{modular_feature_inversion_2026,
title = {Understanding Transformer-Based Vision Models via Modular Feature Inversion},
journal = {Transactions on Machine Learning Research},
year = {2026},
author = {Rathjens, Jan and Reyhanian, Shirin and Kappel, David and Wiskott, Laurenz},
url = {https://openreview.net/forum?id=O5sMv2o3EV}
}This repository is released under the license provided in LICENSE.