Skip to content

Repository files navigation

Uncovering Grounding IDs:
How External Cues Shape Multimodal Binding

Hosein Hasani* · Amirmohammad Izadi* · Fatemeh Askari* · Mobin Bagherian* · Sadegh Mohammadian · Mohammad Izadi · Mahdieh Soleymani Baghshah
Sharif University of Technology

ICML 2026 · Official Code Release

ICML 2026 · OpenReview · arXiv:2509.24072 · MIT License

Release: v0.1 · Status: initial code release


Overview

Large vision-language models benefit from external visual structure such as symbols, row labels, and grid lines, but the internal mechanism behind this improvement is not immediately visible. We find that aligned visual and textual cues induce latent identifiers that bind image regions to their corresponding descriptions. We call these representations Grounding IDs.

Grounding IDs propagate through model representations and attention, strengthen cross-modal alignment, and improve partition-based reasoning. Through activation swapping, layerwise probing, similarity analysis, and natural-image interventions, we show that these identifiers causally control which object a model associates with a query symbol.

Conceptual overview of Grounding IDs

Grounding ID Mechanism

Grounding IDs travel with patched object activations

Activation patching transfers the hidden Grounding ID with the object representation. The query follows the transferred binding even though the visible row labels remain unchanged. MP4 version.

Highlights

  • Emergent multimodal identifiers: shared external cues induce latent Grounding IDs that connect visual partitions with their textual references.
  • Causal control of binding: patching object-region activations transfers symbol-object associations between contexts and changes the model's answer accordingly.
  • Layerwise structure: logit-lens and attention-head analyses trace how partition-specific information develops through the model.
  • Relational geometry: differences between symbol representations align with differences between their corresponding Grounding IDs.
  • Generalization beyond controlled scenes: multi-object and natural-image experiments show that the binding mechanism extends to more realistic visual settings.
  • Practical grounding effects: stronger cue-induced alignment improves visual reasoning and reduces hallucination across multimodal tasks.

Activation Swapping

Activation swapping procedure

Activation swapping log-probability results

The main intervention transfers object-region activations from a source context c' to a target context c. The patched context c* follows the transferred hidden binding rather than the target context's original symbol-object association.

Repository Layout

experiments/
  activation_swapping/
    run_object_region_activation_swap.py    # crossed object-region activation swap
    run_single_object_activation_transfer.py # single source-to-target object transfer
    run_row_metadata_activation_swap.py      # row/label metadata activation swap
    make_pairs_diverse.py                   # diverse source-target pair construction
  layerwise/
    intervene_logitlens_pairs.py            # layerwise intervention logit lens
    run_binding_experiment.py               # attention-head binding measurements
    postprocess_heads.py                    # head-level result aggregation
  similarity/
    save_patch_means.py                     # symbol/object ROI activation extraction
    aggregate_patch_means.py                # activation aggregation by symbol
    compute_means_symbol_permuted.py         # symbol-permutation analysis
    symbol_heatmap.py                       # relational-similarity visualization
  disjoint_symbols/                         # non-overlapping symbol interventions
  multi_object/                             # multi-object generation and evaluation
  natural_patchscope/                       # natural-image PatchScope experiments
  attention_span/                           # generation-attention span analysis

Installation

conda create -n grounding-ids python=3.10 -y
conda activate grounding-ids
pip install -r requirements.txt

The experiments are designed for GPU execution with Qwen2.5-VL. The examples use Qwen/Qwen2.5-VL-7B-Instruct; model arguments also accept a local checkpoint path.

Running Experiments

Object-Region Activation Swap

python experiments/activation_swapping/run_object_region_activation_swap.py \
  --model_dir Qwen/Qwen2.5-VL-7B-Instruct \
  --data_dir /path/to/controlled_images \
  --metadata /path/to/all_samples_metadata.json \
  --source shapes_000_with_symbols.png \
  --target shapes_001_with_symbols.png \
  --row_pairs '1->2,2->1' \
  --out_dir results/activation_swap/example \
  --pad 1 \
  --layers all \
  --greedy

Rows can be addressed by physical position (1->2), absolute metadata row (11->2), or symbol (@->&). Metadata records provide the image filename, canvas and grid dimensions, symbol rows, and object positions.

Patch-Activation Similarity

python experiments/similarity/save_patch_means.py \
  --model_dir Qwen/Qwen2.5-VL-7B-Instruct \
  --data_dir /path/to/controlled_images \
  --metadata /path/to/all_samples_metadata.json \
  --out_dir results/patch_means \
  --pad 1 \
  --layers all

python experiments/similarity/aggregate_patch_means.py \
  --in_dir results/patch_means/pmeans \
  --out_npz results/patch_means/means.npz

Additional Entry Points

python experiments/layerwise/intervene_logitlens_pairs.py --help
python experiments/layerwise/run_binding_experiment.py --help
python experiments/disjoint_symbols/run_inject_symbols_100.py --help
python experiments/multi_object/eval_multiobject_pair_swap_binding.py --help
python experiments/natural_patchscope/row_symbol_activation_patching.py --help
python experiments/attention_span/collect_generation_attention.py --help

All experiment scripts expose command-line arguments for model, dataset, metadata, and output paths.

Citation

@inproceedings{hasani2026grounding,
  title     = {Uncovering Grounding IDs: How External Cues Shape Multimodal Binding},
  author    = {Hasani, Hosein and Izadi, Amirmohammad and Askari, Fatemeh and
               Bagherian, Mobin and Mohammadian, Sadegh and Izadi, Mohammad and
               Soleymani Baghshah, Mahdieh},
  booktitle = {International Conference on Machine Learning},
  year      = {2026}
}

License

This repository is released under the MIT License.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages