PaperGitHubDocs

Can two AI agents talk each other through defusing a bomb?

GPTNT is a benchmark for real-time, asymmetric collaboration between multimodal agents, built on the cooperative game Keep Talking and Nobody Explodes. Against a live, unpausing clock, one agent sees the bomb but not the instructions; the other holds the instructions but never sees the bomb; neither can defuse it alone.

DefuserSees the bomb · not the manual
/// No signal ///Bomb feedNot visible to the Expert
Natural language · the only communication channel
Expert ▸

Describe the bomb to me: how many modules do you see, and what does each one look like?

click(x=0.70, y=0.38)

do_nothing()

do_nothing()

out()

flip()

do_nothing()

do_nothing()

right()

do_nothing()

do_nothing()

right()

flip()

do_nothing()

do_nothing()

right()

do_nothing()

right()

do_nothing()

do_nothing()

flip()

do_nothing()

do_nothing()

right()

do_nothing()

right()

do_nothing()

do_nothing()

down()

do_nothing()

do_nothing()

do_nothing()

up()

do_nothing()

up()

do_nothing()

do_nothing()

down()

do_nothing()

flip()

do_nothing()

do_nothing()

right()

do_nothing()

right()

do_nothing()

down()

do_nothing()

do_nothing()

up()

do_nothing()

◂ Defuser

Front face: timer at 02:51, no strikes. One module: a large white button labeled "ABORT" with a small LED in the top right corner of the module frame (unlit/unsolved). Four copper grill panels. Back face: three modules. 1. Top right: a module with a large "2" displayed in the upper half and "1324" in the lower half. Status LED unlit (unsolved). 2. Bottom left: a module with four symbol buttons — top-left: a reversed-R shape, top-right: lambda (λ), bottom-left: a six-pointed star, bottom-right: omega (ω). Status LED is bright green (already solved). 3. Bottom right: a module showing wavy lines — white/light wavy lines on top half, red wavy lines on bottom half. Status LED unlit (unsolved). Top edge: serial number label (upside down when viewed from top) reads "6PZ398". Also a port panel with two RCA ports (one red, one green) and one DB9 serial port. No batteries or indicator lights found.

do_nothing()

Expert ▸

Let's start with the white ABORT button on the front. Hold the button down and tell me what color the strip on the right side of the module lights up.

do_nothing()

do_nothing()

click(x=0.70, y=0.38)

do_nothing()

click(x=0.49, y=0.52)

do_nothing()

do_nothing()

hold(x=0.49, y=0.52)

do_nothing()

do_nothing()

◂ Defuser

I'm holding the button. The strip on the right side of the module appears unlit/dark. Also, I got a strike earlier from accidentally clicking - there is now 1 red X showing. Timer is at 02:06.

do_nothing()

◂ Defuser

The strip is now lit yellow. Timer at 01:57.

Expert ▸

Keep holding. Release the button when the timer shows a 1 anywhere in it (e.g., 1:XX, X:1X, or X:X1).

Expert ▸

Release when the timer shows a 5 anywhere (e.g., 1:55, 1:50, 0:59, 0:58... any digit is 5).

release()

do_nothing()

do_nothing()

do_nothing()

do_nothing()

do_nothing()

do_nothing()

do_nothing()

do_nothing()

do_nothing()

do_nothing()

do_nothing()

do_nothing()

release()

do_nothing()

do_nothing()

do_nothing()

click(x=0.49, y=0.52)

do_nothing()

do_nothing()

Async self-play · the clock never stops
ExpertReads the whole manual · 23 pages
/// No access ///ManualNot visible to the Defuser
0:00 / 0:00

Amit Parekh*, Sabrina McCallum*, Kareem Al-Hasan*, Malvina Nikandrou, Alessandro Suglia, Ioannis Konstas

Heriot-Watt University · University of Edinburgh

Five frontier models. Zero bombs defused.

No model we test—open or closed—successfully defuses a single bomb. Nine out of ten different human pairs could solve at least one.

02:47

The clock never stops.

Models generate tokens in real time.

Defuse or explode.

No rubric or LLM judge needed because the game provides the ground truth.

The bomb the Defuser sees
Defuser
sees the bomb
The manual the Expert holds
Expert
holds the manual
The only channel between them
is natural language.

The Defuser can point however models want to.

Our KTANE mod uses normalised (x, y) coordinates and outputs segmentation masks, so the Defuser can use coordinates or set-or-marks.

The bomb as the Defuser sees it, addressed by (x, y) coordinates
Coordinates
click {"x":0.51, "y":0.48}
The same Simon module labeled with set-of-marks A, B, C, DSet-of-marks
press "D"
The bomb defusal manual, spread across its pages

The Expert reasons over the full manual.

All twenty-three pages of rules, wiring diagrams, and symbol tables—in context from the first move.

See the manual
No manual

Collaborating,
or just recalling?

Diagnose real collaboration by taking the manual away to see how strong the parametric knowledge is.

It won’t be saturated and retired.

GPTNT runs on Keep Talking and Nobody Explodes—the same bombs, manual, and ticking clock people play against, and nothing is simplified for the models. It inherits the game’s living modding community, so as models improve we add harder modules—and eventually make them do The CenturionThe Centurion — a single bomb covered in roughly 100 modulesOne bomb with ~100 multimodal and multilingual modules.The pinnacle for any player, human or AI..

Wires module
Simon Says module
Keypad module
The Button module
Complicated Wires module
Maze module
Memory module
Morse Code module
Password module
Who's on First module
Wire Sequence module
MODS

Multimodal & Multilingual

The same bomb, rendered in EnglishThe same bomb, rendered in 中文The same bomb, rendered in العربيةThe same bomb, rendered in 한국어The same bomb, rendered in РусскийThe same bomb, rendered in ไทยThe same bomb, rendered in 日本語The same bomb, rendered in עברית

Leaderboard

Zero-shot self-playSame model for expert and defuser, does not share context., pass @1

Click any model for its full results →each dot = 10% solved · more dots is better
#ModelInteract?How did models interact with the gameReal-time(async)?Async: expert and defuser act on independent live clocks — no shared turns.Turn-taking(sync)?Sync: expert and defuser alternate in lockstep turns.
Missions?Full multi-module missions defused end-to-end.Modules?Share of individual bomb modules solved across missions.Any module?Missions where at least one module was solved before failure.MissionsModulesAny module
1AnthropicClaude Sonnet 4.6set-of-marks 0% 15% 50% 10% 29% 50%
2OpenAIGPT-5.2set-of-marks 0% 12% 30% 10% 21% 50%
3GoogleGemini 3 Flash (Preview)set-of-marks 0% 9% 30% 0% 15% 40%
4QwenQwen3.5 (27B)set-of-marks 0% 6% 20% 0% 12% 30%
5InternVLInternVL3.5 (38B)set-of-marks 0% 3% 10% 0% 12% 40%
Human players 25% 60% 93%

Ran GPTNT on your own model?

Submit your run and we'll add it to the board. New models and protocols welcome.

Submit results →

Contamination checks

Ablations surfacing previous exposure to the game during training — no partner, no manual. The more a model scores here, the more it leans on parametric knowledge rather than genuine, in-context collaboration.

each dot = 10% · more dots = more recall, not collaboration
ModelSingle Agent?One model plays alone with no partner and no manual — pure parametric recall of the game.Expert VQA?The expert must read the bomb from the image alone, with no manual to consult.
AnthropicClaude Sonnet 4.6 26% 22%
OpenAIGPT-5.2 24% 20%
GoogleGemini 3 Flash (Preview) 12% 36%
QwenQwen3.5 (27B) 12% 10%
InternVLInternVL3.5 (38B) 11% 22%
Random baseline 3% 21%

Replay

Of actual games played by models we tested

Defuser view of the bombDefuser view
1:15
5:00
Step 6 / 27
Strikes 1 / 3
DefuserExpert

FAQ

What is GPTNT?
GPTNT is a benchmark for real-time collaboration between multimodal models, built on the cooperative game Keep Talking and Nobody Explodes. Two agents must coordinate to defuse procedurally generated bomb puzzles against a live countdown. One agent is the Defuser: it has access to the bomb but not the instructions for defusing it. The other agent is the Expert: it holds the instructions manual but cannot see or manipulate the bomb.
Why is this hard for AI agents?
GPTNT combines conditions that are usually studied in isolation: asymmetric information, real-time asynchronous action, visual grounding, procedural reasoning, and sustained multi-turn communication — all at once. Current models break down on tracking state across turns, acting within the time budget, handling ambiguous descriptions, and recovering from mistakes.
What do you release, and do I need the game?
We release everything needed to run the benchmark: the game mod and microservice framework, the fixed suite of mission configurations, the processed manual, the role-specific system prompts, and the single-agent diagnostic evaluations. You do need your own copy of Keep Talking and Nobody Explodes ($14.99, Windows/macOS/Linux). We do not distribute the game or any of its source, as that would violate its license.
What are the "async" and "sync" modes?
The bomb clock is the whole challenge, and the two modes differ in one thing: whether it keeps running while a model thinks. In async (real-time) mode the clock never stops — every token a model generates burns game time, and both agents act in parallel, so reasoning too long means running out of time mid-move. This is the real game, and the real world: time doesn't wait for your turn. sync (turn-based) mode freezes the clock during generation and advances a fixed step per turn, separating whether a model reasons accurately from whether it reasons efficiently — and lowering the barrier to entry by matching the turn-based loop today's models assume.
asyncboth at once, out of sync
Defuser
Expert
0:18Live
syncone turn at a time
Defuser
Expert
0:18Live
What's the difference between a strikeout and a timeout?
failed mission ends in one of two terminal states: a strikeout occurs when the Defuser makes three mistakes, leading to three strikes, or a timeout when the timer expires before all modules are solved. These are diagnostically different failures: a model can be careful-but-slow or fast-but-reckless.
How do you separate real collaboration from memorised solutions?
Because the game predates current models, parts of its manual and puzzle logic may appear in training data. As a contamination check, we run dedicated single-agent evaluations without providing the manual to expose what a model already knows.
Will the benchmark get solved and retired?
GPTNT is designed to stay ahead of improving models. New missions can be procedurally generated and reconfigured by varying the time limit and the count, type, and placement of modules. New puzzle types from the game's active modding community can be added without rebuilding the framework, keeping the task space ahead of any fixed training distribution.

Citation

GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes

Parekh, McCallum, Al-Hasan, Nikandrou, Suglia, Konstas

arXiv:2606.28514 · Preprint 2026

arXiv
BibTeX
@misc{parekh2026gptnt,
  title = {{{GPTNT}}: {{Benchmarking Real-Time Collaboration Between Multimodal Agents}} on {{Keep Talking And Nobody Explodes}}},
  shorttitle = {{{GPTNT}}},
  author = {Parekh, Amit and McCallum, Sabrina and {Al-Hasan}, Kareem and Nikandrou, Malvina and Suglia, Alessandro and Konstas, Ioannis},
  year = 2026,
  month = jun,
  number = {arXiv:2606.28514},
  eprint = {2606.28514},
  primaryclass = {cs.AI},
  publisher = {arXiv},
  doi = {10.48550/arXiv.2606.28514},
  archiveprefix = {arXiv}
}