Skip to content
@aisa-group

AI Safety and Alignment Group

AI Safety and Alignment Group at the ELLIS Institute Tübingen and Max Planck Institute for Intelligent Systems

ELLIS Institute Tübingen

AI Safety and Alignment Group

ELLIS Institute Tübingen · Max Planck Institute for Intelligent Systems

Group page · Updates · Datasets

AI safety · alignment · evaluation

We develop algorithmic approaches to reduce harms from increasingly capable general-purpose AI systems. Our work focuses on the alignment and evaluation of autonomous language-model agents, frontier-model risks and capabilities, and model generalisation and steerability.

Research projects

Project Paper Dataset Website
ResearchArena arXiv Trajectories research-arena.ai
PostTrainBench arXiv Trajectories posttrainbench.com
InferenceBench arXiv Trajectories inferencebench.ai
Instrumental Choices arXiv Agent traces instrumentalchoices.com
Evaluation awareness arXiv EvalAwareBench
Skill-Inject arXiv skill-inject.com
Prompt injection in agent skills arXiv
QuantSightBench arXiv quantsightbench.com

Teaching

The AI Safety course at the University of Tübingen is openly available.

Browse all repositories or follow the group on Substack.

Popular repositories Loading

  1. PostTrainBench PostTrainBench Public

    Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours

    Python 473 57

  2. tue-ai-safety-course tue-ai-safety-course Public

    AI safety course at the University of Tübingen (Summer Semester 2026)

    HTML 115 9

  3. skill-inject skill-inject Public

    Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks

    Python 88 4

  4. InferenceBench InferenceBench Public

    Benchmarking Open-Ended Inference Optimization by AI Agents

    Python 34 5

  5. promptinject-agent-skills promptinject-agent-skills Public

    Agent Skills Enable a New Class of Realistic and Trivially Simple Prompt Injections

    Python 21 3

  6. decomposing-eval-awareness decomposing-eval-awareness Public

    Decomposing and measuring evaluation awareness in existing benchmarks and our proposed EvalAwareBench.

    Python 19 4

Repositories

Showing 10 of 16 repositories

Top languages

Loading…

Most used topics

Loading…