Lists (2)
Sort Name ascending (A-Z)
Stars
3
stars
written in HTML
Clear filter
This project investigates whether language models remain epistemically consistent when subjected to varying forms of social pressure. While models are generally trained to reject obvious falsehoods…
Doctor-facing benchmark: how often do frontier LLMs cave to a clinician's wrong medical claim? 9 models, 202 scenarios, Design A vs B knowledge control. BlueDot AI Safety sprint.
Alignment research: how honest human-AI dialogue produces measurably better AI outputs without modifying weights or training