Kiana Jafari
Researcher at Stanford University
Researching human-agent teaming and human-centered AI, designing intelligent systems that enhance collaboration and empower human decision-making.
More about meRecent work
All papers- FAccT
The Doctor Will (Still) See You Now: On the Structural Limits of Agentic AI in Healthcare
A qualitative study based on interviews with 20 stakeholders examining how agentic AI is defined, evaluated, and constrained in healthcare, identifying three mutually reinforcing tensions: conceptual fragmentation, an autonomy contradiction, and an evaluation blind spot.
Agentic AIHealthcare AIResponsible AI - FAccT
Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
A mixed-methods study examining inter-rater reliability among three psychiatrists evaluating 360 LLM-generated mental health responses, revealing systematic expert disagreement driven by incompatible clinical frameworks rather than measurement error.
AI SafetyRLHFMental Health AI - arXiv
How Language Models Fail: Token-Level Signatures of Committed and Persistent Reasoning Failures
Characterizing language model reasoning failures through token-level uncertainty signals, and showing they arise through two empirically distinguishable processes: committed failure, where a model locks onto a wrong path early, and persistent uncertainty, where uncertainty accumulates throughout the trace.
LLMReasoningUncertainty QuantificationFailure Detection
News
All news- The Doctor Will (Still) See You Now: On the Structural Limits of Agentic AI in Healthcare
Accepted to FAccT 2026.
- Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
Accepted to FAccT 2026.
- The Measurement Imbalance in Agentic AI Evaluation Undermines Industry Productivity Claims
Accepted to NeurIPS 2025 by our reviewers and the area chair, but REJECTED by NeurIPS.