This is the GitHub account of the NLP Lab led by Prof. Ziyu Yao at George Mason University, Department of Computer Science.
Group Webpage: https://ziyuyao.org/group/
- ICML 2025 Tutorial on Mechanistic Interpretability for Language Models. Website: https://ziyu-yao-nlp-lab.github.io/ICML25-MI-Tutorial.github.io/
- MathVC NSF Project on building an LLM Agent-powered platform for Mathematics Education. Website: https://ziyu-yao-nlp-lab.github.io/MathVC-NSF.github.io/
- Why Do LLM-based Web Agents Fail? A Hierarchical Planning Perspective, ACL 2026. Paper Code
- Autospatial: Visual-Language Reasoning for Social Robot Navigation through Efficient Spatial Reasoning Learning, IROS 2025. Paper
- Revisiting Prompt Optimization with Large Reasoning Models-A Case Study on Event Extraction, Preprint 2025 Paper
- Evaluating Vision-Language Models as Evaluators in Path Planning, IEEE/CVF CVPR, 2025 Paper Code Dataset
- A Survey on Large Language Models for Automated Planning, Preprint 2025 Paper
- Instruction-Tuning LLMs for Event Extraction with Annotation Guidelines, Findings of ACL 2025. PaperCode
- Efficient but Vulnerable: Benchmarking and Defending LLM Batch Prompting Attack, ACL Findings, 2025. Paper Code
- DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search, ICLR, 2025. Paper Code
- Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning, ICLR Workshop on LLM Agents, 2024 Paper Code
- Look Further Ahead: Testing the Limits of GPT-4 in Path Planning, IEEE 20th International Conference on Automation Science and Engineering, 2024 Paper Code
- Instances Need More Care: Rewriting Prompts for Instances with LLMs in the Loop Yields Better Zero-Shot Performance, ACL Findings 2024 Paper
- Large Language Model Cascades with Mixture of Thoughts Representations for Cost-efficient Reasoning, ICLR, 2024. Paper Code
- Improving Generalization in Language Model-based Text-to-SQL Semantic Parsing: Two Simple Semantic Boundary-based Techniques, ACL 2023. Paper Code
- MailEx: Email event and argument extraction, EMNLP 2023. PaperDataset/Code
- Gentopia: A Collaborative Platform for Tool-Augmented LLMs, EMNLP Demo, 2023. Paper Code
- Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?, Preprint 2026. Paper Code
- Data-driven Circuit Discovery for Interpretability of Language Models, Preprint 2026. Paper Code
- Inside-Out: Measuring Generalization in Vision Transformers Through Inner Workings, CVPR 2026 (Highlight, Top 3%). Paper Code
- Failure by Interference: Language Models Make Balanced Parentheses Errors When Faulty Mechanisms Overshadow Sound Ones, NeurIPS 2025. Paper Code
- Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models, EMNLP 2025. Paper
- All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokens, EMNLP 2025. Paper Code
- A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models, Findings of EMNLP 2025. Paper
- A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models, Preprint 2025. Paper
- Mechanistic Understanding of Language Models in Syntactic Code Completion, AAAI KnowFM Workshop 2025. Paper
- Understanding the Effect of Algorithm Transparency of Model Explanations in Text-to-SQL Semantic Parsing, Preprint 2025. Paper
- An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs, ACL 2024. Paper Code
- Navigating the Shortcut Maze: A Comprehensive Analysis of Shortcut Learning in Text Classification by Language Models, Findings of EMNLP 2024. Paper Code
- Explaining Large Language Model-Based Neural Semantic Parsers, AAAI 2023. Student Abstract
- Designing AI Peers for Collaborative Mathematical Problem Solving with Middle School Students: A Participatory Design Study, CHI 2026. Paper
- IntelliExplain: Enhancing Conversational Code Generation for Non-Professional Programmers, Preprint, 2024. Website Paper
- Learning to Simulate Natural Language Feedback for Interactive Semantic Parsing, ACL 2023. Paper Code
- Reassessing Code Authorship Attribution in the Era of Language Models, ACM TOSEM, 2026. Paper
- PeerMathDial: A Middle School Dialogue Dataset for Student Collaborative Math Problem Solving, ACL-BEA Workshop, 2026. Paper Code
- Guiding AI to Fix Its Own Flaws: An Empirical Study on LLM-Driven Secure Code Generation, Preprint 2025. Paper
- Beneath the Surface: How Large Language Models Reflect Hidden Bias, Findings of EMNLP 2025. Paper Code
- Evaluating the Effectiveness of Persona Simulation in Opinion Prediction with GPT-4.1, ICDM 2025 Workshops. Paper
- MathVC: An LLM-Simulated Multi-Character Virtual Classroom for Mathematics Education, AAAI AI4Edu Workshop, 2025. Website Paper Code
- Can LLMs Simulate Personas with Reversed Performance? A Benchmark for Counterfactual Instruction Following, Preprint, 2025. Paper Code
- A Paradigm Shift from "Human Writing" to "Machine Generation" in Personality Test Development: An Application of State-of-the-Art Natural Language Processing, Journal of Business and Psychology, 2023. Paper