| Physics-informed machine learning: A survey on problems, methods and applications Z Hao, S Liu, Y Zhang, C Ying, Y Feng, H Su, J Zhu arXiv preprint arXiv:2211.08064, 2022 | 419 | 2022 |
| How robust is google's bard to adversarial image attacks? Y Dong, H Chen, J Chen, Z Fang, X Yang, Y Zhang, Y Tian, H Su, J Zhu NeurIPS 2023 R0-FoMo Workshop, 2023 | 274 | 2023 |
| Pinnacle: A comprehensive benchmark of physics-informed neural networks for solving pdes H Zhongkai, J Yao, C Su, H Su, Z Wang, F Lu, Z Xia, Y Zhang, S Liu, L Lu, ... NeurIPS 2024, 2024 | 191* | 2024 |
| Rethinking model ensemble in transfer-based adversarial attacks H Chen, Y Zhang, Y Dong, X Yang, H Su, J Zhu ICLR 2024, 2024 | 169 | 2024 |
| Multitrust: A comprehensive benchmark towards trustworthy multimodal large language models Y Zhang, Y Huang, Y Sun, C Liu, Z Zhao, Z Fang, Y Wang, H Chen, ... NeurIPS 2024, 2024 | 145* | 2024 |
| Understanding the robustness of 3D object detection with bird's-eye-view representations in autonomous driving Z Zhu*, Y Zhang*, H Chen, Y Dong, S Zhao, W Ding, J Zhong, S Zheng CVPR 2023, 2023 | 119* | 2023 |
| Stair: Improving safety alignment with introspective reasoning Y Zhang, S Zhang, Y Huang, Z Xia, Z Fang, X Yang, R Duan, D Yan, ... ICML 2025 (Oral), 2025 | 95 | 2025 |
| A survey on autonomy-induced security risks in large model-based agents H Su, J Luo, C Liu, X Yang, Y Zhang, Y Dong, J Zhu IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026 | 63 | 2026 |
| Realsafe-r1: Safety-aligned deepseek-r1 without compromising reasoning capability Y Zhang, Z Zeng, D Li, Y Huang, Z Deng, Y Dong ICML 2025 R2-FM Workshop, 2025 | 58 | 2025 |
| Exploring the transferability of visual prompting for multimodal large language models Y Zhang, Y Dong, S Zhang, T Min, H Su, J Zhu CVPR 2024 (Highlight), 2024 | 39 | 2024 |
| Mitigating Overthinking in Large Reasoning Models via Manifold Steering Y Huang, H Chen, S Ruan, Y Zhang, X Wei, Y Dong NeurIPS 2025, 2025 | 35 | 2025 |
| DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios Y Huang, Y Sun, Y Zhang, R Zhang, Y Dong, X Wei NeurIPS 2025, 2025 | 32 | 2025 |
| Understanding pre-training and fine-tuning from loss landscape perspectives H Chen, Y Dong, Z Wei, Y Huang, Y Zhang, H Su, J Zhu arXiv e-prints, arXiv: 2505.17646, 2025 | 29* | 2025 |
| Breaking the ceiling: Exploring the potential of jailbreak attacks through expanding strategy space Y Huang, Y Sun, S Ruan, Y Zhang, Y Dong, X Wei ACL 2025 Findings, 2025 | 27 | 2025 |
| Oyster-I: Beyond Refusal--Constructive Safety Alignment for Responsible Language Models R Duan, J Liu, X Jia, S Zhao, R Cheng, F Wang, C Wei, Y Xie, C Liu, D Li, ... arXiv preprint arXiv:2509.01909, 2025 | 23 | 2025 |
| Scaling laws for black box adversarial attacks C Liu, H Chen, Y Zhang, J Zhu, Y Dong arXiv preprint arXiv:2411.16782, 2024 | 14 | 2024 |
| To make yourself invisible with adversarial semantic contours Y Zhang, Z Zhu, H Su, J Zhu, S Zheng, Y He, H Xue Computer Vision and Image Understanding 230, 103659, 2023 | 13* | 2023 |
| Unrestricted adversarial attacks on imagenet competition Y Chen, X Mao, Y He, H Xue, C Li, Y Dong, QA Fu, X Yang, W Xiang, ... arXiv preprint arXiv:2110.09903, 2021 | 12 | 2021 |
| Towards Safe Reasoning in Large Reasoning Models via Corrective Intervention Y Zhang, Y Ding, J Yang, T Luo, D Li, R Duan, Q Liu, H Su, Y Dong, J Zhu ICLR 2026, 2026 | 10 | 2026 |
| Unveiling trust in multimodal large language models: Evaluation, analysis, and mitigation Y Zhang, Y Huang, Y Wang, Y Sun, C Liu, Z Zhao, Z Fang, H Chen, ... arXiv preprint arXiv:2508.15370, 2025 | 7 | 2025 |