Jianwei Li

Hi! I am currently a PhD student at North Carolina State University, under the guidance of Prof. JUNG-EUN KIM. Previously, I worked as a machine learning scientist in Moffett AI, and was guided by the chief scientist and co-founder of Moffett.AI: Ian En-Hsu Yen. Before that, I got my master's degree from the CS department of San Jose State University and was advised by Prof. Mark Stamp. I earned my bachelor's degree from Shandong University.

Currently, I am a Research Scientist Intern at the TikTok Foundations & Intelligence Service (TNS) team for Summer 2026, after starting the summer with the NLP Trust & Safety (TNS) team. Previously, in Summer 2025, I interned with TikTok Responsible Recommendation Systems (RRS) team, where I worked on safety-aware recommendation for multimodal LLM.

/ Google Scholar / GitHub / LinkedIn / Twitter

Shadow LLM Guardians

Address: Raleigh, North Carolina
Email: ljw040426 AT gmail DOT com | jli265 AT ncsu DOT edu
Research Interest

My research centers on Secret Alignment — covert alignment behaviors in large language models and the agents built on them — as the technical foundation for their secure and personalized deployment. I build, understand, evaluate, and defend against these behaviors in the models and agents that increasingly act on our behalf.

What is Secret Alignment? Introduced by me and Prof. Jung-Eun Kim in early 2025, Secret Alignment describes a class of hidden, condition-dependent behaviors in aligned models — the model's response is systematically different under a specific trigger, context, or credential, while remaining indistinguishable from a normally aligned model to any observer without that trigger.

Because the same mechanism can serve legitimate control or hostile subversion, Secret Alignment is fundamentally dual-use:

Adversarial forms (defend against)
  • Backdoor attacks
  • Multi-persona injection (attacker trigger)
  • Sandbagging on capability / safety evaluations
  • Deceptive alignment
  • Steganographic collusion between agents
Legitimate forms (build well)
  • Authentication of model origin
  • Copyright / provenance watermarking
  • Personalization (verified user)
  • Access control & permission gating
  • Kill switch / emergency shutdown

Whether Secret Alignment serves defense or attack depends on who holds the trigger — which is why my work treats it as a first-class object of study and evaluation.

News
Scroll down for more news ↓
  • [July, 2026] 📝 Invited to serve as Reviewer for AAAI 2027
  • [June, 2026] 🔄 Transitioned to the Foundations & Intelligence Service (TNS) team at TikTok
  • [June, 2026] 📝 Invited to serve as Reviewer for TMLR
  • [May, 2026] 🥇 Selected as Gold Reviewer of ICML 2026
  • [May, 2026] 💼 Start a research internship program with the NLP Trust & Safety (TNS) team at TikTok
  • [April, 2026] 📝 Invited to serve as Reviewer for NeurIPS 2026
  • [April, 2026] 🎉 One paper is accepted by ICML 2026
  • [March, 2026] 📝 serve as Reviewer for ICML 2026
  • [Jan, 2026] 🎉 Two papers are accepted by ICLR 2026
  • [Oct, 2025] 🎓 Pass PhD written preliminary exam at NCSU
  • [May–Aug, 2025] 💼 Start a research internship program with the Responsible Recommendation Systems (RRS) team at TikTok
  • [May, 2025] 🎉 One paper is accepted by ICML 2025
  • [Feb, 2025] 🎉 One paper is accepted by CPAL 2025
  • [Jul, 2024] 🤝 Start research under the guidance of Prof. Jung-Eun Kim
  • ----------------------------------------
  • [Dec, 2023] 🎤 One paper selected as Oral of FL@FM-NeurIPS 2023
  • [Nov, 2023] 🛡️ Initiated the Shadow-LLM-Guardians Group: https://github.com/Shadow-LLM
  • [Oct, 2023] 🎉 Two papers are accepted by EMNLP 2023
  • [Aug, 2023] 🎓 Start PhD student life at NC State University; Serve as official affiliation of New York University
  • ----------------------------------------
  • [Apr, 2023] 📢 Publicity Chair of the International Workshop onResource-Efficient Learning for Knowledge Discovery @KDD 2023.
  • [Oct, 2022] 📢 Publicity Chair of the first workshop on DL-Hardware Co-Design for AI Acceleration @AAAI 2023.
  • [Sep, 2022] Participated in AI Hardware Summit 2022.
  • [Sep, 2022] 🏆 Moffett S30 Accelerator wins MLPerf V2.1
Preprints
Jianwei Li, Jung-Eun KimSecurity Before Safety: A Backdoor-Centric View of LLM Output Risks in the Private AI Era PDF
Xingli Fang, Jianwei Li, Varun Mulchandani, Jung-Eun KimTrustworthy AI: Safety, Bias, and Privacy — A Survey PDF
Selected Publications
pub
Jianwei Li, Jung-Eun Kim
Position: Retire the “Positive Backdoor” Label—Secret Alignment Requires Strict and Systematic Evaluation
Position PaperSecret Alignment Evaluation
ICML 2026
pub
Jianwei Li, Jung-Eun Kim
Purifying Generative LLMs from Backdoors without Prior Knowledge or Clean Reference
Main PaperDefense Backdoor Attack PDF Code Project page
ICLR 2026
pub
Jianwei Li, Jung-Eun Kim
Safety Alignment Can Be Not Superficial With Explicit Safety Signals
Main PaperDefense Jailbreak Attack PDF Code Project page
ICML 2025
pub
Jianwei Li, Jung-Eun Kim
Superficial Safety Alignment Hypothesis
Main PaperDefense Fine-tuning Attack PDF Code Project page
ICLR 2026 (arXiv 2024)
pub
Jianwei Li, Sheng Liu, Qi Lei
Beyond Gradient and Priors in Privacy Attacks: Leveraging Pooler Layer Inputs of Language Models in Federated Learning
Workshop OralDefense Privacy Leakage
FL@FM NeurIPS 2023
Teaching & Research Assistant
  • 2023.08-2024.05, NCSU, Teachin Assistant
  • 2024.06~present, NCSU, Reserach Assistant
© 2022 jianwei.li All rights reserved
(Last update: May 14, 2026.)