Chaowei Xiao, Johns Hopkins University 30,463Citations 69h-index
Email: chaoweixiao [at] jhu [dot] edu

Chaowei Xiao

Assistant Professor · Johns Hopkins University|Researcher · NVIDIA Research

Bio: I am Chaowei Xiao, an assistant professor at Johns Hopkins University (JHU) and a researcher at NVIDIA Research, currently visiting Stanford University. I obtained my Ph.D. from the University of Michigan, Ann Arbor, and my bachelor's degree from Tsinghua University.

My research aims to build safe AGI, with a focus on AI and agent security. I have also recently become interested in Recursive Self Improvement, Continual Learning, Scientific Discovery and principled methods for diverse application domains, including AI4Science, embodied agents, and computer-use agents.

Longer bio

Chaowei Xiao is an Assistant Professor at Johns Hopkins University and a researcher at NVIDIA Research. His research aims to build safe and secure AI, with a focus on LLM security and AI agent security.

He has received several prestigious awards, including the Schmidt Sciences AI2050 Early Career Fellowship, MIT Technology Review Innovators Under 35 (TR35), the Impact Award from Argonne National Laboratory, and various industry faculty awards. His work has been recognized with several paper awards at top-tier conferences, including the USENIX Security Distinguished Paper Award (2024), the EWSN Best Paper Award (2021), the MobiCom Best Paper Award (2014), ACM Gordon Bell Prize Finalist (2024), and the ACM Gordon Bell Special Prize for HPC-Based COVID-19 Research (2023). His research has been cited more than 30,000 times and has been featured in media outlets such as Nature, Wired, Fortune, and The New York Times. One of his research outputs is on display at the Science Museum in London.

Our group plans to recruit multiple PhD students. I am interested in students in general AI, cybersecurity, or robotics (VLA / world models).

I'm also looking for multiple postdocs with experience in cybersecurity, software engineering, security, RL, or robotics.

🏆 Awards (selected)

🐾 Research all publications →

📰 News

  • [Oct 2026] Four papers accepted to NeurIPS 2026 on agent security and safety: Agent MechSuits (safety steering for multi-turn CLI agents), LPG (latent policy guardrails), MaskForge (jailbreaking diffusion LLMs), and position-weighted self-distillation for reasoning. Congratulations to all authors!
  • [Oct 2026] More 2026 papers on LLM and agent security: ReasoningBomb (CCS 2026), Your Harness is Not Secure (EMNLP 2026), Concept Concentration (ICML 2026), and four papers at ICLR 2026, including ARMOR and A2ASecBench.
  • [Mar 2026] Serving as Senior Area Chair for ACL, EMNLP, and NeurIPS.
  • [Jan 2025] We have 9 papers at ICLR and 3 papers at ACL on model safety and security. Congratulations to all authors!
  • [Dec 2024] We received the Fall Research Competition Award at UW–Madison. Thank you UW–Madison and OVCR.
  • [Nov 2024] Our group recently received funding and donations. Thank you Amazon and Apple!
  • [Nov 2024] Our lab will have a winter break this December. Lab members will enjoy some well-deserved vacation time with their families and loved ones.
  • [Sep 2024] Four papers at the NeurIPS regular track.
  • [Sep 2024] Our study on safety of RLHF alignment is accepted to S&P (Oakland) 2025.
  • [Jul 2024] Our multimodal jailbreak benchmark is accepted to COLM. It is from the interns in my group.
  • [Jul 2024] 4/4 papers accepted to ECCV on trustworthy VLMs and driving. Two of them are from interns in my group.
  • [Jun 2024] Senior Area Chair for the NeurIPS Datasets & Benchmarks track, and AC for the NeurIPS regular track.
  • [May 2024] Our jailbreak paper is accepted to USENIX Security. Congratulations, Zhiyuan!
  • [Mar 2024] Five papers at NAACL on LLM security (4 main, 1 Findings): two on backdoor attacks, one on backdoor defense, one on jailbreak attacks, and one on model fingerprinting.
  • [Mar 2024] PreDa for personalized federated learning is accepted at CVPR 2024.
  • [Jan 2024] Three papers at ICLR and two papers at TMLR.
  • [Oct 2023] Our paper MoleculeSTM is accepted to Nature Machine Intelligence. MoleculeSTM aligns natural language and molecule representations in the same space.
Older news
  • [Oct 2023] Three papers at EMNLP and one at NeurIPS. Our NeurIPS paper studies a new threat to instruction tuning of LLMs by injecting ads — the first work to view LLMs as generative models and attack their generative property.
  • [Oct 2023] Our tutorial on Security and Privacy in the Era of Large Language Models is accepted to NAACL.
  • [May 2023] One paper at ACL. Congratulations to Zhuofeng and Jiazhao! We propose an attention-based method to defend against NLP backdoor attacks.
  • [Apr 2023] Two papers at ICML. Congratulations to Jiachen and Zhiyuan! We propose the first benchmark for code copyright of code generation models.
  • [Feb 2023] Two papers at CVPR. Congratulations to Yiming and Xiaogeng (an intern from my group at ASU).
  • [Feb 2023] I will give a tutorial at CVPR 2023 on trustworthiness in the era of foundation models.
  • [Jan 2023] Impact Award from Argonne National Laboratory.
  • [Jan 2023] One paper accepted to USENIX Security 2023.
  • [Jan 2023] Three papers accepted to ICLR 2023. [a] explains why and how diffusion models improve adversarial robustness, and designs DensePure for state-of-the-art certified robustness. [b] is our first attempt at a retrieval-based framework for AI drug discovery.
  • [Dec 2022] Our team won the ACM Gordon Bell Special Prize for COVID-19 Research.
  • [Sep 2022] One paper accepted to USENIX Security 2023, and two papers accepted to NeurIPS 2022.
  • [Sep 2022] RobustTraj is accepted to CoRL as an oral presentation. We train trajectory prediction models that are robust to adversarial attacks.
  • [Aug 2022] Talk at the virtual seminar series on Challenges and Opportunities for Security & Privacy in Machine Learning.
  • [Jul 2022] Our survey on the challenges and opportunities of machine learning security is accepted to ACM Computing Surveys.
  • [Jul 2022] Two papers accepted to ECCV 2022.
  • [May 2022] Two papers accepted to ICML 2022.
  • [Mar 2022] Talks at the AAAI 2022 Workshop on Practical Deep Learning in the Wild and the AAAI 2022 Workshop on Adversarial Machine Learning and Beyond.
  • [Feb 2022] One paper accepted to ICLR.

🎤 Recent Talks

Earlier talks
  • Invited Talk · CVPR Workshop of Adversarial Machine Learning on Computer Vision: Robustness of Foundation Models · 2024-06
  • Tutorial · NAACL, Combating Security and Privacy Issues in the Era of Large Language Models · 2024-06
  • Invited Talk · ICLR Workshop on Secure and Trustworthy Large Language Models · 2024-05
  • Invited Talk · NeurIPS TDW Workshop · 2023-12

📄 Selected Publications full list →

30,463 citations · h-index 69 · Google Scholar, updated 2026-10-09

* equal contribution

Model / Agent Safety and Robustness Evaluation (Red-teaming)

  • AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
    Xiaogeng Liu*, Peiran Li*, G. Edward Suh, Yevgeniy Vorobeychik, Zhuoqing Mao, Somesh Jha, Patrick McDaniel, Huan Sun, Bo Li, Chaowei Xiao
    ICLR 2025 Spotlight 267 citations [pdf]
  • AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
    Xiaogeng Liu, Nan Xu, Muhao Chen, Chaowei Xiao
    ICLR 2024 1.7k citations [pdf]
  • DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agents
    Zhaorun Chen, Xun Liu, Haibo Tong, Chengquan Guo, Yuzhou Nie, Jiawei Zhang, Mintong Kang, Chejian Xu, Qichang Liu, Xiaogeng Liu, Tianneng Shi, Chaowei Xiao, Sanmi Koyejo, Percy Liang, Wenbo Guo, Dawn Song, Bo Li
    arXiv [pdf]
  • A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
    Fangzhou Wu, Ning Zhang, Somesh Jha, Patrick McDaniel, Chaowei Xiao
    arXiv 147 citations [pdf]
  • Do Not Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models
    Zhiyuan Yu, Xiaogeng Liu, Shuning Liang, Zach Cameron, Chaowei Xiao, Ning Zhang
    USENIX Security 2024 Distinguished Paper Award 256 citations [pdf]

Safety Alignment / Mitigation

  • Safety Mid-Training: Internalizing LLM Safety as a Foundational Capability
    Zhengyue Zhao, Liwei Jiang, Yejin Choi, Chaowei Xiao
    Preprint
  • ARMOR: Aligning Secure and Safe Large Language Models via Meticulous Reasoning
    Zhengyue Zhao, Yingzi Ma, Somesh Jha, Marco Pavone, Patrick McDaniel, Chaowei Xiao
    ICLR 2026 [pdf]
  • Mitigating Fine-tuning Jailbreak Attack with Backdoor Enhanced Alignment
    Jiongxiao Wang, Jiazhao Li, Yiquan Li, Xiangyu Qi, Muhao Chen, Junjie Hu, Yixuan Li, Bo Li, Chaowei Xiao
    NeurIPS 2024 77 citations [pdf]

Agent Security via System-level Solutions

  • DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agents
    Hao Li, Xiaogeng Liu, Hung-Chun Chiu, Dianqi Li, Ning Zhang, Chaowei Xiao
    NeurIPS 2025 72 citations [pdf]
  • AGrail: A Lifelong Agent Guardrail with Effective and Adaptive Safety Detection
    Weidi Luo, Shenghong Dai, Xiaogeng Liu, Suman Banerjee, Huan Sun, Muhao Chen, Chaowei Xiao
    ACL 2025 111 citations [pdf]
  • PIGuard: Prompt Injection Guardrail via Mitigating Overdefense for Free
    Hao Li, Xiaogeng Liu, Ning Zhang, Chaowei Xiao
    ACL 2025 91 citations [pdf]
  • System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
    Fangzhou Wu, Ethan Cecchetti, Chaowei Xiao
    arXiv 102 citations [pdf]

AI4Science (Bio and Math)

  • ProteinDT: A Text-guided Protein Design Framework
    Shengchao Liu, Yanjing Li, Zhuoxinran Li, Anthony Gitter, Yutao Zhu, Jiarui Lu, Zhao Xu, Weili Nie, Arvind Ramanathan, Chaowei Xiao*, Jian Tang*, Hongyu Guo*, Anima Anandkumar*
    Nature Machine Intelligence 2025 142 citations [pdf]
  • Multi-modal Molecule Structure–Text Model for Text-based Retrieval and Editing
    Shengchao Liu, Weili Nie, Chengpeng Wang, Jiarui Lu, Zhuoran Qiao, Ling Liu, Jian Tang*, Chaowei Xiao*, Animashree Anandkumar*
    Nature Machine Intelligence 407 citations [pdf]
  • ChatDrug: ChatGPT-powered Conversational Drug Editing Using Retrieval and Domain Feedback
    Shengchao Liu, Jiongxiao Wang, Yijin Yang, Chengpeng Wang, Ling Liu, Hongyu Guo, Chaowei Xiao
    ICLR 2024 72 citations [pdf]
  • LeanAgent: Lifelong Learning for Formal Theorem Proving
    Adarsh Kumarappan, Mo Tiwari, Peiyang Song, Robert Joseph George, Chaowei Xiao, Anima Anandkumar
    ICLR 2025 46 citations [pdf]

Foundation Models, Agents, Test-time Training

  • Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language Models
    Manli Shu, Weili Nie, De-An Huang, Zhiding Yu, Tom Goldstein, Anima Anandkumar, Chaowei Xiao
    NeurIPS 2022 837 citations [pdf]
  • Dolphins: Multimodal Language Model for Driving
    Yingzi Ma, Yulong Cao, Jiachen Sun, Marco Pavone, Chaowei Xiao
    ECCV 2024 232 citations [pdf]
  • Voyager: An Open-Ended Embodied Agent with Large Language Models
    Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, Anima Anandkumar
    TMLR 2024 3.7k citations [pdf]
  • Prismer: A Vision-Language Model with Multi-Task Experts
    Shikun Liu, Linxi Fan, Edward Johns, Zhiding Yu, Chaowei Xiao, Anima Anandkumar
    TMLR 2024 23 citations [pdf]
  • Shall We Pretrain Autoregressive Language Models with Retrieval? A Comprehensive Study
    Boxin Wang, Wei Ping, Peng Xu, Lawrence McAfee, Zihan Liu, Mohammad Shoeybi, Yi Dong, Oleksii Kuchaiev, Bo Li, Chaowei Xiao, Anima Anandkumar, Bryan Catanzaro
    EMNLP 2023 98 citations [pdf]
  • Re-ViLM: Retrieval-Augmented Visual Language Model for Zero and Few-Shot Image Captioning
    Zhuolin Yang, Wei Ping, Zihan Liu, Vijay Anand Korthikanti, Weili Nie, De-An Huang, Linxi Fan, Zhiding Yu, Shiyi Lan, Bo Li, Mohammad Shoeybi, Ming-Yu Liu, Yuke Zhu, Bryan Catanzaro, Chaowei Xiao*, Anima Anandkumar*
    EMNLP 2023 83 citations [pdf]

Trustworthy LLMs

  • Can Watermarks be Used to Detect LLM IP Infringement For Free? Watermark
    Zhengyue Zhao, Xiaogeng Liu, Somesh Jha, Patrick McDaniel, Bo Li, Chaowei Xiao
    ICLR 2025 [pdf]
  • HaloScope: Harnessing Unlabeled LLM Generations for Hallucination Detection Hallucination
    Xuefeng Du, Chaowei Xiao, Yixuan Li
    NeurIPS 2024 Spotlight 174 citations [pdf]
  • Instructional Fingerprinting of Large Language Models Fingerprinting
    Jiashu Xu, Fei Wang, Mingyu Derek Ma, Pang Wei Koh, Chaowei Xiao, Muhao Chen
    NAACL 2024 168 citations [pdf]
  • AgentPoison: Red-teaming LLM Agents via Memory or Knowledge Base Backdoor Poisoning Memory Poisoning
    Zhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song, Bo Li
    NeurIPS 2024 597 citations [pdf]
  • On the Exploitability of Instruction Tuning Training-time threats
    Manli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping, Chaowei Xiao*, Tom Goldstein*
    NeurIPS 2023 200 citations [pdf]

Adversarial Machine Learning

  • Diffusion Models for Adversarial Purification Defense
    Weili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao, Arash Vahdat, Anima Anandkumar
    ICML 2022 1.2k citations [pdf] [code]
  • DensePure: Understanding Diffusion Models towards Adversarial Robustness Certification
    Chaowei Xiao*, Zhongzhu Chen*, Kun Jin*, Jiongxiao Wang*, Weili Nie, Mingyan Liu, Anima Anandkumar, Bo Li, Dawn Song
    ICLR 2023 137 citations [pdf]
  • Invisible for both Camera and LiDAR: Security of Multi-Sensor Fusion based Perception in Autonomous Driving Under Physical-World Attacks Physical Attacks
    Yulong Cao*, Ningfei Wang*, Chaowei Xiao*, Dawei Yang*, Jin Fang, Ruigang Yang, Qi Alfred Chen, Mingyan Liu, Bo Li
    IEEE S&P (Oakland) 2021 471 citations [pdf]
  • Spatially Transformed Adversarial Examples Attacks
    Chaowei Xiao*, Jun-Yan Zhu*, Bo Li, Warren He, Mingyan Liu, Dawn Song
    ICLR 2018 760 citations [pdf]
  • Generating Adversarial Examples with Adversarial Networks Attacks
    Chaowei Xiao, Bo Li, Jun-Yan Zhu, Warren He, Mingyan Liu, Dawn Song
    IJCAI 2018 1.4k citations [pdf]
  • Robust Physical-World Attacks on Machine Learning Models Physical Attacks
    Kevin Eykholt*, Ivan Evtimov*, Earlence Fernandes, Bo Li, Amir Rahmati, Chaowei Xiao, Atul Prakash, Tadayoshi Kohno, Dawn Song
    CVPR 2018 4.4k citations [pdf]

👥 Current PhD Students

  • Xiaogeng Liu · PhD @ JHU (transferred from UW–Madison)
  • Zhengyue Zhao · PhD @ JHU (transferred from UW–Madison)
  • Yingzi Ma · PhD @ UW–Madison
  • Eddy Luo · PhD @ UGA (co-advised with Zhen Xiang)
  • Hao Li · PhD @ WashU (co-advised with Ning Zhang)
  • Bowen Sun · PhD @ JHU

✉️ Contact

Email: chaoweixiao [at] jhu [dot] edu