Chaowei Xiao
Assistant Professor · Johns Hopkins University|Researcher · NVIDIA Research
Bio: I am Chaowei Xiao, an assistant professor at Johns Hopkins University (JHU) and a researcher at NVIDIA Research, currently visiting Stanford University. I obtained my Ph.D. from the University of Michigan, Ann Arbor, and my bachelor's degree from Tsinghua University.
My research aims to build safe AGI, with a focus on AI and agent security. I have also recently become interested in Recursive Self Improvement, Continual Learning, Scientific Discovery and principled methods for diverse application domains, including AI4Science, embodied agents, and computer-use agents.
Longer bio
Chaowei Xiao is an Assistant Professor at Johns Hopkins University and a researcher at NVIDIA Research. His research aims to build safe and secure AI, with a focus on LLM security and AI agent security.
He has received several prestigious awards, including the Schmidt Sciences AI2050 Early Career Fellowship, MIT Technology Review Innovators Under 35 (TR35), the Impact Award from Argonne National Laboratory, and various industry faculty awards. His work has been recognized with several paper awards at top-tier conferences, including the USENIX Security Distinguished Paper Award (2024), the EWSN Best Paper Award (2021), the MobiCom Best Paper Award (2014), ACM Gordon Bell Prize Finalist (2024), and the ACM Gordon Bell Special Prize for HPC-Based COVID-19 Research (2023). His research has been cited more than 30,000 times and has been featured in media outlets such as Nature, Wired, Fortune, and The New York Times. One of his research outputs is on display at the Science Museum in London.
Our group plans to recruit multiple PhD students. I am interested in students in general AI, cybersecurity, or robotics (VLA / world models).
I'm also looking for multiple postdocs with experience in cybersecurity, software engineering, security, RL, or robotics.
🏆 Awards (selected)
- MIT Technology Review Innovators Under 35 (TR35 China) 2026
- SafeBench Competition, Second Prize (JailbreakV) · 2025
- USENIX Security Distinguished Paper Award · 2024
- Schmidt Sciences AI2050 Early Career Fellow · 2024
- ACM Gordon Bell Prize Finalist · 2024
- NVIDIA Graduate Fellowship (Security track), awarded to my PhD student Xiaogeng Liu · 2024
- Stanford/Elsevier Top 2% Scientists List · 2024
- ACM Gordon Bell Special Prize for HPC-Based COVID-19 Research · 2023
- Impact Award, Argonne National Laboratory · 2023
- EWSN Best Paper Award · 2021
- MobiCom Best Paper Award · 2014
🐾 Research all publications →
-
Evaluating AI & Agent Safety. How and why do AI models and agents fail?
-
How can we evaluate model safety automatically and at scale, even as models recursively self-improve?
-
What threats do AI agents face, and how do we evaluate them?
-
What new risks come with multimodal models, reasoning models, and new architectures such as diffusion language models?
-
How do we build benchmarks that comprehensively evaluate LLM and agent safety and security?
-
What threats arise at training time?
Instruction Backdoors AutoPoison BGMAttack RLHFPoison Preference Poisoning
-
-
Alignment & Guardrails. How do we make AI and agent safety robust, not a thin layer?
-
How can we make models safer through alignment?
ARMOR Safety Mid-Training Concept Concentration Surface Heuristics AdaShield BackdoorAlign LM Detox RePD FIUBench
-
How do we detect and remove backdoors?
-
-
Agent Security. Can we secure agents by design, even when the model can be fooled?
-
Can system-level designs, such as information-flow control and dynamic security-policy enforcement, secure AI agents?
-
Can guardrails adapt to new threats without over-blocking?
-
Can we control what an agent knows and does from the inside?
-
-
Training Models & Building Agents. How do we build AI models and agents that keep improving?
-
How do we build lifelong-learning (self-evolving) agents and multi-agent systems?
-
Can models adapt at test time and keep improving (continual learning)?
-
Can multimodal models perceive and act in the physical world?
VoxFormer Dolphins MuirBench RealGen DreamDrive dVLM-AD CALICO TrafficRLHF
-
Can foundation models be more efficient and versatile?
Transformer-LS RETRO Study Prismer Re-ViLM T-Stitch DataGen HMN RelViT HiCL
-
-
AI4Science. Molecule design, proteins, and AI4Math.
MoleculeSTM ChatDrug RetMol ProteinDT MProt-DPO GenSLMs LeanAgent
-
Trustworthy & Robust ML. How do models behave under worst-case inputs, and what else does trust require?
-
Do adversarial attacks work in the physical world?
-
What do adversarial examples look like beyond pixel noise?
AdvGAN stAdv SemanticAdv Spatial Consistency Deep RL Attacks SMACK Diffusion Attack
-
Can we make models robust, empirically and provably?
DiffPure SA-MDP CROWN-IBP FAN AugMax DensePure Conserved Features AdvIT DiffSmooth AudioPure Consistency Purification SSNI 3D Purification 3D Self-Supervision Shape Robustness
-
How do we protect model IP and data privacy, and catch hallucinations?
HaloScope LLM Fingerprinting LLM Watermark CodeIPPrompt Copyright Detective Private Data Publishing SecretGen Co-Membership Attack DP Video RF Privacy Exploit Discovery
-
📰 News
- [Oct 2026] Four papers accepted to NeurIPS 2026 on agent security and safety: Agent MechSuits (safety steering for multi-turn CLI agents), LPG (latent policy guardrails), MaskForge (jailbreaking diffusion LLMs), and position-weighted self-distillation for reasoning. Congratulations to all authors!
- [Oct 2026] More 2026 papers on LLM and agent security: ReasoningBomb (CCS 2026), Your Harness is Not Secure (EMNLP 2026), Concept Concentration (ICML 2026), and four papers at ICLR 2026, including ARMOR and A2ASecBench.
- [Mar 2026] Serving as Senior Area Chair for ACL, EMNLP, and NeurIPS.
- [Jan 2025] We have 9 papers at ICLR and 3 papers at ACL on model safety and security. Congratulations to all authors!
- [Dec 2024] We received the Fall Research Competition Award at UW–Madison. Thank you UW–Madison and OVCR.
- [Nov 2024] Our group recently received funding and donations. Thank you Amazon and Apple!
- [Nov 2024] Our lab will have a winter break this December. Lab members will enjoy some well-deserved vacation time with their families and loved ones.
- [Sep 2024] Four papers at the NeurIPS regular track.
- [Sep 2024] Our study on safety of RLHF alignment is accepted to S&P (Oakland) 2025.
- [Jul 2024] Our multimodal jailbreak benchmark is accepted to COLM. It is from the interns in my group.
- [Jul 2024] 4/4 papers accepted to ECCV on trustworthy VLMs and driving. Two of them are from interns in my group.
- [Jun 2024] Senior Area Chair for the NeurIPS Datasets & Benchmarks track, and AC for the NeurIPS regular track.
- [May 2024] Our jailbreak paper is accepted to USENIX Security. Congratulations, Zhiyuan!
- [Mar 2024] Five papers at NAACL on LLM security (4 main, 1 Findings): two on backdoor attacks, one on backdoor defense, one on jailbreak attacks, and one on model fingerprinting.
- [Mar 2024] PreDa for personalized federated learning is accepted at CVPR 2024.
- [Jan 2024] Three papers at ICLR and two papers at TMLR.
- [Oct 2023] Our paper MoleculeSTM is accepted to Nature Machine Intelligence. MoleculeSTM aligns natural language and molecule representations in the same space.
Older news
- [Oct 2023] Three papers at EMNLP and one at NeurIPS. Our NeurIPS paper studies a new threat to instruction tuning of LLMs by injecting ads — the first work to view LLMs as generative models and attack their generative property.
- [Oct 2023] Our tutorial on Security and Privacy in the Era of Large Language Models is accepted to NAACL.
- [May 2023] One paper at ACL. Congratulations to Zhuofeng and Jiazhao! We propose an attention-based method to defend against NLP backdoor attacks.
- [Apr 2023] Two papers at ICML. Congratulations to Jiachen and Zhiyuan! We propose the first benchmark for code copyright of code generation models.
- [Feb 2023] Two papers at CVPR. Congratulations to Yiming and Xiaogeng (an intern from my group at ASU).
- [Feb 2023] I will give a tutorial at CVPR 2023 on trustworthiness in the era of foundation models.
- [Jan 2023] Impact Award from Argonne National Laboratory.
- [Jan 2023] One paper accepted to USENIX Security 2023.
- [Jan 2023] Three papers accepted to ICLR 2023. [a] explains why and how diffusion models improve adversarial robustness, and designs DensePure for state-of-the-art certified robustness. [b] is our first attempt at a retrieval-based framework for AI drug discovery.
- [Dec 2022] Our team won the ACM Gordon Bell Special Prize for COVID-19 Research.
- [Sep 2022] One paper accepted to USENIX Security 2023, and two papers accepted to NeurIPS 2022.
- [Sep 2022] RobustTraj is accepted to CoRL as an oral presentation. We train trajectory prediction models that are robust to adversarial attacks.
- [Aug 2022] Talk at the virtual seminar series on Challenges and Opportunities for Security & Privacy in Machine Learning.
- [Jul 2022] Our survey on the challenges and opportunities of machine learning security is accepted to ACM Computing Surveys.
- [Jul 2022] Two papers accepted to ECCV 2022.
- [May 2022] Two papers accepted to ICML 2022.
- [Mar 2022] Talks at the AAAI 2022 Workshop on Practical Deep Learning in the Wild and the AAAI 2022 Workshop on Adversarial Machine Learning and Beyond.
- [Feb 2022] One paper accepted to ICLR.
🎤 Recent Talks
- Discussion lead, How to Secure AI Agents · Alignment Workshop @ San Diego (NeurIPS) · 2025-12
- Invited Talk · 5th Workshop on Adversarial Machine Learning on Computer Vision: Foundation Models + X @ CVPR · 2025-06
- Invited Talk · Secure Generative AI Agents Workshop @ IEEE S&P · 2025-05
- Invited Talk · International Symposium on Trustworthy Foundation Models · 2025-04
- Invited Talk · SFU @ NeurIPS · 2024-12
- Invited Talk · LLM and Agent Safety Competition @ NeurIPS · 2024-12
- Keynote · CCS Workshop on Large AI Systems and Models with Privacy and Safety Analysis · 2024-10
- Invited Talk · Trillion Parameter Consortium (TPC), LLM safety and security · 2024-10
- Invited Talk · NSF Workshop on Large Language Models for Network Security · 2024-10
Earlier talks
- Invited Talk · CVPR Workshop of Adversarial Machine Learning on Computer Vision: Robustness of Foundation Models · 2024-06
- Tutorial · NAACL, Combating Security and Privacy Issues in the Era of Large Language Models · 2024-06
- Invited Talk · ICLR Workshop on Secure and Trustworthy Large Language Models · 2024-05
- Invited Talk · NeurIPS TDW Workshop · 2023-12
📄 Selected Publications full list →
30,463 citations · h-index 69 · Google Scholar, updated 2026-10-09
* equal contribution
Model / Agent Safety and Robustness Evaluation (Red-teaming)
-
AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
-
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
-
DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI AgentsarXiv [pdf]
-
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
-
Do Not Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models
Safety Alignment / Mitigation
-
Safety Mid-Training: Internalizing LLM Safety as a Foundational CapabilityPreprint
-
ARMOR: Aligning Secure and Safe Large Language Models via Meticulous ReasoningICLR 2026 [pdf]
-
Mitigating Fine-tuning Jailbreak Attack with Backdoor Enhanced Alignment
Agent Security via System-level Solutions
-
DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agents
-
AGrail: A Lifelong Agent Guardrail with Effective and Adaptive Safety Detection
-
PIGuard: Prompt Injection Guardrail via Mitigating Overdefense for Free
-
System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
AI4Science (Bio and Math)
-
ProteinDT: A Text-guided Protein Design Framework
-
Multi-modal Molecule Structure–Text Model for Text-based Retrieval and Editing
-
ChatDrug: ChatGPT-powered Conversational Drug Editing Using Retrieval and Domain Feedback
-
LeanAgent: Lifelong Learning for Formal Theorem Proving
Foundation Models, Agents, Test-time Training
-
Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language Models
-
Dolphins: Multimodal Language Model for Driving
-
Voyager: An Open-Ended Embodied Agent with Large Language Models
-
Prismer: A Vision-Language Model with Multi-Task Experts
-
Shall We Pretrain Autoregressive Language Models with Retrieval? A Comprehensive Study
-
Re-ViLM: Retrieval-Augmented Visual Language Model for Zero and Few-Shot Image Captioning
Trustworthy LLMs
-
Can Watermarks be Used to Detect LLM IP Infringement For Free? WatermarkICLR 2025 [pdf]
-
HaloScope: Harnessing Unlabeled LLM Generations for Hallucination Detection Hallucination
-
Instructional Fingerprinting of Large Language Models Fingerprinting
-
AgentPoison: Red-teaming LLM Agents via Memory or Knowledge Base Backdoor Poisoning Memory Poisoning
-
On the Exploitability of Instruction Tuning Training-time threats
Adversarial Machine Learning
-
Diffusion Models for Adversarial Purification Defense
-
DensePure: Understanding Diffusion Models towards Adversarial Robustness Certification
-
Invisible for both Camera and LiDAR: Security of Multi-Sensor Fusion based Perception in Autonomous Driving Under Physical-World Attacks Physical Attacks
-
Spatially Transformed Adversarial Examples Attacks
-
Generating Adversarial Examples with Adversarial Networks Attacks
-
Robust Physical-World Attacks on Machine Learning Models Physical Attacks
👥 Current PhD Students
- Xiaogeng Liu · PhD @ JHU (transferred from UW–Madison)
- Zhengyue Zhao · PhD @ JHU (transferred from UW–Madison)
- Yingzi Ma · PhD @ UW–Madison
- Eddy Luo · PhD @ UGA (co-advised with Zhen Xiang)
- Hao Li · PhD @ WashU (co-advised with Ning Zhang)
- Bowen Sun · PhD @ JHU
✉️ Contact
Email: chaoweixiao [at] jhu [dot] edu