Hi~ I am Peng Ding (丁鹏), receiving my Ph.D. from the School of Computer Science, Nanjing University in September 2026, supervised by Prof. Shujian Huang.
My research interests focus on the safety of large language models (LLMs), including jailbreak attacks, defense mechanisms, and interpretability. I am also interested in other topics related to LLMs, such as reasoning, reinforcement learning, and Agentic AI.
I am currently working at Alibaba Group in Hangzhou, on LLM safety and Agentic AI. Please feel free to contact me via email!
🔥 News
- 2026.05: 🎉🎉 Our paper “Friend or Foe: How LLMs’ Safety Mind Gets Fooled by Intent Shift Attack” is accepted by TACL.
- 2025.08: 🎉🎉 Our paper “SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language Models” is accepted by EMNLP 2025.
- 2025.05: 🎉🎉 Our paper “Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement” is accepted by ACL 2025 (Findings).
- 2024.07: 🎉🎉 Our paper “Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs” is accepted by MM 2024.
- 2024.03: 🎉🎉 Our paper “A Wolf in Sheep’s Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily” is accepted by NAACL 2024 (Oral).
📝 Publications

Friend or Foe: How LLMs’ Safety Mind Gets Fooled by Intent Shift Attack
Peng Ding, Jun Kuang, Wen Sun, Zongyu Wang, Xuezhi Cao, Xunliang Cai, Jiajun Chen, Shujian Huang
💻 [Code]: Link
📄 [Paper]: Link




📖 Educations
- 2019.06 - 2026.09, Ph.D., School of Computer Science, Nanjing University.
- 2016.09 - 2019.06, Master’s degree, School of Information Science and Engineering, Yunnan University.
💻 Experiences
- 2026.08 - now, Alibaba Group, Hangzhou, China.
- 2023.08 - 2026.07, Meituan Inc., Shanghai, China.
🎖 Honors and Awards
- 2018.10 Yunnan Provincial Government Scholarship.