I am a second-year master student at the School of AI, Beihang University (BUAA), supervised by Prof. Lei Sha. I was honored to be advised by Jing Shao, JG Yao, and Junxian He.
My previous research focused on the safety alignment of AI and long-horizon reasoning, and I am now seeking a PhD position for 2027 Fall.
π₯ News
- 2026.08: Β ππ SafeSteer is accepted by EMNLP 2026.
- 2026.04: Β ππ SSP is accepted by ACL 2026 Findings.
- 2026.02: Β ππ ReVeL is accepted by CVPR 2026.
- 2025.08: Β ππ Two papers (LARF and DIffusionAttacker) are accepted by EMNLP 2025 and DIffusionAttacker is selected as Oral Presentation.
- 2025.03: Β ππ Two papers (ActorBreaker and VLSBench) are accepted by ACL 2025 and ActorBreaker is selected as Outstanding Paper.
- 2024.09: Β ππ ASETF is accepted by EMNLP 2024 and selected as Oral Presentation.
π Publications
SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
Hao Li*, Jingkun An*, Zijun Song*, Pengyu Zhu, Rui Li, Hao Wang, Wendi Feng, Yesheng Liu, Lijun Li, Jin-Ge Yao, Lei Sha
TL;DR: Restricting the reverse KL penalty to algorithmically identified safety tokens, distilled from an activation-steering safety instructor, aligns LLMs with only 100 harmful samples and no general-purpose data β cutting the alignment tax on general capabilities.
Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay
Hao Wang, Yanting Wang, Hao Li, Rui Li, Lei Sha
TL;DR: Safety Self-Play (SSP) lets one model be both attacker and defender in a single RL loop, with UCB-sampled reflective experience replay over low-reward failures, so defenses keep evolving instead of overfitting static adversarial datasets.
Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT
Yesheng Liu, Hao Li, Haiyu Xu, Baoqi Pei, Jiahao Wang, Mingxuan Zhao, Jingshu Zheng, Zheqi He, JG Yao, Bowen Qin, Xi Yang, Jiajun Zhang
TL;DR: Multiple-choice options leak exploitable signals that inflate scores and reward guessing during RFT; ReVeL rewrites them into open-form yet still verifiable questions, improving OpenQA accuracy by ~6 points and exposing up to 20 points of MCQA score inflation.
Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment
Hao Li*, Lijun Li*, Zhenghao Lu, Xianyi Wei, Rui Li, Jing Shao, Lei Sha
TL;DR: Seemingly benign finetuning data hides safety-degrading features; LARF locates safety-sensitive layers and uses their representations to filter those samples out, mitigating the alignment degradation that downstream finetuning otherwise causes.
DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak
Hao Wang, Hao Li, Junda Zhu, Xinyuan Wang, Chengwei Pan, Minlie Huang, Lei Sha
TL;DR: A seq2seq text diffusion model rewrites harmful prompts end-to-end, steered by an attack loss and Gumbel-Softmax differentiable sampling, which frees the attack from suffix templates and iterative token search while beating prior methods on success rate, fluency, and diversity.
LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts
Qibing Ren*, Hao Li*, Dongrui Liu, Zhanxu Xie, Xiaoya Lu, Yu Qiao, Lei Sha, Junchi Yan, Lizhuang Ma, Jing Shao
TL;DR: Grounded in actor-network theory, ActorBreaker mines actors tied to a toxic goal within the pre-training distribution and builds multi-turn benign-looking prompts that walk models into unsafe content β revealing that safety training covers too narrow a semantic space.
VLSBench: Unveiling Visual Leakage in Multimodal Safety
Xuhao Hu, Dongrui Liu, Hao Li, Xuanjing Huang, Jing Shao
TL;DR: Existing multimodal safety benchmarks leak the imageβs risky content into the text query, so models can refuse from text alone; the leakage-free VLSBench (2.2k pairs) restores reliable cross-modal evaluation and challenges LLaVA, Qwen2-VL, and GPT-4o alike.
ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings
Hao Wang*, Hao Li*, Minlie Huang, Lei Sha
TL;DR: Translating continuous adversarial suffix embeddings into coherent natural text replaces costly discrete token search, yielding fluent, perplexity-filter-resistant and transferable jailbreak prompts at a fraction of the compute.
π Reports
-
AISafetyLab: A Comprehensive Framework for AI Safety Evaluation and Improvement, CoAI & Lesca Group
-
SafeWork-R1: Coevolving Safety and Intelligence under the AI-45 Law, Shanghai AI Lab
π Educations
- 2024.09 - present, Master, Beihang University, Beijing.
- 2020.09 - 2024.06, Bachelor, Beihang University, Beijing.
π Selected Honors and Awards
- 2025, National Scholarship in China.
- 2023, Special Prize (Top 1) in βChallenge Cupβ Competition of Science Achievement in China.
π§© Academic Services
- Conference Review: ACL, EMNLP, NAACL, AAAI
- Workshop Challenge Organizer: Trustworthy Multi-modal Foundation Models and AI Agents (TiFA) in ICML 2024.
π» Internships
- 2026.03 - 2026.06, Agentic RL, TikTok AI Innovation Center
- 2025.08 - 2026.01, VLM post-training & evaluation, BAAI
- 2024.07 β 2025.07, LLM and VLM safety, Shanghai AI Lab