I am a second-year master student at the School of AI, Beihang University (BUAA), supervised by Prof. Lei Sha. I was honored to be advised by Jing Shao, JG Yao, and Junxian He.

My previous research focused on the safety alignment of AI and long-horizon reasoning, and I am now seeking a PhD position for 2027 Fall.

πŸ”₯ News

  • 2026.08: Β πŸŽ‰πŸŽ‰ SafeSteer is accepted by EMNLP 2026.
  • 2026.04: Β πŸŽ‰πŸŽ‰ SSP is accepted by ACL 2026 Findings.
  • 2026.02: Β πŸŽ‰πŸŽ‰ ReVeL is accepted by CVPR 2026.
  • 2025.08: Β πŸŽ‰πŸŽ‰ Two papers (LARF and DIffusionAttacker) are accepted by EMNLP 2025 and DIffusionAttacker is selected as Oral Presentation.
  • 2025.03: Β πŸŽ‰πŸŽ‰ Two papers (ActorBreaker and VLSBench) are accepted by ACL 2025 and ActorBreaker is selected as Outstanding Paper.
  • 2024.09: Β πŸŽ‰πŸŽ‰ ASETF is accepted by EMNLP 2024 and selected as Oral Presentation.

πŸ“ Publications

EMNLP 2026

SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment

Hao Li*, Jingkun An*, Zijun Song*, Pengyu Zhu, Rui Li, Hao Wang, Wendi Feng, Yesheng Liu, Lijun Li, Jin-Ge Yao, Lei Sha

TL;DR: Restricting the reverse KL penalty to algorithmically identified safety tokens, distilled from an activation-steering safety instructor, aligns LLMs with only 100 harmful samples and no general-purpose data β€” cutting the alignment tax on general capabilities.

ACL 2026 Findings

Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay

Hao Wang, Yanting Wang, Hao Li, Rui Li, Lei Sha

TL;DR: Safety Self-Play (SSP) lets one model be both attacker and defender in a single RL loop, with UCB-sampled reflective experience replay over low-reward failures, so defenses keep evolving instead of overfitting static adversarial datasets.

CVPR 2026

Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT

Yesheng Liu, Hao Li, Haiyu Xu, Baoqi Pei, Jiahao Wang, Mingxuan Zhao, Jingshu Zheng, Zheqi He, JG Yao, Bowen Qin, Xi Yang, Jiajun Zhang

TL;DR: Multiple-choice options leak exploitable signals that inflate scores and reward guessing during RFT; ReVeL rewrites them into open-form yet still verifiable questions, improving OpenQA accuracy by ~6 points and exposing up to 20 points of MCQA score inflation.

EMNLP 2025

Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment

Hao Li*, Lijun Li*, Zhenghao Lu, Xianyi Wei, Rui Li, Jing Shao, Lei Sha

TL;DR: Seemingly benign finetuning data hides safety-degrading features; LARF locates safety-sensitive layers and uses their representations to filter those samples out, mitigating the alignment degradation that downstream finetuning otherwise causes.

EMNLP 2025 Oral

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak

Hao Wang, Hao Li, Junda Zhu, Xinyuan Wang, Chengwei Pan, Minlie Huang, Lei Sha

TL;DR: A seq2seq text diffusion model rewrites harmful prompts end-to-end, steered by an attack loss and Gumbel-Softmax differentiable sampling, which frees the attack from suffix templates and iterative token search while beating prior methods on success rate, fluency, and diversity.

ACL 2025 Outstanding Paper

LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts

Qibing Ren*, Hao Li*, Dongrui Liu, Zhanxu Xie, Xiaoya Lu, Yu Qiao, Lei Sha, Junchi Yan, Lizhuang Ma, Jing Shao

TL;DR: Grounded in actor-network theory, ActorBreaker mines actors tied to a toxic goal within the pre-training distribution and builds multi-turn benign-looking prompts that walk models into unsafe content β€” revealing that safety training covers too narrow a semantic space.

ACL 2025

VLSBench: Unveiling Visual Leakage in Multimodal Safety

Xuhao Hu, Dongrui Liu, Hao Li, Xuanjing Huang, Jing Shao

TL;DR: Existing multimodal safety benchmarks leak the image’s risky content into the text query, so models can refuse from text alone; the leakage-free VLSBench (2.2k pairs) restores reliable cross-modal evaluation and challenges LLaVA, Qwen2-VL, and GPT-4o alike.

EMNLP 2024 Oral

ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings

Hao Wang*, Hao Li*, Minlie Huang, Lei Sha

TL;DR: Translating continuous adversarial suffix embeddings into coherent natural text replaces costly discrete token search, yielding fluent, perplexity-filter-resistant and transferable jailbreak prompts at a fraction of the compute.

πŸ“ Reports

πŸ“– Educations

  • 2024.09 - present, Master, Beihang University, Beijing.
  • 2020.09 - 2024.06, Bachelor, Beihang University, Beijing.

πŸŽ– Selected Honors and Awards

  • 2025, National Scholarship in China.
  • 2023, Special Prize (Top 1) in β€œChallenge Cup” Competition of Science Achievement in China.

🧩 Academic Services

  • Conference Review: ACL, EMNLP, NAACL, AAAI
  • Workshop Challenge Organizer: Trustworthy Multi-modal Foundation Models and AI Agents (TiFA) in ICML 2024.

πŸ’» Internships

  • 2026.03 - 2026.06, Agentic RL, TikTok AI Innovation Center
  • 2025.08 - 2026.01, VLM post-training & evaluation, BAAI
  • 2024.07 – 2025.07, LLM and VLM safety, Shanghai AI Lab