About me

I currently serve as an Assistant Professor with the VCC group at the College of Computer Science and Software Engineering, Shenzhen University. I earned my Ph.D at Xidian University in 2024, and obtained my B.Sc from the same university in 2018. My Ph.D advisor was Prof. Chen Bo.

My long-term research goal is to build reliable AI systems that can perceive, reason, learn from feedback, and eventually act intelligently in the physical world.

  • Embodied AI: Vision-language-action models, world models, agents, and reward models.
  • Multimodal understanding & reasoning: Multimodal large language models, multimodal retrieval, and visual question answering.
  • Machine learning & Bayesian analysis: Variational inference, hierarchical models, and optimal transport.

We aim to develop principled algorithms and practical systems that connect perception, language, decision-making, and action. In particular, we are excited about embodied intelligence: AI agents that can understand human intent, learn from complex environments, and assist people in the real world. We welcome students who are curious, self-motivated, and eager to explore ambitious ideas at the frontier of AGI.

Open Positions

  • We are actively looking for motivated undergraduate and Master students who are excited about embodied AI, multimodal foundation models, Bayesian machine learning, and related areas. If you enjoy asking bold questions and building real systems, please feel free to send me your CV.
  • Students will have opportunities to work on frontier research problems with the goal of publishing at top AI/ML/CV conferences. Please refer to Research and Publications for our recent work. Outstanding students can be recommended to leading universities for Ph.D. studies.
  • I also welcome informal conversations with junior PhD, Master, and undergraduate students. If you would like to discuss research ideas, career plans, or the future of AGI and embodied intelligence, feel free to contact me. I set aside time each week for these conversations.

News

  • [05/2026] One paper is accepted by ICML 2026 (Zero-Shot 3D Question Answering via KeyVT). Our first paper to 3D world. Congratulations to all coauthors!
  • [04/2026] Zhang Li received the Best Presentation Award at ICIGP. Congratulations!
  • [02/2026] Two papers are accepted by CVPR 2026 (SparseSAE and STiTch for Zero-Shot Composed Image Retrieval). Congratulations to all coauthors!
  • [07/2025] One paper is accepted by ICCV 2025 (Dynamic Multimodal Prototype Learning for CLIP). Congratulations to Xinyu!
  • [05/2025] One paper is accepted by IJCAI 2025 (Consistency Alignment for the Compositional Zero-Shot task). Congratulations to Miaoge!
  • [07/2024] We are presenting one paper at ECCV 2024 (Instruction Tuning-free Visual Token Complement for MLLMs), and one paper at UAI 2024 (Patch-Prompt Aligned Bayesian Prompt Tuning for Vision-Language Models). Congratulations to my coauthors!
  • [07/2024] Hello, SZU. I accept the offer from Shenzhen University and join the VCC group at College of Computer Science and Software Engineering as an Assistant Professor. The next level of my academic journey is about to begin.
  • I regularly serve as a reviewer for top-tier conferences and journals, including NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, TPAMI, IJCV, TKDE, and TNNLS.