I am a senior undergraduate student at School of Mechanical, Electronic and Control Engineering, Beijing Jiaotong University (BJTU), expecting to receive my B.Eng. degree in 2027. I will pursue my graduate studies at Institute of Automation, Chinese Academy of Sciences (CASIA) as a master’s student. My current research centers on multimodal perception and embodied agents; at CASIA I plan to work on vision-language-action (VLA) models and world models. If you are interested in my research, please feel free to contact me at 23222002@bjtu.edu.cn.
🔥 News
- 2026.10: 🔥🔥 We release ROMA, an LLM-Based System for Real-World Object-Centric Multi-Sensory Active Perception! One step towards active multi-sensory embodied agents!
- 2026.03: 🎉🎉 Nonholonomic Narrow Dead-end Escape with Reinforcement Learning is accepted to CSAI 2026! The codes are released!
📝 Publications
- Nonholonomic Narrow Dead-end Escape with Reinforcement Learning
Denghan Xiong*, Yanzhe Zhao*, Yutong Chen*, Zichun Wang* (*equal contribution)
CSAI 2026
[Paper] | [Code]
📄 Preprints

💻 Projects
- Visually based automatic docking system for two-wheeled robots — A YOLO-based ROS 2 pipeline that uses an RGB-D camera to detect docking targets and close the loop for autonomous two-wheeled robot docking.
PythonPyTorchROS2
📖 Educations
- 2027.09 - 2030.06 (expected), M.S., Institute of Automation, Chinese Academy of Sciences (CASIA).
- 2023.09 - 2027.06 (expected), B.Eng., School of Mechanical, Electronic and Control Engineering, Beijing Jiaotong University (BJTU).
💼 Internships
- 2025.10 - 2026.10, Research Assistant, Gewu Lab, Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China. Advised by Prof. Di Hu.
Co-first author of ROMA, an LLM-based system for real-world object-centric multi-sensory active perception. ROMA turns passive sensing into an active reasoning–interaction–feedback loop: a multi-sensory LLM (ROMA-7B) decides which evidence is missing and which object, interaction and modality to probe, while a physical interface executes it on a real robot arm and streams back vision, audio, touch and force feedback, reaching 72.9% success versus 53.0% for the strongest baseline.
I was responsible for the following: 1. Grasp pose planning of real-world interaction; 2. Anotation design and data collection
🎖 Honors and Awards
- 2024.10 Shenzhou Railway Scholarship, Beijing Jiaotong University.