Ruiyang Xu 徐瑞阳
M.S. Student in CS
Shanghai Jiao Tong University (SJTU)
Minhang District, Shanghai, China
xry2022 [at] sjtu.edu.cn
About
I am a graduate student in computer science at Shanghai Jiao Tong University, advised by Prof. Xie Chen.
My research focuses on cutting-edge Multimodal Large Language Models (MLLMs), specifically multimodal reasoning and understanding. As a Research Intern at Alibaba Qwen, I have contributed to Qwen3-Omni and Qwen3.5-Omni, focusing on detailed captioning and agents.
I have been involved in several research projects in multimodal reasoning, omni-modal perception, and audio understanding, some of which have appeared at ICML, ICLR, and NeurIPS.
In my spare time, I listen to Britpop and J-rock.
News
- 2026.06 Released MMAE, the first comprehensive benchmark for instruction-based audio editing.
- 2026.06 OmniAgent is now open-sourced — code, paper, and models all available.
- 2026.05 OmniAgent introducing native active perception has been accepted to ICML 2026.
- 2026.03 Released Qwen3.5-Omni, featuring superior performance in audio-visual detailed captioning.
- 2026.03 Check out Omni-Cloze, a novel benchmark dedicated to multimodal detailed captioning.
- 2026.01 Omni-Captioner accepted at ICLR 2026.
- 2026.01 Join our Audio Reasoning Challenge at Interspeech 2026. Check the project page.
- 2025.09 Qwen3-Omni is released. Check the blog.
- 2025.09 MMAR benchmark is accepted at NeurIPS 2025.
- 2025.05 SLAM-Omni is accepted to the ACL 2025 Findings.
Selected Publications
* Equal contribution
* Equal contribution
Education
M.S. in Computer Science
B.S. in Computer Science