|
Research
My current research focuses on multimodal large language models, with a particular emphasis on improving their perceptual and spatial reasoning capabilities.
I am also interested in developing efficient online video models and enhancing the applicability of vision-language models in real-world scenarios.
Previously, my research centered on geometric graph neural networks, especially for simulating complex physical systems.
Representative papers are highlighted.
|
Internship
Rayneo.AI, Remote Intern (May. 2025 ~ March. 2026)
Mentor: Sijie Cheng
Multimodal Large Language Models, Egocentric Video Understanding
Alibaba Taobao and Tmall, Content Understanding Group, Research Intern (Aug. 2026 ~ Present)
Multimodal Large Language Models, Agent
|
Last updated: August 15, 2026
Template from Jon Barron. Big thanks!
|
|