7 papers · 1 filter
InterSketch: An Interleaved Reasoning Model with Self-correcting Visual Sketch and Stepwise Reward
Zhiwei Ning, Wenwen Tong, Xiangli Kong +12
While vision-language models (VLMs) have exhibited multi-turn visual reasoning capabilities, their reasoning trajectories remain relatively shallow and are dominated by a text-cent…
Stage-adaptive Token Selection for Efficient Omni-modal LLMs
Zijie Xin, Jie Yang, Ruixiang Zhao +4
Omni-modal large language models (om-LLMs) achieve unified audio-visual understanding by encoding video and audio into temporally aligned token sequences interleaved at the window…
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
Ruixiang Zhao, Jie Yang, Zijie Xin +4
Omni-proactive streaming video understanding, i.e., autonomously deciding when to speak and what to say from continuous audio-visual streams, is an emerging capability of omni-moda…
KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model
Jie Yang, Wang Zeng, Sheng Jin +5
The emergence of Multimodal Large Language Models (MLLMs) has revolutionized image understanding by bridging textual and visual modalities. However, these models often struggle wit…
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
Jie Yang, Feipeng Ma, Zitian Wang +4
Building on the success of text-based reasoning models like DeepSeek-R1, extending these capabilities to multimodal reasoning holds great promise. While recent works have attempted…
Unlock the Power of Unlabeled Data in Language Driving Model
Chaoqun Wang, Jie Yang, Xiaobin Hong +1
Recent Vision-based Large Language Models~(VisionLLMs) for autonomous driving have seen rapid advancements. However, such promotion is extremely dependent on large-scale high-quali…