6 papers
AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models
Guiyu Zhao, Longteng Guo, Yanghong Mei +7
While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon task…
TimeThink: Reasoning with Time for Video LLMs
Handong Li, Longteng Guo, Zikang Liu +8
Video reasoning requires models to identify and verify temporally localized evidence within long video sequences. Recent Video Large Language Models (Video-LLMs) have shown promisi…
VeriSpace: Spatially Grounded Action Verification for Vision-Language-Action Models
Guiyu Zhao, Longteng Guo, Junyou Zhu +6
Vision-language-action (VLA) models have shown strong promise for robotic manipulation, but their reliability at test time remains limited by one-shot action prediction, where even…
S1-MMAlign: A Large-Scale, Multi-Disciplinary Dataset for Scientific Figure-Text Understanding
He Wang, Longteng Guo, Pengkang Huo +4
Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap between complex scientific imagery…
BrainGPT: Unleashing the Potential of EEG Generalist Foundation Model by Autoregressive Pre-training
Tongtian Yue, Xuange Gao, Shuning Xue +4
Electroencephalogram (EEG) signals are pivotal in providing insights into spontaneous brain activity, highlighting their significant importance in neuroscience research. However, t…
Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question Answering
Dongze Hao, Qunbo Wang, Longteng Guo +2
While large visual-language models (LVLM) have shown promising results on traditional visual question answering benchmarks, it is still challenging for them to answer complex VQA p…