6 papers
Advantage-Guided Gate: Reshaping Open-Ended Reasoning for Vision-Based Spatial Intelligence
Ling Lin, Yang Bai, Congcong Zhu +6
Multimodal large language models (MLLMs) have demonstrated significant potential in complex spatial scene understanding and reasoning tasks. However, their open-ended reasoning pro…
MeshTok: Efficient Multi-Scale Tokenization for Scalable PDE Transformers
Yanshun Zhao, Xiaoyu Peng, Jiamin Jiang +2
Conventional patchified Transformers operate on uniform spatial partitions, distributing computational effort evenly across the domain irrespective of local features. This inflexib…
OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models
Ling Lin, Yang Bai, Heng Su +5
Existing Visual-Language Models (VLMs) have achieved significant progress by being trained on massive-scale datasets, typically under the assumption that data are independent and i…
Physics-Informed Deformable Gaussian Splatting: Towards Unified Constitutive Laws for Time-Evolving Material Field
Haoqin Hong, Ding Fan, Fubin Dou +4
Recently, 3D Gaussian Splatting (3DGS), an explicit scene representation technique, has shown significant promise for dynamic novel-view synthesis from monocular video input. Howev…
Physics-informed Temporal Alignment for Auto-regressive PDE Foundation Models
Congcong Zhu, Xiaoyan Xu, Jiayue Han +1
Auto-regressive partial differential equation (PDE) foundation models have shown great potential in handling time-dependent data. However, these models suffer from the shortcut pro…
STSA: Spatial-Temporal Semantic Alignment for Visual Dubbing
Zijun Ding, Mingdie Xiong, Congcong Zhu +1
Existing audio-driven visual dubbing methods have achieved great success. Despite this, we observe that the semantic ambiguity between spatial and temporal domains significantly de…