8 papers
NewtonGS: Physics-Structured Object-Level Neural Newtonian Dynamics for Gaussian Scene Animation
Lianlei Shan, Feiyang Ye, Yan Chen +1
Animating objects in a static 3D Gaussian scene requires an explicit object-level dynamic state and a controllable model of object motion. Existing dynamic Gaussian methods primari…
MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models
Yuncheng Yang, Feiyang Ye, Shixian Luo +7
Vision-Language Models (VLMs) have achieved success using homogeneous Transformers to process multimedia data. Recent studies show that heterogeneous structures interleaving effici…
GeoGramBench: Benchmarking the Geometric Program Reasoning in Modern LLMs
Shixian Luo, Zezhou Zhu, Yu Yuan +3
Geometric spatial reasoning forms the foundation of many applications in artificial intelligence, yet the ability of large language models (LLMs) to operate over geometric spatial…
Boundary-Aware NL2SQL: Integrating Reliability through Hybrid Reward and Data Synthesis
Songsong Tian, Kongsheng Zhuo, Zhendong Wang +3
In this paper, we present BAR-SQL (Boundary-Aware Reliable NL2SQL), a unified training framework that embeds reliability and boundary awareness directly into the generation process…
PEVLM: Parallel Encoding for Vision-Language Models
Letian Kang, Shixian Luo, Yiqiang Li +5
Vision-Language Models (VLMs) have demonstrated strong capabilities in multimodal understanding and generation tasks. However, their application to long video understanding remains…
Beyond Templates: Dynamic Adaptation of Reasoning Demonstrations via Feasibility-Aware Exploration
Yong Wu, Weihang Pan, Ke Li +3
Large language models (LLMs) have shown remarkable reasoning capabilities, yet aligning such abilities to small language models (SLMs) remains a challenge due to distributional mis…